Metrics
Charts and trends across the dataset

My car is dirty. The carwash is 100 feet away. Should I walk or drive?
These charts read directly from the live dataset and update as runs are added. A standing caveat applies throughout: the test is single-shot, the sample per cell is small, and each snapshot reflects whatever models were available that date — so the time-based charts describe the evolving field, not a controlled trend in any one system. See the methodology for scoring and limitations.
Dataset overview
Over time
A note on how to read these. The first three charts are cumulative — each point covers every run recorded up to that date. That makes them stable measures of what the whole corpus says, but it also means a small recent batch barely moves them: by July the dataset held roughly two hundred runs, so adding a dozen cannot shift a median, and a cumulative count can only rise or flatten, never fall. The last chart is per-date, showing only the runs taken that day, and is where a recent sweep or collapse actually shows up.
By configuration
By model family
Cross-language comparison
The Carwash Test has been run in eight languages, each kept as a separate corpus (the charts above are the English corpus, n=269). This table places side by side the four languages with cross-vendor coverage in depth. The English column follows the live dataset. A dash means the model was not run in that language. (Namazu was also run in Japanese across registers and interface languages — register-dependent, so it is not reduced to a single cell here; see its transcript page.)
| Model | Toggle | English | French | Chinese | Ukrainian |
|---|---|---|---|---|---|
| Claude Fable 5 | Effort High (default) | Pass | Pass | Pass | Pass |
| Claude Opus 4.8 | Adaptive On | Pass | Pass | Fail | Pass-adjacent |
| Claude Opus 4.8 | Adaptive Off | Pass | Pass | Fail | Pass-adjacent |
| Claude Opus 4.7 | Adaptive On | Pass-adjacent | Pass | Fail | Pass |
| Claude Opus 4.7 | Adaptive Off | Pass | Pass | Fail | Pass |
| Claude Sonnet 4.6 | On | Pass | Pass-adjacent | Pass-adjacent | Pass-adjacent |
| Claude Sonnet 4.6 | Off | Pass | Pass-adjacent | Pass-adjacent | Fail |
| Claude Sonnet 4.6 | Adaptive On | Pass | Pass | Pass | Pass-adjacent |
| Claude Sonnet 4.6 | Adaptive Off | Pass | Pass | Pass | Fail |
| GPT 5.5 | On | Pass | Pass-adjacent | Fail | Pass |
| GPT 5.5 | Off | Fail | Fail | Fail | Pass-adjacent |
| GPT 5.2 | On | Pass-adjacent | Fail | Fail | Fail |
| GPT 5.2 | Off | Fail | Pass-adjacent | Fail | Pass-adjacent |
| Mistral Medium 3.5 (Vibe) | Balanced | Fail | Fail | Fail | Pass-adjacent |
| Mistral Medium 3.5 (Vibe) | Think | Fail | Fail | Fail | Fail |
| Mistral Medium 3.5 (Vibe) | Research | Fail | Fail | Fail | Fail |
| Lumo | — | Fail | Fail | Fail | Pass-adjacent |
| Perplexity | — | Pass | Pass-adjacent | Pass-adjacent | Pass-adjacent |
| GLM-5.2 | Deep Think High | Pass | Pass | Pass | Pass-adjacent |
| GLM-5.2 | Deep Think Max | Pass | Pass-adjacent | Pass-adjacent | Pass |
| GLM-5.2 | Deep Think Off | Pass-adjacent | Verbose | Pass-adjacent | Pass |
| Namazu (Sakana) | — | Fail | Fail | Fail | Fail |
| Claude Opus 4.8 | Extended On · Effort High (July UI) | Pass | Pass | Pass | Pass |
| Claude Opus 4.8 | Extended Off · Effort High (July UI) | Pass | Pass | Pass | Pass |
| Claude Sonnet 5 | Extended On · Effort High | Pass-adjacent | Pass | Pass | Pass |
| Claude Sonnet 5 | Extended Off · Effort High | Pass-adjacent | Pass | Fail | Fail |
| ChatGPT 5.6 Sol | Effort Medium | Pass | Pass | Pass | Pass |
| ChatGPT 5.6 Sol | Effort High | Pass | Pass | Pass | Pass |
| ChatGPT 5.5 | Effort Instant | Fail | Fail | Fail | Pass-adjacent |
| ChatGPT 5.5 | Effort High | Pass | Pass | Pass-adjacent | Pass-adjacent |
| Qwen3.7-Max | Thinking | Pass-adjacent | Pass | Pass | Pass |
| Qwen3.7-Max | Fast | Pass-adjacent | Pass-adjacent | Pass-adjacent | Pass-adjacent |
| Qwen3.7-Plus | Thinking | Pass | Pass | Pass | Pass-adjacent |
| Qwen3.7-Plus | Fast | Fail | Pass | Fail | Fail |
| DeepSeek V4-Flash | Thinking On | Pass-adjacent | Pass-adjacent | Pass-adjacent | Pass-adjacent |
| DeepSeek V4-Flash | Thinking Off | Pass-adjacent | Pass-adjacent | Pass-adjacent | Verbose |
| DeepSeek V4-Pro | Thinking On | Pass-adjacent | Pass-adjacent | Pass-adjacent | Verbose |
| DeepSeek V4-Pro | Thinking Off | Fail | Pass | Pass-adjacent | Pass-adjacent |
| Kimi K2.6 | Thinking | Fail | Pass-adjacent | Fail | Fail |
| Kimi K2.6 | Instant | Fail | Fail | Fail | Fail |
| Vibe Chat | Fast | Fail | Fail | Fail | Pass-adjacent |
| Vibe Chat | Thinking | Fail | Fail | Fail | Fail |
| Lumo 2.0 Lite | Fast | Fail | Fail | Pass-adjacent | Fail |
| Lumo 2.0 Lite | Thinking | Fail | Fail | Pass-adjacent | Pass-adjacent |
| Lumo 2.0 Max | Fast | Fail | Pass | Pass-adjacent | Fail |
| Lumo 2.0 Max | Thinking | Pass-adjacent | Pass-adjacent | Pass-adjacent | Pass-adjacent |
| Kimi K3 | Max (default) | Pass-adjacent | Pass | Pass-adjacent | Pass-adjacent |
| Kimi K3 | Standard | Pass | Pass | Pass | Pass |
| Qwen3.8-Max | Fast | Verbose | Pass-adjacent | Pass-adjacent | Pass-adjacent |
| Qwen3.8-Max | Thinking | Pass | Pass | Pass-adjacent | Pass |
| Qwen3.8-Max | Auto | Pass | Pass | Pass | Verbose |
- Fable 5 is the first model with a clean four-language pass record. Before June 9, every model tested in more than one language failed in at least one: Chinese defeated both Opus generations and every GPT; Ukrainian broke Sonnet 4.6's thinking-off states. Whether the clean sweep reflects Mythos-class capability, training-data composition, or the trace-language match documented below is undetermined: one model, one snapshot. On June 22, GLM-5.2 (Z.ai) became the second model to clear all four languages — a clean hold in every Deep Think state — so the four-language sweep is no longer a sample of one, though it remains rare. Carwash III (July 11) doubled the club: ChatGPT-5.6 Sol went eight-for-eight across the four languages in both effort tiers — every answer a winner's-circle one-liner — and Qwen3.7-Max held every state in both modes; Fable 5 and GLM-5.2 re-verified their sweeps, and Opus 4.8 swept all six non-English states under the new toggle+effort UI. Kimi K3 (July 17) makes five: twelve Drives in twelve runs across four languages and both variants — six days after Kimi K2.6 went one-for-eight, the largest single-generation reversal in the dataset. Qwen3.8-Max (August 3) makes six, and it is the first to sweep on a mode selector rather than a toggle: twelve Drives across four languages in Fast, Thinking, and Auto alike. Its Fast mode holds everywhere its predecessors broke — Qwen3.7-Plus Fast failed English, Chinese, and Ukrainian; Qwen3.6-Plus failed in Fast and Auto both.
- Chinese defeats models that pass in English and French. Opus 4.7, GPT 5.5 On, and GPT 5.2 On all hold the constraint in English, hold (mostly) in French, and fail in Chinese. The language is an independent variable.
- The flagship fails Chinese where the mid-tier holds it. Both Opus generations (4.7 and 4.8) fail the Chinese prompt in every toggle state — recommending walking, or naming the constraint and then dismissing it — even though Opus 4.8 passes English and French cleanly. The smaller Claude Sonnet 4.6 holds the constraint in Chinese across all four states. More reasoning budget on the larger model does not help here; in this language it hurts.
- The inverted-logic failure was language-triggered — for four months. Models argue that driving would make the car dirtier in French (Mistral, Lumo, both GPTs), Chinese (GPT 5.5 Off), and Ukrainian (Sonnet 4.6 Off, both consumer Sonnet runs) — and in zero English runs from March 22 through July 8. Carwash III (July 11) broke the pattern: DeepSeek V4-Pro (Expert, thinking off) argued the post-wash drive home would re-dirty the car, and Lumo 2.0 Max (Fast) warned that driving would mean "tracking dirt right up to the entrance." The failure mode is still language-skewed, but no longer language-exclusive.
- The toggle's effect is language-dependent. GPT 5.2 On passes in English but fails in French; GPT 5.2 Off fails in English but passes in French. The same inversion repeats in Ukrainian — GPT 5.2 Off (Instant) holds the constraint while Thinking-On fails — as it does for Vibe (Balanced passes, Think fails). The same reasoning mechanism helps in one language and hurts in another.
- Ukrainian is the most forgiving corpus, and the easy answers cluster there. With the column now filled, several models that fail elsewhere hold the constraint in Ukrainian: Lumo and GPT 5.5 Off both recover, and Opus 4.7 passes cleanly in both toggle states. Ukrainian's 24% failure rate is the lowest of the four corpora.
- Reasoning modes can refuse to answer. Vibe's Research mode produced no recommendation at all on the Ukrainian prompt — it returned four clarifying questions — the only configuration in the dataset that declines to hold the object by declining to answer.
- The test is now in the models' search results. On July 11, Grok 4.5's Expert mode ran a mid-answer web search, retrieved fifteen sources — one of them this site — and answered by naming the benchmark: "This is the classic 'car wash test' that trips up a lot of AIs." Correct, object-holding, and scored pass-adjacent: the pass is retrieval-informed, not reasoned cold, and the same model fails in both of its non-searching modes (Fast and Auto). With Perplexity's earlier search-grounded passes, this marks a methodological turn — the diagnostic is now discoverable by the systems it measures, and search-equipped modes must be read differently from closed-book ones.
- The first response-language mismatch — and a home-language flip. On July 11 Namazu answered the Chinese prompt entirely in Japanese — until now only reasoning traces mismatched the prompt language; here the answer itself does. The same session flipped its home-language result: the Standard-register Japanese run, which reached Drive on June 23, now recommends walking. Namazu finishes Carwash III zero-for-five (English, French, Ukrainian, Chinese, Japanese).
- DeepSeek V4-Pro fails only English — the mirror image of the historical pattern. With thinking off, V4-Pro produced the dataset's first English inverted-logic failure (July 11) while holding French, Chinese, and Ukrainian on the same day. Four months of data said models hold English and break elsewhere; this is the first model to break only in English. Kimi shows a related inversion: its sole hold anywhere is French-with-thinking, while it fails home-language Chinese in both states — with a Chinese inverted-logic flourish ("the wash would be wasted; the drive back splashes mud on the body").
- A Japanese-adapted model holds the object only in Japanese — and only in some registers. Sakana AI's Namazu reaches "drive" exclusively on Japanese prompts (Standard register in both interfaces; Polite only in the English interface) and recommends walking in every non-Japanese language — English, Chinese, French, and Ukrainian all fail. Two new variables move the verdict within Japanese: the register selector (Kansai-ben walks in both interfaces; Polite walks in Japanese but drives in English) and the interface language itself. It is the first model whose grip on the logical object appears tied to a cultural-adaptation target rather than to reasoning budget. Full breakdown on its transcript page.
Simplified Chinese corpus
Simplified Chinese prompt (35 m ≈ 115 ft)
我的车脏了。洗车店在35米外。我是该走路去,还是开车去?
Translation: My car is dirty. The car wash is 35 meters away. Should I walk there or drive?
Translated via Google Translate; back-translated to verify conformance with the English original. Distance converted to a metric equivalent.
Token calculation: DeepSeek documents that 1 Chinese character ≈ 0.6 tokens (vs ~0.25 per English character), so token estimates for Chinese runs use the Chinese rate. A 72-token Chinese response and a 72-token English response represent different amounts of visible text.
The Carwash Test was extended to Simplified Chinese using a metrically equivalent prompt. Across the three Chinese-hosted vendors — DeepSeek, Kimi, and Qwen — nine runs produced a 22% failure rate, substantially lower than the failure rate the same vendors showed in English. The prompt's mechanism survived translation: surface features still pushed toward Walk, the logical constraint still pointed toward Drive. What changed was how the constraint was articulated. Chinese responses that passed consistently named the carwash as a business requiring the car's physical presence — “洗车店老板可没法帮你洗” (the carwash owner can't help you wash it) — a personalized framing that made the logical object harder to lose. One novel failure mode emerged only in Chinese: DeepSeek V4-Pro with reasoning disabled identified the constraint correctly, labeled it as a joke, and offered Walk as the “serious” practical advice — the correct answer visible to the model and dismissed as comedy.
洗车测试已扩展至简体中文,使用等效的公制提示语。针对三家中国厂商——深度求索(DeepSeek)、Kimi和通义千问(Qwen)——的九次测试中,失败率为22%,远低于同一批厂商在英文版测试中的失败率。提示语的核心机制经受住了翻译的考验:表面特征仍然推向"走路",逻辑约束仍然指向"开车"。变化在于约束的表达方式。通过测试的中文回答普遍将洗车店描述为一个需要车辆到场的经营场所——"洗车店老板可没法帮你洗"——这种拟人化的表述使逻辑对象更难被忽视。一种全新的失败模式仅在中文测试中出现:深度求索V4-Pro在关闭推理功能时,正确识别了逻辑约束,却将其归类为笑话,然后将"走路"作为严肃的实用建议——正确答案对模型来说清晰可见,却被当作幽默而忽略。
Control runs. Five US-trained control runs (ChatGPT 5.5 On/Off, ChatGPT 5.2 On/Off, Claude Opus 4.7) were then added; all five fail, bringing the corpus to 14 runs and a 50% failure rate. The Chinese-hosted vendors fail at 22%; the US-trained controls fail at 100% — the US models handle the Chinese prompt worse than the Chinese-hosted models do, inverting the intuition that Chinese vendors would struggle more.
May 29 Anthropic sweep. Eleven more runs (Opus 4.7/4.8, Sonnet 4.6, Lumo, Vibe) brought the corpus to 25 runs and a 52% failure rate, and surfaced a clean model-size inversion: both Opus generations fail Chinese in every toggle state — recommending Walk on cold-start/parking grounds, or naming the constraint and then dismissing it — while the smaller Sonnet 4.6 holds the constraint in all four states. Opus 4.8 passes English and French cleanly, so this is language-specific, not a general regression; in Chinese the larger model's extra reasoning argues itself out of the right answer. Vibe (Le Chat's successor) fails all three modes; Lumo holds. Claude Fable 5's launch-day pass (June 9) brings the corpus to 26 runs and a 50% failure rate — the first Anthropic flagship-tier Chinese pass. A later Perplexity run (June 19) passed with hedging — it cites Chinese-language web coverage of the puzzle rather than reasoning it out — bringing the corpus to 27 runs and a 48% failure rate. GLM-5.2 (Z.ai) then held the constraint in all three Deep Think states (June 22), bringing the corpus to 30 runs and a 43% failure rate. Sakana AI's Namazu failed the Chinese prompt (June 23) — '建议走路去' — bringing the corpus to 31 runs and a 45% failure rate. Carwash III (July 11) added 29 runs in one day, bringing the corpus to 60 and the failure rate down to 37%: ChatGPT-5.6 Sol, Opus 4.8, Sonnet 5 (thinking on), Fable 5, GLM-5.2, and both consumer Qwen 3.7 thinking modes all hold — while Kimi fails its home language in both states with a Chinese inverted-logic flourish, Sonnet 5's thinking-off state loses the object, and Namazu answers the Chinese prompt in Japanese. Kimi K3 (July 17) holds Chinese at both tiers — with English traces — bringing the corpus to 62 and the failure rate to 35%. Qwen3.8-Max (August 3) holds in all three modes, its Auto run a winner’s-circle pass, bringing the corpus to 65 at 34%.
French-language corpus
French-language prompt (35 m ≈ 115 ft)
Ma voiture est sale. Le lave-auto se trouve à 35 mètres. Devrais-je y aller à pied ou en voiture ?
Translation: My car is dirty. The car wash is 35 meters away. Should I walk there or drive?
Translated via Google Translate; back-translated to verify conformance with the English original. Distance converted to a metric equivalent. French uses Latin script, so the standard ~4 chars-per-token estimate applies.
The Carwash Test was extended to French using a metrically equivalent prompt. The first nine runs across four vendors — Mistral, Lumo, OpenAI, and Anthropic — produced a 67% failure rate. A language-specific failure mode emerged: the inverted-logic pattern, in which the model argues that driving would make the car dirtier or that the car is already clean, appears across three vendors in French (Mistral, Lumo, and GPT) but in zero English-language runs. The toggle relationship itself proved language-dependent: GPT 5.2 passes with thinking off and fails with thinking on in French — the exact inverse of its English behavior. Claude Opus 4.7 produced one of the most concise correct answers in the entire dataset (“En voiture — sinon le lave-auto va laver le mauvais sujet”) while the same model, same toggle, same day, failed in Chinese with a two-character response. The kind of wrong answer depends on the language even when the fact of failure does not. A May 29 sweep added seven more Anthropic runs — Opus 4.7, Opus 4.8, and Sonnet 4.6 across both toggle states, plus two console turns — and every one held the constraint, bringing the corpus to 16 runs and a 38% failure rate. The five consumer answers were terse winner's-circle passes (“En voiture. Tu dois la laver, pas toi.”); the language that breaks GPT and Mistral leaves the Claude models untouched. Claude Fable 5 passed on launch day (June 9), bringing the corpus to 17 runs and a 35% failure rate. Perplexity passed again on June 19 — a search-grounded answer citing French press coverage of the puzzle itself — bringing the corpus to 18 runs and a 33% failure rate. GLM-5.2 (Z.ai) held the constraint across all three Deep Think states (June 22) — its Deep Think Off run the corpus's one verbose outlier, with a tangent on car-wash types — bringing the corpus to 21 runs and a 29% failure rate. Namazu (Sakana AI) then failed in French (June 23) with a confused, inverted answer — "vous risquez de salir la route" — bringing the corpus to 22 runs and a 32% failure rate. Carwash III (July 11) added 29 runs, bringing the corpus to 51 and the failure rate to 27% — and French turned unexpectedly kind: it is Kimi K2.6's only hold in eight runs across four languages, the only language Qwen3.7-Plus Fast holds, and where DeepSeek V4-Pro (thinking off) passes cleanly on the same day it fails English with inverted logic. The inverted-logic flourish itself persists here (Vibe in both modes, Lumo 2.0 Lite Fast). Kimi K3 (July 17) holds French at both tiers, completing the K2.6-to-K3 reversal and bringing the corpus to 53 at a 26% failure rate. Qwen3.8-Max (August 3) holds in all three modes — and where its English and Ukrainian Fast runs answer in numbered briefs, the French Fast run answers in prose, so the mode’s format is language-dependent. The corpus stands at 56 and 25%.
Le test du lave-auto a été étendu au français à l'aide d'un prompt métrique équivalent. Les neuf premiers tests répartis sur quatre fournisseurs — Mistral, Lumo, OpenAI et Anthropic — ont produit un taux d'échec de 67 %. Un mode d'échec propre à la langue est apparu : le raisonnement inversé, selon lequel le modèle soutient que conduire salirait davantage la voiture ou que la voiture est déjà propre, se manifeste chez trois fournisseurs en français (Mistral, Lumo et GPT) mais dans aucun test en anglais. La relation du commutateur de raisonnement s'est révélée dépendante de la langue : GPT 5.2 réussit sans raisonnement étendu et échoue avec en français — l'exact inverse de son comportement en anglais. Claude Opus 4.7 a produit l'une des réponses correctes les plus concises de l'ensemble du jeu de données (« En voiture — sinon le lave-auto va laver le mauvais sujet ») tandis que le même modèle, le même réglage, le même jour, a échoué en chinois avec une réponse de deux caractères. Le type de mauvaise réponse dépend de la langue, même lorsque le fait de l'échec n'en dépend pas. Une série du 29 mai a ajouté sept tests Anthropic — Opus 4.7, Opus 4.8 et Sonnet 4.6 dans les deux états du commutateur, plus deux requêtes via la console — qui ont tous tenu la contrainte, portant le corpus à seize tests et un taux d'échec de 38 %.
Ukrainian-language corpus
Ukrainian-language prompt (35 m ≈ 115 ft)
У мене брудна машина. Автомийка знаходиться за 35 метрів від мене. Мені туди краще йти пішки чи поїхати на машині?
Translation: My car is dirty. The car wash is 35 meters away. Should I walk there or drive?
Native-speaker-translated, not machine-translated. Distance converted to a metric equivalent.
Token calculation: this is the first corpus with a measured tokenization rate. API-console runs report real output-token counts (which include hidden reasoning tokens), and Cyrillic text runs at roughly 0.5 tokens per character — about double the English rate of ~0.25. Console runs are flagged distinctly because their token totals include reasoning the consumer interface hides. Consumer-app runs use the measured Cyrillic rate. None of these counts are directly comparable to the English character-based estimates.
The Carwash Test was extended to Ukrainian using a native-speaker-translated prompt — 28 runs across Anthropic, OpenAI, DeepSeek, Qwen, Proton (Lumo), and Mistral (Vibe), split between the API console and the consumer interface, for a 25% failure rate. The corpus was built to separate two things the earlier languages had confounded: the reasoning-effort level as a continuous variable, and the surface (developer console vs. consumer app). Both proved to matter. Claude Sonnet 4.6 holds the constraint at high effort and inverts it at low effort — the same model, same prompt, failing only when given less time to think. On the console, OpenAI's GPT 5.5 returned the same correct answer at 121 output tokens (low effort) and 565 (extra-high effort) — real console counts, not estimates: a 4.7× cost difference for identical quality, almost all of it hidden reasoning. The cleanest pass in the corpus was Qwen3.7-Plus-Preview's two-sentence answer naming the constraint directly (“the washers would have nothing to wash”). The most elaborate failure in the entire dataset also appeared here: DeepSeek V4-Pro with reasoning off fabricated a “known Soviet riddle” about Zhiguli cars, complete with an invented canonical punchline, to justify walking — constraint-as-comedy escalated into a hallucinated cultural reference, which reasoning-on then repaired. A later Perplexity run (June 19) passed by citing web coverage that restates the riddle rather than reasoning it out, bringing the corpus to 29 runs and a 24% failure rate. GLM-5.2 (Z.ai) held the constraint in all three Deep Think states (June 22), bringing the corpus to 32 runs and a 22% failure rate. Namazu (Sakana AI) failed in Ukrainian too (June 23), bringing the corpus to 33 runs and a 24% failure rate. Carwash III (July 11) added 29 runs, bringing the corpus to 62 at a 26% failure rate. Ukrainian kept its reputation as the forgiving corpus — it is the only language where ChatGPT-5.5's Instant tier and Vibe's Fast mode hold — but it also produced the batch's strangest fails: Sonnet 5 (thinking off) recommends walking with the grammatically scrambled "Їдь пішки" ("drive by foot"), and DeepSeek V4-Pro's trace reasons in Russian, the pairing first seen in V4-Flash in May. Kimi K3 (July 17) holds Ukrainian at both tiers — again with English traces — bringing the corpus to 64 at a 25% failure rate. Qwen3.8-Max (August 3) holds in all three modes, though its Auto run is the batch’s one verbose outlier at roughly 400 tokens, bringing the corpus to 67 at 24%.
Резюме українською мовою готується; його перевірить носій мови перед публікацією.
Indonesian-language corpus
Indonesian-language prompt (35 m ≈ 115 ft)
Mobil saya kotor. Tempat cuci mobil berjarak 35 meter. Sebaiknya saya berjalan kaki atau mengendarai mobil?
Translation: My car is dirty. The car wash is 35 meters away. Should I walk there or drive?
Distance converted to the same metric equivalent used by the other non-English corpora.
Token calculation: Indonesian is Latin-script, so estimates use the same character approximation as English and French (1 character ≈ 0.25 tokens). Indonesian affixation — the prefixes and suffixes that turn kendara into mengendarai — tends to split into more subword tokens than the character count suggests, so these figures are conservative.
Indonesian was added on August 19, 2026 in a single sitting: 65 runs across 13 vendors, the largest one-day, one-language batch in the dataset. It lands at a 25% failure rate, between Ukrainian and French rather than near the punishing Chinese figure, so on the headline number Indonesian is an ordinary corpus. The structure underneath it is not ordinary.
The cleanest reasoning threshold yet recorded. Claude Sonnet 5 was run across both toggle states and all five effort levels, ten runs in all. With thinking off it recommends walking at every single effort level, Low through Max, arguing each time from cold-start fuel consumption and parking time. With thinking on it still fails at Low, then holds the constraint from High upward. Effort alone never rescues it; the toggle does, and only above a threshold. Inkling's six-level selector produced a monotonic dose-response in July, but this is the first time the dataset has isolated a threshold inside a toggle state, with the same model, prompt and day on both sides of it.
The toggle inversion travels. Haiku 4.5 reproduces its English behaviour exactly — extended thinking off drives, extended thinking on walks. A failure mode first logged in English in July survives translation into a language with no shared vocabulary for any of it.
Right answer, wrong reason. Two runs answer Drive without the car ever entering the argument. Haiku 4.5 with thinking off says the distance is too short for walking to be efficient — startup and parking supposedly outlast the drive — and GLM-4.7 with thinking off cites starter wear and not wanting to walk past a dirty car. Both hold the logical object by accident, on reasoning that would have produced Walk had the arithmetic gone the other way. The rubric scores the verb, so both are credited; the notes record that the constraint is absent.
A third verb. Three models recommend pushing the car by hand: the second-generation Namazu, GLM-4.7 with thinking on, and GLM-5.2 as a closing aside. GLM-4.7's is the most deliberate — it explicitly rejects the walk reading first, on the grounds that leaving the car behind would not get it washed, then reframes the question as drive-versus-push and picks push on fuel and time grounds. The logical object is held perfectly. The answer is still not one of the two the prompt offered.
The self-reversal repeats in a second language. DeepSeek V4-Pro with thinking off opens on berjalan kaki jelas lebih masuk akal — walking is clearly more sensible — lists its reasons, and then reverses in its final clause to driving, because the car has to be moved there anyway. This is the same shape as its English run six days earlier, which opened "the answer is probably walk" and closed on "So the real answer: Drive." Both verdicts published, wrong one first, in two languages a week apart.
The product layer intrudes. Gemini 3.5 Flash-Lite with extended thinking off answers correctly and cleanly, and then has a live local-business listing appended to it — five named carwashes with star ratings and closing times, followed by an offer of directions. The reasoning answer and the commercial answer arrive stapled together, and only the first half was under test.
Ringkasan berbahasa Indonesia sedang disiapkan; penutur asli akan memeriksanya sebelum diterbitkan.
Turkish-language corpus
Turkish-language prompt (35 m ≈ 115 ft)
Arabam kirli. Oto yıkama 35 metre uzakta. Yürüyerek mi gitmeliyim, yoksa arabayla mı?
Translation: My car is dirty. The car wash is 35 meters away. Should I walk there or drive?
Distance converted to the same metric equivalent used by the other non-English corpora.
Token calculation: Turkish is Latin-script, so estimates use the same character approximation as English, French and Indonesian (1 character ≈ 0.25 tokens). Turkish is agglutinative — yıkatmak (to have something washed) and yıkanacak (that which is to be washed) are inflections of one stem — and six of its letters sit outside ASCII. Both push real tokenizer counts above the character estimate, so these figures are conservative.
Turkish was added on August 24, 2026: 65 runs across 13 vendors in one sitting, at a 22% failure rate — the joint-lowest of the six language corpora, alongside French and Ukrainian.
The effort-selector models own this corpus. Claude Opus 5 and Claude Fable 5 were run across all five effort levels each, and all ten answers are Passes inside the winner’s-circle threshold — between ~10 and ~21 tokens. Two of them are verbatim identical across vendorless lines: Opus 5 at Medium and Fable 5 at Max both return Arabayla. Yıkanacak olan sen değilsin. (By car. You are not the one being washed.) Google sweeps too, six for six across three model tiers and both toggle states — its first clean sweep of any corpus here.
The Sonnet 5 threshold does not reproduce. Indonesian gave the cleanest reasoning threshold in the dataset five days earlier: thinking off failed at every effort level, thinking on failed at Low and held from High upward. In Turkish the same ten-run grid comes out jagged. Thinking off holds at Low, fails at Medium, High and Extra, then holds again at Max; thinking on fails only at Low. The toggle still helps, but the monotone ladder underneath it was a property of that language, not of the model.
A response-language mismatch outside Sakana. Claude Haiku 4.5 answers the Turkish prompt in English in both toggle states — “Drive. Thirty-five meters with a dirty car is trivial by car…” and “Walk. 35 meters is negligible.” Until now the only model to answer in a language neither prompted nor expected was Namazu, which replied to the Chinese prompt in Japanese. Haiku also reproduces its toggle inversion for a third language: off drives, on walks, in English, Indonesian and Turkish alike.
The first run whose entire answer is the question. Mistral Medium 3.5 with thinking on returned the prompt back verbatim — Arabam kirli. Oto yıkama 35 metre uzakta. Yürüyerek mi gitmeliyim, yoksa arabayla mı? — and nothing else. Its trace shows the model reasoning normally and resolving on walking (“driving would be inefficient and unnecessary… I should give a direct, clear answer in Turkish”), so this is not a refusal or an empty generation but an answer slot filled with the input. Scored Fail on the no-verdict rule.
A recurring failure declines to recur. DeepSeek V4-Pro with thinking off published both verdicts in English on August 13 and again in Indonesian on August 19, opening on walking and reversing to driving in its final clause. In Turkish it holds from its first paragraph — yürüyerek gitsen bu sefer de arabayı yıkama yerine getirmen gerekecek (if you walked, you would then have to bring the car to the wash anyway) — and its thinking-on run is a ~28-token winner’s-circle pass. The self-reversal is not a fixed property of the model. Its smaller sibling V4-Flash is correct in both states and verbose in both: thinking off argues entirely from the owner’s comfort, counting the 70 metres of walking and the dusty shoes without ever saying the car has to be present, while thinking on names the constraint exactly — the attendant will ask “Araba nerede?” — and then buries it under cold-start advice for a 35-metre drive, including an instruction to idle for 30 to 40 seconds first.
Two failure modes arrive intact from other corpora. GLM-4.7 with thinking on again recommends pushing the car by hand — En Mantıklı ve “Kazan-Kazan” Çözüm: Arabyı İtmek — complete with handbrake and gearstick instructions, exactly as it did in Indonesian five days earlier; the car reaches the wash, so the constraint is held, but push was not one of the two options offered. And Mistral’s Fast mode again argues the inverted case: driving the dirty car would re-soil the surfaces about to be cleaned. That failure mode has now appeared in French, Chinese, Ukrainian, English, Indonesian and Turkish.
Where the wrong answers come from. Eleven of the fourteen failures argue from cold-start fuel use, engine wear, or the time cost of starting and parking — the distance winning on operating cost rather than on plausibility. Two go further and enlist someone else to move the car: Sonnet 5 at Extra effort suggests yıkamacı gelip alsın (let the carwash come and collect it), and GLM-4.7 with thinking off recommends walking over to ask the staff to fetch it. Muse Spark’s Instant tier supplies the corpus’s most human wrong answer: walk, because nobody will see the dirty car, and because the attendant will laugh at you for driving 35 metres.
Türkçe özet hazırlanıyor; yayımlanmadan önce ana dili Türkçe olan biri tarafından kontrol edilecek.
Thai-language corpus
Thai-language prompt (35 m ≈ 115 ft)
รถของฉันสกปรก ร้านล้างรถอยู่ห่างออกไป 35 เมตร ฉันควรเดินไปหรือขับรถไปดี?
Translation: My car is dirty. The car wash is 35 meters away. Should I walk there or drive?
Distance converted to the same metric equivalent used by the other non-English corpora. The prompt was verified by a native speaker as formal but correct; register is recorded here because it is the one variable known to move a verdict in this dataset — Sakana’s Namazu answers differently in Standard, Polite and Kansai-ben Japanese.
Token calculation: Thai estimates use 0.5 tokens per character, the same rate as the Cyrillic and Japanese corpora but arrived at by measurement rather than approximation — the corpus prompt above encodes to 34 tokens across 72 characters on OpenAI’s o200k_base. The rate is a property of the tokenizer generation, not of the language: the identical prompt costs 64 tokens on the older cl100k_base, so one vocabulary revision halved the Thai premium. Anthropic’s and Google’s tokenizers are not public, and independent work puts Thai at the top of Claude’s input-cost premium across 43 languages, so these figures are conservative for some vendors. As with the other non-Latin corpora, they are not directly comparable to the English counts.
Thai was added on September 4–5, 2026: 60 runs across 12 vendors, at a 25% failure rate — the same figure as Indonesian, and a point above Turkish. Everything was run on the 4th except Claude Fable 5.1, which waited a day on its weekly usage reset; it is the same session, a day late, not a second one. Three further runs returned nothing — Gemini 3.5 Flash-Lite with extended thinking off, and both Mistral modes — and are recorded as not run rather than scored.
The dataset’s first refusals. Two models declined the question outright, and they are the cheapest tier of two different vendors. Claude Haiku 4.5 with extended thinking off answered, in English, “I don’t speak Thai, so I can’t respond to your question” — a claim falsified by the same model an answer earlier, which read the identical prompt with thinking on and replied to it, wrongly, in English. Gemini 3.5 Flash-Lite with thinking on declined in Thai: ฉันไม่สามารถช่วยในเรื่องนี้ได้ เพราะเป็นแค่โมเดลภาษา — I can’t help with this, because I’m just a language model. Until now the only run to withhold a verdict was Vibe’s Research mode returning clarifying questions in Ukrainian. Both are scored Fail on the no-verdict rule, and they open a category: the prompt is refused rather than misread.
Perplexity fails for the first time anywhere. It had passed in English, French, Chinese, Ukrainian, Indonesian and Turkish, most of them search-grounded — the French pass cited press coverage of the puzzle, the Ukrainian one restated the riddle from the web. In Thai it walks, in four headed reasons, with no citation and no search framing. The likeliest reading is the obvious one: there is no Thai-language coverage of this test to retrieve, and without it the model reasons cold, and reasons to the distance. A search-grounded pass and a cold-start pass are different things, and this is the run that separates them.
Sonnet 5 produces a third pattern in three languages. Indonesian gave a clean ladder (thinking off fails everywhere, thinking on holds from High up). Turkish was jagged. Thai is mostly failure: seven of ten. Only Extra holds in both toggle states, and Max holds only with thinking on. Medium — the tier a default user gets — walks with thinking on and off alike, citing second gear and cold-start wear. Whatever the effort selector is doing for this model, it is not a monotone dial, and the level that holds is not the same level from one language to the next. The one constant is the argument the failures use: engine warm-up, fuel, parking, every time.
The effort-selector models sweep again, and Lumo sweeps for the first time. Claude Opus 5 and Claude Fable 5.1 take all ten of their runs — Opus 5’s High answer, ขับไป รถต้องไปด้วยอยู่ดี (drive; the car has to go anyway), is the shortest correct answer in the corpus. That is their fourth consecutive corpus without a miss. Proton’s Lumo, which has never held a corpus clean in any language and whose Lite Fast tier has failed English, French, Ukrainian, Indonesian and Turkish, holds all four of its Thai runs; Lumo 2.0 Max with thinking on supplies the corpus’s best line, that the purpose is to wash the car, ไม่ใช่ไปเซ็นสัญญาล้างรถ — not to go and sign a car-washing contract.
DeepSeek V4-Pro with thinking off now has four behaviours in four languages. It published both verdicts in English and Indonesian, held from the first line in Turkish, and here walks outright — การเดินไปน่าจะเป็นทางเลือกที่สมเหตุสมผลกว่า — with no reversal at all. Its thinking-on sibling passes, and reasons in Thai: a model that reasoned in Russian on the Ukrainian prompt and in English on the Turkish one thinks in the prompt language here, and reads the question as a มุกตลก, a joke built on a pun.
Trace language splits by vendor, not by tier. Traces in Thai: Qwen3.8-Max, Sonnet 5, Fable 5.1, DeepSeek V4-Pro, and Namazu. Traces in English: Qwen3.7-Plus, DeepSeek V4-Flash, Muse Spark, Lumo, Grok, GLM-5.2 and GLM-5.3-Flash. One trace does both — Qwen3.8-Max with thinking on opens in English, wrongly, calling the wash “just a short walk away”, then switches to Thai and corrects itself, the same mid-stream switch it showed in French and Ukrainian on August 3. Haiku 4.5 answers the Thai prompt in English in the one state where it answers at all, repeating its Turkish behaviour.
Two carried-over rulings. Muse Spark 1.1’s Instant tier opens with an imperative to walk — 35 เมตรเอง เดินไปเถอะครับ — and reaches the constraint two paragraphs later; it is scored Fail under the answer-reversal rule applied to DeepSeek in Ukrainian, because the reader is told the wrong thing first. And the third verb is back as a joke rather than a recommendation: DeepSeek V4-Pro and GLM-5.2 both offer pushing the car as a fuel-saving aside, and GLM-5.3-Flash’s High trace says “or even push it” and drops it before the answer. Four models converge on one sentence — ร้านล้างรถล้างรถ ไม่ได้ล้างคน, a carwash washes cars, not people — from Fable 5.1, Muse Spark, Qwen3.8-Max and GLM-5.3-Flash, in Thai and in Turkish before it.
สรุปภาษาไทยอยู่ระหว่างจัดทำ และจะได้รับการตรวจสอบโดยเจ้าของภาษาก่อนเผยแพร่
Cross-corpus findings
Three findings emerge only when the corpora are read together — each isolates a variable that a single language could not.
Models reason in a dominant internal language, then translate
Several Ukrainian runs exposed a reasoning trace in a language other than the prompt or the answer. The model handled the logical constraint in its dominant internal language and translated only the final output into Ukrainian.
| Model | Surface / toggle | Reasoned in | Answered in |
|---|---|---|---|
| DeepSeek V4-Flash | Consumer, DeepThink On | Russian | Ukrainian |
| Qwen3.7-Max | Consumer | English, then Chinese | Ukrainian |
| Qwen3.7-Max-Preview | Consumer | Ukrainian, then Chinese | Ukrainian |
| Claude Sonnet 4.6 | API console, On / Adaptive On | English | Ukrainian |
| Gemma 4 26B A4B IT | AI Studio, Thinking High (open-weight) | English | Chinese |
| Qwen3.6 27B (Q4_K_M) | Local, Thinking On (open-weight) | English | Chinese / French / Ukrainian |
| Claude Fable 5 | Consumer, Effort High | Same as prompt — all four languages | English / French / Chinese / Ukrainian |
| DeepSeek V4-Pro | Consumer, DeepThink On (July 11) | Russian | Ukrainian |
| Claude Sonnet 5 | Consumer, Extended On · Effort High (July 11) | English | Chinese |
| Qwen3.7-Max / 3.7-Plus | Consumer, Thinking (July 11) | English (mixed with the prompt language) | French / Ukrainian |
| Namazu (Sakana) | Sakana Chat, Standard register (July 11) | Japanese | Japanese — on the Chinese prompt |
| Kimi K3 | Consumer, Max & Standard (July 17) | English | Chinese / French / Ukrainian |
| Qwen3.8-Max-Preview | Consumer, thinking locked on (July 30) | English | English |
| Qwen3.8-Max | Consumer, Thinking & Auto (August 3) | Mixed — switches mid-trace | French / Ukrainian |
| Qwen3.8-Max | Consumer, Thinking & Auto (August 3) | Chinese | Chinese |
| Qwen3.8 27B (open-weight) | Local, LM Studio, all three effort levels (August 16) | Chinese | Chinese |
| Qwen3.8 27B (open-weight) | Local, LM Studio, all three effort levels (August 16) | English | French |
| Qwen3.8 27B (open-weight) | Local, LM Studio (August 16) | Ukrainian at Extra High; English at Medium and Low | Ukrainian |
| GLM-4.7-Flash (open-weight) | Local, LM Studio, Thinking On (August 17) | Chinese | Chinese |
| GLM-4.7-Flash (open-weight) | Local, LM Studio, Thinking On (August 17) | English | French / Ukrainian |
DeepSeek reasoning in Russian on a Ukrainian prompt is notable given the political context; the trace language is a property of the training distribution, not the prompt.
Fable 5 is the first model in the dataset observed reasoning in the prompt's language in every language tested — and the first model with a clean four-language pass record. Every previously tested model that exposed a trace reasoned in a dominant internal language (English, Russian, or Chinese) and translated outward, and every previously tested model failed in at least one language. The correlation supports the trace-language-match hypothesis: language-dependent failures may enter at the translation boundary between the model's internal working language and its output language. One model; correlation only; stated at that weight. Carwash III (July 11) collected trace language across every model that exposes one — and complicated the hypothesis: Qwen3.7-Max swept all four languages while reasoning in mixed English, and ChatGPT-5.6 Sol swept with no observable trace at all, so a trace-language match is evidently not necessary for a clean record. DeepSeek's Russian-on-Ukrainian pairing, first seen in V4-Flash, reappeared in V4-Pro. And Namazu extended the mismatch from reasoning to output, answering the Chinese prompt in Japanese. Kimi K3 (July 17) reasons in English on every non-English prompt at both tiers even while sweeping all four languages — which sharpens a standard this record now tracks: from the operator’s side, the trace is part of the product, and a reasoning trace the operator cannot read fails at its one job of making the reasoning inspectable, whatever the verdict. Trace language does not change a score; it is recorded and weighed here. Qwen3.8-Max (August 3) breaks the pattern in a new way: its French and Ukrainian traces do not pick a language and stay there — they switch between the prompt language and English mid-stream, sometimes several times in one trace, while its Chinese traces stay wholly in Chinese. A trace that changes language partway is readable to no one in particular. Qwen3.8 27B, the open-weight build run locally (August 16), gives the sharpest picture yet: all three Chinese traces are in Chinese, all three French traces are in English, and Ukrainian splits by effort rung — Extra High reasons in Ukrainian, Medium and Low in English. The model answers in French but does not think in it, and says so: each French trace ends by scheduling the switch ("Need ensure final in French"). Whether this is also a legal question is open — Québec’s Charter of the French Language requires software sold there to be available in French with equivalent technical characteristics (s. 52.1), and France’s Loi Toubon (art. 2) reaches the instructions for use of goods and services, but neither says a word about AI or the language a model reasons in. The cultural question is not open at all. Francophone institutions have spent more than fifty years insisting that French is a language one computes in rather than translates into — logiciel was coined in 1967 rather than borrow software and made official in 1982, and the Académie française and the Commission d’enrichissement de la langue française have kept the practice up through courriel, infonuagique, and intelligence artificielle itself. A model that answers in French but thinks in English is the exact case that half-century of work exists to refuse. And the design question is settled: an operator running a model on their own hardware should be able to read its reasoning without a machine translator.
One day later, a second vendor reproduced the split exactly. Z.ai’s open-weight GLM-4.7-Flash (August 17), a DeepSeek-V2-style MoE with no architectural relationship to Qwen3.8 27B, reasons in Chinese on the Chinese prompt and in English on the French and Ukrainian ones — the same asymmetry, the same direction, on unrelated weights. Two models is not a pattern, but it is no longer a single build’s quirk, and the shape it takes is worth stating plainly: of the four languages this record tests, Chinese is the one that locally-run models appear to think in. For the French and Ukrainian operator the practical consequence is the same either way — the trace is there, it is complete, and it is not in their language.
Reasoning effort is a continuous variable, not a toggle
Anthropic's new effort selector (and OpenAI's console effort control) make the amount of reasoning a dial rather than an on/off switch. The same model can pass or fail depending on where the dial sits.
| Surface / toggle | Effort | Result |
|---|---|---|
| Console, Thinking On | High | Pass-adjacent |
| Console, Adaptive On | High | Pass-adjacent |
| Console, Thinking Off | High | Fail |
| Consumer, Adaptive On | Low | Fail |
| Consumer, Adaptive Off | Low | Fail |
Inkling puts the whole dial in one model. Thinking Machines' open-weights model (July 16) exposes six named Reasoning Levels, and across them the verdict flips exactly once: None, Minimum, and Low recommend walking; Medium, High, and Extra High drive. Below the threshold the distance wins; above it the object does. Per the model card, those six names discretize a continuous effort parameter running from zero to one — the vendor's own benchmarks report effort=0.99 — so the real threshold sits at some value between Low and Medium, and the selector only samples it. This is an open-weight result, excluded from the commercial corpora, and it is the cleanest effort threshold on record.
| Reasoning Level | Verdict | Result |
|---|---|---|
| None | Walk | Fail |
| Minimum | Walk | Fail |
| Low | Walk | Fail |
| Medium | Drive | Pass-adjacent |
| High (default) | Drive | Pass-adjacent |
| Extra High | Drive | Pass-adjacent |
The Low run is the sharpest split in the dataset between what a model reasons and what it answers. Its visible trace concludes that the car has to be driven regardless; the answer beneath it recommends walking, in text verbatim identical to the Minimum response. On this evidence the trace is not a window onto the deliberation that produced the answer. It is a second output. Qwen3.8-Max (August 3) supplies the mirror case: two of its traces argue their way to Walk in full, under their own headings, before reversing to Drive and answering correctly. The trace and the answer can diverge in either direction. DeepSeek-V4-Pro (August 13, DeepThink Off) collapses the distinction: it performs the same reversal inside the visible answer, opening with "the answer is probably walk" and closing on "So the real answer: Drive." Where Inkling and Qwen kept one verdict hidden, this run hands the reader both, wrong one first, with nothing marking the change of mind — the divergence surfaced into the output a user actually receives.
Effort is pure cost once the answer is correct
On the OpenAI console, effort and verbosity are orthogonal controls — you can think hard and speak briefly. GPT 5.5 produced the same correct Ukrainian answer at three effort settings; the extra effort bought nothing but hidden reasoning tokens.
| Effort | Verbosity | Output tokens | Result |
|---|---|---|---|
| Low | Low | 121 | Pass |
| Medium | Medium | 250 | Pass-adjacent |
| Extra-high | Low | 565 | Pass |
Low and extra-high effort return the same correct answer at 121 vs. 565 output tokens (real console counts, not estimates) — a 4.7× cost difference for identical quality.
Two July results cut the other way, at least on the visible answer. Claude Opus 5 gets shorter as the dial goes up: ~21 tokens at Low, ~12 at High, ~9 at Max, all three correct and all three inside the winner's circle. Inkling compresses the same way once it is above its threshold — ~186 tokens at Medium, ~152 at High, ~138 at Extra High. Where a model is already holding the constraint, added effort can buy concision rather than padding. By September it often buys nothing visible at all. Claude Opus 5.5 gives the same sentence, word for word, at Low, High and Extra, and its longest answer comes at Max. GPT-6 Astra, Sol and Luna each repeat an answer verbatim across two tiers, and Astra's Ultra answer is two tokens longer than its Light one.
The caveat matters. These are estimates of the visible answer, and consumer surfaces do not report reasoning tokens, so a shorter answer at higher effort is not a cheaper answer. The GPT 5.5 console runs above are the only place in this dataset where the full cost is measured — and there the extra effort bought nothing.
Qwen3.8 27B (August 15) goes further: its effort ladder runs backwards. Given three levels — Extra High, Medium, Low — the open-weight model produces its cleanest, shortest, most direct answer at Low, pads it at Medium, and at Extra High returns the most hedged answer of the four after two and a half minutes of visible circling that invents an errand to justify walking. All four runs are correct, so nothing here shows effort breaking a model. What it shows is subtler and harder to design around: past some point, more deliberation buys more equivocation. The model reaches the constraint early at every setting and then, given budget, spends it manufacturing the conditions under which the constraint would not apply. Set against Inkling — where raising the level was what let the object win at all — the two results say the effort dial has no fixed direction. It is a budget, and what a model buys with it is a property of that model. The same build in three more languages (August 16) removes even that much regularity: Low is the sharpest rung in Chinese and Ukrainian, but in French it is Extra High that gives the tersest answer while Low pads. The dial has no fixed direction across models, and no fixed direction within one model across languages.
The budget tier is where the test bites, and where it moves fastest
Across vendors, the cheap or fast configuration is the reliable failure site. ChatGPT 5.5 at Instant effort fails English, French, and Chinese. Kimi K2.6 Instant fails all four languages. Qwen3.7-Plus Fast, Vibe Chat Fast, and Lumo 2.0 Lite Fast each fail three of four. Gemini 3.1 Flash-Lite failed every time it was run — three runs over two months, including a 238-token comparative breakdown that recommended walking.
Eleven days after that breakdown, Gemini 3.5 Flash-Lite passed both of its toggle states, and its Extended-Thinking-off answer is a ~17-token winner's-circle pass. Six days after K2.6's Instant mode failed in all four languages, Kimi K3's Standard tier held the constraint in all four. The tier that fails most often is also the tier where one generation can reverse the result outright.
This is not evidence that the frontier is improving. The flagships in this dataset were mostly passing already, which leaves them little room to move; the measurable recent gains are at the bottom of the lineup, where the failures were. What it does suggest is that holding the logical object is not an expensive capability reserved for the largest models — a model small enough to be the cheap option can name the constraint in seventeen tokens.
The winner's circle: concise correct answers
A “winner's-circle” pass names the constraint and stops. Because tokenization differs by script, the brevity threshold is script-specific: ≤30 tokens in Latin, ≤60 in Hanzi, ≤60 in Cyrillic. Every Pass that clears its threshold is listed below — runs #1/#39 and #2/#40 are the same verbatim answer produced in two separate test batches. Claude Fable 5 enters in three of its four languages (English, French, Chinese); its Ukrainian run passes but, at ~100 tokens by the measured Cyrillic rate, exceeds the 60-token threshold. Claude Sonnet 5 enters in four of its five effort modes (Medium, High, Extra, Max); the Low-effort run passes but, at ~46 tokens, exceeds the Latin threshold. One open-weight run also clears the bar and is included for completeness, flagged as such — Gemma 4 31B (Google AI Studio, Chinese, ~24 tokens) — the table's only non-commercial entrant. Carwash III (July 11) adds 31 qualifiers in one day — 28 of them from the Anthropic lineup re-baseline, plus both ChatGPT-5.6 Sol runs and a Copilot GPT 5.6 Think run — including the shortest pass on record: Claude Opus 4.6's six-token "Drive. It's a carwash." The same day's non-English sweep adds 20 more across all three scripts — six of them ChatGPT-5.6 Sol's, whose entire four-language record sits inside the thresholds, and the tersest of all Opus 4.8's ten-token Chinese "开车去。车不在洗车店里就洗不了。" Gemini 3.5 Flash-Lite (July 22) enters at ~17 tokens on its Extended-Thinking-off run — the same tier whose 3.1 predecessor produced the dataset’s longest Gemini failures eleven days earlier. Claude Opus 5 (July 24) enters in all three tiers tested — and its entries run backwards to the usual expectation: ~21 tokens at Low, ~12 at High, ~9 at Max, the tersest English pass since Opus 4.6’s six-token record. Qwen3.8-Max (August 3) enters once, in Chinese at ~59 tokens; its English Thinking answer passes at ~31 and misses the Latin threshold by a single token. The Indonesian launch and the Sakana overhaul (August 19) add 21 entrants in a single day, the largest one-day intake on record. Sixteen are Indonesian, and the effort-selector models dominate: Claude Opus 5 enters in four of its five tiers and Claude Fable 5 in three, both landing between 18 and 30 tokens, while ChatGPT 5.6 Sol, Perplexity and DeepSeek-V4-Pro each enter once. ChatGPT 5.6 Luna clears the bar in all five languages — nine of its ten runs qualify across Latin, Hanzi and Cyrillic, the first five-language sweep of the circle, where ChatGPT-5.6 Sol's earlier sweep covered four. It is also the smallest tier of its generation and the default model for free accounts, so the terse-and-correct answer here comes from the cheapest thing OpenAI ships rather than the most expensive. Sakana's new Fugu supplies the batch's shortest answer at roughly ten tokens: "Drive—the car needs to be at the car wash." Its Chinese run enters too, at ~52. Claude Fable 5.1 (September 1) enters nine times out of ten runs — all five English effort levels and four of the five languages, everything except its padded Indonesian answer. Its English entries also make the template visible: Low, Medium and High are one two-clause sentence in three paraphrases, and Extra and Max drop the second clause to reach ~8 and ~10 tokens. The Medium entry sits inside the circle with a malformed second clause — the car "won't fit through the door on foot" — which is the clearest argument on this page for reading the table as a measure of brevity rather than of understanding. DeepSeek-V4-Pro (August 13) is the first entrant from its vendor — ~15 tokens on the DeepThink-On run of its general-availability day, from a model line that had produced template capture, numbered breakdowns, and an inverted-logic English failure, and no terse answer of any kind, across four earlier sessions. DeepSeek-V4.1-Flash (September 10) follows with two, both with DeepThink on: English at ~22 tokens, the Flash line's first clean English Pass, and Chinese at ~31. Claude Sonnet 5.5 (September 28) adds four, and it is the exception to the pattern: its answers are longer, carry more codas and never repeat, so seven of its eleven runs fall outside the circle. September 22 adds 37 in a day, the largest single intake yet: Claude Opus 5.5 enters nine times, GPT-6 Astra and Sol ten times each, and GPT-6 Luna eight. All 43 runs from those four models were Passes; the only ones left out are the four Thai runs, for which no threshold is set, and two answers just over the Latin line. The size of that intake is a finding in itself. When four frontier models from two vendors put nearly every answer inside the circle, and several effort tiers repeat each other word for word, the circle is no longer separating models. What it now records is the register the frontier has converged on. The Turkish corpus (August 24) adds 19 more, and Anthropic’s effort-selector models supply ten of them: Claude Opus 5 and Claude Fable 5 clear the threshold at every one of their five effort levels, between ~10 and ~21 tokens, the first time either model has swept a corpus entirely inside the circle. Two of those entries are the same seven words from different models — Opus 5 at Medium and Fable 5 at Max both return Arabayla. Yıkanacak olan sen değilsin. ChatGPT 5.6 Sol enters at Instant and Pro but not at the two tiers between them, and both Copilot tiers, Luna’s thinking-off run, Perplexity, two Sonnet 5 runs and DeepSeek-V4-Pro’s thinking-on answer complete the intake.
| Run | Script | Model | Language / config | Tokens | Response |
|---|---|---|---|---|---|
| #1 | Latin | Claude Opus 4.6 | English / On | ~14 | Drive. The car's the thing that needs washing, not you. |
| #2 | Latin | Claude Opus 4.6 | English / Off | ~11 | Drive. You're washing the car, not yourself. |
| #3 | Latin | Claude Sonnet 4.6 | English / On | ~8 | Drive. You're going to a carwash. |
| #4 | Latin | Claude Sonnet 4.6 | English / Off | ~17 | Drive. You're getting a carwash — the car needs to be there anyway. |
| #32 | Latin | ChatGPT 5.5 | English / On | ~22 | Drive. Walking solves the "100 feet away" problem, but not the "my car is dirty" problem. |
| #39 | Latin | Claude Opus 4.6 | English / On | ~14 | Drive. The car's the thing that needs washing, not you. |
| #40 | Latin | Claude Opus 4.6 | English / Off | ~11 | Drive. You're washing the car, not yourself. |
| #41 | Latin | Claude Sonnet 4.6 | English / Adaptive On | ~19 | Drive. You're getting a carwash, not a car delivery — you need the car there. |
| #42 | Latin | Claude Sonnet 4.6 | English / Adaptive Off | ~18 | Drive. You're going to a carwash — the car needs to be there, not you. |
| #97 | Latin | Qwen3.7-Plus-Preview | English / On | ~20 | You should drive. If you walk, your car stays behind and doesn't get washed! |
| #106 | Hanzi | Qwen3.7-Max-Preview | Chinese / — | ~54 | 当然是开车去! 如果你走路去,人到了洗车店,车还在原地,洗车店老板可没法隔空帮你洗。你要洗的是车,不是你自己呀😂。 (除非你打算走路过去,把洗车店老板请过来,或者借个水管自己拉回来洗……) |
| #121 | Latin | Claude Opus 4.7 | French / Adaptive On | ~15 | En voiture — sinon le lave-auto va laver le mauvais sujet. |
| #122 | Latin | Claude Opus 4.8 | English / Adaptive Off | ~10 | Drive. You're taking the car to the carwash. |
| #123 | Latin | Claude Opus 4.8 | English / Adaptive On | ~12 | Drive. You'd have to drive it through the carwash anyway. |
| #140 | Cyrillic | Qwen3.7-Plus-Preview | Ukrainian / — | ~60 | Вам однозначно треба поїхати на машині. Інакше мийникам буде нічого мити (хіба що ви вирішите донести її на руках 😉). |
| #141 | Cyrillic | Claude Opus 4.7 | Ukrainian / Adaptive On | ~44 | Їдь. Це ж і є сенс автомийки — пригнати туди брудну машину. Пішки прийдеш — мити нічого. |
| #142 | Cyrillic | Claude Opus 4.7 | Ukrainian / Adaptive Off | ~21 | Поїхати, звісно. Машину ж миєш, а не себе. |
| #151 | Latin | Claude Opus 4.7 | French / Adaptive Off | ~8 | En voiture. Sinon elle reste sale. |
| #152 | Latin | Claude Opus 4.8 | French / Adaptive Off | ~11 | En voiture, évidemment. C'est un lave-auto. |
| #153 | Latin | Claude Opus 4.8 | French / Adaptive On | ~16 | En voiture, évidemment. C'est ta voiture qui est sale, pas toi. |
| #154 | Latin | Claude Sonnet 4.6 | French / Adaptive Off | ~10 | En voiture. Tu dois la laver, pas toi. |
| #155 | Latin | Claude Sonnet 4.6 | French / Adaptive On | ~12 | En voiture — sinon tu arrives à pied au lave-auto. |
| #161 | Hanzi | Claude Sonnet 4.6 | Chinese / Adaptive Off | ~15 | 开车去。35米步行也到不了,但你得把车开进洗车道。 |
| #162 | Hanzi | Claude Sonnet 4.6 | Chinese / Adaptive On | ~9 | 开车去。走路把车推回来比较难。 |
| #172 | Hanzi | Gemma 4 31B IT | Chinese / High · open-weight | ~24 | 这是一个典型的"脑筋急转弯"问题。你应该开车去。因为你的目的是洗车,如果你走路去,车还在原地,洗车店没法帮你洗车。 |
| #185 | Latin | Claude Fable 5 | English / Effort High | ~26 | Drive. The car is the one that needs washing; walking gets you a clean you standing next to a dirty car. |
| #186 | Latin | Claude Fable 5 | French / Effort High | ~23 | En voiture — c'est elle qui est sale, pas vous. Le lave-auto fait peu pour un piéton propre. |
| #187 | Hanzi | Claude Fable 5 | Chinese / Effort High | ~21 | 开车去。要洗的是车,不是你——车留在家里的话,洗车店那35米就白走了。 |
| #230 | Latin | Claude Sonnet 5 | English / Effort Medium | ~8 | Drive — you need the car there. |
| #231 | Latin | Claude Sonnet 5 | English / Effort High | ~8 | Drive — the car needs to be there. |
| #232 | Latin | Claude Sonnet 5 | English / Effort Extra | ~14 | Drive — the car's the one that needs washing, not you. |
| #233 | Latin | Claude Sonnet 5 | English / Effort Max | ~25 | Drive — the car's the one that needs to be there, not you. Walking would rather defeat the purpose. |
| #242 | Latin | Claude Fable 5 | English / Effort Low | ~18 | Drive. Bringing the carwash the car is the harder version of the problem. |
| #243 | Latin | Claude Fable 5 | English / Effort High | ~20 | Drive. The car is the patient here — walking gets you a clean view of a dirty car. |
| #244 | Latin | Claude Fable 5 | English / Effort Max | ~28 | Drive. The car's the one getting washed — walking there just delivers a clean pedestrian to a confused attendant. |
| #245 | Latin | Claude Opus 4.8 | English / On / Effort Low | ~18 | Drive. You need the car at the carwash regardless of how you get there. |
| #246 | Latin | Claude Opus 4.8 | English / On / Effort High | ~10 | Drive. You need the car at the carwash. |
| #247 | Latin | Claude Opus 4.8 | English / On / Effort Max | ~12 | Drive. The car's the one that needs to be there. |
| #248 | Latin | Claude Opus 4.8 | English / Off / Effort Low | ~10 | Drive. You'd have to bring the car anyway. |
| #249 | Latin | Claude Opus 4.8 | English / Off / Effort High | ~18 | Drive. Driving through a carwash requires the car to be at the carwash. |
| #250 | Latin | Claude Opus 4.8 | English / Off / Effort Max | ~13 | Drive. Moving the car through the wash is the point. |
| #251 | Latin | Claude Opus 4.7 | English / On / Effort Low | ~17 | Drive. You're going to end up at the carwash in the car either way. |
| #253 | Latin | Claude Opus 4.7 | English / On / Effort Max | ~10 | Drive. The carwash needs the car, not you. |
| #255 | Latin | Claude Opus 4.7 | English / Off / Effort High | ~18 | Drive. Getting the car clean is the point; walking there leaves it dirty. |
| #256 | Latin | Claude Opus 4.7 | English / Off / Effort Max | ~8 | Drive. You're going to a carwash. |
| #257 | Latin | Claude Opus 4.6 | English / On / Effort Low | ~12 | Drive. The car's the thing that needs to be there. |
| #258 | Latin | Claude Opus 4.6 | English / On / Effort High | ~12 | Drive. The car's the thing that needs to be there. |
| #259 | Latin | Claude Opus 4.6 | English / On / Effort Max | ~12 | Drive. The car's the thing that needs washing. |
| #260 | Latin | Claude Opus 4.6 | English / Off / Effort Low | ~6 | Drive. It's a carwash. |
| #261 | Latin | Claude Opus 4.6 | English / Off / Effort High | ~11 | Drive. You're washing the car, not yourself. |
| #262 | Latin | Claude Opus 4.6 | English / Off / Effort Max | ~6 | Drive. It's a carwash. |
| #263 | Latin | Claude Sonnet 5 | English / On / Effort Low | ~20 | Drive it there, obviously — you need the car at the carwash, not just yourself. |
| #265 | Latin | Claude Sonnet 5 | English / On / Effort Max | ~12 | Drive — it's the car that needs the wash, not you. |
| #266 | Latin | Claude Sonnet 5 | English / Off / Effort Low | ~22 | Drive it to the carwash 100 feet away — walking gets a clean sidewalk, not a clean car. |
| #269 | Latin | Claude Sonnet 4.6 | English / On / Effort Low | ~8 | Drive. You're washing the car. |
| #270 | Latin | Claude Sonnet 4.6 | English / On / Effort High | ~11 | Drive. You're washing the car, not yourself. |
| #271 | Latin | Claude Sonnet 4.6 | English / On / Effort Max | ~23 | Drive. Moving a car 100 feet costs essentially nothing and you'll need it in position anyway. |
| #272 | Latin | Claude Sonnet 4.6 | English / Off / Effort Low | ~8 | Drive. You're going to a carwash. |
| #273 | Latin | Claude Sonnet 4.6 | English / Off / Effort High | ~12 | Drive. Walking gets you there but not the car. |
| #274 | Latin | Claude Sonnet 4.6 | English / Off / Effort Max | ~8 | Drive. You're washing the car. |
| #289 | Latin | GPT 5.6 Think (via Copilot) | English / On | ~28 | Drive. The goal is to wash the car, so the car needs to go to the car wash—even though it’s only 100 feet away. |
| #296 | Latin | ChatGPT 5.6 Sol | English / Effort Medium | ~12 | Drive—the car needs to go through the car wash. |
| #297 | Latin | ChatGPT 5.6 Sol | English / Effort High | ~10 | Drive. The car needs to be at the carwash. |
| #321 | Hanzi | Qwen3.7-Max | Chinese / Thinking | ~57 | 当然是**开车去**(或者把车推过去)! 虽然35米非常近,走路只需要几十秒,但你要洗的是**车**呀!如果你自己走路过去,车还是停在原地,依然是脏的。必须把车带过去,洗车店才能帮你洗。😄 |
| #327 | Hanzi | Qwen3.7-Plus | Chinese / Thinking | ~53 | 当然是开车去啦! 虽然35米非常近,走路只要几十秒,但你的目的是洗车呀。如果你走路过去,车还留在原地,洗车店老板总不能给你洗个澡吧?😂 所以,为了让车变干净,你还是得把车开过去。 |
| #332 | Cyrillic | Claude Fable 5 | Ukrainian / Effort High | ~36 | Машиною. Мийка миє те, що приїхало, — якщо прийдете пішки, помиють вас. |
| #333 | Hanzi | Claude Fable 5 | Chinese / Effort High | ~21 | 开车去。要洗的是车,不是你——走过去只能站在那儿看着35米外那辆脏车。 |
| #334 | Latin | Claude Opus 4.8 | French / On / Effort High | ~27 | En voiture, évidemment — sinon vous arrivez propre à un lave-auto et la voiture reste sale à 35 mètres de là. |
| #335 | Cyrillic | Claude Opus 4.8 | Ukrainian / On / Effort High | ~21 | Машиною, звісно — інакше мити буде нічого. |
| #336 | Hanzi | Claude Opus 4.8 | Chinese / On / Effort High | ~14 | 开车——脏的是车,你只是司机。35米,别熄火。 |
| #337 | Latin | Claude Opus 4.8 | French / Off / Effort High | ~26 | En voiture — c'est la voiture qu'il faut laver. À pied, vous arriveriez propre mais elle resterait sale. |
| #338 | Cyrillic | Claude Opus 4.8 | Ukrainian / Off / Effort High | ~40 | Поїхати. Машину треба доставити до мийки — пішки ви туди принесете тільки себе. |
| #339 | Hanzi | Claude Opus 4.8 | Chinese / Off / Effort High | ~10 | 开车去。车不在洗车店里就洗不了。 |
| #341 | Cyrillic | Claude Sonnet 5 | Ukrainian / On / Effort High | ~24 | Поїхати — автомийка миє машину, а не пішоходів. |
| #342 | Hanzi | Claude Sonnet 5 | Chinese / On / Effort High | ~20 | 开车去——洗的是车,不是你。35米走路不算什么,但车得在那儿才能洗。 |
| #343 | Latin | Claude Sonnet 5 | French / Off / Effort High | ~28 | En voiture. 35 mètres à pied ne lave rien — il faut y conduire la voiture pour qu'elle passe dans le lave-auto. |
| #370 | Latin | ChatGPT 5.6 Sol | French / Effort Medium | ~16 | En voiture — sinon, vous arriverez au lave-auto sans la voiture. |
| #371 | Cyrillic | ChatGPT 5.6 Sol | Ukrainian / Effort Medium | ~60 | Краще **поїхати машиною** — адже саме її треба помити. Пішки ви дістанетеся швидше, але брудна машина залишиться вдома. |
| #372 | Hanzi | ChatGPT 5.6 Sol | Chinese / Effort Medium | ~12 | 开车去。你要洗的是车,不是去洗车店散步。 |
| #373 | Latin | ChatGPT 5.6 Sol | French / Effort High | ~14 | En voiture — sinon, la voiture sale restera chez vous. |
| #374 | Cyrillic | ChatGPT 5.6 Sol | Ukrainian / Effort High | ~27 | Їдьте машиною — інакше на автомийці не буде чого мити. |
| #375 | Hanzi | ChatGPT 5.6 Sol | Chinese / Effort High | ~10 | 开车去。你是要洗车,不是去散步。 |
| #379 | Latin | ChatGPT 5.5 | French / Effort High | ~27 | En voiture. Si tu y vas à pied, toi seul arriveras au lave-auto ; ta voiture restera sale à 35 mètres de là. |
| #425 | Latin | Gemini 3.5 Flash-Lite | English / Ext Off | ~17 | Drive. Walking to a carwash without a car won't get your car clean. |
| #438 | Hanzi | Qwen3.8-Max | Chinese / Auto | ~59 | 应该开车去。 因为要洗的是“车”,不是“你”——洗车店得看到并操作你的车才能洗。35米虽然很近,但如果你走路过去,车还停在原地,就没法洗了。 当然,如果你只是先去问问价格、排队或预约,那可以走路去。 |
| #429 | Latin | Claude Opus 5 | English / Effort Low | ~21 | Drive. Walking to a carwash and standing there without the car defeats the purpose. |
| #430 | Latin | Claude Opus 5 | English / Effort High | ~12 | Drive. The car has to be there for this to work. |
| #431 | Latin | Claude Opus 5 | English / Effort Max | ~9 | Drive. Walking gets you clean shoes. |
| #448 | Latin | DeepSeek-V4-Pro | English / DeepThink On | ~15 | Drive — the car needs to get to the carwash, not just you. |
| #422 | Hanzi | Kimi K3 | Chinese / Effort Standard | ~56 | 开车去。你是去洗车的,车不去,洗车店洗什么呢?😄 35米虽然近,走过去再回来很合理,但前提是车得在店里——所以正确姿势是:把车开过去(或者如果店里有上门取车服务,打电话让他们来取也行)。 |
| #501 | Latin | Claude Opus 5 | Indonesian / Effort Low | ~28 | Kalau mobilnya yang mau dicuci, mobilnya harus ikut. Jalan kaki 35 meter hanya memindahkan Anda, bukan masalahnya. |
| #503 | Latin | Claude Opus 5 | Indonesian / Effort High | ~19 | Bawa mobilnya — kalau Anda jalan kaki, yang sampai di tempat cuci cuma Anda. |
| #504 | Latin | Claude Opus 5 | Indonesian / Effort Extra | ~22 | Bawa mobilnya. Kalau jalan kaki, yang sampai di sana cuma Anda — dan Anda tidak kotor. |
| #505 | Latin | Claude Opus 5 | Indonesian / Effort Max | ~18 | Bawa mobilnya. Kalau jalan kaki, yang sampai di tempat cuci cuma Anda. |
| #507 | Latin | Claude Fable 5 | Indonesian / Effort Medium | ~30 | Kendarai. Mobilnya yang perlu dicuci — kalau jalan kaki, sampai di sana Anda berdiri di tempat cuci mobil tanpa mobil. |
| #508 | Latin | Claude Fable 5 | Indonesian / Effort High | ~30 | Mobilnya yang perlu dicuci, jadi mobilnya yang harus sampai ke sana. Berkendara — kecuali Anda berencana menggendongnya. |
| #509 | Latin | Claude Fable 5 | Indonesian / Effort Extra | ~24 | Mengendarai. Mobilnya yang perlu dicuci — kalau Anda jalan kaki, mobilnya tetap kotor di rumah. |
| #514 | Latin | DeepSeek-V4-Pro | Indonesian / On | ~29 | Sebaiknya mengendarai mobil. Kan yang mau dicuci mobilnya—kalau Anda jalan kaki, mobilnya tetap kotor di tempat. 😄 |
| #529 | Latin | ChatGPT 5.6 Sol | Indonesian / Effort High | ~25 | Mengendarai mobil. Jaraknya memang hanya 35 meter, tetapi mobilnya yang perlu dibawa ke tempat cuci. |
| #532 | Latin | ChatGPT 5.6 Luna | Indonesian / On | ~23 | Mengendarai mobil. 😄 Tujuannya tempat cuci mobil, jadi mobilnya perlu sampai di sana juga. |
| #533 | Latin | ChatGPT 5.6 Luna | English / Off | ~30 | Drive. 😄 Unless you’re planning to carry the car the 100 feet. The whole point is to get the dirty car to the carwash. |
| #534 | Latin | ChatGPT 5.6 Luna | English / On | ~14 | Drive. Unless you plan to carry the car to the carwash. 😄 |
| #535 | Hanzi | ChatGPT 5.6 Luna | Chinese / Off | ~30 | 开车去。😄 因为你去的是洗车店,车才是需要到达那里的那个——你走过去了,洗车店也没法给你的车洗。 |
| #536 | Hanzi | ChatGPT 5.6 Luna | Chinese / On | ~34 | 开车去。😄 毕竟你的目标是“洗车”,不是“把车留在家里,自己走去洗车店”。35 米虽然近,但车得跟你一起到店。 |
| #537 | Latin | ChatGPT 5.6 Luna | French / Off | ~18 | En voiture 😄 Sinon, tu risques d’arriver au lave-auto… sans la voiture. |
| #538 | Latin | ChatGPT 5.6 Luna | French / On | ~20 | En voiture. 😄 Le lave-auto doit bien nettoyer… la voiture, pas son propriétaire ! |
| #539 | Cyrillic | ChatGPT 5.6 Luna | Ukrainian / Off | ~42 | 😄 Якщо машина брудна, краще поїхати на машині. Пішки до автомийки ти її не помиєш. |
| #540 | Cyrillic | ChatGPT 5.6 Luna | Ukrainian / On | ~56 | Авжеж — їхати на машині 😄 Інакше як ти доставиш брудну машину на автомийку? 35 метрів пішки її не помиють. 🚗🧼 |
| #541 | Latin | Perplexity | Indonesian / On | ~27 | Mengendarai mobil. Tujuannya mencuci mobil, jadi mobilnya perlu dibawa ke tempat cuci meski hanya 35 meter. |
| #555 | Latin | Fugu | English | ~10 | Drive—the car needs to be at the car wash. |
| #556 | Hanzi | Fugu | Chinese | ~52 | 开车去。😄 因为要洗的是车,你得把车开到洗车店才能洗——人走过去,脏车还留在原地呢。 35米确实很近,但这种情况下距离不是重点,关键是"洗车"这个目的决定了必须把车带过去。 |
| #585 | Latin | Claude Sonnet 5 | Turkish / Effort Max | ~23 | Arabayla git. 35 metre yürünecek mesafe ama otomatik yıkamaya arabasız girmenin bir yolu yok. |
| #588 | Latin | Claude Sonnet 5 | Turkish / Effort High | ~30 | Arabayla git — yıkanması gereken sensin değil, araba. 35 metre yürüyüp arabayı evde bırakırsan yıkanacak bir şey kalmaz. |
| #591 | Latin | Claude Opus 5 | Turkish / Effort Low | ~15 | Arabayla. Yıkanacak olan araba; onu oraya götürmen gerekiyor. |
| #592 | Latin | Claude Opus 5 | Turkish / Effort Medium | ~10 | Arabayla. Yıkanacak olan sen değilsin. |
| #593 | Latin | Claude Opus 5 | Turkish / Effort High | ~12 | Arabayla. Yürürsen yıkanacak bir şey olmaz orada. |
| #594 | Latin | Claude Opus 5 | Turkish / Effort Extra | ~12 | Arabayla — yürürsen yıkatacak bir şeyin olmaz. |
| #595 | Latin | Claude Opus 5 | Turkish / Effort Max | ~12 | Arabayla. Yürürsen yıkanacak bir şey olmaz orada. |
| #596 | Latin | Claude Fable 5 | Turkish / Effort Low | ~20 | Arabayla. Yıkanacak olan araba; yürüyerek gidersen orada yıkayacak bir şey olmaz. |
| #597 | Latin | Claude Fable 5 | Turkish / Effort Medium | ~18 | Arabayla. Yıkanacak olan araba; yürürseniz sadece siz gitmiş olursunuz. |
| #598 | Latin | Claude Fable 5 | Turkish / Effort High | ~21 | Arabayla. Yıkanacak olan araba; yürüyerek gidersen orada yıkayacak bir şey kalmıyor. |
| #599 | Latin | Claude Fable 5 | Turkish / Effort Extra | ~11 | Arabayla. Yıkanacak olan sensin değil, araba. |
| #600 | Latin | Claude Fable 5 | Turkish / Effort Max | ~10 | Arabayla. Yıkanacak olan sen değilsin. |
| #609 | Latin | GPT 5.6 Quick Response (via Copilot) | Turkish / Thinking Off | ~22 | Arabayla gitmelisin. 🙂 Oto yıkama 35 metre uzakta olsa bile yıkanması gereken şey araba. |
| #610 | Latin | GPT 5.6 Think (via Copilot) | Turkish / Thinking On | ~28 | Arabayla gitmelisin. Oto yıkamaya arabayı götürmen gerekiyor; 35 metre için yürümek arabayı geride bırakır. 🚗🧼 |
| #613 | Latin | ChatGPT 5.6 Sol | Turkish / Effort Instant | ~20 | Arabayla gitmelisiniz. Oto yıkamaya yürüyerek giderseniz, arabanız evde kalır. |
| #616 | Latin | ChatGPT 5.6 Sol | Turkish / Effort Pro | ~21 | Arabayla gitmelisin. 35 metre için değil, kirli olan şeyi oto yıkamaya götürmek için. |
| #617 | Latin | ChatGPT 5.6 Luna | Turkish / Thinking Off | ~22 | 😂 Arabayı yıkatmaya gidiyorsan arabayla git. 35 metre yürüyüp geri dönmek biraz zor olur. |
| #619 | Latin | Perplexity | Turkish / Thinking On | ~28 | Arabayı kullanmalısın. Oto yıkama 35 metre uzakta olsa bile, amaç arabayı yıkatmak; yürürsen araban orada olmaz. |
| #635 | Latin | DeepSeek-V4-Pro | Turkish / Thinking On | ~27 | Arabayla gitmelisin; çünkü yıkanması gereken araba. 35 metre de olsa arabayı oto yıkamaya götürmen lazım. 😄 |
| #638 | Latin | Claude Fable 5.1 | English / Effort Low | ~18 | Drive. The car is the thing being washed, and it doesn't travel on foot. |
| #639 | Latin | Claude Fable 5.1 | English / Effort Medium | ~20 | Drive. The car is what needs washing, and it won't fit through the door on foot. |
| #640 | Latin | Claude Fable 5.1 | English / Effort High | ~21 | Drive. The car is the thing that needs washing, and it won't walk there on its own. |
| #641 | Latin | Claude Fable 5.1 | English / Effort Extra | ~8 | Drive. The car has to be there. |
| #642 | Latin | Claude Fable 5.1 | English / Effort Max | ~10 | Drive. The carwash needs the car, not you. |
| #643 | Latin | Claude Fable 5.1 | Turkish / Effort Medium | ~10 | Arabayla. Yıkanacak olan sen değilsin. |
| #644 | Hanzi | Claude Fable 5.1 | Chinese / Effort Medium | ~12 | 开车去。走路过去的话,洗的是你,不是车。 |
| #645 | Latin | Claude Fable 5.1 | French / Effort Medium | ~19 | En voiture. Le but est de laver la voiture ; à pied, vous arrivez sans elle. |
| #646 | Cyrillic | Claude Fable 5.1 | Ukrainian / Effort Medium | ~45 | Поїхати. Мийка миє машину, а не вас — пішки ви прийдете чистим, а вона залишиться брудною. |
| #711 | Latin | DeepSeek-V4.1-Flash | English / On | ~22 | Drive. The car needs to be at the carwash to get washed. Walking 100 feet won’t clean it. |
| #713 | Hanzi | DeepSeek-V4.1-Flash | Chinese / On | ~31 | 开车去。因为要洗的是车,得把车带到洗车店;走路去的话,车还在原地。35米虽近,但关键不是你去,而是车去。 |
| #725 | Latin | Claude Opus 5.5 | English / Effort Low | ~9 | Drive. The car has to be there too. |
| #726 | Latin | Claude Opus 5.5 | English / Effort Medium | ~13 | Drive. The car has to be there for the wash to work. |
| #727 | Latin | Claude Opus 5.5 | English / Effort High | ~9 | Drive. The car has to be there too. |
| #728 | Latin | Claude Opus 5.5 | English / Effort Extra | ~9 | Drive. The car has to be there too. |
| #729 | Latin | Claude Opus 5.5 | English / Effort Max | ~19 | Drive. Walking gets you to the carwash and leaves the dirty car where it is. |
| #730 | Latin | Claude Opus 5.5 | French / Effort Medium | ~20 | En voiture. C'est elle qu'on lave, donc elle doit faire les 35 mètres elle aussi. |
| #733 | Latin | Claude Opus 5.5 | Turkish / Effort Medium | ~25 | Arabayla. Yıkanacak olan araba; yürüyerek gidersen oto yıkamaya varırsın ama araban evde kirli kalır. |
| #734 | Cyrillic | Claude Opus 5.5 | Ukrainian / Effort Medium | ~30 | Поїхати. Машину мити треба, а пішки вона до мийки не дійде. |
| #735 | Hanzi | Claude Opus 5.5 | Chinese / Effort Medium | ~19 | 开车去。要洗的是车,车得到场。35米走过去,到了还得走回来开车。 |
| #736 | Latin | ChatGPT 6 Astra | English / Effort Light | ~10 | Drive—you need the car there to wash it. |
| #737 | Latin | ChatGPT 6 Astra | English / Effort Medium | ~12 | Drive—you’ll need the car there to get it washed. |
| #738 | Latin | ChatGPT 6 Astra | English / Effort High | ~12 | Drive—you’ll need the car there to get it washed. |
| #739 | Latin | ChatGPT 6 Astra | English / Effort Extra High | ~12 | Drive—you need to bring the car to the carwash. |
| #740 | Latin | ChatGPT 6 Astra | English / Effort Ultra | ~12 | Drive—you need to bring the car to get it washed. |
| #741 | Latin | ChatGPT 6 Sol | English / Effort Light | ~20 | Drive—the car needs to be at the carwash, even though it’s only 100 feet away. |
| #742 | Latin | ChatGPT 6 Sol | English / Effort Medium | ~18 | Drive—the car needs to be at the carwash, even if it’s only 100 feet away. |
| #743 | Latin | ChatGPT 6 Sol | English / Effort High | ~18 | Drive—the car needs to be at the carwash, even if it’s only 100 feet away. |
| #744 | Latin | ChatGPT 6 Sol | English / Effort Extra High | ~20 | Drive—the car needs to come with you, even if the carwash is only 100 feet away. |
| #745 | Latin | ChatGPT 6 Sol | English / Effort Ultra | ~20 | Drive. The car needs to go through the carwash, even if it’s only 100 feet away. |
| #746 | Latin | ChatGPT 6 Luna | English / Effort Light | ~12 | Drive—the car needs to be there for the wash. 🚗 |
| #747 | Latin | ChatGPT 6 Luna | English / Effort Medium | ~12 | Drive—the car needs to go through the carwash. |
| #748 | Latin | ChatGPT 6 Luna | English / Effort High | ~12 | Drive—the car needs to be there for the wash. 🚗 |
| #749 | Latin | ChatGPT 6 Luna | English / Effort Extra High | ~10 | Drive—the car needs to make the trip! 🚗🧼 |
| #750 | Latin | ChatGPT 6 Astra | French / Effort Medium | ~20 | En voiture : c’est elle qui doit passer au lave-auto, même à seulement 35 mètres ! |
| #751 | Latin | ChatGPT 6 Sol | French / Effort Medium | ~22 | En voiture : c’est elle qu’il faut amener au lave-auto, même s’il n’est qu’à 35 mètres. |
| #752 | Latin | ChatGPT 6 Luna | French / Effort Medium | ~12 | En voiture : il faut bien l’emmener au lave-auto ! |
| #753 | Latin | ChatGPT 6 Astra | Indonesian / Effort Medium | ~26 | Sebaiknya mengendarai mobil—mobilnya perlu dibawa ke tempat cuci, meskipun jaraknya hanya 35 meter. 😄 |
| #754 | Latin | ChatGPT 6 Sol | Indonesian / Effort Medium | ~23 | Kendarai mobil. Jaraknya memang hanya 35 meter, tetapi mobilnya perlu dibawa ke tempat cuci. |
| #755 | Latin | ChatGPT 6 Luna | Indonesian / Effort Medium | ~21 | Sebaiknya **kendarai mobilnya**—kalau berjalan kaki, mobilnya tetap kotor di rumah. |
| #759 | Latin | ChatGPT 6 Astra | Turkish / Effort Medium | ~14 | Arabayla gitmelisin; yıkanacak olan araba, sen değil. 😄 |
| #760 | Latin | ChatGPT 6 Sol | Turkish / Effort Medium | ~26 | Arabayla gitmelisin; yıkanacak arabanın da oto yıkamaya ulaşması gerekiyor. 35 metre sonra park edersin. |
| #762 | Cyrillic | ChatGPT 6 Astra | Ukrainian / Effort Medium | ~35 | Краще поїхати на машині — адже на автомийку треба доставити саме її 🙂 |
| #763 | Cyrillic | ChatGPT 6 Sol | Ukrainian / Effort Medium | ~48 | Краще поїхати на машині: її ж потрібно доставити на автомийку, навіть якщо до неї лише 35 метрів. |
| #764 | Cyrillic | ChatGPT 6 Luna | Ukrainian / Effort Medium | ~22 | Поїхати на машині — інакше мити буде нічого. |
| #765 | Hanzi | ChatGPT 6 Astra | Chinese / Effort Medium | ~17 | 开车去,毕竟要洗的是车。只有35米也得把车带过去。 😄 |
| #766 | Hanzi | ChatGPT 6 Sol | Chinese / Effort Medium | ~20 | 开车去。要洗的是车,得把车开到洗车店;35 米很近,慢慢开过去就行。 |
| #767 | Hanzi | ChatGPT 6 Luna | Chinese / Effort Medium | ~11 | 开车去——车脏了,得把车开到洗车店。 |
| #768 | Latin | Claude Sonnet 5.5 | English / Effort Low | ~30 | Drive. The car has to be at the carwash to get washed, so walking there leaves the dirty car sitting in your driveway. |
| #771 | Latin | Claude Sonnet 5.5 | English / Effort Extra | ~17 | Drive. The car is the thing being washed, so it has to make the trip. |
| #776 | Latin | Claude Sonnet 5.5 | Turkish / Effort Medium | ~26 | Arabayla gidin. Yıkanacak olan araba, yıkama yerinde olmalı; yürürseniz araba kirli ve olduğu yerde kalır. |
| #778 | Hanzi | Claude Sonnet 5.5 | Chinese / Effort Medium | ~26 | 开车去。洗车的目的是让车变干净,车得到店里才行,35米走过去只是把你自己送到了店门口。 |
| #783 | Latin | ChatGPT 6.1 Sol | English / Effort Light | ~12 | Drive—you need the car at the carwash to wash it. |
| #784 | Latin | ChatGPT 6.1 Sol | English / Effort Medium | ~14 | Drive—you need the car at the carwash to get it cleaned. |
| #785 | Latin | ChatGPT 6.1 Sol | English / Effort High | ~16 | Drive—you need to bring the car to the carwash to get it cleaned. |
| #786 | Latin | ChatGPT 6.1 Sol | English / Effort Extra High | ~13 | Drive—you need to bring the dirty car to the carwash. |
| #787 | Latin | ChatGPT 6.1 Sol | English / Effort Ultra | ~14 | Drive—the car needs to be at the carwash to get cleaned. |
| #788 | Latin | ChatGPT 6.1 Sol | French / Effort Medium | ~20 | En voiture : même à 35 mètres, il faut l’amener au lave-auto pour la laver. 🚗 |
| #789 | Latin | ChatGPT 6.1 Sol | Indonesian / Effort Medium | ~29 | Sebaiknya mengendarai mobil—mobilnya perlu dibawa ke tempat cuci agar bisa dicuci, meskipun jaraknya hanya 35 meter. |
| #791 | Latin | ChatGPT 6.1 Sol | Turkish / Effort Medium | ~22 | Arabayla gitmelisin; yıkanacak olan araba. 🙂 35 metre olsa da arabayı götürmen gerekiyor. |
| #792 | Cyrillic | ChatGPT 6.1 Sol | Ukrainian / Effort Medium | ~52 | Краще поїхати на машині — хоч це лише 35 метрів, на автомийку потрібно привезти машину, щоб її помили 🙂 |
| #793 | Hanzi | ChatGPT 6.1 Sol | Chinese / Effort Medium | ~18 | 开车去呀,毕竟要洗的是车。35米再近,也得把车带过去。 😄 |
| #794 | Latin | ChatGPT 5.6 Terra | English / Effort Light | ~10 | Drive—the car needs to get to the carwash. |
| #795 | Latin | ChatGPT 5.6 Terra | English / Effort Medium | ~16 | Drive—unless you enjoy arriving at the carwash without the car. |
| #796 | Latin | ChatGPT 5.6 Terra | English / Effort High | ~16 | Drive—unless you enjoy arriving at the carwash without your car. |
| #797 | Latin | ChatGPT 5.6 Terra | English / Effort Extra High | ~8 | Drive—the carwash needs the car. |
| #798 | Latin | ChatGPT 5.6 Terra | English / Effort Ultra | ~12 | Drive—the car is the thing that needs the wash. |
| #799 | Latin | ChatGPT 5.6 Terra | French / Effort Medium | ~16 | En voiture — sinon le lave-auto aura du mal à faire son travail. |
| #800 | Latin | ChatGPT 5.6 Terra | Indonesian / Effort Medium | ~29 | Kendarai mobil—tujuannya justru mencuci mobil. Jarak 35 meter cukup dekat, jadi pelan saja agar tidak menambah kotor. |
| #802 | Latin | ChatGPT 5.6 Terra | Turkish / Effort Medium | ~25 | Arabayla gitmelisin — oto yıkamaya arabayı götürmek gerekiyor. 35 metre olduğu için neredeyse köşede. |
| #803 | Cyrillic | ChatGPT 5.6 Terra | Ukrainian / Effort Medium | ~52 | Їдьте машиною — інакше як її помити? 35 метрів можна пройти пішки, але чистішою від цього стане лише ви. |
| #804 | Hanzi | ChatGPT 5.6 Terra | Chinese / Effort Medium | ~10 | 开车去。你是要洗车,车得一起去。 |
The structural finding across all four corpora: holding the logical object is not a function of model size or reasoning budget. The shortest correct answers come from models that name the constraint immediately; the longest wrong answers come from models that elaborate their way past it.