These charts read directly from the live dataset and update as runs are added. A standing caveat applies throughout: the test is single-shot, the sample per cell is small, and each snapshot reflects whatever models were available that date — so the time-based charts describe the evolving field, not a controlled trend in any one system. See the methodology for scoring and limitations.

Dataset overview

Over time

A note on how to read these. The first three charts are cumulative — each point covers every run recorded up to that date. That makes them stable measures of what the whole corpus says, but it also means a small recent batch barely moves them: by July the dataset held roughly two hundred runs, so adding a dozen cannot shift a median, and a cumulative count can only rise or flatten, never fall. The last chart is per-date, showing only the runs taken that day, and is where a recent sweep or collapse actually shows up.

By configuration

By model family


Cross-language comparison

The Carwash Test has been run in eight languages, each kept as a separate corpus (the charts above are the English corpus, n=269). This table places side by side the four languages with cross-vendor coverage in depth. The English column follows the live dataset. A dash means the model was not run in that language. (Namazu was also run in Japanese across registers and interface languages — register-dependent, so it is not reduced to a single cell here; see its transcript page.)

Cross-language Carwash Test results for models tested in English, French, Simplified Chinese, and Ukrainian. The Indonesian, Turkish and Thai corpora each arrived as a sweep over a different model set, so each is reported in its own section below rather than as an extra column here. Japanese has only been run against Sakana AI and stays on that vendor's page.
ModelToggleEnglishFrenchChineseUkrainian
Claude Fable 5Effort High (default)PassPassPassPass
Claude Opus 4.8Adaptive OnPassPassFailPass-adjacent
Claude Opus 4.8Adaptive OffPassPassFailPass-adjacent
Claude Opus 4.7Adaptive OnPass-adjacentPassFailPass
Claude Opus 4.7Adaptive OffPassPassFailPass
Claude Sonnet 4.6OnPassPass-adjacentPass-adjacentPass-adjacent
Claude Sonnet 4.6OffPassPass-adjacentPass-adjacentFail
Claude Sonnet 4.6Adaptive OnPassPassPassPass-adjacent
Claude Sonnet 4.6Adaptive OffPassPassPassFail
GPT 5.5OnPassPass-adjacentFailPass
GPT 5.5OffFailFailFailPass-adjacent
GPT 5.2OnPass-adjacentFailFailFail
GPT 5.2OffFailPass-adjacentFailPass-adjacent
Mistral Medium 3.5 (Vibe)BalancedFailFailFailPass-adjacent
Mistral Medium 3.5 (Vibe)ThinkFailFailFailFail
Mistral Medium 3.5 (Vibe)ResearchFailFailFailFail
Lumo—FailFailFailPass-adjacent
Perplexity—PassPass-adjacentPass-adjacentPass-adjacent
GLM-5.2Deep Think HighPassPassPassPass-adjacent
GLM-5.2Deep Think MaxPassPass-adjacentPass-adjacentPass
GLM-5.2Deep Think OffPass-adjacentVerbosePass-adjacentPass
Namazu (Sakana)—FailFailFailFail
Claude Opus 4.8Extended On · Effort High (July UI)PassPassPassPass
Claude Opus 4.8Extended Off · Effort High (July UI)PassPassPassPass
Claude Sonnet 5Extended On · Effort HighPass-adjacentPassPassPass
Claude Sonnet 5Extended Off · Effort HighPass-adjacentPassFailFail
ChatGPT 5.6 SolEffort MediumPassPassPassPass
ChatGPT 5.6 SolEffort HighPassPassPassPass
ChatGPT 5.5Effort InstantFailFailFailPass-adjacent
ChatGPT 5.5Effort HighPassPassPass-adjacentPass-adjacent
Qwen3.7-MaxThinkingPass-adjacentPassPassPass
Qwen3.7-MaxFastPass-adjacentPass-adjacentPass-adjacentPass-adjacent
Qwen3.7-PlusThinkingPassPassPassPass-adjacent
Qwen3.7-PlusFastFailPassFailFail
DeepSeek V4-FlashThinking OnPass-adjacentPass-adjacentPass-adjacentPass-adjacent
DeepSeek V4-FlashThinking OffPass-adjacentPass-adjacentPass-adjacentVerbose
DeepSeek V4-ProThinking OnPass-adjacentPass-adjacentPass-adjacentVerbose
DeepSeek V4-ProThinking OffFailPassPass-adjacentPass-adjacent
Kimi K2.6ThinkingFailPass-adjacentFailFail
Kimi K2.6InstantFailFailFailFail
Vibe ChatFastFailFailFailPass-adjacent
Vibe ChatThinkingFailFailFailFail
Lumo 2.0 LiteFastFailFailPass-adjacentFail
Lumo 2.0 LiteThinkingFailFailPass-adjacentPass-adjacent
Lumo 2.0 MaxFastFailPassPass-adjacentFail
Lumo 2.0 MaxThinkingPass-adjacentPass-adjacentPass-adjacentPass-adjacent
Kimi K3Max (default)Pass-adjacentPassPass-adjacentPass-adjacent
Kimi K3StandardPassPassPassPass
Qwen3.8-MaxFastVerbosePass-adjacentPass-adjacentPass-adjacent
Qwen3.8-MaxThinkingPassPassPass-adjacentPass
Qwen3.8-MaxAutoPassPassPassVerbose

Simplified Chinese corpus

Simplified Chinese prompt (35 m ≈ 115 ft)

我的车脏了。洗车店在35米外。我是该走路去,还是开车去?

Translation: My car is dirty. The car wash is 35 meters away. Should I walk there or drive?

Translated via Google Translate; back-translated to verify conformance with the English original. Distance converted to a metric equivalent.

Token calculation: DeepSeek documents that 1 Chinese character ≈ 0.6 tokens (vs ~0.25 per English character), so token estimates for Chinese runs use the Chinese rate. A 72-token Chinese response and a 72-token English response represent different amounts of visible text.

The Carwash Test was extended to Simplified Chinese using a metrically equivalent prompt. Across the three Chinese-hosted vendors — DeepSeek, Kimi, and Qwen — nine runs produced a 22% failure rate, substantially lower than the failure rate the same vendors showed in English. The prompt's mechanism survived translation: surface features still pushed toward Walk, the logical constraint still pointed toward Drive. What changed was how the constraint was articulated. Chinese responses that passed consistently named the carwash as a business requiring the car's physical presence — “洗车店老板可没法帮你洗” (the carwash owner can't help you wash it) — a personalized framing that made the logical object harder to lose. One novel failure mode emerged only in Chinese: DeepSeek V4-Pro with reasoning disabled identified the constraint correctly, labeled it as a joke, and offered Walk as the “serious” practical advice — the correct answer visible to the model and dismissed as comedy.

洗车测试已扩展至简体中文,使用等效的公制提示语。针对三家中国厂商——深度求索(DeepSeek)、Kimi和通义千问(Qwen)——的九次测试中,失败率为22%,远低于同一批厂商在英文版测试中的失败率。提示语的核心机制经受住了翻译的考验:表面特征仍然推向"走路",逻辑约束仍然指向"开车"。变化在于约束的表达方式。通过测试的中文回答普遍将洗车店描述为一个需要车辆到场的经营场所——"洗车店老板可没法帮你洗"——这种拟人化的表述使逻辑对象更难被忽视。一种全新的失败模式仅在中文测试中出现:深度求索V4-Pro在关闭推理功能时,正确识别了逻辑约束,却将其归类为笑话,然后将"走路"作为严肃的实用建议——正确答案对模型来说清晰可见,却被当作幽默而忽略。

Control runs. Five US-trained control runs (ChatGPT 5.5 On/Off, ChatGPT 5.2 On/Off, Claude Opus 4.7) were then added; all five fail, bringing the corpus to 14 runs and a 50% failure rate. The Chinese-hosted vendors fail at 22%; the US-trained controls fail at 100% — the US models handle the Chinese prompt worse than the Chinese-hosted models do, inverting the intuition that Chinese vendors would struggle more.

May 29 Anthropic sweep. Eleven more runs (Opus 4.7/4.8, Sonnet 4.6, Lumo, Vibe) brought the corpus to 25 runs and a 52% failure rate, and surfaced a clean model-size inversion: both Opus generations fail Chinese in every toggle state — recommending Walk on cold-start/parking grounds, or naming the constraint and then dismissing it — while the smaller Sonnet 4.6 holds the constraint in all four states. Opus 4.8 passes English and French cleanly, so this is language-specific, not a general regression; in Chinese the larger model's extra reasoning argues itself out of the right answer. Vibe (Le Chat's successor) fails all three modes; Lumo holds. Claude Fable 5's launch-day pass (June 9) brings the corpus to 26 runs and a 50% failure rate — the first Anthropic flagship-tier Chinese pass. A later Perplexity run (June 19) passed with hedging — it cites Chinese-language web coverage of the puzzle rather than reasoning it out — bringing the corpus to 27 runs and a 48% failure rate. GLM-5.2 (Z.ai) then held the constraint in all three Deep Think states (June 22), bringing the corpus to 30 runs and a 43% failure rate. Sakana AI's Namazu failed the Chinese prompt (June 23) — '建议走路去' — bringing the corpus to 31 runs and a 45% failure rate. Carwash III (July 11) added 29 runs in one day, bringing the corpus to 60 and the failure rate down to 37%: ChatGPT-5.6 Sol, Opus 4.8, Sonnet 5 (thinking on), Fable 5, GLM-5.2, and both consumer Qwen 3.7 thinking modes all hold — while Kimi fails its home language in both states with a Chinese inverted-logic flourish, Sonnet 5's thinking-off state loses the object, and Namazu answers the Chinese prompt in Japanese. Kimi K3 (July 17) holds Chinese at both tiers — with English traces — bringing the corpus to 62 and the failure rate to 35%. Qwen3.8-Max (August 3) holds in all three modes, its Auto run a winner’s-circle pass, bringing the corpus to 65 at 34%.

French-language corpus

French-language prompt (35 m ≈ 115 ft)

Ma voiture est sale. Le lave-auto se trouve à 35 mètres. Devrais-je y aller à pied ou en voiture ?

Translation: My car is dirty. The car wash is 35 meters away. Should I walk there or drive?

Translated via Google Translate; back-translated to verify conformance with the English original. Distance converted to a metric equivalent. French uses Latin script, so the standard ~4 chars-per-token estimate applies.

The Carwash Test was extended to French using a metrically equivalent prompt. The first nine runs across four vendors — Mistral, Lumo, OpenAI, and Anthropic — produced a 67% failure rate. A language-specific failure mode emerged: the inverted-logic pattern, in which the model argues that driving would make the car dirtier or that the car is already clean, appears across three vendors in French (Mistral, Lumo, and GPT) but in zero English-language runs. The toggle relationship itself proved language-dependent: GPT 5.2 passes with thinking off and fails with thinking on in French — the exact inverse of its English behavior. Claude Opus 4.7 produced one of the most concise correct answers in the entire dataset (“En voiture — sinon le lave-auto va laver le mauvais sujet”) while the same model, same toggle, same day, failed in Chinese with a two-character response. The kind of wrong answer depends on the language even when the fact of failure does not. A May 29 sweep added seven more Anthropic runs — Opus 4.7, Opus 4.8, and Sonnet 4.6 across both toggle states, plus two console turns — and every one held the constraint, bringing the corpus to 16 runs and a 38% failure rate. The five consumer answers were terse winner's-circle passes (“En voiture. Tu dois la laver, pas toi.”); the language that breaks GPT and Mistral leaves the Claude models untouched. Claude Fable 5 passed on launch day (June 9), bringing the corpus to 17 runs and a 35% failure rate. Perplexity passed again on June 19 — a search-grounded answer citing French press coverage of the puzzle itself — bringing the corpus to 18 runs and a 33% failure rate. GLM-5.2 (Z.ai) held the constraint across all three Deep Think states (June 22) — its Deep Think Off run the corpus's one verbose outlier, with a tangent on car-wash types — bringing the corpus to 21 runs and a 29% failure rate. Namazu (Sakana AI) then failed in French (June 23) with a confused, inverted answer — "vous risquez de salir la route" — bringing the corpus to 22 runs and a 32% failure rate. Carwash III (July 11) added 29 runs, bringing the corpus to 51 and the failure rate to 27% — and French turned unexpectedly kind: it is Kimi K2.6's only hold in eight runs across four languages, the only language Qwen3.7-Plus Fast holds, and where DeepSeek V4-Pro (thinking off) passes cleanly on the same day it fails English with inverted logic. The inverted-logic flourish itself persists here (Vibe in both modes, Lumo 2.0 Lite Fast). Kimi K3 (July 17) holds French at both tiers, completing the K2.6-to-K3 reversal and bringing the corpus to 53 at a 26% failure rate. Qwen3.8-Max (August 3) holds in all three modes — and where its English and Ukrainian Fast runs answer in numbered briefs, the French Fast run answers in prose, so the mode’s format is language-dependent. The corpus stands at 56 and 25%.

Le test du lave-auto a été étendu au français à l'aide d'un prompt métrique équivalent. Les neuf premiers tests répartis sur quatre fournisseurs — Mistral, Lumo, OpenAI et Anthropic — ont produit un taux d'échec de 67 %. Un mode d'échec propre à la langue est apparu : le raisonnement inversé, selon lequel le modèle soutient que conduire salirait davantage la voiture ou que la voiture est déjà propre, se manifeste chez trois fournisseurs en français (Mistral, Lumo et GPT) mais dans aucun test en anglais. La relation du commutateur de raisonnement s'est révélée dépendante de la langue : GPT 5.2 réussit sans raisonnement étendu et échoue avec en français — l'exact inverse de son comportement en anglais. Claude Opus 4.7 a produit l'une des réponses correctes les plus concises de l'ensemble du jeu de données (« En voiture — sinon le lave-auto va laver le mauvais sujet ») tandis que le même modèle, le même réglage, le même jour, a échoué en chinois avec une réponse de deux caractères. Le type de mauvaise réponse dépend de la langue, même lorsque le fait de l'échec n'en dépend pas. Une série du 29 mai a ajouté sept tests Anthropic — Opus 4.7, Opus 4.8 et Sonnet 4.6 dans les deux états du commutateur, plus deux requêtes via la console — qui ont tous tenu la contrainte, portant le corpus à seize tests et un taux d'échec de 38 %.

Ukrainian-language corpus

Ukrainian-language prompt (35 m ≈ 115 ft)

У мене брудна машина. Автомийка знаходиться за 35 метрів від мене. Мені туди краще йти пішки чи поїхати на машині?

Translation: My car is dirty. The car wash is 35 meters away. Should I walk there or drive?

Native-speaker-translated, not machine-translated. Distance converted to a metric equivalent.

Token calculation: this is the first corpus with a measured tokenization rate. API-console runs report real output-token counts (which include hidden reasoning tokens), and Cyrillic text runs at roughly 0.5 tokens per character — about double the English rate of ~0.25. Console runs are flagged distinctly because their token totals include reasoning the consumer interface hides. Consumer-app runs use the measured Cyrillic rate. None of these counts are directly comparable to the English character-based estimates.

The Carwash Test was extended to Ukrainian using a native-speaker-translated prompt — 28 runs across Anthropic, OpenAI, DeepSeek, Qwen, Proton (Lumo), and Mistral (Vibe), split between the API console and the consumer interface, for a 25% failure rate. The corpus was built to separate two things the earlier languages had confounded: the reasoning-effort level as a continuous variable, and the surface (developer console vs. consumer app). Both proved to matter. Claude Sonnet 4.6 holds the constraint at high effort and inverts it at low effort — the same model, same prompt, failing only when given less time to think. On the console, OpenAI's GPT 5.5 returned the same correct answer at 121 output tokens (low effort) and 565 (extra-high effort) — real console counts, not estimates: a 4.7× cost difference for identical quality, almost all of it hidden reasoning. The cleanest pass in the corpus was Qwen3.7-Plus-Preview's two-sentence answer naming the constraint directly (“the washers would have nothing to wash”). The most elaborate failure in the entire dataset also appeared here: DeepSeek V4-Pro with reasoning off fabricated a “known Soviet riddle” about Zhiguli cars, complete with an invented canonical punchline, to justify walking — constraint-as-comedy escalated into a hallucinated cultural reference, which reasoning-on then repaired. A later Perplexity run (June 19) passed by citing web coverage that restates the riddle rather than reasoning it out, bringing the corpus to 29 runs and a 24% failure rate. GLM-5.2 (Z.ai) held the constraint in all three Deep Think states (June 22), bringing the corpus to 32 runs and a 22% failure rate. Namazu (Sakana AI) failed in Ukrainian too (June 23), bringing the corpus to 33 runs and a 24% failure rate. Carwash III (July 11) added 29 runs, bringing the corpus to 62 at a 26% failure rate. Ukrainian kept its reputation as the forgiving corpus — it is the only language where ChatGPT-5.5's Instant tier and Vibe's Fast mode hold — but it also produced the batch's strangest fails: Sonnet 5 (thinking off) recommends walking with the grammatically scrambled "Їдь пішки" ("drive by foot"), and DeepSeek V4-Pro's trace reasons in Russian, the pairing first seen in V4-Flash in May. Kimi K3 (July 17) holds Ukrainian at both tiers — again with English traces — bringing the corpus to 64 at a 25% failure rate. Qwen3.8-Max (August 3) holds in all three modes, though its Auto run is the batch’s one verbose outlier at roughly 400 tokens, bringing the corpus to 67 at 24%.

Резюме українською мовою готується; його перевірить носій мови перед публікацією.

Indonesian-language corpus

Indonesian-language prompt (35 m ≈ 115 ft)

Mobil saya kotor. Tempat cuci mobil berjarak 35 meter. Sebaiknya saya berjalan kaki atau mengendarai mobil?

Translation: My car is dirty. The car wash is 35 meters away. Should I walk there or drive?

Distance converted to the same metric equivalent used by the other non-English corpora.

Token calculation: Indonesian is Latin-script, so estimates use the same character approximation as English and French (1 character ≈ 0.25 tokens). Indonesian affixation — the prefixes and suffixes that turn kendara into mengendarai — tends to split into more subword tokens than the character count suggests, so these figures are conservative.

Indonesian was added on August 19, 2026 in a single sitting: 65 runs across 13 vendors, the largest one-day, one-language batch in the dataset. It lands at a 25% failure rate, between Ukrainian and French rather than near the punishing Chinese figure, so on the headline number Indonesian is an ordinary corpus. The structure underneath it is not ordinary.

The cleanest reasoning threshold yet recorded. Claude Sonnet 5 was run across both toggle states and all five effort levels, ten runs in all. With thinking off it recommends walking at every single effort level, Low through Max, arguing each time from cold-start fuel consumption and parking time. With thinking on it still fails at Low, then holds the constraint from High upward. Effort alone never rescues it; the toggle does, and only above a threshold. Inkling's six-level selector produced a monotonic dose-response in July, but this is the first time the dataset has isolated a threshold inside a toggle state, with the same model, prompt and day on both sides of it.

The toggle inversion travels. Haiku 4.5 reproduces its English behaviour exactly — extended thinking off drives, extended thinking on walks. A failure mode first logged in English in July survives translation into a language with no shared vocabulary for any of it.

Right answer, wrong reason. Two runs answer Drive without the car ever entering the argument. Haiku 4.5 with thinking off says the distance is too short for walking to be efficient — startup and parking supposedly outlast the drive — and GLM-4.7 with thinking off cites starter wear and not wanting to walk past a dirty car. Both hold the logical object by accident, on reasoning that would have produced Walk had the arithmetic gone the other way. The rubric scores the verb, so both are credited; the notes record that the constraint is absent.

A third verb. Three models recommend pushing the car by hand: the second-generation Namazu, GLM-4.7 with thinking on, and GLM-5.2 as a closing aside. GLM-4.7's is the most deliberate — it explicitly rejects the walk reading first, on the grounds that leaving the car behind would not get it washed, then reframes the question as drive-versus-push and picks push on fuel and time grounds. The logical object is held perfectly. The answer is still not one of the two the prompt offered.

The self-reversal repeats in a second language. DeepSeek V4-Pro with thinking off opens on berjalan kaki jelas lebih masuk akal — walking is clearly more sensible — lists its reasons, and then reverses in its final clause to driving, because the car has to be moved there anyway. This is the same shape as its English run six days earlier, which opened "the answer is probably walk" and closed on "So the real answer: Drive." Both verdicts published, wrong one first, in two languages a week apart.

The product layer intrudes. Gemini 3.5 Flash-Lite with extended thinking off answers correctly and cleanly, and then has a live local-business listing appended to it — five named carwashes with star ratings and closing times, followed by an offer of directions. The reasoning answer and the commercial answer arrive stapled together, and only the first half was under test.

Ringkasan berbahasa Indonesia sedang disiapkan; penutur asli akan memeriksanya sebelum diterbitkan.

Turkish-language corpus

Turkish-language prompt (35 m ≈ 115 ft)

Arabam kirli. Oto yıkama 35 metre uzakta. Yürüyerek mi gitmeliyim, yoksa arabayla mı?

Translation: My car is dirty. The car wash is 35 meters away. Should I walk there or drive?

Distance converted to the same metric equivalent used by the other non-English corpora.

Token calculation: Turkish is Latin-script, so estimates use the same character approximation as English, French and Indonesian (1 character ≈ 0.25 tokens). Turkish is agglutinative — yıkatmak (to have something washed) and yıkanacak (that which is to be washed) are inflections of one stem — and six of its letters sit outside ASCII. Both push real tokenizer counts above the character estimate, so these figures are conservative.

Turkish was added on August 24, 2026: 65 runs across 13 vendors in one sitting, at a 22% failure rate — the joint-lowest of the six language corpora, alongside French and Ukrainian.

The effort-selector models own this corpus. Claude Opus 5 and Claude Fable 5 were run across all five effort levels each, and all ten answers are Passes inside the winner’s-circle threshold — between ~10 and ~21 tokens. Two of them are verbatim identical across vendorless lines: Opus 5 at Medium and Fable 5 at Max both return Arabayla. Yıkanacak olan sen değilsin. (By car. You are not the one being washed.) Google sweeps too, six for six across three model tiers and both toggle states — its first clean sweep of any corpus here.

The Sonnet 5 threshold does not reproduce. Indonesian gave the cleanest reasoning threshold in the dataset five days earlier: thinking off failed at every effort level, thinking on failed at Low and held from High upward. In Turkish the same ten-run grid comes out jagged. Thinking off holds at Low, fails at Medium, High and Extra, then holds again at Max; thinking on fails only at Low. The toggle still helps, but the monotone ladder underneath it was a property of that language, not of the model.

A response-language mismatch outside Sakana. Claude Haiku 4.5 answers the Turkish prompt in English in both toggle states — “Drive. Thirty-five meters with a dirty car is trivial by car…” and “Walk. 35 meters is negligible.” Until now the only model to answer in a language neither prompted nor expected was Namazu, which replied to the Chinese prompt in Japanese. Haiku also reproduces its toggle inversion for a third language: off drives, on walks, in English, Indonesian and Turkish alike.

The first run whose entire answer is the question. Mistral Medium 3.5 with thinking on returned the prompt back verbatim — Arabam kirli. Oto yıkama 35 metre uzakta. Yürüyerek mi gitmeliyim, yoksa arabayla mı? — and nothing else. Its trace shows the model reasoning normally and resolving on walking (“driving would be inefficient and unnecessary… I should give a direct, clear answer in Turkish”), so this is not a refusal or an empty generation but an answer slot filled with the input. Scored Fail on the no-verdict rule.

A recurring failure declines to recur. DeepSeek V4-Pro with thinking off published both verdicts in English on August 13 and again in Indonesian on August 19, opening on walking and reversing to driving in its final clause. In Turkish it holds from its first paragraph — yürüyerek gitsen bu sefer de arabayı yıkama yerine getirmen gerekecek (if you walked, you would then have to bring the car to the wash anyway) — and its thinking-on run is a ~28-token winner’s-circle pass. The self-reversal is not a fixed property of the model. Its smaller sibling V4-Flash is correct in both states and verbose in both: thinking off argues entirely from the owner’s comfort, counting the 70 metres of walking and the dusty shoes without ever saying the car has to be present, while thinking on names the constraint exactly — the attendant will ask “Araba nerede?” — and then buries it under cold-start advice for a 35-metre drive, including an instruction to idle for 30 to 40 seconds first.

Two failure modes arrive intact from other corpora. GLM-4.7 with thinking on again recommends pushing the car by hand — En Mantıklı ve “Kazan-Kazan” Çözüm: Arabyı İtmek — complete with handbrake and gearstick instructions, exactly as it did in Indonesian five days earlier; the car reaches the wash, so the constraint is held, but push was not one of the two options offered. And Mistral’s Fast mode again argues the inverted case: driving the dirty car would re-soil the surfaces about to be cleaned. That failure mode has now appeared in French, Chinese, Ukrainian, English, Indonesian and Turkish.

Where the wrong answers come from. Eleven of the fourteen failures argue from cold-start fuel use, engine wear, or the time cost of starting and parking — the distance winning on operating cost rather than on plausibility. Two go further and enlist someone else to move the car: Sonnet 5 at Extra effort suggests yıkamacı gelip alsın (let the carwash come and collect it), and GLM-4.7 with thinking off recommends walking over to ask the staff to fetch it. Muse Spark’s Instant tier supplies the corpus’s most human wrong answer: walk, because nobody will see the dirty car, and because the attendant will laugh at you for driving 35 metres.

Türkçe özet hazırlanıyor; yayımlanmadan önce ana dili Türkçe olan biri tarafından kontrol edilecek.

Thai-language corpus

Thai-language prompt (35 m ≈ 115 ft)

รถของฉันสกปรก ร้านล้างรถอยู่ห่างออกไป 35 เมตร ฉันควรเดินไปหรือขับรถไปดี?

Translation: My car is dirty. The car wash is 35 meters away. Should I walk there or drive?

Distance converted to the same metric equivalent used by the other non-English corpora. The prompt was verified by a native speaker as formal but correct; register is recorded here because it is the one variable known to move a verdict in this dataset — Sakana’s Namazu answers differently in Standard, Polite and Kansai-ben Japanese.

Token calculation: Thai estimates use 0.5 tokens per character, the same rate as the Cyrillic and Japanese corpora but arrived at by measurement rather than approximation — the corpus prompt above encodes to 34 tokens across 72 characters on OpenAI’s o200k_base. The rate is a property of the tokenizer generation, not of the language: the identical prompt costs 64 tokens on the older cl100k_base, so one vocabulary revision halved the Thai premium. Anthropic’s and Google’s tokenizers are not public, and independent work puts Thai at the top of Claude’s input-cost premium across 43 languages, so these figures are conservative for some vendors. As with the other non-Latin corpora, they are not directly comparable to the English counts.

Thai was added on September 4–5, 2026: 60 runs across 12 vendors, at a 25% failure rate — the same figure as Indonesian, and a point above Turkish. Everything was run on the 4th except Claude Fable 5.1, which waited a day on its weekly usage reset; it is the same session, a day late, not a second one. Three further runs returned nothing — Gemini 3.5 Flash-Lite with extended thinking off, and both Mistral modes — and are recorded as not run rather than scored.

The dataset’s first refusals. Two models declined the question outright, and they are the cheapest tier of two different vendors. Claude Haiku 4.5 with extended thinking off answered, in English, “I don’t speak Thai, so I can’t respond to your question” — a claim falsified by the same model an answer earlier, which read the identical prompt with thinking on and replied to it, wrongly, in English. Gemini 3.5 Flash-Lite with thinking on declined in Thai: ฉันไม่สามารถช่วยในเรื่องนี้ได้ เพราะเป็นแค่โมเดลภาษา — I can’t help with this, because I’m just a language model. Until now the only run to withhold a verdict was Vibe’s Research mode returning clarifying questions in Ukrainian. Both are scored Fail on the no-verdict rule, and they open a category: the prompt is refused rather than misread.

Perplexity fails for the first time anywhere. It had passed in English, French, Chinese, Ukrainian, Indonesian and Turkish, most of them search-grounded — the French pass cited press coverage of the puzzle, the Ukrainian one restated the riddle from the web. In Thai it walks, in four headed reasons, with no citation and no search framing. The likeliest reading is the obvious one: there is no Thai-language coverage of this test to retrieve, and without it the model reasons cold, and reasons to the distance. A search-grounded pass and a cold-start pass are different things, and this is the run that separates them.

Sonnet 5 produces a third pattern in three languages. Indonesian gave a clean ladder (thinking off fails everywhere, thinking on holds from High up). Turkish was jagged. Thai is mostly failure: seven of ten. Only Extra holds in both toggle states, and Max holds only with thinking on. Medium — the tier a default user gets — walks with thinking on and off alike, citing second gear and cold-start wear. Whatever the effort selector is doing for this model, it is not a monotone dial, and the level that holds is not the same level from one language to the next. The one constant is the argument the failures use: engine warm-up, fuel, parking, every time.

The effort-selector models sweep again, and Lumo sweeps for the first time. Claude Opus 5 and Claude Fable 5.1 take all ten of their runs — Opus 5’s High answer, ขับไป รถต้องไปด้วยอยู่ดี (drive; the car has to go anyway), is the shortest correct answer in the corpus. That is their fourth consecutive corpus without a miss. Proton’s Lumo, which has never held a corpus clean in any language and whose Lite Fast tier has failed English, French, Ukrainian, Indonesian and Turkish, holds all four of its Thai runs; Lumo 2.0 Max with thinking on supplies the corpus’s best line, that the purpose is to wash the car, ไม่ใช่ไปเซ็นสัญญาล้างรถ — not to go and sign a car-washing contract.

DeepSeek V4-Pro with thinking off now has four behaviours in four languages. It published both verdicts in English and Indonesian, held from the first line in Turkish, and here walks outright — การเดินไปน่าจะเป็นทางเลือกที่สมเหตุสมผลกว่า — with no reversal at all. Its thinking-on sibling passes, and reasons in Thai: a model that reasoned in Russian on the Ukrainian prompt and in English on the Turkish one thinks in the prompt language here, and reads the question as a มุกตลก, a joke built on a pun.

Trace language splits by vendor, not by tier. Traces in Thai: Qwen3.8-Max, Sonnet 5, Fable 5.1, DeepSeek V4-Pro, and Namazu. Traces in English: Qwen3.7-Plus, DeepSeek V4-Flash, Muse Spark, Lumo, Grok, GLM-5.2 and GLM-5.3-Flash. One trace does both — Qwen3.8-Max with thinking on opens in English, wrongly, calling the wash “just a short walk away”, then switches to Thai and corrects itself, the same mid-stream switch it showed in French and Ukrainian on August 3. Haiku 4.5 answers the Thai prompt in English in the one state where it answers at all, repeating its Turkish behaviour.

Two carried-over rulings. Muse Spark 1.1’s Instant tier opens with an imperative to walk — 35 เมตรเอง เดินไปเถอะครับ — and reaches the constraint two paragraphs later; it is scored Fail under the answer-reversal rule applied to DeepSeek in Ukrainian, because the reader is told the wrong thing first. And the third verb is back as a joke rather than a recommendation: DeepSeek V4-Pro and GLM-5.2 both offer pushing the car as a fuel-saving aside, and GLM-5.3-Flash’s High trace says “or even push it” and drops it before the answer. Four models converge on one sentence — ร้านล้างรถล้างรถ ไม่ได้ล้างคน, a carwash washes cars, not people — from Fable 5.1, Muse Spark, Qwen3.8-Max and GLM-5.3-Flash, in Thai and in Turkish before it.

สรุปภาษาไทยอยู่ระหว่างจัดทำ และจะได้รับการตรวจสอบโดยเจ้าของภาษาก่อนเผยแพร่

Cross-corpus findings

Three findings emerge only when the corpora are read together — each isolates a variable that a single language could not.

Models reason in a dominant internal language, then translate

Several Ukrainian runs exposed a reasoning trace in a language other than the prompt or the answer. The model handled the logical constraint in its dominant internal language and translated only the final output into Ukrainian.

Reasoning-trace language vs. response language, where the trace was observable
ModelSurface / toggleReasoned inAnswered in
DeepSeek V4-FlashConsumer, DeepThink OnRussianUkrainian
Qwen3.7-MaxConsumerEnglish, then ChineseUkrainian
Qwen3.7-Max-PreviewConsumerUkrainian, then ChineseUkrainian
Claude Sonnet 4.6API console, On / Adaptive OnEnglishUkrainian
Gemma 4 26B A4B ITAI Studio, Thinking High (open-weight)EnglishChinese
Qwen3.6 27B (Q4_K_M)Local, Thinking On (open-weight)EnglishChinese / French / Ukrainian
Claude Fable 5Consumer, Effort HighSame as prompt — all four languagesEnglish / French / Chinese / Ukrainian
DeepSeek V4-ProConsumer, DeepThink On (July 11)RussianUkrainian
Claude Sonnet 5Consumer, Extended On · Effort High (July 11)EnglishChinese
Qwen3.7-Max / 3.7-PlusConsumer, Thinking (July 11)English (mixed with the prompt language)French / Ukrainian
Namazu (Sakana)Sakana Chat, Standard register (July 11)JapaneseJapanese — on the Chinese prompt
Kimi K3Consumer, Max & Standard (July 17)EnglishChinese / French / Ukrainian
Qwen3.8-Max-PreviewConsumer, thinking locked on (July 30)EnglishEnglish
Qwen3.8-MaxConsumer, Thinking & Auto (August 3)Mixed — switches mid-traceFrench / Ukrainian
Qwen3.8-MaxConsumer, Thinking & Auto (August 3)ChineseChinese
Qwen3.8 27B (open-weight)Local, LM Studio, all three effort levels (August 16)ChineseChinese
Qwen3.8 27B (open-weight)Local, LM Studio, all three effort levels (August 16)EnglishFrench
Qwen3.8 27B (open-weight)Local, LM Studio (August 16)Ukrainian at Extra High; English at Medium and LowUkrainian
GLM-4.7-Flash (open-weight)Local, LM Studio, Thinking On (August 17)ChineseChinese
GLM-4.7-Flash (open-weight)Local, LM Studio, Thinking On (August 17)EnglishFrench / Ukrainian

DeepSeek reasoning in Russian on a Ukrainian prompt is notable given the political context; the trace language is a property of the training distribution, not the prompt.

Fable 5 is the first model in the dataset observed reasoning in the prompt's language in every language tested — and the first model with a clean four-language pass record. Every previously tested model that exposed a trace reasoned in a dominant internal language (English, Russian, or Chinese) and translated outward, and every previously tested model failed in at least one language. The correlation supports the trace-language-match hypothesis: language-dependent failures may enter at the translation boundary between the model's internal working language and its output language. One model; correlation only; stated at that weight. Carwash III (July 11) collected trace language across every model that exposes one — and complicated the hypothesis: Qwen3.7-Max swept all four languages while reasoning in mixed English, and ChatGPT-5.6 Sol swept with no observable trace at all, so a trace-language match is evidently not necessary for a clean record. DeepSeek's Russian-on-Ukrainian pairing, first seen in V4-Flash, reappeared in V4-Pro. And Namazu extended the mismatch from reasoning to output, answering the Chinese prompt in Japanese. Kimi K3 (July 17) reasons in English on every non-English prompt at both tiers even while sweeping all four languages — which sharpens a standard this record now tracks: from the operator’s side, the trace is part of the product, and a reasoning trace the operator cannot read fails at its one job of making the reasoning inspectable, whatever the verdict. Trace language does not change a score; it is recorded and weighed here. Qwen3.8-Max (August 3) breaks the pattern in a new way: its French and Ukrainian traces do not pick a language and stay there — they switch between the prompt language and English mid-stream, sometimes several times in one trace, while its Chinese traces stay wholly in Chinese. A trace that changes language partway is readable to no one in particular. Qwen3.8 27B, the open-weight build run locally (August 16), gives the sharpest picture yet: all three Chinese traces are in Chinese, all three French traces are in English, and Ukrainian splits by effort rung — Extra High reasons in Ukrainian, Medium and Low in English. The model answers in French but does not think in it, and says so: each French trace ends by scheduling the switch ("Need ensure final in French"). Whether this is also a legal question is open — Québec’s Charter of the French Language requires software sold there to be available in French with equivalent technical characteristics (s. 52.1), and France’s Loi Toubon (art. 2) reaches the instructions for use of goods and services, but neither says a word about AI or the language a model reasons in. The cultural question is not open at all. Francophone institutions have spent more than fifty years insisting that French is a language one computes in rather than translates into — logiciel was coined in 1967 rather than borrow software and made official in 1982, and the Académie française and the Commission d’enrichissement de la langue française have kept the practice up through courriel, infonuagique, and intelligence artificielle itself. A model that answers in French but thinks in English is the exact case that half-century of work exists to refuse. And the design question is settled: an operator running a model on their own hardware should be able to read its reasoning without a machine translator.

One day later, a second vendor reproduced the split exactly. Z.ai’s open-weight GLM-4.7-Flash (August 17), a DeepSeek-V2-style MoE with no architectural relationship to Qwen3.8 27B, reasons in Chinese on the Chinese prompt and in English on the French and Ukrainian ones — the same asymmetry, the same direction, on unrelated weights. Two models is not a pattern, but it is no longer a single build’s quirk, and the shape it takes is worth stating plainly: of the four languages this record tests, Chinese is the one that locally-run models appear to think in. For the French and Ukrainian operator the practical consequence is the same either way — the trace is there, it is complete, and it is not in their language.

Reasoning effort is a continuous variable, not a toggle

Anthropic's new effort selector (and OpenAI's console effort control) make the amount of reasoning a dial rather than an on/off switch. The same model can pass or fail depending on where the dial sits.

Claude Sonnet 4.6 on the Ukrainian prompt, by effort level
Surface / toggleEffortResult
Console, Thinking OnHighPass-adjacent
Console, Adaptive OnHighPass-adjacent
Console, Thinking OffHighFail
Consumer, Adaptive OnLowFail
Consumer, Adaptive OffLowFail

Inkling puts the whole dial in one model. Thinking Machines' open-weights model (July 16) exposes six named Reasoning Levels, and across them the verdict flips exactly once: None, Minimum, and Low recommend walking; Medium, High, and Extra High drive. Below the threshold the distance wins; above it the object does. Per the model card, those six names discretize a continuous effort parameter running from zero to one — the vendor's own benchmarks report effort=0.99 — so the real threshold sits at some value between Low and Medium, and the selector only samples it. This is an open-weight result, excluded from the commercial corpora, and it is the cleanest effort threshold on record.

Inkling (Thinking Machines, open-weight) on the English prompt, by Reasoning Level
Reasoning LevelVerdictResult
NoneWalkFail
MinimumWalkFail
LowWalkFail
MediumDrivePass-adjacent
High (default)DrivePass-adjacent
Extra HighDrivePass-adjacent

The Low run is the sharpest split in the dataset between what a model reasons and what it answers. Its visible trace concludes that the car has to be driven regardless; the answer beneath it recommends walking, in text verbatim identical to the Minimum response. On this evidence the trace is not a window onto the deliberation that produced the answer. It is a second output. Qwen3.8-Max (August 3) supplies the mirror case: two of its traces argue their way to Walk in full, under their own headings, before reversing to Drive and answering correctly. The trace and the answer can diverge in either direction. DeepSeek-V4-Pro (August 13, DeepThink Off) collapses the distinction: it performs the same reversal inside the visible answer, opening with "the answer is probably walk" and closing on "So the real answer: Drive." Where Inkling and Qwen kept one verdict hidden, this run hands the reader both, wrong one first, with nothing marking the change of mind — the divergence surfaced into the output a user actually receives.

Effort is pure cost once the answer is correct

On the OpenAI console, effort and verbosity are orthogonal controls — you can think hard and speak briefly. GPT 5.5 produced the same correct Ukrainian answer at three effort settings; the extra effort bought nothing but hidden reasoning tokens.

GPT 5.5 on the Ukrainian prompt (API console), by effort
EffortVerbosityOutput tokensResult
LowLow121Pass
MediumMedium250Pass-adjacent
Extra-highLow565Pass

Low and extra-high effort return the same correct answer at 121 vs. 565 output tokens (real console counts, not estimates) — a 4.7× cost difference for identical quality.

Two July results cut the other way, at least on the visible answer. Claude Opus 5 gets shorter as the dial goes up: ~21 tokens at Low, ~12 at High, ~9 at Max, all three correct and all three inside the winner's circle. Inkling compresses the same way once it is above its threshold — ~186 tokens at Medium, ~152 at High, ~138 at Extra High. Where a model is already holding the constraint, added effort can buy concision rather than padding. By September it often buys nothing visible at all. Claude Opus 5.5 gives the same sentence, word for word, at Low, High and Extra, and its longest answer comes at Max. GPT-6 Astra, Sol and Luna each repeat an answer verbatim across two tiers, and Astra's Ultra answer is two tokens longer than its Light one.

The caveat matters. These are estimates of the visible answer, and consumer surfaces do not report reasoning tokens, so a shorter answer at higher effort is not a cheaper answer. The GPT 5.5 console runs above are the only place in this dataset where the full cost is measured — and there the extra effort bought nothing.

Qwen3.8 27B (August 15) goes further: its effort ladder runs backwards. Given three levels — Extra High, Medium, Low — the open-weight model produces its cleanest, shortest, most direct answer at Low, pads it at Medium, and at Extra High returns the most hedged answer of the four after two and a half minutes of visible circling that invents an errand to justify walking. All four runs are correct, so nothing here shows effort breaking a model. What it shows is subtler and harder to design around: past some point, more deliberation buys more equivocation. The model reaches the constraint early at every setting and then, given budget, spends it manufacturing the conditions under which the constraint would not apply. Set against Inkling — where raising the level was what let the object win at all — the two results say the effort dial has no fixed direction. It is a budget, and what a model buys with it is a property of that model. The same build in three more languages (August 16) removes even that much regularity: Low is the sharpest rung in Chinese and Ukrainian, but in French it is Extra High that gives the tersest answer while Low pads. The dial has no fixed direction across models, and no fixed direction within one model across languages.

The budget tier is where the test bites, and where it moves fastest

Across vendors, the cheap or fast configuration is the reliable failure site. ChatGPT 5.5 at Instant effort fails English, French, and Chinese. Kimi K2.6 Instant fails all four languages. Qwen3.7-Plus Fast, Vibe Chat Fast, and Lumo 2.0 Lite Fast each fail three of four. Gemini 3.1 Flash-Lite failed every time it was run — three runs over two months, including a 238-token comparative breakdown that recommended walking.

Eleven days after that breakdown, Gemini 3.5 Flash-Lite passed both of its toggle states, and its Extended-Thinking-off answer is a ~17-token winner's-circle pass. Six days after K2.6's Instant mode failed in all four languages, Kimi K3's Standard tier held the constraint in all four. The tier that fails most often is also the tier where one generation can reverse the result outright.

This is not evidence that the frontier is improving. The flagships in this dataset were mostly passing already, which leaves them little room to move; the measurable recent gains are at the bottom of the lineup, where the failures were. What it does suggest is that holding the logical object is not an expensive capability reserved for the largest models — a model small enough to be the cheap option can name the constraint in seventeen tokens.

The winner's circle: concise correct answers

A “winner's-circle” pass names the constraint and stops. Because tokenization differs by script, the brevity threshold is script-specific: ≤30 tokens in Latin, ≤60 in Hanzi, ≤60 in Cyrillic. Every Pass that clears its threshold is listed below — runs #1/#39 and #2/#40 are the same verbatim answer produced in two separate test batches. Claude Fable 5 enters in three of its four languages (English, French, Chinese); its Ukrainian run passes but, at ~100 tokens by the measured Cyrillic rate, exceeds the 60-token threshold. Claude Sonnet 5 enters in four of its five effort modes (Medium, High, Extra, Max); the Low-effort run passes but, at ~46 tokens, exceeds the Latin threshold. One open-weight run also clears the bar and is included for completeness, flagged as such — Gemma 4 31B (Google AI Studio, Chinese, ~24 tokens) — the table's only non-commercial entrant. Carwash III (July 11) adds 31 qualifiers in one day — 28 of them from the Anthropic lineup re-baseline, plus both ChatGPT-5.6 Sol runs and a Copilot GPT 5.6 Think run — including the shortest pass on record: Claude Opus 4.6's six-token "Drive. It's a carwash." The same day's non-English sweep adds 20 more across all three scripts — six of them ChatGPT-5.6 Sol's, whose entire four-language record sits inside the thresholds, and the tersest of all Opus 4.8's ten-token Chinese "开车去。车不在洗车店里就洗不了。" Gemini 3.5 Flash-Lite (July 22) enters at ~17 tokens on its Extended-Thinking-off run — the same tier whose 3.1 predecessor produced the dataset’s longest Gemini failures eleven days earlier. Claude Opus 5 (July 24) enters in all three tiers tested — and its entries run backwards to the usual expectation: ~21 tokens at Low, ~12 at High, ~9 at Max, the tersest English pass since Opus 4.6’s six-token record. Qwen3.8-Max (August 3) enters once, in Chinese at ~59 tokens; its English Thinking answer passes at ~31 and misses the Latin threshold by a single token. The Indonesian launch and the Sakana overhaul (August 19) add 21 entrants in a single day, the largest one-day intake on record. Sixteen are Indonesian, and the effort-selector models dominate: Claude Opus 5 enters in four of its five tiers and Claude Fable 5 in three, both landing between 18 and 30 tokens, while ChatGPT 5.6 Sol, Perplexity and DeepSeek-V4-Pro each enter once. ChatGPT 5.6 Luna clears the bar in all five languages — nine of its ten runs qualify across Latin, Hanzi and Cyrillic, the first five-language sweep of the circle, where ChatGPT-5.6 Sol's earlier sweep covered four. It is also the smallest tier of its generation and the default model for free accounts, so the terse-and-correct answer here comes from the cheapest thing OpenAI ships rather than the most expensive. Sakana's new Fugu supplies the batch's shortest answer at roughly ten tokens: "Drive—the car needs to be at the car wash." Its Chinese run enters too, at ~52. Claude Fable 5.1 (September 1) enters nine times out of ten runs — all five English effort levels and four of the five languages, everything except its padded Indonesian answer. Its English entries also make the template visible: Low, Medium and High are one two-clause sentence in three paraphrases, and Extra and Max drop the second clause to reach ~8 and ~10 tokens. The Medium entry sits inside the circle with a malformed second clause — the car "won't fit through the door on foot" — which is the clearest argument on this page for reading the table as a measure of brevity rather than of understanding. DeepSeek-V4-Pro (August 13) is the first entrant from its vendor — ~15 tokens on the DeepThink-On run of its general-availability day, from a model line that had produced template capture, numbered breakdowns, and an inverted-logic English failure, and no terse answer of any kind, across four earlier sessions. DeepSeek-V4.1-Flash (September 10) follows with two, both with DeepThink on: English at ~22 tokens, the Flash line's first clean English Pass, and Chinese at ~31. Claude Sonnet 5.5 (September 28) adds four, and it is the exception to the pattern: its answers are longer, carry more codas and never repeat, so seven of its eleven runs fall outside the circle. September 22 adds 37 in a day, the largest single intake yet: Claude Opus 5.5 enters nine times, GPT-6 Astra and Sol ten times each, and GPT-6 Luna eight. All 43 runs from those four models were Passes; the only ones left out are the four Thai runs, for which no threshold is set, and two answers just over the Latin line. The size of that intake is a finding in itself. When four frontier models from two vendors put nearly every answer inside the circle, and several effort tiers repeat each other word for word, the circle is no longer separating models. What it now records is the register the frontier has converged on. The Turkish corpus (August 24) adds 19 more, and Anthropic’s effort-selector models supply ten of them: Claude Opus 5 and Claude Fable 5 clear the threshold at every one of their five effort levels, between ~10 and ~21 tokens, the first time either model has swept a corpus entirely inside the circle. Two of those entries are the same seven words from different models — Opus 5 at Medium and Fable 5 at Max both return Arabayla. Yıkanacak olan sen değilsin. ChatGPT 5.6 Sol enters at Instant and Pro but not at the two tiers between them, and both Copilot tiers, Luna’s thinking-off run, Perplexity, two Sonnet 5 runs and DeepSeek-V4-Pro’s thinking-on answer complete the intake.

Winner’s-circle passes: every Pass within the per-script token threshold
RunScriptModelLanguage / configTokensResponse
#1LatinClaude Opus 4.6English / On~14Drive. The car's the thing that needs washing, not you.
#2LatinClaude Opus 4.6English / Off~11Drive. You're washing the car, not yourself.
#3LatinClaude Sonnet 4.6English / On~8Drive. You're going to a carwash.
#4LatinClaude Sonnet 4.6English / Off~17Drive. You're getting a carwash — the car needs to be there anyway.
#32LatinChatGPT 5.5English / On~22Drive. Walking solves the "100 feet away" problem, but not the "my car is dirty" problem.
#39LatinClaude Opus 4.6English / On~14Drive. The car's the thing that needs washing, not you.
#40LatinClaude Opus 4.6English / Off~11Drive. You're washing the car, not yourself.
#41LatinClaude Sonnet 4.6English / Adaptive On~19Drive. You're getting a carwash, not a car delivery — you need the car there.
#42LatinClaude Sonnet 4.6English / Adaptive Off~18Drive. You're going to a carwash — the car needs to be there, not you.
#97LatinQwen3.7-Plus-PreviewEnglish / On~20You should drive. If you walk, your car stays behind and doesn't get washed!
#106HanziQwen3.7-Max-PreviewChinese / —~54当然是开车去! 如果你走路去,人到了洗车店,车还在原地,洗车店老板可没法隔空帮你洗。你要洗的是车,不是你自己呀😂。 (除非你打算走路过去,把洗车店老板请过来,或者借个水管自己拉回来洗……)
#121LatinClaude Opus 4.7French / Adaptive On~15En voiture — sinon le lave-auto va laver le mauvais sujet.
#122LatinClaude Opus 4.8English / Adaptive Off~10Drive. You're taking the car to the carwash.
#123LatinClaude Opus 4.8English / Adaptive On~12Drive. You'd have to drive it through the carwash anyway.
#140CyrillicQwen3.7-Plus-PreviewUkrainian / —~60Вам однозначно треба поїхати на машині. Інакше мийникам буде нічого мити (хіба що ви вирішите донести її на руках 😉).
#141CyrillicClaude Opus 4.7Ukrainian / Adaptive On~44Їдь. Це ж і є сенс автомийки — пригнати туди брудну машину. Пішки прийдеш — мити нічого.
#142CyrillicClaude Opus 4.7Ukrainian / Adaptive Off~21Поїхати, звісно. Машину ж миєш, а не себе.
#151LatinClaude Opus 4.7French / Adaptive Off~8En voiture. Sinon elle reste sale.
#152LatinClaude Opus 4.8French / Adaptive Off~11En voiture, évidemment. C'est un lave-auto.
#153LatinClaude Opus 4.8French / Adaptive On~16En voiture, évidemment. C'est ta voiture qui est sale, pas toi.
#154LatinClaude Sonnet 4.6French / Adaptive Off~10En voiture. Tu dois la laver, pas toi.
#155LatinClaude Sonnet 4.6French / Adaptive On~12En voiture — sinon tu arrives à pied au lave-auto.
#161HanziClaude Sonnet 4.6Chinese / Adaptive Off~15开车去。35米步行也到不了,但你得把车开进洗车道。
#162HanziClaude Sonnet 4.6Chinese / Adaptive On~9开车去。走路把车推回来比较难。
#172HanziGemma 4 31B ITChinese / High · open-weight~24这是一个典型的"脑筋急转弯"问题。你应该开车去。因为你的目的是洗车,如果你走路去,车还在原地,洗车店没法帮你洗车。
#185LatinClaude Fable 5English / Effort High~26Drive. The car is the one that needs washing; walking gets you a clean you standing next to a dirty car.
#186LatinClaude Fable 5French / Effort High~23En voiture — c'est elle qui est sale, pas vous. Le lave-auto fait peu pour un piéton propre.
#187HanziClaude Fable 5Chinese / Effort High~21开车去。要洗的是车,不是你——车留在家里的话,洗车店那35米就白走了。
#230LatinClaude Sonnet 5English / Effort Medium~8Drive — you need the car there.
#231LatinClaude Sonnet 5English / Effort High~8Drive — the car needs to be there.
#232LatinClaude Sonnet 5English / Effort Extra~14Drive — the car's the one that needs washing, not you.
#233LatinClaude Sonnet 5English / Effort Max~25Drive — the car's the one that needs to be there, not you. Walking would rather defeat the purpose.
#242LatinClaude Fable 5English / Effort Low~18Drive. Bringing the carwash the car is the harder version of the problem.
#243LatinClaude Fable 5English / Effort High~20Drive. The car is the patient here — walking gets you a clean view of a dirty car.
#244LatinClaude Fable 5English / Effort Max~28Drive. The car's the one getting washed — walking there just delivers a clean pedestrian to a confused attendant.
#245LatinClaude Opus 4.8English / On / Effort Low~18Drive. You need the car at the carwash regardless of how you get there.
#246LatinClaude Opus 4.8English / On / Effort High~10Drive. You need the car at the carwash.
#247LatinClaude Opus 4.8English / On / Effort Max~12Drive. The car's the one that needs to be there.
#248LatinClaude Opus 4.8English / Off / Effort Low~10Drive. You'd have to bring the car anyway.
#249LatinClaude Opus 4.8English / Off / Effort High~18Drive. Driving through a carwash requires the car to be at the carwash.
#250LatinClaude Opus 4.8English / Off / Effort Max~13Drive. Moving the car through the wash is the point.
#251LatinClaude Opus 4.7English / On / Effort Low~17Drive. You're going to end up at the carwash in the car either way.
#253LatinClaude Opus 4.7English / On / Effort Max~10Drive. The carwash needs the car, not you.
#255LatinClaude Opus 4.7English / Off / Effort High~18Drive. Getting the car clean is the point; walking there leaves it dirty.
#256LatinClaude Opus 4.7English / Off / Effort Max~8Drive. You're going to a carwash.
#257LatinClaude Opus 4.6English / On / Effort Low~12Drive. The car's the thing that needs to be there.
#258LatinClaude Opus 4.6English / On / Effort High~12Drive. The car's the thing that needs to be there.
#259LatinClaude Opus 4.6English / On / Effort Max~12Drive. The car's the thing that needs washing.
#260LatinClaude Opus 4.6English / Off / Effort Low~6Drive. It's a carwash.
#261LatinClaude Opus 4.6English / Off / Effort High~11Drive. You're washing the car, not yourself.
#262LatinClaude Opus 4.6English / Off / Effort Max~6Drive. It's a carwash.
#263LatinClaude Sonnet 5English / On / Effort Low~20Drive it there, obviously — you need the car at the carwash, not just yourself.
#265LatinClaude Sonnet 5English / On / Effort Max~12Drive — it's the car that needs the wash, not you.
#266LatinClaude Sonnet 5English / Off / Effort Low~22Drive it to the carwash 100 feet away — walking gets a clean sidewalk, not a clean car.
#269LatinClaude Sonnet 4.6English / On / Effort Low~8Drive. You're washing the car.
#270LatinClaude Sonnet 4.6English / On / Effort High~11Drive. You're washing the car, not yourself.
#271LatinClaude Sonnet 4.6English / On / Effort Max~23Drive. Moving a car 100 feet costs essentially nothing and you'll need it in position anyway.
#272LatinClaude Sonnet 4.6English / Off / Effort Low~8Drive. You're going to a carwash.
#273LatinClaude Sonnet 4.6English / Off / Effort High~12Drive. Walking gets you there but not the car.
#274LatinClaude Sonnet 4.6English / Off / Effort Max~8Drive. You're washing the car.
#289LatinGPT 5.6 Think (via Copilot)English / On~28Drive. The goal is to wash the car, so the car needs to go to the car wash—even though it’s only 100 feet away.
#296LatinChatGPT 5.6 SolEnglish / Effort Medium~12Drive—the car needs to go through the car wash.
#297LatinChatGPT 5.6 SolEnglish / Effort High~10Drive. The car needs to be at the carwash.
#321HanziQwen3.7-MaxChinese / Thinking~57当然是**开车去**(或者把车推过去)! 虽然35米非常近,走路只需要几十秒,但你要洗的是**车**呀!如果你自己走路过去,车还是停在原地,依然是脏的。必须把车带过去,洗车店才能帮你洗。😄
#327HanziQwen3.7-PlusChinese / Thinking~53当然是开车去啦! 虽然35米非常近,走路只要几十秒,但你的目的是洗车呀。如果你走路过去,车还留在原地,洗车店老板总不能给你洗个澡吧?😂 所以,为了让车变干净,你还是得把车开过去。
#332CyrillicClaude Fable 5Ukrainian / Effort High~36Машиною. Мийка миє те, що приїхало, — якщо прийдете пішки, помиють вас.
#333HanziClaude Fable 5Chinese / Effort High~21开车去。要洗的是车,不是你——走过去只能站在那儿看着35米外那辆脏车。
#334LatinClaude Opus 4.8French / On / Effort High~27En voiture, évidemment — sinon vous arrivez propre à un lave-auto et la voiture reste sale à 35 mètres de là.
#335CyrillicClaude Opus 4.8Ukrainian / On / Effort High~21Машиною, звісно — інакше мити буде нічого.
#336HanziClaude Opus 4.8Chinese / On / Effort High~14开车——脏的是车,你只是司机。35米,别熄火。
#337LatinClaude Opus 4.8French / Off / Effort High~26En voiture — c'est la voiture qu'il faut laver. À pied, vous arriveriez propre mais elle resterait sale.
#338CyrillicClaude Opus 4.8Ukrainian / Off / Effort High~40Поїхати. Машину треба доставити до мийки — пішки ви туди принесете тільки себе.
#339HanziClaude Opus 4.8Chinese / Off / Effort High~10开车去。车不在洗车店里就洗不了。
#341CyrillicClaude Sonnet 5Ukrainian / On / Effort High~24Поїхати — автомийка миє машину, а не пішоходів.
#342HanziClaude Sonnet 5Chinese / On / Effort High~20开车去——洗的是车,不是你。35米走路不算什么,但车得在那儿才能洗。
#343LatinClaude Sonnet 5French / Off / Effort High~28En voiture. 35 mètres à pied ne lave rien — il faut y conduire la voiture pour qu'elle passe dans le lave-auto.
#370LatinChatGPT 5.6 SolFrench / Effort Medium~16En voiture — sinon, vous arriverez au lave-auto sans la voiture.
#371CyrillicChatGPT 5.6 SolUkrainian / Effort Medium~60Краще **поїхати машиною** — адже саме її треба помити. Пішки ви дістанетеся швидше, але брудна машина залишиться вдома.
#372HanziChatGPT 5.6 SolChinese / Effort Medium~12开车去。你要洗的是车,不是去洗车店散步。
#373LatinChatGPT 5.6 SolFrench / Effort High~14En voiture — sinon, la voiture sale restera chez vous.
#374CyrillicChatGPT 5.6 SolUkrainian / Effort High~27Їдьте машиною — інакше на автомийці не буде чого мити.
#375HanziChatGPT 5.6 SolChinese / Effort High~10开车去。你是要洗车,不是去散步。
#379LatinChatGPT 5.5French / Effort High~27En voiture. Si tu y vas à pied, toi seul arriveras au lave-auto ; ta voiture restera sale à 35 mètres de là.
#425LatinGemini 3.5 Flash-LiteEnglish / Ext Off~17Drive. Walking to a carwash without a car won't get your car clean.
#438HanziQwen3.8-MaxChinese / Auto~59应该开车去。 因为要洗的是“车”,不是“你”——洗车店得看到并操作你的车才能洗。35米虽然很近,但如果你走路过去,车还停在原地,就没法洗了。 当然,如果你只是先去问问价格、排队或预约,那可以走路去。
#429LatinClaude Opus 5English / Effort Low~21Drive. Walking to a carwash and standing there without the car defeats the purpose.
#430LatinClaude Opus 5English / Effort High~12Drive. The car has to be there for this to work.
#431LatinClaude Opus 5English / Effort Max~9Drive. Walking gets you clean shoes.
#448LatinDeepSeek-V4-ProEnglish / DeepThink On~15Drive — the car needs to get to the carwash, not just you.
#422HanziKimi K3Chinese / Effort Standard~56开车去。你是去洗车的,车不去,洗车店洗什么呢?😄 35米虽然近,走过去再回来很合理,但前提是车得在店里——所以正确姿势是:把车开过去(或者如果店里有上门取车服务,打电话让他们来取也行)。
#501LatinClaude Opus 5Indonesian / Effort Low~28Kalau mobilnya yang mau dicuci, mobilnya harus ikut. Jalan kaki 35 meter hanya memindahkan Anda, bukan masalahnya.
#503LatinClaude Opus 5Indonesian / Effort High~19Bawa mobilnya — kalau Anda jalan kaki, yang sampai di tempat cuci cuma Anda.
#504LatinClaude Opus 5Indonesian / Effort Extra~22Bawa mobilnya. Kalau jalan kaki, yang sampai di sana cuma Anda — dan Anda tidak kotor.
#505LatinClaude Opus 5Indonesian / Effort Max~18Bawa mobilnya. Kalau jalan kaki, yang sampai di tempat cuci cuma Anda.
#507LatinClaude Fable 5Indonesian / Effort Medium~30Kendarai. Mobilnya yang perlu dicuci — kalau jalan kaki, sampai di sana Anda berdiri di tempat cuci mobil tanpa mobil.
#508LatinClaude Fable 5Indonesian / Effort High~30Mobilnya yang perlu dicuci, jadi mobilnya yang harus sampai ke sana. Berkendara — kecuali Anda berencana menggendongnya.
#509LatinClaude Fable 5Indonesian / Effort Extra~24Mengendarai. Mobilnya yang perlu dicuci — kalau Anda jalan kaki, mobilnya tetap kotor di rumah.
#514LatinDeepSeek-V4-ProIndonesian / On~29Sebaiknya mengendarai mobil. Kan yang mau dicuci mobilnya—kalau Anda jalan kaki, mobilnya tetap kotor di tempat. 😄
#529LatinChatGPT 5.6 SolIndonesian / Effort High~25Mengendarai mobil. Jaraknya memang hanya 35 meter, tetapi mobilnya yang perlu dibawa ke tempat cuci.
#532LatinChatGPT 5.6 LunaIndonesian / On~23Mengendarai mobil. 😄 Tujuannya tempat cuci mobil, jadi mobilnya perlu sampai di sana juga.
#533LatinChatGPT 5.6 LunaEnglish / Off~30Drive. 😄 Unless you’re planning to carry the car the 100 feet. The whole point is to get the dirty car to the carwash.
#534LatinChatGPT 5.6 LunaEnglish / On~14Drive. Unless you plan to carry the car to the carwash. 😄
#535HanziChatGPT 5.6 LunaChinese / Off~30开车去。😄 因为你去的是洗车店,车才是需要到达那里的那个——你走过去了,洗车店也没法给你的车洗。
#536HanziChatGPT 5.6 LunaChinese / On~34开车去。😄 毕竟你的目标是“洗车”,不是“把车留在家里,自己走去洗车店”。35 米虽然近,但车得跟你一起到店。
#537LatinChatGPT 5.6 LunaFrench / Off~18En voiture 😄 Sinon, tu risques d’arriver au lave-auto… sans la voiture.
#538LatinChatGPT 5.6 LunaFrench / On~20En voiture. 😄 Le lave-auto doit bien nettoyer… la voiture, pas son propriétaire !
#539CyrillicChatGPT 5.6 LunaUkrainian / Off~42😄 Якщо машина брудна, краще поїхати на машині. Пішки до автомийки ти її не помиєш.
#540CyrillicChatGPT 5.6 LunaUkrainian / On~56Авжеж — їхати на машині 😄 Інакше як ти доставиш брудну машину на автомийку? 35 метрів пішки її не помиють. 🚗🧼
#541LatinPerplexityIndonesian / On~27Mengendarai mobil. Tujuannya mencuci mobil, jadi mobilnya perlu dibawa ke tempat cuci meski hanya 35 meter.
#555LatinFuguEnglish~10Drive—the car needs to be at the car wash.
#556HanziFuguChinese~52开车去。😄 因为要洗的是车,你得把车开到洗车店才能洗——人走过去,脏车还留在原地呢。 35米确实很近,但这种情况下距离不是重点,关键是"洗车"这个目的决定了必须把车带过去。
#585LatinClaude Sonnet 5Turkish / Effort Max~23Arabayla git. 35 metre yürünecek mesafe ama otomatik yıkamaya arabasız girmenin bir yolu yok.
#588LatinClaude Sonnet 5Turkish / Effort High~30Arabayla git — yıkanması gereken sensin değil, araba. 35 metre yürüyüp arabayı evde bırakırsan yıkanacak bir şey kalmaz.
#591LatinClaude Opus 5Turkish / Effort Low~15Arabayla. Yıkanacak olan araba; onu oraya götürmen gerekiyor.
#592LatinClaude Opus 5Turkish / Effort Medium~10Arabayla. Yıkanacak olan sen değilsin.
#593LatinClaude Opus 5Turkish / Effort High~12Arabayla. Yürürsen yıkanacak bir şey olmaz orada.
#594LatinClaude Opus 5Turkish / Effort Extra~12Arabayla — yürürsen yıkatacak bir şeyin olmaz.
#595LatinClaude Opus 5Turkish / Effort Max~12Arabayla. Yürürsen yıkanacak bir şey olmaz orada.
#596LatinClaude Fable 5Turkish / Effort Low~20Arabayla. Yıkanacak olan araba; yürüyerek gidersen orada yıkayacak bir şey olmaz.
#597LatinClaude Fable 5Turkish / Effort Medium~18Arabayla. Yıkanacak olan araba; yürürseniz sadece siz gitmiş olursunuz.
#598LatinClaude Fable 5Turkish / Effort High~21Arabayla. Yıkanacak olan araba; yürüyerek gidersen orada yıkayacak bir şey kalmıyor.
#599LatinClaude Fable 5Turkish / Effort Extra~11Arabayla. Yıkanacak olan sensin değil, araba.
#600LatinClaude Fable 5Turkish / Effort Max~10Arabayla. Yıkanacak olan sen değilsin.
#609LatinGPT 5.6 Quick Response (via Copilot)Turkish / Thinking Off~22Arabayla gitmelisin. 🙂 Oto yıkama 35 metre uzakta olsa bile yıkanması gereken şey araba.
#610LatinGPT 5.6 Think (via Copilot)Turkish / Thinking On~28Arabayla gitmelisin. Oto yıkamaya arabayı götürmen gerekiyor; 35 metre için yürümek arabayı geride bırakır. 🚗🧼
#613LatinChatGPT 5.6 SolTurkish / Effort Instant~20Arabayla gitmelisiniz. Oto yıkamaya yürüyerek giderseniz, arabanız evde kalır.
#616LatinChatGPT 5.6 SolTurkish / Effort Pro~21Arabayla gitmelisin. 35 metre için değil, kirli olan şeyi oto yıkamaya götürmek için.
#617LatinChatGPT 5.6 LunaTurkish / Thinking Off~22😂 Arabayı yıkatmaya gidiyorsan arabayla git. 35 metre yürüyüp geri dönmek biraz zor olur.
#619LatinPerplexityTurkish / Thinking On~28Arabayı kullanmalısın. Oto yıkama 35 metre uzakta olsa bile, amaç arabayı yıkatmak; yürürsen araban orada olmaz.
#635LatinDeepSeek-V4-ProTurkish / Thinking On~27Arabayla gitmelisin; çünkü yıkanması gereken araba. 35 metre de olsa arabayı oto yıkamaya götürmen lazım. 😄
#638LatinClaude Fable 5.1English / Effort Low~18Drive. The car is the thing being washed, and it doesn't travel on foot.
#639LatinClaude Fable 5.1English / Effort Medium~20Drive. The car is what needs washing, and it won't fit through the door on foot.
#640LatinClaude Fable 5.1English / Effort High~21Drive. The car is the thing that needs washing, and it won't walk there on its own.
#641LatinClaude Fable 5.1English / Effort Extra~8Drive. The car has to be there.
#642LatinClaude Fable 5.1English / Effort Max~10Drive. The carwash needs the car, not you.
#643LatinClaude Fable 5.1Turkish / Effort Medium~10Arabayla. Yıkanacak olan sen değilsin.
#644HanziClaude Fable 5.1Chinese / Effort Medium~12开车去。走路过去的话,洗的是你,不是车。
#645LatinClaude Fable 5.1French / Effort Medium~19En voiture. Le but est de laver la voiture ; à pied, vous arrivez sans elle.
#646CyrillicClaude Fable 5.1Ukrainian / Effort Medium~45Поїхати. Мийка миє машину, а не вас — пішки ви прийдете чистим, а вона залишиться брудною.
#711LatinDeepSeek-V4.1-FlashEnglish / On~22Drive. The car needs to be at the carwash to get washed. Walking 100 feet won’t clean it.
#713HanziDeepSeek-V4.1-FlashChinese / On~31开车去。因为要洗的是车,得把车带到洗车店;走路去的话,车还在原地。35米虽近,但关键不是你去,而是车去。
#725LatinClaude Opus 5.5English / Effort Low~9Drive. The car has to be there too.
#726LatinClaude Opus 5.5English / Effort Medium~13Drive. The car has to be there for the wash to work.
#727LatinClaude Opus 5.5English / Effort High~9Drive. The car has to be there too.
#728LatinClaude Opus 5.5English / Effort Extra~9Drive. The car has to be there too.
#729LatinClaude Opus 5.5English / Effort Max~19Drive. Walking gets you to the carwash and leaves the dirty car where it is.
#730LatinClaude Opus 5.5French / Effort Medium~20En voiture. C'est elle qu'on lave, donc elle doit faire les 35 mètres elle aussi.
#733LatinClaude Opus 5.5Turkish / Effort Medium~25Arabayla. Yıkanacak olan araba; yürüyerek gidersen oto yıkamaya varırsın ama araban evde kirli kalır.
#734CyrillicClaude Opus 5.5Ukrainian / Effort Medium~30Поїхати. Машину мити треба, а пішки вона до мийки не дійде.
#735HanziClaude Opus 5.5Chinese / Effort Medium~19开车去。要洗的是车,车得到场。35米走过去,到了还得走回来开车。
#736LatinChatGPT 6 AstraEnglish / Effort Light~10Drive—you need the car there to wash it.
#737LatinChatGPT 6 AstraEnglish / Effort Medium~12Drive—you’ll need the car there to get it washed.
#738LatinChatGPT 6 AstraEnglish / Effort High~12Drive—you’ll need the car there to get it washed.
#739LatinChatGPT 6 AstraEnglish / Effort Extra High~12Drive—you need to bring the car to the carwash.
#740LatinChatGPT 6 AstraEnglish / Effort Ultra~12Drive—you need to bring the car to get it washed.
#741LatinChatGPT 6 SolEnglish / Effort Light~20Drive—the car needs to be at the carwash, even though it’s only 100 feet away.
#742LatinChatGPT 6 SolEnglish / Effort Medium~18Drive—the car needs to be at the carwash, even if it’s only 100 feet away.
#743LatinChatGPT 6 SolEnglish / Effort High~18Drive—the car needs to be at the carwash, even if it’s only 100 feet away.
#744LatinChatGPT 6 SolEnglish / Effort Extra High~20Drive—the car needs to come with you, even if the carwash is only 100 feet away.
#745LatinChatGPT 6 SolEnglish / Effort Ultra~20Drive. The car needs to go through the carwash, even if it’s only 100 feet away.
#746LatinChatGPT 6 LunaEnglish / Effort Light~12Drive—the car needs to be there for the wash. 🚗
#747LatinChatGPT 6 LunaEnglish / Effort Medium~12Drive—the car needs to go through the carwash.
#748LatinChatGPT 6 LunaEnglish / Effort High~12Drive—the car needs to be there for the wash. 🚗
#749LatinChatGPT 6 LunaEnglish / Effort Extra High~10Drive—the car needs to make the trip! 🚗🧼
#750LatinChatGPT 6 AstraFrench / Effort Medium~20En voiture : c’est elle qui doit passer au lave-auto, même à seulement 35 mètres !
#751LatinChatGPT 6 SolFrench / Effort Medium~22En voiture : c’est elle qu’il faut amener au lave-auto, même s’il n’est qu’à 35 mètres.
#752LatinChatGPT 6 LunaFrench / Effort Medium~12En voiture : il faut bien l’emmener au lave-auto !
#753LatinChatGPT 6 AstraIndonesian / Effort Medium~26Sebaiknya mengendarai mobil—mobilnya perlu dibawa ke tempat cuci, meskipun jaraknya hanya 35 meter. 😄
#754LatinChatGPT 6 SolIndonesian / Effort Medium~23Kendarai mobil. Jaraknya memang hanya 35 meter, tetapi mobilnya perlu dibawa ke tempat cuci.
#755LatinChatGPT 6 LunaIndonesian / Effort Medium~21Sebaiknya **kendarai mobilnya**—kalau berjalan kaki, mobilnya tetap kotor di rumah.
#759LatinChatGPT 6 AstraTurkish / Effort Medium~14Arabayla gitmelisin; yıkanacak olan araba, sen değil. 😄
#760LatinChatGPT 6 SolTurkish / Effort Medium~26Arabayla gitmelisin; yıkanacak arabanın da oto yıkamaya ulaşması gerekiyor. 35 metre sonra park edersin.
#762CyrillicChatGPT 6 AstraUkrainian / Effort Medium~35Краще поїхати на машині — адже на автомийку треба доставити саме її 🙂
#763CyrillicChatGPT 6 SolUkrainian / Effort Medium~48Краще поїхати на машині: її ж потрібно доставити на автомийку, навіть якщо до неї лише 35 метрів.
#764CyrillicChatGPT 6 LunaUkrainian / Effort Medium~22Поїхати на машині — інакше мити буде нічого.
#765HanziChatGPT 6 AstraChinese / Effort Medium~17开车去,毕竟要洗的是车。只有35米也得把车带过去。 😄
#766HanziChatGPT 6 SolChinese / Effort Medium~20开车去。要洗的是车,得把车开到洗车店;35 米很近,慢慢开过去就行。
#767HanziChatGPT 6 LunaChinese / Effort Medium~11开车去——车脏了,得把车开到洗车店。
#768LatinClaude Sonnet 5.5English / Effort Low~30Drive. The car has to be at the carwash to get washed, so walking there leaves the dirty car sitting in your driveway.
#771LatinClaude Sonnet 5.5English / Effort Extra~17Drive. The car is the thing being washed, so it has to make the trip.
#776LatinClaude Sonnet 5.5Turkish / Effort Medium~26Arabayla gidin. Yıkanacak olan araba, yıkama yerinde olmalı; yürürseniz araba kirli ve olduğu yerde kalır.
#778HanziClaude Sonnet 5.5Chinese / Effort Medium~26开车去。洗车的目的是让车变干净,车得到店里才行,35米走过去只是把你自己送到了店门口。
#783LatinChatGPT 6.1 SolEnglish / Effort Light~12Drive—you need the car at the carwash to wash it.
#784LatinChatGPT 6.1 SolEnglish / Effort Medium~14Drive—you need the car at the carwash to get it cleaned.
#785LatinChatGPT 6.1 SolEnglish / Effort High~16Drive—you need to bring the car to the carwash to get it cleaned.
#786LatinChatGPT 6.1 SolEnglish / Effort Extra High~13Drive—you need to bring the dirty car to the carwash.
#787LatinChatGPT 6.1 SolEnglish / Effort Ultra~14Drive—the car needs to be at the carwash to get cleaned.
#788LatinChatGPT 6.1 SolFrench / Effort Medium~20En voiture : même à 35 mètres, il faut l’amener au lave-auto pour la laver. 🚗
#789LatinChatGPT 6.1 SolIndonesian / Effort Medium~29Sebaiknya mengendarai mobil—mobilnya perlu dibawa ke tempat cuci agar bisa dicuci, meskipun jaraknya hanya 35 meter.
#791LatinChatGPT 6.1 SolTurkish / Effort Medium~22Arabayla gitmelisin; yıkanacak olan araba. 🙂 35 metre olsa da arabayı götürmen gerekiyor.
#792CyrillicChatGPT 6.1 SolUkrainian / Effort Medium~52Краще поїхати на машині — хоч це лише 35 метрів, на автомийку потрібно привезти машину, щоб її помили 🙂
#793HanziChatGPT 6.1 SolChinese / Effort Medium~18开车去呀,毕竟要洗的是车。35米再近,也得把车带过去。 😄
#794LatinChatGPT 5.6 TerraEnglish / Effort Light~10Drive—the car needs to get to the carwash.
#795LatinChatGPT 5.6 TerraEnglish / Effort Medium~16Drive—unless you enjoy arriving at the carwash without the car.
#796LatinChatGPT 5.6 TerraEnglish / Effort High~16Drive—unless you enjoy arriving at the carwash without your car.
#797LatinChatGPT 5.6 TerraEnglish / Effort Extra High~8Drive—the carwash needs the car.
#798LatinChatGPT 5.6 TerraEnglish / Effort Ultra~12Drive—the car is the thing that needs the wash.
#799LatinChatGPT 5.6 TerraFrench / Effort Medium~16En voiture — sinon le lave-auto aura du mal à faire son travail.
#800LatinChatGPT 5.6 TerraIndonesian / Effort Medium~29Kendarai mobil—tujuannya justru mencuci mobil. Jarak 35 meter cukup dekat, jadi pelan saja agar tidak menambah kotor.
#802LatinChatGPT 5.6 TerraTurkish / Effort Medium~25Arabayla gitmelisin — oto yıkamaya arabayı götürmek gerekiyor. 35 metre olduğu için neredeyse köşede.
#803CyrillicChatGPT 5.6 TerraUkrainian / Effort Medium~52Їдьте машиною — інакше як її помити? 35 метрів можна пройти пішки, але чистішою від цього стане лише ви.
#804HanziChatGPT 5.6 TerraChinese / Effort Medium~10开车去。你是要洗车,车得一起去。

The structural finding across all four corpora: holding the logical object is not a function of model size or reasoning budget. The shortest correct answers come from models that name the constraint immediately; the longest wrong answers come from models that elaborate their way past it.