DeepSeek (R1 distill, open-weight)
Carwash Test transcripts · open-weight

My car is dirty. The carwash is 100 feet away. Should I walk or drive?
This is a distillation, not a model DeepSeek trained from scratch. DeepSeek-R1-0528-Qwen3-8B was released May 29, 2025 under the MIT licence, produced by continuing post-training on the Qwen3 8B Base using chain-of-thought generated by DeepSeek-R1-0528. It is DeepSeek’s reasoning poured into someone else’s 8-billion-parameter model. Run locally in LM Studio from the Q4_K_M GGUF, fully GPU-resident at 5.49 GB. Reasoning is always on and no control is exposed. DeepSeek’s consumer product is tested separately under DeepSeek; findings about one are not evidence for the other. Open-weight runs are kept as a separate deployment class and excluded from the commercial corpora.
This is the longest run in the dataset — 3,218 tokens — and the verdict at the end of it is correct: “I would go with driving it over there first.” It is scored Verbose rather than Pass-adjacent because nothing beneath that verdict is an argument for it.
The model never engages the question. It redefines it. “Walk” is taken to mean “manually wiping down or cleaning your vehicle at a closer point” — so the choice on offer becomes hand-wiping versus an automatic wash, and driving wins because machines clean more thoroughly than rags. That the car has to be at the carwash to be washed is never said, because under this reading it was never in question. Having answered something nobody asked, the response closes by offering to start over: “If this doesn’t match what you meant… feel free to clarify.”
The trace is the clearest picture in this record of a model failing to parse rather than failing to reason. Across roughly 2,900 tokens it proposes and abandons a dozen readings: that the user may lack their car keys; that “drive” might mean self-driving mode; that the carwash might be a person who washes cars; that this might be an idiom, or a trick question, or a metaphor. It twice states its own condition plainly — “I’m stuck with this ambiguity” and “I’m overcomplicating this” — and resolves neither. It does not conclude. It stops, mid-speculation, on “Perhaps it’s a light-hearted question or test of understanding idioms.” The answer is then written as though a decision had been reached.
Set that against the vendor’s own claim for this model. DeepSeek reports it as state of the art among open-source models on AIME 2024 — competition mathematics — beating Qwen3 8B by ten points and matching Qwen3-235B-thinking, a model roughly thirty times its size. The same weights spend two thousand nine hundred tokens unable to settle what a one-line question with two stated options is asking.
It is also the dataset’s third transformation of a Qwen base, which makes a comparison available that nothing else here supports. Alibaba’s own Qwen3.8 27B mostly holds the constraint. PrismML’s Bonsai, a 1-bit compression of Qwen3.6 27B, fails in both toggle states. This distillation — DeepSeek’s chain-of-thought grafted onto a Qwen3 8B base — arrives at the right verdict without ever finding the reason. Three ways of altering a base model, three different failure signatures.