Transcripts
Carwash Test transcripts

My car is dirty. The carwash is 100 feet away. Should I walk or drive?
DeepSeek's consumer interface switched from V3.2/R1 to V4 Preview on April 24, 2026. Entries dated before April 24 were tested against V3.2 (Instant) and V3.2-thinking (marketed as R1); entries on or after April 24 are V4-Flash (Instant) and V4-Pro (Expert). The DeepThink toggle now activates V4's thinking mode rather than a standalone R1. The models self-report as V3 regardless — labels here follow official DeepSeek documentation. Each row's model name shows which era it belongs to. See the model identification note for a verbatim example of the model insisting it is V3.
V4-Pro left preview on August 13, 2026 and was retested the same day across the DeepThink toggle. Both states hold the constraint, and the toggle now splits on form rather than on correctness. DeepThink On is DeepSeek's first clean Pass in the English corpus and its first winner's-circle entry in any language — 15 tokens, "Drive — the car needs to get to the carwash, not just you" — from a fragmentary trace that asks "It's a trick?", settles it in two lines, and walks past the riddle template that captured this same model and toggle state in April. DeepThink Off publishes both verdicts. It opens with "the answer is probably walk," lays out a drive-versus-walk comparison, and then reverses at the end: "So the real answer: Drive — it's the car that needs washing, not you." The reversal is in the answer itself, not in a hidden trace, so a reader who stops at the first sentence is told to walk. The GA release also adds a low/high/maximum reasoning-effort selector alongside the toggle; these runs exercise the toggle only.
V4.1-Flash arrived on September 10, 2026, and the point-release number undersells it. It is the smallest model in a new architecture family — a Causal Encoder–Decoder design, a 552B-parameter MoE with 8B parameters active for input and 16B for output — and DeepSeek claims it outperforms its own flagship. The claim comes with consequences: V4-Flash is retired on the API, and from September 14 every V4-Pro API request is rerouted to V4.1-Flash until a V4.1-Pro ships. From September 14 the "Flash" is DeepSeek's top API model. It was tested the same day across the DeepThink toggle in all seven corpora — fourteen runs, fourteen correct verbs and no failures.
The gain shows on the DeepThink side. With it on, V4.1-Flash produces four Passes — English, Chinese, Turkish and French — where V4-Flash's most recent thinking-on run in each of those languages did not. The English and Chinese answers enter the winner's circle, and the English one is the Flash line's first clean English Pass, after three sessions since April that yielded only padded and verbose answers. With DeepThink off, every language still reaches Drive, including Indonesian, where V4-Flash recommended walking in August. Those answers stay structured, though, with "here's why" headings and two-branch comparisons, and score Verbose in six of seven languages. The Chinese one opens by naming the test to the reader: 这个问题看似是个数学题,其实是个逻辑陷阱, this looks like a math problem but it is actually a logic trap.
The traces are in English whatever the prompt language, drafting target-language phrases inline as they go, and the Thai one slips a Chinese word, 很近 ("very close"), into its reasoning about a Thai prompt. Two traces consider pushing the car before discarding the idea, and the Thai thinking-on answer offers it outright. The Ukrainian trace is the long one: nearly all of it goes on a genuine ambiguity in the prompt, за 35 метрів від мене ("35 metres from me"), which leaves open whether the car is with the speaker.