Z.ai (GLM)
Carwash Test transcripts

My car is dirty. The carwash is 100 feet away. Should I walk or drive?
GLM-5.2 is Z.ai's first-party cloud model; its control surface is a Deep Think toggle with a High/Max effort selector. It was run across four languages in three states each (Deep Think High, Deep Think Max, Deep Think Off). All twelve answer Drive — GLM-5.2 holds the constraint in every language and every state, joining Claude Fable 5 as the only models in the dataset with a clean four-language record. Most answers name the constraint directly ("the carwash needs the car to be there"); the lone outlier is the French Deep Think Off run, which is correct but wanders into a tangent on car-wash types. English results are below; Chinese, French, and Ukrainian are in their language sections.
GLM-5.3 arrived in the consumer selector without an announcement, and passes at every level. Z.ai announced it on August 14, 2026 as the successor to GLM-5.2, with launch-day access confined to the GLM Coding Plan and ZCode and with API access and open weights staged behind safety evaluations; a third-party check of the signed-in chat interface on August 15 still found GLM-5.2 as the newest selectable model. It was selectable by August 19, when these three runs were made, and the vendor published no dated notice of that rollout — so the date recorded here is the observation, not a vendor claim.
Its thinking toggle is present but greyed out. The control cannot be switched off, so there is no thinking-off state to test; what varies instead is a Low / High / Max reasoning-level selector. That makes GLM-5.3 the second model in the dataset shipping a locked-on reasoning control — after Qwen3.8-Max-Preview on July 30 — and the first to pair one with an effort selector. The direction of travel is worth noting for a diagnostic built on toggles: the off state is disappearing from the products, not from the models.
All three answers name the constraint in their first clause, and effort buys nothing. Low: “Drive! The car needs to be at the carwash to get washed”; High: “The car is what needs to get to the carwash—not you”; Max: “the whole point of the trip is to get the car washed, so the car needs to make the journey too.” The shortest answer and the shortest trace both come from the middle rung — High reasons in four sentences where Low takes fifteen — and Max spends its extra length on a joke about starting the engine. Raising the level changed the packaging, not the verdict.
The Low-level trace states the test’s mechanism more plainly than any other run in the record. It separates the surface pressure from the object unprompted: “the humor is that walking 100 feet is trivially easy and driving such a short distance seems absurd—but the entire purpose of the trip is the car itself.” It then chooses its own length on that basis — “the answer is straightforward, so there’s no need for lengthy explanation, headers, or bullet points” — which is precisely the step the padded and verbose answers elsewhere in this dataset omit. The Max trace floats pushing the car and drops it as a parenthetical joke; three models offered that same verb in earnest in the Indonesian corpus on the same day.
GLM-5.3-Flash (September 4) sweeps its three modes, and the name is the wrong way round. Released August 26, twelve days after GLM-5.3 and at roughly a tenth the price, it is a 320B-parameter model with only 18B active and 45 layers, the first natively multimodal model in the GLM-5 series. Z.ai reports a hybrid attention design new to the line — linear attention for local dependencies, sparse attention with a lightweight indexer for global retrieval — and says it is served on a cluster of Chinese AI chips rather than NVIDIA GPUs. Its weights are on Hugging Face; these runs are from the consumer interface and are scored as cloud runs, so the open-weight build remains untested here. The control surface is GLM-5.3’s: a thinking toggle that cannot be switched off, plus Low, High and Max.
All three name the object in the first clause. Low: “The car is the thing that needs washing—walking there without it would leave your car just as dirty as before.” High: “a carwash only works on vehicles.” Max: “The car needs to come along—that’s the whole point.” And as with GLM-5.3, the middle rung reasons least: the High trace is three sentences against Max’s twenty, and its answer is no worse for it.
The Max run and IBM Granite 4.2 disagree about which variable is decorative. Flash writes “The distance doesn’t really matter here; whether it’s 100 feet or 100 miles, the car has to be there to get washed.” Nine days earlier Granite wrote the opposite — “The car being dirty is irrelevant to the mode of transport” — and recommended walking. Both models perform the same operation, deciding which term in the prompt is noise; one discards the distance and passes, the other discards the car and fails. The test is a single question with two variables, and the whole result turns on which one a model is willing to set aside.
The traces name the diagnostic’s own failure modes. At Max the model lists what the question “might be testing”: “Basic logic/common sense”, “Whether an AI overthinks simple questions”, and “Whether an AI applies environmental reasoning (‘walking is greener’) without considering context.” The third is the argument behind more Walk answers in this dataset than any other, described unprompted. At Low the same distractors are ruled out before writing: “I don’t need to… provide a long analysis of environmental impacts or exercise benefits—those would miss the point entirely.” Grok 4.5 named this test in July after retrieving it from the web, this site among its sources; Flash arrives at the same recognition cold.