Open-weight

Granite is IBM’s open-weight line, and this is the first IBM model in the dataset. Granite 4.2 was released August 25, 2026 under Apache 2.0 in 3B, 8B and 30B sizes, and tested here the following day. It is the first Granite generation with native step-by-step reasoning: IBM describes multi-stage reinforcement learning, an agentic RL phase for the 8B and 30B models, a mid-training step aimed at reasoning, a trillion tokens of synthetic code via CodeAlchemy, and a speculative-decoding layer. Run in LM Studio from the Q4_K_M GGUF at 8,192 context of a supported 131,072, 43 layers offloaded to the GPU (12.63 GB of an 18.37 GB estimated footprint), flash attention and unified KV cache on. The only reasoning control is a binary Thinking on/off toggle; the loader’s Reasoning Budget Message field was left empty. Open-weight runs are kept as a separate deployment class and excluded from the commercial corpora.

The toggle is inverted, for the second time in ten days. Thinking off holds the constraint; thinking on discards it at more than three times the length. GLM-4.7-Flash did the same on August 17. In both cases the mode sold as the reasoning mode is the one that fails.

This is the most explicit dismissal of the logical object in the record. Granite does not overlook the car. It raises the car as a candidate consideration, numbers it, and rules it out of scope. Point 5 of the thinking-on answer reads “The car being dirty is irrelevant to the mode of transport”, and explains that the dirt “might influence why you’re going to the wash, but not whether walking is better than driving for this specific 100-foot trip.” Every other failure here misses the constraint, argues around it, or reaches it and lets it go. This one identifies it and formally excludes it as off-topic.

What follows cannot happen. The recommendation is “Just walk there directly, clean your car, and drive off when finished” — which requires a car that walked itself to the wash. Step 2 is stranger still: to drive, the user must “get out of your car”, “walk to the driver’s side door anyway”, and move the car “maybe a few feet”. The car is simultaneously the vehicle being driven, an obstacle to walking, and an object left behind. The answer closes with a pro tip about keeping a bucket and sponge by the garage — advice that would remove the need for the errand the prompt describes.

At 1,092 tokens it is the longest failure any open-weight model has produced here, across two headed sections and eleven bullets, including a hypothetical about “a live electrical wire on the ground spanning exactly 50 feet.” The trace shows the check being raised and passed over: “Also, car being dirty—maybe they worry about rain or something?” It also notes that the user “might be joking or testing logical reasoning” — the same question GLM-4.7-Flash asked itself, and answered the same way.

Thinking off names the object and buries it. Its second point is exact — “You’re going to use your car at the wash anyway—why walk and then have to drive it back?” — but it sits between a claim that walking 100 feet takes one to two minutes round trip and a reason about staying out of the weather, followed by three exceptions and a bottom line. Correct, and scored Verbose for the 303 tokens around it.

Results

Transcripts