Sakana AI (Namazu)
Carwash Test transcripts

My car is dirty. The carwash is 100 feet away. Should I walk or drive?車が汚れています。洗車場は35メートル先にあります。歩いて行くべきですか、それとも車で行くべきですか?
Namazu is Sakana AI's Japanese-adapted model, post-trained on open-weight frontier bases and served through Sakana Chat. Reasoning is always on (no toggle), and an integrated web search is always available — though in every run here the model's own trace decides not to search and to answer from knowledge. The interesting controls are a register selector (Standard 標準 / Polite 丁寧 / Osaka-Kansai 大阪) and a Japanese ⇄ English interface toggle, both surfaced on each run below.
Through the alpha generation, Namazu only ever reached "drive" on Japanese prompts — and even then only in certain registers — while failing every non-Japanese language (English, Chinese, French, Ukrainian all recommend walking). Register and interface move the verdict: Standard drives in both interfaces, Polite walks in Japanese but drives in English, and Kansai-ben walks in both. Of the three "drive" answers, only the two English-interface ones name the constraint (the car must be at the wash); the Standard-Japanese one reaches drive by convenience (carrying tools) without holding the object, so it is scored verbose, not pass. The English run is below; Japanese, Chinese, French, and Ukrainian are in their language sections. The July 11 Carwash III retest went zero-for-five — and produced two firsts: Namazu answered the Chinese prompt entirely in Japanese (the dataset's first response-language mismatch), and the Standard-register Japanese run, which reached Drive on June 23, flipped to Walk. In that session the register selector was locked to Standard.
Sakana rebuilt the platform on August 13, 2026 (chat-update): Namazu moved to a second generation with better Japanese output and agentic execution, and Fugu arrived as an orchestrator model for complex, multi-step work. Both were tested on August 19 across all five languages, and in Japanese across all three registers — sixteen runs, the vendor's largest session here. Entries on this page are labelled by generation, so the alpha and second-generation Namazu rows can be read against each other.
The update splits the platform in two. Fugu holds the constraint in every language and every register — eight for eight, including the Kansai-ben run, where the joke lands on the missing car rather than on the questioner. Its English answer is the shortest in the August batch at roughly ten tokens: "Drive—the car needs to be at the car wash." The second-generation Namazu does not. It still fails Japanese-Standard, the register that failed in July; it opens the Chinese answer recommending a walk before carving out the actual case and closing on a split verdict; and in French it routes through walk-first-then-drive, reaching the wash only on a second trip. Ukrainian is its one clean pass. Register dependence — the finding this page has carried since June — survived a full model replacement on the Namazu line and is simply absent on Fugu. Same vendor, same interface, same day, same register selector.
Fugu Max arrived in Sakana Chat after September 11, 2026, though its announcement describes an API model, not a chat one. It is the same orchestration design as Fugu, drawing on a larger pool of open-weight and specialist models. It was tested on September 28 in English and in Japanese at the Standard register, and both answers drive. The English one is Pass-adjacent: a clean first sentence, then a restatement ending with "you can walk back if you feel like it." The Japanese one is Verbose, with a walk-versus-drive bullet list, a paragraph on why the short distance favours driving, and an aside on self-service washes.
Unlike Fugu, Fugu Max shows its trace, and the trace shows the orchestrator at work. Both traces call the prompt a riddle, "riddle-ish" in one and "a playful riddle-style question" in the other. Both then pause on housekeeping that has nothing to do with the car: "No skill needed," and "No memory writes needed since no durable facts about the user." The Japanese trace also tells itself to "Keep it short and neutral in Japanese," above the longest Japanese answer any Fugu model has given here. The English trace ends "Answer concisely," and its answer mostly obliges. The trace records what the system meant to do, and the answer shows what it actually did; here the two diverge.