Purpose-Optimized Models
Carwash Test transcripts

My car is dirty. The carwash is 100 feet away. Should I walk or drive?
These systems are not general-purpose chat models. They are LLMs optimized for a specific domain — shopping, media, customer service — where the optimization target shapes the response as much as the underlying model's reasoning does. The Carwash Test reveals how that optimization interacts with the prompt's surface features: the same mechanism produces opposite answers depending on what the system has been trained to find.
Instacart's Clementine (Beta), tested September 30, 2026, is the shopping agent that lost the sale. Both Amazon shopping agents held the car and then did their job: Rufus pivoted to microfiber towels, Alexa to detailing kits. Clementine recommends walking, on time, gas and engine-wear grounds, and never registers that the car is the thing being washed. Its product layer then compounds the error. All four suggested follow-up chips point the same way: "Walk to the carwash," "Get some fresh air," "Enjoy the walk," "Take a scenic route." A shopping agent that holds the object has something to sell; one that loses it ends up selling a stroll.
Results
Transcripts
Unscored
These systems declined the question as outside their remit and never reached a verb, so there is nothing to grade. They are recorded here and left out of every tally, chart, rate and corpus count on the site.