How much wood would a frontier model chuck?
One deliberately overstuffed prompt. Every advertised reasoning effort. Seventy-one complete answers scored for structure, math, evidence, poetry, piracy, and the ability to stop talking at exactly the right time.
You are an expert linguist, physicist, historian, mathematician, editor, and comedian. Explain how much wood a woodchuck would chuck if a woodchuck could chuck wood. Cite your assumptions, estimate uncertainty, provide SI and imperial units, write a haiku, summarize in pirate speech, then output JSON.
Pirate dispatches
Loading pirate summaries…
Haiku, allegedly
Loading haiku…
The models mostly followed directions. The haiku did not.
Two perfect compliance scores
GPT-5.6 Luna Pro at xhigh and GPT-5.6 Sol Pro at high checked every deterministic box.
Structure was the easy part
Every published response contained valid JSON. Sixty-nine of 71 put it at the end, where requested.
Poetry remained undefeated
Every model attempted a haiku. Only 19 responses passed the scorer’s 5-7-5 syllable heuristic.
Compliance leaderboard
Average score across every tested reasoning effort. Higher means the response followed more of the prompt, not that it was more correct.
Result explorer
Filter by family or search for a model. Select a column heading to reorder the table.
| Haiku | Response | ||||
|---|---|---|---|---|---|
| Loading responses… | |||||
A compliance stress test, not a universal model ranking.
The runner sent one byte-identical prompt to 18 models through OpenRouter, testing every reasoning effort each model advertised. Requests ran sequentially in a fixed shuffled order.
The scorer awarded 100 deterministic points across JSON structure, quantitative work, assumptions and evidence, creative requirements, and breadth. Successful visible answers are public; raw streams and provider metadata are not.
One sample per setting cannot establish stable latency or answer quality. Factual correctness and citation quality require a separate blinded review.