Experiment 001 · July 15, 2026

How much wood would a frontier model chuck?

One deliberately overstuffed prompt. Every advertised reasoning effort. Seventy-one complete answers scored for structure, math, evidence, poetry, piracy, and the ability to stop talking at exactly the right time.

The exact prompt
You are an expert linguist, physicist, historian, mathematician, editor, and comedian. Explain how much wood a woodchuck would chuck if a woodchuck could chuck wood. Cite your assumptions, estimate uncertainty, provide SI and imperial units, write a haiku, summarize in pirate speech, then output JSON.
71complete responses
18models tested
92.4mean compliance
$8.16core model cost
What happened

The models mostly followed directions. The haiku did not.

02

Structure was the easy part

Every published response contained valid JSON. Sixty-nine of 71 put it at the end, where requested.

03

Poetry remained undefeated

Every model attempted a haiku. Only 19 responses passed the scorer’s 5-7-5 syllable heuristic.

Model averages

Compliance leaderboard

Average score across every tested reasoning effort. Higher means the response followed more of the prompt, not that it was more correct.

Loading model scores…

All 71 responses

Result explorer

Filter by family or search for a model. Select a column heading to reorder the table.

71 results
Haiku Response
Loading responses…
Method, briefly

A compliance stress test, not a universal model ranking.

The runner sent one byte-identical prompt to 18 models through OpenRouter, testing every reasoning effort each model advertised. Requests ran sequentially in a fixed shuffled order.

The scorer awarded 100 deterministic points across JSON structure, quantitative work, assumptions and evidence, creative requirements, and breadth. Successful visible answers are public; raw streams and provider metadata are not.

One sample per setting cannot establish stable latency or answer quality. Factual correctness and citation quality require a separate blinded review.