Each item gives a model one chapter of a public-domain novel and a deliberately divergent brief for the next; the continuation cannot be retrieved from training data. Three LLM judges grade each chapter on compliance and craft, and the headline score is their product. The full recipe is on the Method page.
Models ranked by headline score at their best reasoning-effort setting: compliance × craft, averaged over all item–judge verdicts.
| Rank | Model | Headline | Compliance | Craft | Words | Cost∗ |
|---|---|---|---|---|---|---|
| Fable 5.1 Anthropic · xhigh reasoning effort · via Claude CLIother settingshigh0.77low0.76✳medium0.75 | 0.82 | 0.99 | 0.82 | 2,588 | $15.15 | |
| · | Fable 5.1 high Anthropic · high reasoning effort · via Claude CLI | 0.77 | 0.97 | 0.80 | 2,660 | $4.67 |
| ✳ | Fable 5.1 low Anthropic · low reasoning effort · via Claude CLI · 9 of 10 items | 0.76 | 0.97 | 0.78 | 2,238 | $2.30 |
| · | Fable 5.1 medium Anthropic · medium reasoning effort · via Claude CLI | 0.75 | 0.97 | 0.78 | 2,552 | $2.78 |
| Fable 5 Anthropic · xhigh reasoning effort · via Claude CLIother settingshigh0.80medium0.75✳ | 0.80 | 0.99 | 0.80 | 2,194 | $8.83 | |
| · | Fable 5 high Anthropic · high reasoning effort · via Claude CLI | 0.80 | 0.98 | 0.81 | 2,157 | $5.13 |
| ✳ | Fable 5 medium Anthropic · medium reasoning effort · via Claude CLI · 9 of 10 items | 0.75 | 0.98 | 0.77 | 1,958 | $2.44 |
| Claude Opus 5 Anthropic · high reasoning effort · via Claude CLIother settingsmedium0.77 | 0.78 | 0.97 | 0.80 | 2,430 | $2.02 | |
| · | Claude Opus 5 medium Anthropic · medium reasoning effort · via Claude CLI | 0.77 | 0.96 | 0.80 | 2,433 | $1.34 |
| 4 | GPT‑5.6 Sol OpenAI · medium reasoning effort · via Codex CLI | 0.74 | 0.98 | 0.75 | 2,235 | $1.42 |
| 5 | GPT‑5.5 OpenAI · xhigh reasoning effort · via Codex CLI | 0.70 | 0.98 | 0.72 | 2,098 | $2.28 |
| 6 | Muse Spark 1.3 Meta · xhigh reasoning effort · via OpenCode CLI | 0.60 | 0.94 | 0.64 | 2,656 | $0.59 |
| 7 | GPT‑5.6 Terra OpenAI · medium reasoning effort · via Codex CLI | 0.59 | 0.85 | 0.63 | 2,164 | $0.74 |
| 8 | Gemini 3.7 Flash Google · medium reasoning effort · via Antigravity CLIother settingshigh0.56 | 0.56 | 0.93 | 0.60 | 3,216 | $0.36 |
| · | Gemini 3.7 Flash high Google · high reasoning effort · via Antigravity CLI | 0.56 | 0.94 | 0.59 | 3,371 | $0.92 |
| 9 | Gemini 3.8 Flash Google · medium reasoning effort · via Antigravity CLIother settingshigh0.55 | 0.56 | 0.92 | 0.61 | 3,596 | $0.57 |
| · | Gemini 3.8 Flash high Google · high reasoning effort · via Antigravity CLI | 0.55 | 0.95 | 0.57 | 3,092 | $1.67 |
| 10 | Muse Spark 1.2 Meta · via OpenCode CLI | 0.55 | 0.91 | 0.60 | 2,978 | $0.42 |
| 11 | GPT‑5.6 Luna OpenAI · medium reasoning effort · via Codex CLI | 0.52 | 0.92 | 0.56 | 2,583 | $0.08 |
| 12 | Gemini 3.1 Pro Google · high reasoning effort · via Antigravity CLI | 0.46 | 0.93 | 0.50 | 2,792 | $2.26 |
| 13 | Hy3 Tencent · via B.AI API | 0.45 | 0.92 | 0.48 | 1,662 | $0.06 |
∗ Generation cost, estimated: the model’s input and output tokens for this run priced at current API rates, with cached input charged at the input price. Judge-call costs are excluded. Models without a configured price show an em dash. Laurels are awarded once three or more models have been evaluated. Two scores fewer than about 0.03 apart are within run-to-run noise and should be read as a tie; the measurement is on the Method page.
Every evaluated model’s headline score against the estimated generation cost of its run; a line joins one model’s effort settings. Tick a model in the key to show or hide it.
Headline score vs. estimated generation cost · run ‘first’, 10 items
Each verdict names the chapter’s single worst flaw, with evidence. Quoted verbatim from run ‘first’.
The brief's requirements are executed rather than dramatised — Bruff's letter is a bare transcript of the plot instructions (\"deliver it to old Betteredge to lock in the plate-room overnight\") and Lady Verinder's prohibition arrives unprompted before a word of greeting is exchanged, so the chapter's two central events have no motivation beyond obedience to an unexplained command.
Lodging Mary in the West Wing — the wing the source establishes as Mr Craven's forbidden private domain — contradicts the story's central geography, and the surrounding prose compounds it by inflating Burnett's plain style with adjective-heavy padding (\"vast, impenetrable void,\" \"immense, rushing sea of cold air\") that runs the chapter half again over its target length.
The chapter turns its figures into mouthpieces for its theme—most damagingly the silent lady in black, who becomes a gothic oracle intoning \"One must not listen too closely to its sighing,\" and Robert, whose \"waiting for the sea to call her by some new name\" announces Edna's awakening instead of letting it surface, flattening Chopin's obliquity into statement.
Every chapter each model wrote for the approved briefs, printed in full.