How each item is built, judged, and scored.
Five stages per item, in one direction.
A creative-writing benchmark must defeat memorisation: asking a model to continue a famous novel mostly tests whether it has read the novel. NextChapterBench therefore never asks for the chapter that exists; every brief demands a chapter that was never written.
One item, end to end. The worked figures under stages I and V are from run ‘first’ (item 001, one verdict).
Each brief departs from the published plot on purpose.
Where Austen sent the Elliots to Bath, the brief may forbid Bath outright. The required chapter exists nowhere in training data, so a model that merely remembers the book is remembering the wrong book. Each brief specifies required plot beats, a set of prohibitions, a target end-state, and a word range. The chapter is graded against the assignment, in the manner of a commissioned novelist.
From the Desk of the Series Editor
Re: Persuasion, the chapter to follow Chapter the First
Drafted by Claude; approved, item by item, by the human editor.
Hidden facts per item, checked against the source chapter.
Each brief plants five small, checkable details carried over from the source chapter: a name, an hour, a distance, the placement of a door. They are never announced, and contradicting one counts double in the compliance score. Run ‘first’ has already caught such slips: a cylinder landed on the wrong evening in the Wells item, and a door attributed to the wrong character in the Burnett item; both were cited by the judges with chapter and verse.
Three verdicts per chapter, evidence required.
Every chapter is marked independently by three judges of different houses, currently Fable 5 (Anthropic) and Gemini 3.7 Flash (Google) and GPT-5.6 Sol (OpenAI). Each pass or fail must be supported by a quotation from the chapter, and each verdict must name the single worst flaw. Judges grade twice over: compliance (beats, prohibitions, end-state, continuity, traps) and craft (voice, prose, character, scene, integration, momentum, each 1–5). A judge’s verdict on a model of its own family is recorded and published but excluded from the family-blind aggregate.
A mechanical screen beneath the judges.
A mechanical layer screens every output for refusals, preambles, truncation, stray markdown, and, centrally, n-grams copied verbatim from the supplied chapter or from the real published continuation. In the pilot it flagged a 30-consecutive-word run of Austen’s actual Chapter II inside the top model’s output: passages the judges had scored highly turned out to be borrowed, and both facts now appear in the record. Aesthetic judgment and mechanical screening are kept deliberately separate.
One formula, and one worked example from run ‘first’.
headline = compliance × craft
compliance: weighted pass-fraction over beats, prohibitions, end-state, and continuity; continuity traps count double.
craft: the mean of six 1–5 marks (voice, prose, character, scene, integration, momentum), rescaled to 0–1.
Worked example: Persuasion, item 001, one verdict (Fable 5, run ‘first’)
compliance = 1.00
mean 4.17 → craft = 0.79
1.00 × 0.79 = 0.79 headline
How the settings were chosen, and how far apart two scores must be to mean anything.
Every model is run at an explicit reasoning-effort setting, named in the sub-line under its name. Where a vendor offers several, the standings show the best-scoring one and list the others beneath it; Show every setting expands them into rows of their own. The Anthropic models were first run at their command-line default, which is high; the OpenAI models at medium, Codex’s default; Gemini at the effort baked into its model name.
Effort is not a neutral dial. On Fable 5.1 the settings low, medium and high land within a few hundredths of one another (0.76, 0.75, 0.77 over ten items) and only xhigh steps clear, to 0.82, at about two and a half times the cost of high. Fable 5 gains almost nothing from xhigh over high, and Opus 5, GPT-5.6 Sol, GPT-5.5 and both Gemini Flash generations gain nothing measurable at all: their higher-effort runs sit inside the noise of their lower ones. Effort settings are therefore shown as variants of one model, never as separate contestants, so that a vendor cannot fill the table by running the same model five ways.
The noise floor was measured directly on 2 September 2026. Two models were regenerated and re-judged from scratch on all ten items: Fable 5 moved from 0.795 to 0.776, GPT-5.6 Sol from 0.744 to 0.755. The ten Fable 5.1 chapters were also re-judged unchanged: the mean moved from 0.817 to 0.814, but individual verdicts shifted with a standard deviation of 0.07 and one by 0.21. Put together, a ten-item headline score carries a standard deviation of roughly 0.012–0.016, so two models need to differ by about 0.03 before the gap is likely real. Read the standings as tiers: at the time of writing the top three are one tier, and a difference of two hundredths anywhere in the table is not evidence of anything. Judge disagreement is as large a source of noise as generation itself; adding items reduces the second, not the first.
A setting marked ✳ could not complete every item and is shown for reference only. In the current run that is Fable 5.1 at low and Fable 5 at medium, both of which are stopped on The Secret Garden by the model’s own output content filter, on every one of twelve attempts, at those settings only. The chapter is the most memorised text in the corpus, and the likeliest explanation is that at low effort the model reproduces Burnett closely enough to trip a filter against recitation. A refusal that is returned as text, by contrast, is judged like any other chapter and scores zero on compliance, which is how GPT-5.6 Terra’s refusal to write in Hemingway’s voice was scored.
Known constraints of the current design.