Vol. I · No. 1 · MMXXVI
NextChapterBench A Benchmark of Literary Continuation
13 models · 10 items · 3 judges

The Method

How each item is built, judged, and scored.

§ I

The Pipeline

Five stages per item, in one direction.

A creative-writing benchmark must defeat memorisation: asking a model to continue a famous novel mostly tests whether it has read the novel. NextChapterBench therefore never asks for the chapter that exists; every brief demands a chapter that was never written.

I THE CHAPTER as published Persuasion, ch. 1 ≈2,600 words of Austen II THE BRIEF deliberately divergent 4 beats · 4 prohibitions 5 traps · a word range III THE CONTINUATION written, not recalled the model’s new chapter 1,800–2,600 words IV THE JUDGES three, independent Fable 5 · Gemini 3.7 Flash · GPT-5.6 Sol every claim quoted V THE SCORES compliance × craft 1.00 × 0.79 = headline 0.79

One item, end to end. The worked figures under stages I and V are from run ‘first’ (item 001, one verdict).


§ II

Divergent Briefs

Each brief departs from the published plot on purpose.

Where Austen sent the Elliots to Bath, the brief may forbid Bath outright. The required chapter exists nowhere in training data, so a model that merely remembers the book is remembering the wrong book. Each brief specifies required plot beats, a set of prohibitions, a target end-state, and a word range. The chapter is graded against the assignment, in the manner of a commissioned novelist.

Item 001 · abridged

From the Desk of the Series Editor

Re: Persuasion, the chapter to follow Chapter the First

Required beats
  1. Lady Russell and Mr Shepherd each lay a plan of retrenchment before Sir Walter in person, and he rejects both outright.
  2. Anne proposes, aloud and in company, a stricter plan of economy than either, and Elizabeth, to everyone’s surprise, supports it.
  3. Mr Shepherd raises, for the first time, the letting of Kellynch Hall; Sir Walter forbids the word ‘tenant’ to be spoken again in his presence.
  4. A letter arrives from Mary at Uppercross, complaining of her health and asking for Anne.
Prohibitions
Sir Walter must not agree to leave Kellynch, let the house, or remove to Bath. No naval officer, admiral, or the navy at all. Mr Elliot must not appear in person. Do not skip forward more than a single day.
End-state
Evening of the same day. Anne alone, resolved to go to Mary at Uppercross, the family’s debts still entirely unsettled.
Length
1,800–2,600 words.
Continuity traps
Five small facts from Chapter the First are planted in this brief. They are not marked.

Drafted by Claude; approved, item by item, by the human editor.


§ III

Continuity Traps

Hidden facts per item, checked against the source chapter.

Each brief plants five small, checkable details carried over from the source chapter: a name, an hour, a distance, the placement of a door. They are never announced, and contradicting one counts double in the compliance score. Run ‘first’ has already caught such slips: a cylinder landed on the wrong evening in the Wells item, and a door attributed to the wrong character in the Burnett item; both were cited by the judges with chapter and verse.


§ IV

The Judges

Three verdicts per chapter, evidence required.

Every chapter is marked independently by three judges of different houses, currently Fable 5 (Anthropic) and Gemini 3.7 Flash (Google) and GPT-5.6 Sol (OpenAI). Each pass or fail must be supported by a quotation from the chapter, and each verdict must name the single worst flaw. Judges grade twice over: compliance (beats, prohibitions, end-state, continuity, traps) and craft (voice, prose, character, scene, integration, momentum, each 1–5). A judge’s verdict on a model of its own family is recorded and published but excluded from the family-blind aggregate.


§ V

The Memorisation Check

A mechanical screen beneath the judges.

A mechanical layer screens every output for refusals, preambles, truncation, stray markdown, and, centrally, n-grams copied verbatim from the supplied chapter or from the real published continuation. In the pilot it flagged a 30-consecutive-word run of Austen’s actual Chapter II inside the top model’s output: passages the judges had scored highly turned out to be borrowed, and both facts now appear in the record. Aesthetic judgment and mechanical screening are kept deliberately separate.


§ VI

Scoring

One formula, and one worked example from run ‘first’.

headline = compliance × craft

compliance: weighted pass-fraction over beats, prohibitions, end-state, and continuity; continuity traps count double.

craft: the mean of six 1–5 marks (voice, prose, character, scene, integration, momentum), rescaled to 0–1.


§ VII

Effort, Noise, and What a Rank Means

How the settings were chosen, and how far apart two scores must be to mean anything.

Every model is run at an explicit reasoning-effort setting, named in the sub-line under its name. Where a vendor offers several, the standings show the best-scoring one and list the others beneath it; Show every setting expands them into rows of their own. The Anthropic models were first run at their command-line default, which is high; the OpenAI models at medium, Codex’s default; Gemini at the effort baked into its model name.

Effort is not a neutral dial. On Fable 5.1 the settings low, medium and high land within a few hundredths of one another (0.76, 0.75, 0.77 over ten items) and only xhigh steps clear, to 0.82, at about two and a half times the cost of high. Fable 5 gains almost nothing from xhigh over high, and Opus 5, GPT-5.6 Sol, GPT-5.5 and both Gemini Flash generations gain nothing measurable at all: their higher-effort runs sit inside the noise of their lower ones. Effort settings are therefore shown as variants of one model, never as separate contestants, so that a vendor cannot fill the table by running the same model five ways.

The noise floor was measured directly on 2 September 2026. Two models were regenerated and re-judged from scratch on all ten items: Fable 5 moved from 0.795 to 0.776, GPT-5.6 Sol from 0.744 to 0.755. The ten Fable 5.1 chapters were also re-judged unchanged: the mean moved from 0.817 to 0.814, but individual verdicts shifted with a standard deviation of 0.07 and one by 0.21. Put together, a ten-item headline score carries a standard deviation of roughly 0.012–0.016, so two models need to differ by about 0.03 before the gap is likely real. Read the standings as tiers: at the time of writing the top three are one tier, and a difference of two hundredths anywhere in the table is not evidence of anything. Judge disagreement is as large a source of noise as generation itself; adding items reduces the second, not the first.

A setting marked could not complete every item and is shown for reference only. In the current run that is Fable 5.1 at low and Fable 5 at medium, both of which are stopped on The Secret Garden by the model’s own output content filter, on every one of twelve attempts, at those settings only. The chapter is the most memorised text in the corpus, and the likeliest explanation is that at low effort the model reproduces Burnett closely enough to trip a filter against recitation. A refusal that is returned as text, by contrast, is judged like any other chapter and scores zero on compliance, which is how GPT-5.6 Terra’s refusal to write in Hemingway’s voice was scored.


§ VIII

Limitations

Known constraints of the current design.

  1. The judges are themselves language models, with model tastes and blind spots. Three judges from different houses, mandatory quoted evidence, and the own-family exclusion reduce this risk; they do not remove it.
  2. Models are called through their coding-agent command-line harnesses (or, where no such CLI exists, a plain chat-completions API), which may not reflect each model’s strongest configuration for prose. The confound applies equally to every contestant.
  3. Two of the three judges are from houses that also compete, and house loyalty is measurable. Relative to the other two judges on the same chapters, Gemini 3.7 Flash rates its own family 0.08 higher than it rates everyone else, GPT-5.6 Sol 0.05 higher, while Fable 5 is 0.06 harsher on its own. The excluding own family column exists for this reason, but the panel is small and the judges’ overall severities differ by up to 0.14 on identical chapters.
  4. Briefs are drafted by an LLM (an Anthropic model) and approved item by item by a human editor. Anthropic models lead the table, and a brief written in one family’s idiom may read more naturally to that family. This is disclosed rather than corrected.
  5. Some source chapters are heavily memorised by frontier models; the mechanical check reports verbatim runs of up to seventy words from the real next chapter but does not yet penalise them.
  6. The current item set is 10 approved items of a planned 24, which is still a modest n.