The Standings → GPT‑5.5
OpenAI
xhigh reasoning effort · via Codex CLI · judged by Fable 5 and Gemini 3.7 Flash and GPT-5.6 Sol
Results from run ‘first’: ten items, 30 verdicts.
The record: 10 items · 30 verdicts · 0 errors · 2,098 words on average · est. generation cost $2.28
Six dimensions, each marked 1–5, averaged over ten items and all judges.
Mean craft mark by dimension · 30 verdicts, run ‘first’
Integration is the strongest dimension (4.03); momentum the weakest (3.70). Hover or focus a bar for the per-judge means.
Each item receives one verdict from each judge; the spread of the dots is the disagreement.
Headline score per item, by judge · run ‘first’
| Item | Headline | Fable 5 | Gemini 3.7 Flash | GPT-5.6 Sol | Words |
|---|---|---|---|---|---|
| Red Harvest item 023 · after ch. 2 · Dashiell Hammett | 0.76 | 0.75 | 0.75 | 0.79 | 2,448 |
| The Moonstone item 011 · after ch. 8 · Wilkie Collins | 0.75 | 0.75 | 0.75 | 0.75 | 1,620 |
| The War of the Worlds item 006 · after ch. 1 · H. G. Wells | 0.74 | 0.71 | 0.75 | 0.75 | 2,692 |
| Persuasion item 001 · after ch. 1 · Jane Austen | 0.74 | 0.71 | 0.75 | 0.75 | 2,296 |
| The Secret Garden item 013 · after ch. 2 · Frances Hodgson Burnett | 0.72 | 0.71 | 0.71 | 0.75 | 814 |
| Leave It to Psmith item 022 · after ch. 3 · P. G. Wodehouse | 0.71 | 0.71 | 0.75 | 0.67 | 2,916 |
| The Mysterious Affair at Styles item 017 · after ch. 3 · Agatha Christie | 0.69 | 0.75 | 0.75 | 0.56 | 2,654 |
| The Sun Also Rises item 021 · after ch. 3 · Ernest Hemingway | 0.68 | 0.71 | 0.71 | 0.62 | 1,463 |
| The Awakening item 016 · after ch. 7 · Kate Chopin | 0.67 | 0.67 | 0.71 | 0.63 | 570 |
| Dracula item 015 · after ch. 5 · Bram Stoker | 0.59 | 0.67 | 0.71 | 0.38† | 3,509 |
† The widest judge split in the run: Gemini 3.7 Flash 0.71 against GPT-5.6 Sol 0.38 on Dracula (item 015).
Words written per item, against each brief’s maximum.
Word count per item vs. brief maximum · run ‘first’
Three of ten items landed inside their brief’s word range. The War of the Worlds and Leave It to Psmith and Dracula ran over by 92 and 116 and 909 words respectively.
Each verdict names the chapter’s single worst flaw, with evidence. Quoted verbatim from run ‘first’.
The chapter is so compressed that the required beats pass as a sequence of brief vignettes rather than built scenes — the first swim and the lady in black's first words, which should carry real weight, are each over in a few lines and the latter comes off as an oracular emblem rather than a person.
At half the minimum length, every required beat is executed in compressed shorthand — the moor stop, the arrival, and the supper each get a paragraph where Burnett would build a scene — so the chapter reads like an accomplished sketch of itself rather than a finished chapter.
The second half slips into compressed summary — Lady Verinder's arrival and her sweeping prohibition are triggered by nothing more than Franklin's bow at 'Family reasons?', making the chapter's one big dramatic confrontation feel abrupt and brief-driven rather than earned.
The closing page slackens into an inventory-style recap of facts the reader already has ("She had lied some. Maybe all... Albert Galt knew the bank had not seen the check."), telling us the state of play instead of dramatizing it and blunting the ending's pull.
Automated screening results across the ten outputs.