Study AD · Confirmations & machinery
The Opus confirmation
Studies F through U validated the core editing stack on sonnet-4.5 and gemini-3.5-flash, but the downstream template-editing surface runs claude-opus-4.8, a tier whose benchmark data began at Study W. Study Q had already shown that serializer advice inverts between models, so nothing could be assumed.
Study AD re-ran the stack on opus with corpora, prompts, and graders reused verbatim: 1,160 cells, gates anchored to the weakest previously passing tier. Every gate passed, mostly at or above the top of the prior bands.
The anchored-patch dialect scored 194/200 on the main corpus (the best condition-F number ever measured, and every miss was a tree-reading count question rather than an editing failure), HTML focused views swept 90/90 across 300 to 1000 nodes, search grounding landed exactly on sonnet's oracle bound (43/45, median one call), and all session policies were perfect.
The surprise is the ablation: bare stateless sessions, no history and no worked examples, also scored 240/240 steps with 20/20 intact end states. The examples block that rescued sonnet (13/20 without it) is not load-bearing on the frontier tier; it is cheap insurance for every tier below it, and it stays shipped.
Fan-out remains the honest boundary, now with a third mitigation profile: opus raised the floor to 80 to 89% but still left a third of the 7-plus-target view tasks partially covered, so the app-side decomposition fence stands unchanged at every tier.