Study V · Qualitative (Track 2)
Views carry values, memos carry goals
Qualitative rewrites; judge-graded (Track 2)
Study Overview
The First Judge-Graded Study
The series' first judge-graded study, run under a pre-registered pairwise protocol: both judges (gpt-5.4 primary, haiku-4.5 sensitivity; neither an editor) first had to pass a 50-pair calibration gate with planted verdicts, identity probes, and length probes, and both aced it 50/50.
The Task
The task: "rewrite this paragraph to focus on our central thesis," with the thesis planted in a mission node and the target paragraph deliberately assembled off-thesis. Five arms for where the goal lives.
Blind Arms Polish Obliviously
Blind arms lost all 120 comparisons by obliviously polishing the off-topic paragraph; fluent, valid, useless.
The Memo Carries the Goal
The T memo carried the goal at full parity with an explicit instruction (sonnet's memo arm actually won 10–2 under the primary judge).
The Surprise That Failed the Gate
The surprise that failed the gate: models with the thesis-bearing node IN THE VIEW demonstrably read it (keyword coverage +0.75 vs +0.00 blind) yet lost 117/120; their rewrites orbit the topic where told-outright rewrites anchor the thesis. Study U's view fix moves data perfectly; it moves intent only partway.
The Verdict
Judge agreement 83.3% (kappa 0.53); the keyword proxy ranks the arms identically. Guidance: put the nodes an edit must read in the view; put the goal in the instruction or memo, restated outright.