Study AI · Asking & calibration
The multiplicity hatch
Study AE left an obvious fix on the table: the shipped ask sentence covers what the model cannot see, so ambiguous references (visible, but two of them) slipped through on the mid tiers. One added sentence covering requests that match more than one node looks like it must work, and this series has now caught two obvious-looking clauses doing nothing (the priority meta-rule in AA, the restate ceremony in AF), so the clause was measured before anyone shipped it: amended sentence versus the shipped one, re-run contemporaneously, across the full five-level ladder on three models.
The results split three ways. Sonnet is completely rescued: 3 of 15 asks became 15 of 15, every ask naming both candidate ids, with its silent edit-both habit gone. Gemini improved enormously, 0 of 15 to 11 of 15, but stopped one ask short of the registered detection bar, and its four residual failures all silently edited BOTH matches; the clause converted its coin-flips into asks or edit-boths, never back into correct restraint.
And the clause taxed nothing anywhere: zero false asks in 90 clear-request cells, solve rates identical to control, hard boundaries undisturbed.
The one cost signal lives on the tier we actually ship: opus, already perfect on ambiguity with the shipped sentence alone, started interrogating discretionary requests slightly more often under the amendment (asks on "make it punchier" cells rose from 2 of 15 to 5 of 15, not significant, but the only movement anywhere).
So the registered gate fails by a single ask, and the practical verdict is sharper than a pass: the clause stays unshipped on frontier surfaces, where it buys nothing and whispers a cost; it is a large, free, but partial improvement for sub-frontier deployments; and app-side disambiguation remains the actual contract at every tier. Knowing when a question is the right answer keeps measuring as a capability that prompting approaches but does not close.