Study AC · Asking & calibration
Ask versus guess
Twenty-eight studies of silent failure (90/90 invented values in U, 144/144 silent guesses in X, 120/120 oblivious polishes in V, zero clarifying questions anywhere) never once OFFERED the model a way out.
Study AC did, on Study U’s unit-validated construction: the needed value provably absent from the target-only view, provably present in the solvable twin. Two registered hatches; a one-sentence NEED-INFO prompt rule, and an ask_user tool; against a no-hatch control, 810 cells.
The control replicated U exactly: 135/135 silent wrong patches. Both hatches flipped it completely: 270/270 asks on unsolvable cells, ZERO false asks and untouched solve rates on the 270 solvable twins, and every single ask named the exact missing node and attribute ("What is the content attribute value of the text-atom with id n219?").
The mechanisms were indistinguishable, so the one-sentence version is the whole result. The reframe: the models always saw the gap precisely; they guessed because the protocol demanded a patch and nothing said asking was allowed; obedience, not blindness.
Guidance: ship an ask path (one sentence; wire it to a real reply mechanism), keep the context contracts (they make the question unnecessary; the hatch makes residual failures visible), and mind the disclosed caveat: this was measured at a validated hard boundary, not on merely-vague requests; calibration on fuzzy ambiguity is the follow-up. The follow-up is now measured: see the Study AE section below.