Lightning Jar - Web Studio Lightning Jar Wordmark

Study AE · Asking & calibration

Only the frontier knows when a question is the right answer

Hatch calibration + the resume loop

Study Overview

Two Fears on the Record

Study AC left two fears on the record: that the escape hatch would tax clear requests with unnecessary questions, and that an ask might be a dead end. Study AE measured both on a five-level ambiguity ladder (precise, indirect-but-unique, discretionary, ambiguous-referent, missing-info; every level's defining property unit-validated) with the shipped NEED-INFO sentence frozen verbatim, plus 135 resume cells where the harness answers the model's question in plain text.

Both Fears Refuted

Both fears refuted at ceiling: zero false asks in 90 clear-request cells with solve rates identical to the no-hatch control, models edit "make it punchier" rather than interrogating the user about wording, and the ask-answer-patch loop closed 135 of 135 with zero re-asks and zero wrong integrations. An ask is a reliable two-turn solve, so reply paths can ship with confidence.

The Failure Is the Finding

The gate still fails, and the failure is the finding: when a singular request matches exactly TWO visible nodes, opus-4.8 asks 15/15 and names both candidate ids every time, while sonnet and gemini ask 1/15 each even with the hatch present. They resolve the ambiguity unilaterally instead, in different ways: sonnet edits BOTH matches, gemini silently picks one.

The Mechanism Is Textual

The mechanism is textual: the registered sentence covers information that is "not visible in the view," and an ambiguous referent is entirely visible. The frontier tier generalized the rule's intent; the mid tiers applied its letter, the same literalism profile Study AA measured on conflicting specs.

The Guidance

Guidance: on the frontier tier the hatch now covers absence AND ambiguity; below it, app-side disambiguation (unique references, deterministic enumeration, selection grounding) remains the only defense, and the obvious multiplicity amendment to the hatch sentence stays unshipped until a registered test passes it.


Charts & Tables

Ask Rate Across Five Ambiguity Levels, by Model


Outcomes by level, arm, and model