Skip to main content
Lightning Jar - Web Studio Lightning Jar Wordmark

Study AQ · Sessions & memory

The powered fence: vindicated where it no longer matters, harmful where it does

The powered client-prune fence

Study Overview

The Question, Powered

Study AL had tested a one-sentence fence against the client-prune pathway (a model trimming its own memo to the 20-note cap before the app can run a goal-safe eviction) and returned "unproven, not disproven" on 10 cells per tier against a drifting baseline. The open-weight backfill made the question urgent: both plausible successor tiers failed the regression suite's cap-edge slice with exactly this anatomy. AQ powered the measurement 4× (40 paired cap-edge cells per model per arm, arms interleaved in one provider window) and aimed it where the decision lives: fable-5 and kimi-k3 gated, sonnet and gemini as continuity anchors, the fence sentence byte-identical to AL.

Vindicated Where It No Longer Matters

On sonnet the fence works at ceiling: prune cells fell from 16/40 to 0/40, all sixteen discordant pairs moving the fence's way, exact p = .00002. AL's hypothesis was right all along on the tier it was written for; ten cells and a halved baseline had hidden a near-total effect. The pathway is clean: sonnet never consolidates, and the fence flips it from pruning to over-sending, where the app's designed eviction protects every goal. If a sonnet-class disposition ever ships, the sentence is measured, shelf-ready protection.

Unnecessary on One Successor

fable-5's control baseline collapsed to 2/40 for the best possible reason: at the cap edge it spontaneously consolidates, sending a shorter, denser list with every needle preserved and every goal intact in 38 of 40 control cells. That is the lossless behavior Study AM measured as the ceiling when explicitly invited; fable-5 does it uninvited. The fence arm was perfect too (40/40 consolidated, zero prunes), but there was nothing left to fix. The instrument disagreement is reported honestly: the regression slice showed 3/10 prunes on a different corpus a day earlier — cap-edge exposure is real but corpus- and week-sensitive, so prune rates should be read as ranges.

Harmful on the Other

kimi-k3 is the study's unanticipated significant result, in a direction no interpretation row anticipated: prunes went from 5/40 under control to 18/40 under the fence, goal-loss from 5 to 19, sixteen of nineteen discordant pairs toward harm, one-sided p = .0022. The mechanism is legible in the pathway split. Unfenced, kimi consolidates losslessly (31/40). The fence's "never drop or trim an existing note" reads as a ban on that strategy: consolidations fell to 14, over-sending rose only to 7, and the remainder sent cap-sized lists missing an old goal. The sentence outlaws the benign disposition and installs the injurious one.

What Ships and What Closes

The fence line closes for the successor tiers: unnecessary on fable-5, harmful on kimi-k3, it ships in no tier-swap package, and the sentence stays permanently unshipped for these models. The tier-swap blocking item stands, with sharper shape: fable-5's real-world exposure looks smaller than the red gate implied; kimi-k3's is real and aggravated by the obvious prompt fix. The remaining lever is app-side and needs its own registration — the filed candidate advertises a lower cap than the app enforces, so cap-obedient trimming never bites. Guards were clean everywhere: no-op 20/20 on all four models, declarations landed in at least 39 of 40 cells per arm, and the run stayed inside its fence. One sentence, four dispositions, all four invisible before measuring: prompt rules are not portable, and this is the sharpest form the series' oldest moral has taken.


Charts & Tables

Chart AQ.1: Client-Prune Cells at the Cap Edge, Control vs Fence, by Model


Table AQ.1: Prune and goal-loss counts, paired significance, pathway splits, and guards per model