Lightning Jar - Web Studio Lightning Jar Wordmark

Study S · Sessions & memory

36 edits later, nothing broke; except the cost curve

Long sessions

Study Overview

Both Recipes Hold

Every prior session study ran 12 edits. Study S ran the two surviving recipes through 36 (10 new sessions, seed committed) and both held: full-history S-view went 719/720 steps across both models, stateless-plus-worked-examples S-system went 713/720 with no late-third decay (98–99% at steps 25–36; step 36 is taught as well as step 1), McNemar parity everywhere, zero context ceilings, zero blocked steps.

Cost Is What Diverges

What diverges is cost: keep-history's median per-step input grows linearly to ~24k tokens by step 36 while the stateless recipe stays flat at ~2.1k, which compounds to 449k vs 81k input per session; 5.4–5.6×, against a pre-registered prediction of ≥3×.

The Verdict

The gate passed on both models; the worked-examples recipe is the measured default for long sessions. All seven S-system step failures were placement-class (insert/move), the class Studies M/O/P mapped. Solid: sonnet-4.5; dashed: gemini-3.5-flash.


Charts & Tables

Chart S.1: Median Input Tokens per Step Across a 36-Edit Session


Table S.1: Per-third success, end-state, and tokens per session by recipe