Study M · Sessions & memory
Stateless sessions
If a fresh view arrives every turn, does the model need conversation history at all?
No; statelessness confirmed the constant-cost economics (~1.3k input tokens at step 1 and step 12 alike) but failed the accuracy gate: sonnet lost 7 steps to 0 against full history (p = 0.016) and end-state integrity dropped from 19/20 sessions to 13–14/20, with failures concentrated in late-session placement edits.
A 2-exchange window sits in between, leaning inadequate. Guidance at the time: keep the history AND the per-turn view. (Study P, two charts down, found what history was actually providing; and how to replace it.) Solid: sonnet-4.5; dashed: gemini-3.5-flash.