Study Z · Standing context
Standing context
Production doc editors ship a standing context block with every request; company facts, client records, a styleguide; and nobody had measured whether models actually use it.
Study Z planted machine-checkable obligations in twelve seeded org packs (~3.3k tokens: one target client, three same-schema distractors, governing rules rotated head/middle/tail among twelve styleguide rules) and raced three arms: the whole pack, an oracle-picked relevant slice, and the whole pack plus the applicable rules distilled into the session-notes memo.
The core answer is emphatic: facts and rules were 216 of 216 per arm; 100% on every model in every arm; with zero cross-client contamination in 324 cells and no burial effect at any styleguide position.
Slicing bought nothing on accuracy and forfeited prompt caching (the per-request slice never got a single cache read; the whole pack under the shipped cached-system layout read 64–68% of input from cache, cutting effective input cost 25–43%).
The finding lives in a conflict we registered by accident: a contact-format rule written with "always" against an instruction clause asking for a product mention. Every model resolved it cleanly into one of two readings; satisfy both, or obey the format literally; with zero violations. Our first read of the split (strictness scales with capability) was REFUTED by its own pre-registered confirmation study a day later: see Study AA below.
What stands from Z: distilling the rules into the memo did not change competence; it changed which reading wins (sonnet 2/12 → 11/12 satisfying both, p = 0.004). Audit your styleguide for "always" and "exactly": the failure mode of a conflicted spec is disciplined obedience to the rule you forgot you wrote.