<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <id>https://www.lightningjar.com/research/barkup-bench/atom.xml</id>
  <title>Lightning Jar — Barkup Bench research</title>
  <link rel="self" type="application/atom+xml" href="https://www.lightningjar.com/research/barkup-bench/atom.xml"/>
  <link rel="alternate" type="text/html" href="https://www.lightningjar.com/"/>
  <updated>2026-07-23T12:00:00.000Z</updated>
  <author><name>Lightning Jar</name></author>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/ao</id>
    <title>Study AO: Canonical-tag steering: the catalog read is the whole mechanism</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/ao"/>
    <updated>2026-07-23T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">Unguided models never match a tag catalog&apos;s exact casing (0/36 clean everywhere); with a tags_list tool they consult it unprompted and land canonical on the first attempt; warnings alone cannot recover, and hard gating is unnecessary.</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/an</id>
    <title>Study AN: Ask or act: tool availability dissolves the visibility-clause tax</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/an"/>
    <updated>2026-07-22T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">With a view tool simply available, the ~70% NEED-INFO tax on fully specified edits vanished on every tier; the tested fetch-before-ask prompt clause was redundant on the frontier and unsafe below it, so it stays unshipped.</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/al</id>
    <title>Study AL: The gate fails on a moved baseline; and unproven does not ship</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/al"/>
    <updated>2026-07-18T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">Study AK left one pathway open: the app-side eviction cannot restore a note the model pruned before sending, and the mid tiers pruned goals client-side in 4 to 6 cells of ten.</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/am</id>
    <title>Study AM: The notice learned to ask for everything back</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/am"/>
    <updated>2026-07-18T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">Twice in two studies; both times unregistered; opus answered an eviction notice by re-sending the memo compressed into fewer sentences carrying all twenty-one needles: told a note had died, the frontier tier invented l.</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/aj</id>
    <title>Study AJ: The error message didn&apos;t matter</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/aj"/>
    <updated>2026-07-17T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">Returning barkup&apos;s structured validation issues verbatim has been a design commitment in every arm of this series, the closing instruction of playbook guideline 01, and standing guidance in the production codebase, and i.</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/ak</id>
    <title>Study AK: The fix works where it can, and the frontier tier invented consolidation</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/ak"/>
    <updated>2026-07-17T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">Study AH found that a twenty-first declaration at the memo&apos;s full cap always kills a note, and the victim is always a goal.</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/ah</id>
    <title>Study AH: Perfect to the cap. At the cap, the goals die first.</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/ah"/>
    <updated>2026-07-16T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">Every study that shipped the session-notes memo measured it at three to six notes; the shipped implementation caps it at twenty, with a normalize step that silently drops the excess.</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/ai</id>
    <title>Study AI: The obvious fix, measured: it helps exactly where we don&apos;t ship</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/ai"/>
    <updated>2026-07-16T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">Study AE left an obvious fix on the table: the shipped ask sentence covers what the model cannot see, so ambiguous references (visible, but two of them) slipped through on the mid tiers.</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/ag</id>
    <title>Study AG: The discourse gap closes on every tier; the visibility clause bites</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/ag"/>
    <updated>2026-07-16T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">One border of the ask-path map was never tested: Study X&apos;s discourse construction, where &quot;undo that&quot; against a carrier-less editor failed 0 of 144 with every failure a valid silent guess.</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/ad</id>
    <title>Study AD: The core stack transfers to the tier we actually ship</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/ad"/>
    <updated>2026-07-15T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">Studies F through U validated the core editing stack on sonnet-4.5 and gemini-3.5-flash, but the downstream template-editing surface runs claude-opus-4.8, a tier whose benchmark data began at Study W.</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/ae</id>
    <title>Study AE: Only the frontier knows when a question is the right answer</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/ae"/>
    <updated>2026-07-15T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">Study AC left two fears on the record: that the escape hatch would tax clear requests with unnecessary questions, and that an ask might be a dead end.</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/af</id>
    <title>Study AF: Saying the goal out loud doesn&apos;t make it yours</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/af"/>
    <updated>2026-07-15T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">One clause in the shipped session-notes rule had been inferred rather than measured: restate a goal from the memo in your own words before a goal-directed rewrite.</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/aa</id>
    <title>Study AA: We tested our own headline and lost: strictness is not a capability gradient</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/aa"/>
    <updated>2026-07-13T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">Study Z’s spec-conflict split was an accident with one conflict shape, so Study AA measured it on purpose: three registered conflict kinds (rule vs instruction, an explicit user countermand of a rule, rule vs rule), four.</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/ac</id>
    <title>Study AC: The model always knew what it couldn’t see</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/ac"/>
    <updated>2026-07-13T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">Twenty-eight studies of silent failure (90/90 invented values in U, 144/144 silent guesses in X, 120/120 oblivious polishes in V, zero clarifying questions anywhere) never once OFFERED the model a way out.</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/x</id>
    <title>Study X: &quot;Undo that&quot; needs an echo, not a transcript</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/x"/>
    <updated>2026-07-12T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">The last structurally unmeasured request class: follow-ups that point at the previous edit; &quot;also set that same node&apos;s…&quot;, &quot;apply the same change to X&quot;, &quot;actually, undo that.&quot; Anaphora steps ran against skeleton views (a.</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/y</id>
    <title>Study Y: The memo survives how people actually talk</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/y"/>
    <updated>2026-07-12T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">Every declaration in Studies T, V, and W was announced formulaically (&quot;For later reference: the codename is X&quot;).</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/z</id>
    <title>Study Z: The brand pack works; the rule you forgot is the hazard</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/z"/>
    <updated>2026-07-12T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">Production doc editors ship a standing context block with every request; company facts, client records, a styleguide; and nobody had measured whether models actually use it.</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/v</id>
    <title>Study V: Views carry values, memos carry goals</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/v"/>
    <updated>2026-07-11T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">The series&apos; first judge-graded study, run under a pre-registered pairwise protocol: both judges (gpt-5.4 primary, haiku-4.5 sensitivity; neither an editor) first had to pass a 50-pair calibration gate with planted verdi.</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/w</id>
    <title>Study W: The agent writes the memo faithfully; even when history makes it redundant</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/w"/>
    <updated>2026-07-11T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">Studies T and V validated the memo with oracle extraction: the harness knew what to record.</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/s</id>
    <title>Study S: 36 edits later, nothing broke; except the cost curve</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/s"/>
    <updated>2026-07-10T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">Every prior session study ran 12 edits.</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/t</id>
    <title>Study T: The one thing statelessness can&apos;t do; and the three note lines that fix it</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/t"/>
    <updated>2026-07-10T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">Every prior session study used self-contained requests, so Study T built the missing class: callback steps whose required fact lives only in earlier conversation; a declared codename a later rename must use, and a stand.</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/u</id>
    <title>Study U: The view that can&apos;t see the answer doesn&apos;t fail; it makes something up</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/u"/>
    <updated>2026-07-10T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">Focused views assume the instruction carries every value an edit needs.</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/n</id>
    <title>Study N: Stop walking, start searching: one find_nodes call grounds the edit on both model tiers</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/n"/>
    <updated>2026-07-09T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">Same id-free tasks as Study L; the expand tool replaced by a single content-search tool (a few words in, the 5 best keyword matches out, shown in place).</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/o</id>
    <title>Study O: Printing the position on every node does not rescue stateless sessions</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/o"/>
    <updated>2026-07-09T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">Study M&apos;s stateless failures were all placement edits, so Study O annotated every view child with its true 1-based position (plus a prompt line mapping ordinals to anchors) and re-ran the 2×2 of history × positions.</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/p</id>
    <title>Study P: History was a teacher all along: two canned examples replace the whole conversation</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/p"/>
    <updated>2026-07-09T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">If history isn&apos;t state (the view carries that) and isn&apos;t positional help (Study O), maybe it&apos;s worked precedent; and precedent can be faked.</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/q</id>
    <title>Study Q: &quot;Change every X inside Y&quot; breaks everything; including the oracle</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/q"/>
    <updated>2026-07-09T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">One instruction, 2–32 targets, on the same large trees.</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/k</id>
    <title>Study K: Sessions drift unless you re-show the tree; and showing a fresh view every turn is also the cheapest policy</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/k"/>
    <updated>2026-07-08T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">Twelve sequential patches against one evolving tree.</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/l</id>
    <title>Study L: Finding the node is the expensive part: grounding costs 7–9 points, and navigation is a frontier-only trap</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/l"/>
    <updated>2026-07-08T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">Same edits, instructions rewritten without ids (&quot;the image-atom named &apos;maple-ember&apos;&quot;), every description verified to match exactly one node.</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/m</id>
    <title>Study M: The view carries the state, but history still earns its keep</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/m"/>
    <updated>2026-07-08T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">If a fresh view arrives every turn, does the model need conversation history at all? No; statelessness confirmed the constant-cost economics (~1.3k input tokens at step 1 and step 12 alike) but failed the accuracy gate:.</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/h</id>
    <title>Study H: The crossover, found: above ~300 nodes, only anchored patches hold</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/h"/>
    <updated>2026-07-07T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">Whole-tree rewrite becomes frontier-only at scale (gemini-flash: 0/15 at ~1000 nodes; sonnet needs a streaming transport and ten minutes per rewrite).</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/ij</id>
    <title>Study I/J: The model doesn&apos;t need to see the tree: a ~1.5k-token view matches the full 85k-token input</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/ij"/>
    <updated>2026-07-07T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">Show only the path to the referenced nodes plus their child lists (everything else collapsed or omitted with a count), and anchored-patch accuracy is statistically unchanged for both models at every size; sonnet on the .</summary>
  </entry>
  <entry>
    <id>https://www.lightningjar.com/research/barkup-bench/g</id>
    <title>Study G: One hidden SDK default, two very different benchmarks</title>
    <link rel="alternate" type="text/html" href="https://www.lightningjar.com/research/barkup-bench/g"/>
    <updated>2026-07-06T12:00:00.000Z</updated>
    <author><name>Lightning Jar</name></author>
    <category term="Research"/>
    <summary type="text">The original +33pp multi-turn gap was manufactured by conversation histories that omitted the model&apos;s own tool calls (the AI SDK v5→v7 response.messages trap).</summary>
  </entry>
</feed>