Lightning Jar - Web Studio Lightning Jar Wordmark

Study N · Views & grounding

Stop walking, start searching: one find_nodes call grounds the edit on both model tiers

The retrieval ladder

Study Overview

Search Replaces Navigation

Same id-free tasks as Study L; the expand tool replaced by a single content-search tool (a few words in, the 5 best keyword matches out, shown in place).

Both Tiers Recover

The frontier model matches its id-oracle bound exactly; same 43/45, same two failed tasks; at a median of one search call and ~90% less input; gemini jumps from 23/45 (navigating) to 39/45 (16–0 paired, p < 0.001), its full-tree score at 3% of the cost.

Two More Rungs

Two more rungs: swapping the keyword matcher for text-embedding-3-small changed nothing (target coverage 23/45 vs 24/45; embeddings can't resolve "the 3rd block inside atlas"), and two-stage grounding (gemini reads the tree, names the ids, sonnet patches a focused view) holds 41/45 with the frontier model's median input at 1,484 tokens (−97.4%). The Study L gate, re-tested, passes twice.


Charts & Tables

Chart N.1: Grounded-Instruction Success Across the Retrieval Ladder, by Model


Table N.1: Success, input, and mechanism by condition (45 tasks per cell)