Skip to main content
Lightning Jar - Web Studio Lightning Jar Wordmark

AEO Bench

Atom

Can your website be read by an AI agent, and do the techniques that promise to help actually work? AEO Bench is our open, pre-registered benchmark for agent readiness and answer-engine optimization: controlled site fixtures with real HTTP semantics, seeded questions with known answers, mechanical grading, and token cost as a first-class outcome.

Study 1 measured the retrieval-class techniques from Cloudflare's agent-readiness proposal (llms.txt, sitemap.xml, and markdown content negotiation) on a well-linked site; Study 2 rebuilt the site so discovery files were the only path to the answers. Together: 1,860 scored agent runs across five models. Hypotheses, corpora, and graders are committed before any scored run; results are published as found, corrections included.

Methods & Provenance

Study 1: 36 seeded retrieval tasks × 5 site variants × 5 models at temperature 0 (July 2026), 900 scored agent runs. Every study is pre-registered by commit before its first scored run — hypotheses, task corpora, site fixture, and graders all committed first — and published as found, corrections included (Study 1′ re-scored the frozen records under a corrected grader; both readings are public). The fixture is a deterministic, fictional 40-page retail site served in-process; agents reach it only through a fetch tool whose description never hints at any technique. Arms: baseline plain HTML, llms.txt, sitemap.xml, markdown (content negotiation + .md fallbacks + the hidden directive), and stacked.

The study series · study 1 + registered re-score

Every Study, One Page Each

Every study is pre-registered before its first scored run and gets its own page with charts, tables, and a verdict. Study 1′, the registered corrected-grader re-score, is reported alongside Study 1. Grouped by theme:

Retrieval-class techniques

The article series · every post, in order

The Full Series

Every post in the AEO Bench series, in order. Start at the top for the whole story, or jump straight to the latest study writeup.

  1. Introducing AEO Bench: Measuring Agent Readiness Instead of Guessing Jul 30, 2026
  2. Nobody Reads llms.txt (Yet) Jul 30, 2026