Research
We believe good engineering decisions come from evidence, not vibes. So
we run our own research: pre-registered studies whose hypotheses,
corpora, and analysis plans are committed publicly before a single
model is called, with results published as found, corrections included.
The goal is practical guidance for developers of LLM applications:
claims you can check, recipes you can ship, and open-source packages
that carry the findings into production.
AEO Bench
Active An open, pre-registered benchmark measuring whether agent-readiness and answer-engine-optimization techniques (llms.txt, sitemap.xml, markdown content negotiation) measurably help AI agents use websites. Study 1: 900 scored agent runs across five models against a controlled 40-page fixture, with token cost as a first-class outcome. Results published as found, corrections included.
View the ProjectBarkup Bench
Active An open, pre-registered benchmark series measuring how large language models read and edit structured document trees. Forty-five studies, more than 36,198 scored model runs, eight models, trees from 5 to 1,000 nodes, sessions up to 36 edits. Results published as found, corrections included; every finding shipped into the open-source barkup library.
View the Project