Agent Lab
We build AI agents for private equity and regulated firms, where an improperly configured agent can incur significant liability. AVP, our open observability protocol, benchmarks agents beyond just the model, including the prompts, tools, and verification. This way we can see which build actually fits the job.
Request your own benchmarkFinishing a source-scan and posting the digest. Three frontier models, three runs each.
Rebuilding a PDF page as HTML, scored on structural fidelity. Haiku 4.5, ParseBench, 10 pages.
four-step prompt, double-check
the four-step prompt, nothing added
baseline plus a PDF cheat sheet
one line of prompt plus a tools server
a single-sentence instruction
baseline plus a worked table
Answering product questions right, by how the agent gets the docs. Haiku 4.5, 6 questions.
llms.txt in the system prompt
a tool call to read the docs
the whole pile, in context
the docs behind an MCP server
the docs as an inline skill
no docs at all, the control
Accuracy and Context are voyages from our public captain’s log at agentvoyagerproject.com, on Claude Haiku 4.5. Deliverability is our Code Mode case study, across three frontier models. Open any setup to read the run behind it.
Have a workflow you would trust an agent with? Tell us the task. We put a build through it, on documents like yours, and send you the numbers with the run behind them.