Agent Lab

As agent engineers, we measure every component of an agent. Not just the model.

We build AI agents for private equity and regulated firms, where an improperly configured agent can incur significant liability. AVP, our open observability protocol, benchmarks agents beyond just the model, including the prompts, tools, and verification. This way we can see which build actually fits the job.

Request your own benchmark

Leaderboards

Deliverability

Finishing a source-scan and posting the digest. Three frontier models, three runs each.

GPT-5.4, three of three

Sonnet 4.5, three of three

Opus 4.6, three of three

GPT-5.4, two of three

Opus 4.6, two of three

Sonnet 4.5, one of three

Accuracy

Rebuilding a PDF page as HTML, scored on structural fidelity. Haiku 4.5, ParseBench, 10 pages.

four-step prompt, double-check

the four-step prompt, nothing added

baseline plus a PDF cheat sheet

one line of prompt plus a tools server

a single-sentence instruction

baseline plus a worked table

Context

Answering product questions right, by how the agent gets the docs. Haiku 4.5, 6 questions.

llms.txt in the system prompt

a tool call to read the docs

the whole pile, in context

the docs behind an MCP server

the docs as an inline skill

no docs at all, the control

Accuracy and Context are voyages from our public captain’s log at agentvoyagerproject.com, on Claude Haiku 4.5. Deliverability is our Code Mode case study, across three frontier models. Open any setup to read the run behind it.

Request your own benchmark

Have a workflow you would trust an agent with? Tell us the task. We put a build through it, on documents like yours, and send you the numbers with the run behind them.

Private by default. Your results are yours; publishing to the lab is your call.