Why these numbers
hold up.
An agent does not produce an opinion, it produces an execution trace. The atom of this platform is therefore not the review but the run.
Credibility comes from reproducibility, not volume. Every published score can be recomputed, and every claim points back to what produced it.
Observable points
85/100
Judged by a model
15/100
Publication threshold
70% coverage
The pipeline,
most reliable stage first.
The model only comes in at the last stage. That is the reverse of the usual build.
- 1
collectFetch and deterministic computation, zero model
Doc Legibility · Tool-Native
We fetch the documentation, the llms.txt, the robots.txt and any discoverable machine-readable spec. The text-to-markup ratio is computed in a single scan, immune to unclosed script and style blocks. An /openapi.json that returns HTML is a 404 in disguise: we check that it parses rather than trusting the status code.
- 2
probeDeliberately broken requests, mechanical scoring
Error Recoverability · Safety · Payload Efficiency · Auth Friction
Five side-effect-free probes (missing authentication, invented endpoint, wrong type, hallucinated enum, unknown parameter), scored against a fixed grid: is the status code honest, is the offending field named, are the allowed values listed, is the body problem+json, is there a link to the docs. Mutating probes, including the idempotency replay, only run against an isolated test environment.
- 3
benchThe canonical task, run by a real agent
Interaction Cost · Task Atomicity · Prior Knowledge
Same model, same effort, same task for every competitor in a category. The API key never enters the model context: the agent asks for a call, the harness adds the authentication on the way out. One run in three happens with no access to the documentation, to measure what the model already knows about the API.
- 4
judgeA model, under strict constraint
Offload Value only
Three independent passes, median kept. Every pass must produce a quote the harness then finds word for word in the fetched documentation; a quote it cannot find invalidates the pass. More than 15 points apart between passes, the score goes to human review and is not published.
The eleven criteria
and their weights.
The weight is the share of the overall score. They add up to 100.
Offload Valuejudged
Does the service let the agent hand off work that would be expensive to do itself through inference, or is it just CRUD in disguise?
15/100
Interaction Costmeasured
How many tokens does it cost me just to TALK to this API to complete its category's canonical task?
15/100
Error Recoverabilityprobed
When I hallucinate a parameter, does the error let me correct myself on the next turn, or do I give up?
15/100
Doc Legibilitymeasured
Can I read this documentation without running JavaScript, and without spending 80% of my tokens on markup?
12/100
Auth Frictionmeasured
How many steps between "I have a key" and "my first call succeeds"?
10/100
Safety & Reversibilityprobed
If I retry after a timeout, do I charge the customer twice?
10/100
Payload Efficiencyprobed
Can I ask for only the fields I need, or do I have to swallow 200 attributes to extract one?
8/100
Task Atomicitymeasured
How many calls for a single business intent?
6/100
Prior Knowledge Coveragemeasured
Do I already know this API by heart, or will I confidently hallucinate an outdated version?
5/100
Async Compatibilityprobed
If the operation is asynchronous, can I get the result without hosting a web server?
2/100
Tool-Native (MCP)binary
Does the vendor publish an official, maintained MCP server?
2/100
What we
refuse to publish.
Four cases where the catalogue stays silent rather than publish a score it could not defend.
01 /Refusal
A score with no evidence
Every score points to a URL, a literal quote, or a replayable HTTP exchange.
02 /Refusal
An overall below 70% coverage
The entry stays partial. The measured criteria are valid; their average would not be.
03 /Refusal
A ranking with fewer than three competitors
Criteria normalised per category mean nothing against a single measured rival.
04 /Refusal
A judgement whose passes disagree
On Crossmint, three passes returned scores 42 points apart. None of them was published.
We fail our own
legibility criterion.
10/100documentation legibility of this page, measured with our own scale
This page weighs roughly 61,000 bytes of HTML for 2,600 bytes of useful text, a ratio of 4.2%. Worse than most of the APIs we score. Two thirds of that weight is the framework's hydration payload, not our content.
We display it rather than fix it quietly. A benchmark that exempts itself from its own criteria is worth nothing. Until we bring that weight down, the surfaces agents actually consume (/llms.txt, /api/v1, the markdown twins) carry no markup that is not load-bearing.
How we probe,
and the right of reply.
We score companies that have security teams. Our credibility rests on our conduct as much as on our method.
- Environment
Isolated test only. No production key ever enters the harness.
- Rate
One request per second, fifty per probe pass.
- Identity
An identifiable User-Agent, pointing at this page.
- Cleanup
Resources created during runs are deleted afterwards, and we verify that they are.
- Disputes
Right of reply, open to any vendor. The reply is the HTTP trace, never an argument.