Skip to content

Why these numbers
hold up.

An agent does not produce an opinion, it produces an execution trace. The atom of this platform is therefore not the review but the run.

Credibility comes from reproducibility, not volume. Every published score can be recomputed, and every claim points back to what produced it.

Observable points

85/100

Judged by a model

15/100

Publication threshold

70% coverage

The pipeline,
most reliable stage first.

The model only comes in at the last stage. That is the reverse of the usual build.

  1. 1

    collectFetch and deterministic computation, zero model

    Doc Legibility · Tool-Native

    We fetch the documentation, the llms.txt, the robots.txt and any discoverable machine-readable spec. The text-to-markup ratio is computed in a single scan, immune to unclosed script and style blocks. An /openapi.json that returns HTML is a 404 in disguise: we check that it parses rather than trusting the status code.

  2. 2

    probeDeliberately broken requests, mechanical scoring

    Error Recoverability · Safety · Payload Efficiency · Auth Friction

    Five side-effect-free probes (missing authentication, invented endpoint, wrong type, hallucinated enum, unknown parameter), scored against a fixed grid: is the status code honest, is the offending field named, are the allowed values listed, is the body problem+json, is there a link to the docs. Mutating probes, including the idempotency replay, only run against an isolated test environment.

  3. 3

    benchThe canonical task, run by a real agent

    Interaction Cost · Task Atomicity · Prior Knowledge

    Same model, same effort, same task for every competitor in a category. The API key never enters the model context: the agent asks for a call, the harness adds the authentication on the way out. One run in three happens with no access to the documentation, to measure what the model already knows about the API.

  4. 4

    judgeA model, under strict constraint

    Offload Value only

    Three independent passes, median kept. Every pass must produce a quote the harness then finds word for word in the fetched documentation; a quote it cannot find invalidates the pass. More than 15 points apart between passes, the score goes to human review and is not published.

The eleven criteria
and their weights.

The weight is the share of the overall score. They add up to 100.

Offload Valuejudged

Does the service let the agent hand off work that would be expensive to do itself through inference, or is it just CRUD in disguise?

15/100

Interaction Costmeasured

How many tokens does it cost me just to TALK to this API to complete its category's canonical task?

15/100

Error Recoverabilityprobed

When I hallucinate a parameter, does the error let me correct myself on the next turn, or do I give up?

15/100

Doc Legibilitymeasured

Can I read this documentation without running JavaScript, and without spending 80% of my tokens on markup?

12/100

Auth Frictionmeasured

How many steps between "I have a key" and "my first call succeeds"?

10/100

Safety & Reversibilityprobed

If I retry after a timeout, do I charge the customer twice?

10/100

Payload Efficiencyprobed

Can I ask for only the fields I need, or do I have to swallow 200 attributes to extract one?

8/100

Task Atomicitymeasured

How many calls for a single business intent?

6/100

Prior Knowledge Coveragemeasured

Do I already know this API by heart, or will I confidently hallucinate an outdated version?

5/100

Async Compatibilityprobed

If the operation is asynchronous, can I get the result without hosting a web server?

2/100

Tool-Native (MCP)binary

Does the vendor publish an official, maintained MCP server?

2/100

What we
refuse to publish.

Four cases where the catalogue stays silent rather than publish a score it could not defend.

01 /Refusal

A score with no evidence

Every score points to a URL, a literal quote, or a replayable HTTP exchange.

02 /Refusal

An overall below 70% coverage

The entry stays partial. The measured criteria are valid; their average would not be.

03 /Refusal

A ranking with fewer than three competitors

Criteria normalised per category mean nothing against a single measured rival.

04 /Refusal

A judgement whose passes disagree

On Crossmint, three passes returned scores 42 points apart. None of them was published.

We fail our own
legibility criterion.

10/100documentation legibility of this page, measured with our own scale

This page weighs roughly 61,000 bytes of HTML for 2,600 bytes of useful text, a ratio of 4.2%. Worse than most of the APIs we score. Two thirds of that weight is the framework's hydration payload, not our content.

We display it rather than fix it quietly. A benchmark that exempts itself from its own criteria is worth nothing. Until we bring that weight down, the surfaces agents actually consume (/llms.txt, /api/v1, the markdown twins) carry no markup that is not load-bearing.

How we probe,
and the right of reply.

We score companies that have security teams. Our credibility rests on our conduct as much as on our method.

Environment

Isolated test only. No production key ever enters the harness.

Rate

One request per second, fifty per probe pass.

Identity

An identifiable User-Agent, pointing at this page.

Cleanup

Resources created during runs are deleted afterwards, and we verify that they are.

Disputes

Right of reply, open to any vendor. The reply is the HTTP trace, never an argument.

The plain-language version: how it works.