Skip to content

Software trying to use
other software.

An agent is a program driven by a language model that is handed a goal rather than a list of instructions. To reach it, the agent has to read documentation and call APIs, exactly like a developer, with three differences: it moves fast, it never tires, and everything it reads is billed to it.

An API can be excellent for a human and miserable for an agent. Beautiful documentation that only exists as images, an error message that says nothing but “invalid request”, a checkout that runs through a button in a browser: none of that bothers a person, and all of it stops an agent. AgentScore measures that gap.

The protocol,
in five steps.

One task, one agent, one public documentation. Everything else is measured, never assumed.

  1. 1

    We give an agent a task, not a questionnaire

    1 task per category

    “Charge €10 to a new customer.” The agent gets nothing but the public documentation and a test key. Nobody tells it which endpoints to call or in what order: this is the situation of an agent discovering a service.

  2. 2

    We count what it costs the agent

    tokens · calls · errors

    Every word read or written is a billed token. We measure how many it spends, how many calls it needs, how often it gets it wrong. An API that takes nine calls where another takes two costs four times as much, for the same service.

  3. 3

    We check the result ourselves

    server-side read-back

    The agent says it created the payment? We go read it back through the API. If it is not there, the attempt counts as a failure whatever the agent claims. A confident model is not a model that succeeded.

  4. 4

    We send deliberately broken requests

    5 probes per surface

    We invent a parameter, we put text where a number belongs, we call an endpoint that does not exist. The question is not whether the API refuses, it is what it says. “Bad Request” leaves the agent stuck; “the limit parameter expects an integer” lets it fix itself.

  5. 5

    We publish the evidence, not an opinion

    228 attached pieces of evidence

    Every score points to what produced it: the URL we read, the exact sentence from the documentation, or the full HTTP exchange. A vendor disputing its score gets the trace, not an argument.

The same task, twice the price,
or no price at all.

Write a record to a database and read it back, take ten euros. Identical service from one competitor to the next, unrelated cost. And sometimes no cost at all, because the task is simply out of an agent’s reach.

Firebase2 calls

6 500tokens

One document per POST. That is the floor of the catalogue on this task.

Appwrite9.5 calls

57 300tokens

Create the database, then the collection, then every attribute, before writing a single row. Eight and a half times the price, for the same service.

Paddlea wall

0successes out of 3

Customer, product and price created without a single error, then the same refusal three times over: taking the payment requires a hosted checkout no headless agent gets through.

What we
refuse to do.

Four rules that cost the catalogue entries, and without which the rest would be worth nothing.

01 /Rule

No user reviews

An agent does not produce an opinion, it produces an execution trace.

02 /Rule

No cell filled in by default

What has not been measured stays empty, and never turns into a zero.

03 /Rule

No comparison across jobs

The tokens of sending an email do not compare to the tokens of taking a payment.

04 /Rule

No score when our own measurements disagree

Three passes more than 15 points apart go to human review.

The technical detail of the pipeline lives on the methodology and architecture page.