Firebase2 calls
6 500tokens
One document per POST. That is the floor of the catalogue on this task.
An agent is a program driven by a language model that is handed a goal rather than a list of instructions. To reach it, the agent has to read documentation and call APIs, exactly like a developer, with three differences: it moves fast, it never tires, and everything it reads is billed to it.
An API can be excellent for a human and miserable for an agent. Beautiful documentation that only exists as images, an error message that says nothing but “invalid request”, a checkout that runs through a button in a browser: none of that bothers a person, and all of it stops an agent. AgentScore measures that gap.
One task, one agent, one public documentation. Everything else is measured, never assumed.
1 task per category
“Charge €10 to a new customer.” The agent gets nothing but the public documentation and a test key. Nobody tells it which endpoints to call or in what order: this is the situation of an agent discovering a service.
tokens · calls · errors
Every word read or written is a billed token. We measure how many it spends, how many calls it needs, how often it gets it wrong. An API that takes nine calls where another takes two costs four times as much, for the same service.
server-side read-back
The agent says it created the payment? We go read it back through the API. If it is not there, the attempt counts as a failure whatever the agent claims. A confident model is not a model that succeeded.
5 probes per surface
We invent a parameter, we put text where a number belongs, we call an endpoint that does not exist. The question is not whether the API refuses, it is what it says. “Bad Request” leaves the agent stuck; “the limit parameter expects an integer” lets it fix itself.
228 attached pieces of evidence
Every score points to what produced it: the URL we read, the exact sentence from the documentation, or the full HTTP exchange. A vendor disputing its score gets the trace, not an argument.
Write a record to a database and read it back, take ten euros. Identical service from one competitor to the next, unrelated cost. And sometimes no cost at all, because the task is simply out of an agent’s reach.
Firebase2 calls
6 500tokens
One document per POST. That is the floor of the catalogue on this task.
Appwrite9.5 calls
57 300tokens
Create the database, then the collection, then every attribute, before writing a single row. Eight and a half times the price, for the same service.
Paddlea wall
0successes out of 3
Customer, product and price created without a single error, then the same refusal three times over: taking the payment requires a hosted checkout no headless agent gets through.
Four rules that cost the catalogue entries, and without which the rest would be worth nothing.
01 /Rule
An agent does not produce an opinion, it produces an execution trace.
02 /Rule
What has not been measured stays empty, and never turns into a zero.
03 /Rule
The tokens of sending an email do not compare to the tokens of taking a payment.
04 /Rule
Three passes more than 15 points apart go to human review.