Skip to content

Twilio SendGrid

EmailREST · sendgrid-rest-v3

Email delivery platform at scale.

Overall score

52.6/100

sendgrid-rest-v3

Measured Oct 5, 2026

Status

Ranked

Complete benchmark. `overall` is a weighted average of the applicable criteria, comparable to other surfaces in the same category and the same rubric version.

Coverage

98% of the rubric

Surfaces
evaluated.

We score a surface, not a product. One product often exposes several, and they are not worth the same.

sendgrid-rest-v3restprimary

98% coverage

52.6/100

Criterion
by criterion.

Measured on sendgrid-rest-v3, rubric v0.2.0. A criterion with no score is not counted as zero: it is excluded and the weights are renormalised.

Offload Valuejudged · weight 15 · Oct 5, 2026

93/100

Email sending is a genuine external side effect an agent cannot perform via inference: SMTP delivery, deliverability infrastructure, IP warmup, domain authentication/DKIM, bounce and spam suppression lists, and email address validation. These involve proprietary reputation data and compliance (unsubscribe groups, enforced TLS) impossible to replicate with reasoning tokens. Much of the surface area is admin CRUD (templates, contacts, settings), which caps it slightly below the top.

Interaction Costmeasured · weight 15 · Oct 5, 2026

51/100

Median over 2 successful run(s): 10015 tokens, 2 calls. Ratio to the best in "transactional-email" (3 surfaces measured, v0.2 scale). Cold, without documentation: 1/1 succeeded.

Error Recoverabilityprobed · weight 15 · Oct 5, 2026

39/100

unknown_endpoint → HTTP 404, clarity 30/100 (field named: no, allowed values: no, RFC 9457: no) · wrong_type → HTTP 400, clarity 30/100 (field named: no, allowed values: no, RFC 9457: no) · hallucinated_enum → HTTP 400, clarity 50/100 (field named: no, allowed values: yes, RFC 9457: no) · unknown_field → HTTP 200, clarity 25/100 (field named: no, allowed values: no, RFC 9457: no) · missing_required_param → HTTP 400, clarity 60/100 (field named: no, allowed values: yes, RFC 9457: no) · idempotency_replay → HTTP 201, clarity 30/100 (field named: no, allowed values: no, RFC 9457: no)

Doc Legibilitymeasured · weight 12 · Oct 5, 2026

17/100

text/markup ratio 3.6% (weight 30) · rendering without JavaScript undetermined (excluded from the calculation) (weight 25) · machine-readable spec: none found (weight 25) · llms.txt: https://www.twilio.com/llms.txt (weight 10) · 0 runnable example(s) (weight 10). Weighted average over 75 points of available sub-criteria.

Auth Frictionmeasured · weight 10 · Oct 5, 2026

95/100

Authorization: Bearer <key>, verified by a successful call. One header, no ceremony.

Safety & Reversibilityprobed · weight 10 · Oct 5, 2026

42/100

Two identical POSTs with the same Idempotency-Key header created TWO distinct resources (d-05d889cf52d94f55b4f11bed57417c4a then d-b08c4a7b6cdd423390786de55677bb66). Every network timeout becomes a duplicate. (weight 35) · Isolated test environment available. (weight 25). Weighted average over 60 points of measured sub-criteria.

Payload Efficiencyprobed · weight 8 · Oct 5, 2026

5/100

No discoverable field selection (tested fields, select, _fields, expand, properties): the agent swallows the whole resource. (weight 35) · No cursor pagination marker in the response. (weight 20) · Nested envelope or non-list shape. (weight 10) · No ETag. (Compression is not measurable: the runtime decompresses it and strips the header.) (weight 10). Weighted average over 75 points of measured sub-criteria.

Task Atomicitymeasured · weight 6 · Oct 5, 2026

50/100

Median over 2 successful run(s): 10015 tokens, 2 calls. Ratio to the best in "transactional-email" (3 surfaces measured, v0.2 scale). Cold, without documentation: 1/1 succeeded.

Prior Knowledge Coveragemeasured · weight 5 · Oct 5, 2026

100/100

Median over 2 successful run(s): 10015 tokens, 2 calls. Ratio to the best in "transactional-email" (3 surfaces measured, v0.2 scale). Cold, without documentation: 1/1 succeeded.

Async Compatibilityprobed · weight 2 · Oct 5, 2026

0/100

POST /mail/send returns no trackable identifier: an agent can neither re-read nor wait for the result without a webhook. Long-polling / SSE not probed, excluded from the weighting.

Tool-Native (MCP)binary · weight 2

—/100

Does the vendor publish an official, maintained MCP server?

Runs, verified
never on the agent’s word.

Every row is a real attempt at the category’s canonical task. Success is verified by reading the created resource back.

TaskTokensCallsErrorsDocsResult
send-transactional-email3,81410withoutverified
send-transactional-email15,97330withverified
send-transactional-email4,05710withverified