Skip to content

Tavily

Web searchREST · tavily-rest

Search, extract and crawl API for AI agents.

Overall score

55.7/100

tavily-rest

Measured Oct 5, 2026

Status

Ranked

Complete benchmark. `overall` is a weighted average of the applicable criteria, comparable to other surfaces in the same category and the same rubric version.

Coverage

78% of the rubric

Surfaces
evaluated.

We score a surface, not a product. One product often exposes several, and they are not worth the same.

tavily-restrestprimary

78% coverage

55.7/100

Criterion
by criterion.

Measured on tavily-rest, rubric v0.2.0. A criterion with no score is not counted as zero: it is excluded and the weights are renormalised.

Offload Valuejudged · weight 15 · Oct 5, 2026

88/100

Tavily Search retrieves live web content the model cannot generate from weights: real-time indexed results, publish-date estimates, domain/country/language filtering, per-source semantically relevant chunk extraction, cleaned/parsed HTML, image extraction with descriptions, and an optional LLM-generated synthesized answer. This is genuine external data acquisition plus retrieval ranking and content distillation, all of which would be impossible or extremely costly for an agent to simulate through inference (and a naive agent would otherwise fetch and read whole pages, burning huge context). It falls short of the top band only because it performs no irreversible real-world side effect (no funds, signing, or filing) and because fetch-plus-summarize is approximable by a resourceful agent with raw HTTP access, albeit at much higher token cost.

Interaction Costmeasured · weight 15 · Oct 5, 2026

6/100

Median over 2 successful run(s): 183629 tokens, 9.5 calls. Ratio to the best in "web-search" (3 surfaces measured, v0.2 scale). Cold, without documentation: 1/1 succeeded.

Error Recoverabilityprobed · weight 15 · Oct 5, 2026

54/100

unknown_endpoint → HTTP 404, clarity 25/100 (field named: no, allowed values: no, RFC 9457: no) · wrong_type → HTTP 422, clarity 75/100 (field named: yes, allowed values: yes, RFC 9457: no) · hallucinated_enum → HTTP 400, clarity 75/100 (field named: yes, allowed values: yes, RFC 9457: no) · unknown_field → HTTP 200, clarity 40/100 (field named: no, allowed values: no, RFC 9457: no)

Doc Legibilitymeasured · weight 12 · Oct 5, 2026

48/100

text/markup ratio 4.8% (weight 30) · renders without JavaScript: yes (weight 25) · machine-readable spec: none found (weight 25) · llms.txt: https://docs.tavily.com/llms.txt (weight 10) · 12 runnable example(s) (weight 10). Weighted average over 100 points of available sub-criteria.

Auth Frictionmeasured · weight 10 · Oct 5, 2026

95/100

Authorization: Bearer <key>, verified by a successful call. One header, no ceremony.

Safety & Reversibilityprobed · weight 10

—/100

If I retry after a timeout, do I charge the customer twice?

Payload Efficiencyprobed · weight 8

—/100

Can I ask for only the fields I need, or do I have to swallow 200 attributes to extract one?

Task Atomicitymeasured · weight 6 · Oct 5, 2026

16/100

Median over 2 successful run(s): 183629 tokens, 9.5 calls. Ratio to the best in "web-search" (3 surfaces measured, v0.2 scale). Cold, without documentation: 1/1 succeeded.

Prior Knowledge Coveragemeasured · weight 5 · Oct 5, 2026

100/100

Median over 2 successful run(s): 183629 tokens, 9.5 calls. Ratio to the best in "web-search" (3 surfaces measured, v0.2 scale). Cold, without documentation: 1/1 succeeded.

Async Compatibilityprobed · weight 2

—/100

If the operation is asynchronous, can I get the result without hosting a web server?

Tool-Native (MCP)binary · weight 2

—/100

Does the vendor publish an official, maintained MCP server?

Runs, verified
never on the agent’s word.

Every row is a real attempt at the category’s canonical task. Success is verified by reading the created resource back.

TaskTokensCallsErrorsDocsResult
find-official-source77,44280withoutverified
find-official-source83,31180withverified
find-official-source283,947110withverified