Tavily
Search, extract and crawl API for AI agents.
Overall score
55.7/100
tavily-rest
Measured Oct 5, 2026
Status
Ranked
Complete benchmark. `overall` is a weighted average of the applicable criteria, comparable to other surfaces in the same category and the same rubric version.
Coverage
78% of the rubric
Surfaces
evaluated.
We score a surface, not a product. One product often exposes several, and they are not worth the same.
tavily-restrestprimary
78% coverage
55.7/100
Criterion
by criterion.
Measured on tavily-rest, rubric v0.2.0. A criterion with no score is not counted as zero: it is excluded and the weights are renormalised.
Offload Valuejudged · weight 15 · Oct 5, 2026
88/100
Tavily Search retrieves live web content the model cannot generate from weights: real-time indexed results, publish-date estimates, domain/country/language filtering, per-source semantically relevant chunk extraction, cleaned/parsed HTML, image extraction with descriptions, and an optional LLM-generated synthesized answer. This is genuine external data acquisition plus retrieval ranking and content distillation, all of which would be impossible or extremely costly for an agent to simulate through inference (and a naive agent would otherwise fetch and read whole pages, burning huge context). It falls short of the top band only because it performs no irreversible real-world side effect (no funds, signing, or filing) and because fetch-plus-summarize is approximable by a resourceful agent with raw HTTP access, albeit at much higher token cost.
- https://docs.tavily.com/documentation/api-reference/endpoint/search
Chunks are short content snippets (maximum 500 characters each) pulled directly from the source. Use chunks_per_source to define the maximum number of relevant chunks returned per source and to control the content length.
- https://docs.tavily.com/documentation/api-reference/endpoint/search
Highest relevance with increased latency. Best for detailed, high-precision queries. Returns multiple semantically relevant snippets per URL
- https://docs.tavily.com/documentation/api-reference/endpoint/search
advanced : Highest relevance with increased latency. Best for detailed, high-precision queries. Returns multiple semantically relevant snippets per URL
Interaction Costmeasured · weight 15 · Oct 5, 2026
6/100
Median over 2 successful run(s): 183629 tokens, 9.5 calls. Ratio to the best in "web-search" (3 surfaces measured, v0.2 scale). Cold, without documentation: 1/1 succeeded.
Error Recoverabilityprobed · weight 15 · Oct 5, 2026
54/100
unknown_endpoint → HTTP 404, clarity 25/100 (field named: no, allowed values: no, RFC 9457: no) · wrong_type → HTTP 422, clarity 75/100 (field named: yes, allowed values: yes, RFC 9457: no) · hallucinated_enum → HTTP 400, clarity 75/100 (field named: yes, allowed values: yes, RFC 9457: no) · unknown_field → HTTP 200, clarity 40/100 (field named: no, allowed values: no, RFC 9457: no)
- https://api.tavily.com/search
HTTP 401: { "detail": { "error": "Unauthorized: missing or invalid API key." } }
- https://api.tavily.com/nonexistent_resource_xyz
HTTP 404:
- https://api.tavily.com/search
HTTP 422: {"detail":[{"type":"literal_error","loc":["body","max_results","literal['auto']"],"msg":"Input should be 'auto'","input":"ten","ctx":{"expected":"'auto'"}},{"type":"int_parsing","loc":["body","max_res
- https://api.tavily.com/search
HTTP 400: {"detail":{"error":"Invalid search depth. Must be 'ultra-fast', 'fast', 'basic' or 'advanced'."}}
Doc Legibilitymeasured · weight 12 · Oct 5, 2026
48/100
text/markup ratio 4.8% (weight 30) · renders without JavaScript: yes (weight 25) · machine-readable spec: none found (weight 25) · llms.txt: https://docs.tavily.com/llms.txt (weight 10) · 12 runnable example(s) (weight 10). Weighted average over 100 points of available sub-criteria.
Auth Frictionmeasured · weight 10 · Oct 5, 2026
95/100
Authorization: Bearer <key>, verified by a successful call. One header, no ceremony.
- https://api.tavily.com/search
HTTP 401: { "detail": { "error": "Unauthorized: missing or invalid API key." } }
- https://api.tavily.com/nonexistent_resource_xyz
HTTP 404:
- https://api.tavily.com/search
HTTP 422: {"detail":[{"type":"literal_error","loc":["body","max_results","literal['auto']"],"msg":"Input should be 'auto'","input":"ten","ctx":{"expected":"'auto'"}},{"type":"int_parsing","loc":["body","max_res
- https://api.tavily.com/search
HTTP 400: {"detail":{"error":"Invalid search depth. Must be 'ultra-fast', 'fast', 'basic' or 'advanced'."}}
Safety & Reversibilityprobed · weight 10
—/100
If I retry after a timeout, do I charge the customer twice?
Payload Efficiencyprobed · weight 8
—/100
Can I ask for only the fields I need, or do I have to swallow 200 attributes to extract one?
Task Atomicitymeasured · weight 6 · Oct 5, 2026
16/100
Median over 2 successful run(s): 183629 tokens, 9.5 calls. Ratio to the best in "web-search" (3 surfaces measured, v0.2 scale). Cold, without documentation: 1/1 succeeded.
Prior Knowledge Coveragemeasured · weight 5 · Oct 5, 2026
100/100
Median over 2 successful run(s): 183629 tokens, 9.5 calls. Ratio to the best in "web-search" (3 surfaces measured, v0.2 scale). Cold, without documentation: 1/1 succeeded.
Async Compatibilityprobed · weight 2
—/100
If the operation is asynchronous, can I get the result without hosting a web server?
Tool-Native (MCP)binary · weight 2
—/100
Does the vendor publish an official, maintained MCP server?
Runs, verified
never on the agent’s word.
Every row is a real attempt at the category’s canonical task. Success is verified by reading the created resource back.
| Task | Tokens | Calls | Errors | Docs | Result |
|---|---|---|---|---|---|
| find-official-source | 77,442 | 8 | 0 | without | verified |
| find-official-source | 83,311 | 8 | 0 | with | verified |
| find-official-source | 283,947 | 11 | 0 | with | verified |
The same entry,
without the markup.
You publish this API and dispute a score? The methodology is public and every score above points to the evidence that produced it.