Firecrawl Benchmarks: open retrieval evaluations
Developer Retrieval Benchmark
Firecrawl runs these too. They have no page of their own yet: the numbers and the method are published on the pages that cite them.
Scrape coverage and quality, run Jan 13, 2026
1,000 URLs from ten categories of public web pages, scored on whether each tool returned the core page text.
- Coverage (success rate)
- 96%
- Extraction accuracy (F1)
- 0.638
- Content recall
- 0.639
- Latency (P95)
- 3,387 ms
Dataset and scoring
1,000 URLs drawn from diverse public web domains (news, documentation, e-commerce, finance, and more). A URL counts as covered when the tool retrieved at least 10% of the expected content, defined as core page text excluding navigation, ads, and footers.
Benchmarks Firecrawl appears in but did not run. The figures, the method and the scoring belong to whoever published them.
Openbenchmarks token efficiency (search only), Sep 15, 2026 snapshot
The study's search-only token board: 14 configurations ranked by median LLM tokens per task. Firecrawl places first.
- Median task tokens
- 7,456
- Task completion
- 70.3%
Scope
Search-only: the agent works from search results and opens no page. The same coding agent, model, and prompts run against every API, and a task passes only with a grounded source URL from that run.
Openbenchmarks token-efficiency ranking · openbenchmarks-labs/web-search-for-coding-agents
Openbenchmarks token efficiency (search and fetch), Sep 15, 2026 snapshot
The study's search-and-fetch token board: 14 configurations ranked by median LLM tokens per task. Firecrawl places third behind TinyFish and Nimble.
- Median task tokens
- 17,379
- Task completion
- 76.0%
Scope
Search and fetch: the agent may open pages after searching, with each vendor's own fetch or extract endpoint. Same tasks, agent, model, and prompts as the search-only board; the ranking differs.
Openbenchmarks token-efficiency ranking · openbenchmarks-labs/web-search-for-coding-agents
Openbenchmarks coding tickets, Sep 1, 2026 snapshot
An independent search-only comparison of 11 web search APIs on 100 held-out documentation tickets. Firecrawl places second behind Perplexity.
- Grounded task completion
- 70.3%
Scope
Search-only: every vendor returns snippets and nothing opens a page. The study publishes a separate search and fetch board on which the ranking differs.
Openbenchmarks web search study · openbenchmarks-labs/web-search-for-coding-agents
Openbenchmarks multi-hop discovery, Sep 1, 2026 snapshot
The agent half of the same study: 45 multi-constraint discovery questions, search-only. Firecrawl places twelfth of sixteen configurations, well behind Parallel.
- Multi-hop F1
- 30.4%
Scope
Search-only, with no page opened, on questions that need several searches to assemble a complete set. It is a different task set from the study's coding tickets, so the two rows are not comparable to each other.
Openbenchmarks web search study · openbenchmarks-labs/multi-turn-company-search
AIMultiple agentic search, Dec 2025 snapshot
A third-party comparison of 8 search APIs over 100 AI and LLM queries. Read the ordering as a ranking, not a proven gap: the top intervals overlap.
- Agent Score
- 14.58
- Mean relevant results
- 4.30
- Quality
- 3.39
Same input for everyone
Every system gets the same inputs in the same order with default settings. No per-system tuning and no dropping the cases a system struggles with.
Scoring you can check
Metrics are computed against human-annotated answers, not judged by a model, so the same run produces the same score. Every metric a benchmark page carries is defined in full on that page.
Our own numbers included
Firecrawl runs these benchmarks and appears in them. Every result is published, dated, and changelogged so the comparison can be argued with rather than taken on faith.