Web data benchmarks,
in the open.
Benchmarks for choosing a web data API to search, scrape, interact, and format data for AI.
Published benchmarks
Developer Retrieval Benchmark
An agent runs the same developer questions through each system with one search tool and ten results per call, then scores whether the right answer is cited across three tracks: the repository behind a described capability, the pull request that fixed a bug, and the docs page that answers a how-to. Compare Recall@10 per track and MRR@10 overall.
How we benchmark
Same input for everyone
Every system gets the same inputs in the same order with default settings. No per-system tuning and no dropping the cases a system struggles with.
Scoring you can check
Metrics are computed against human-annotated answers, not judged by a model, so the same run produces the same score. Every metric is defined in full on the benchmark page.
Our own numbers included
Firecrawl runs these benchmarks and appears in them. Every result is published, dated, and changelogged so the comparison can be argued with rather than taken on faith.