Benchmark · company data
Company Search API Benchmark
When you ask a data API for “companies in {country} in {industry} with {N} employees”, how many of the results actually match — and how fast, how deep, and how accurate is the headcount? We asked each provider the same 12 frozen firmographic queries and blind-judged every company they returned.
Updated 2026-07-22 · judge openai/gpt-4.1-mini · 12 queries · K=25· every raw request & judge call is published for verification.
Leaderboard
| Provider | Precision (of returned) | Precision@10 | Size accuracy | Depth (avg/25) | Coverage | Latency |
|---|---|---|---|---|---|---|
| #1Crustdata | 36% | 35.5% | 97.5% | 23 | 11/12 | 5,868ms |
| Coresignal | 16.1% | 10% | 51.1% | 7 | 9/12 | 191ms |
Precision of returned results
Share of returned companies that genuinely match the requested industry and HQ country (blind LLM judge).
Verdict — route by workflow
- Crustdata leads on result quality (36% precision), headcount accuracy (97.5%), and depth (23 companies/query).
- Coresignal is dramatically faster (191ms) — best when latency dominates — but returns fewer, looser matches.
Quota-limited this run
Explorium, People Data Labs returned no results because the test account hit a free-tier / credit limit (Explorium: 12/12 errored; People Data Labs: 12/12 errored). They are excluded from the leaderboard until re-run on a paid tier — not scored as zero.
Methodology
A frozen set of 12 firmographic queries fixes a country, an industry, and a headcount band. Each provider is asked for up to 25 matching companies via its native company-search API. Industry names are translated to each provider’s taxonomy so the query is expressed as faithfully as each API allows.
Relevance is graded blind: each returned company (name + domain + description, with the provider’s identity stripped) is judged by openai/gpt-4.1-mini for whether it is genuinely in the requested industry and headquartered in the requested country. Size accuracy is deterministic — the provider’s own reported headcount is checked against the requested band.
Every HTTP request/response and every judge call is cached under runs/ so the entire leaderboard can be independently re-graded. Judge cost for this run: $0.0255.
Per-query results
| Query | Crustdata | Coresignal |
|---|---|---|
| us-fintech-smb | 30% · 25 | 0% · 7 |
| us-software-startup | 0% · 25 | 0% · 20 |
| uk-adtech-mid | 20% · 25 | 30% · 10 |
| de-machinery-mid | 40% · 25 | 10% · 5 |
| us-health-large | — | — |
| ca-itservices-startup | 50% · 25 | 10% · 20 |
| us-retail-mid | 50% · 25 | — |
| us-realestate-smb | 40% · 25 | 0% · 4 |
| au-fintech-startup | 50% · 25 | 30% · 20 |
| us-biotech-smb | 50% · 25 | 10% · 2 |
| fr-software-smb | 40% · 25 | — |
| us-logistics-mid | 20% · 25 | 0% · 1 |
cell = Precision@10 · results returned
Part of the CompanyDataBenchmarks open benchmark suite · dataset released CC BY 4.0.