Benchmarks

pre-registered comparisons of deep research apis. task sets, provider configurations and judging protocols are frozen and published before any run, and every raw run is public.

LBC-30aug 2026
Claude Code vs Codex CLI under three search backends (builtin, Exa MCP, Valyu MCP) on a frozen 30-task LiveBrowseComp subset. claude-exa wins at 60%; the search backend moves accuracy up to 33 points, the agent at most 7.
30 questions · 2 agents × 3 search backends · 180 runs, 0 failures · blinded BrowseComp grading
RB-30aug 2026
Five deep-research APIs plus a Perplexity calibration anchor on a frozen 30-task ResearcherBench subset. Coverage is a three-way race; faithfulness splits the field in two; parallel has the best overall profile.
30 questions · 6 systems · 180 runs · ~32k claims verified · 2 judge tracks
DRB2-20aug 2026
Nine deep research APIs on a frozen 20-task subset of DeepResearch-Bench-II. Exa effort tiers against each other, then one flagship tier per provider.
20 tasks · 20 domains · 1,504 rubrics · 180 runs