All benchmarksCode ↗LiveBrowseComp ↗Paper ↗

Coding-Agent CLIs × Search Backends on LBC-30

When an agent has to find a hard fact on the live web, what decides whether it succeeds: the agent, or the search tool plugged into it? Most benchmarks test an agent with whatever search it ships with, so a score never separates the model from its retrieval. This study isolates the two.

#Systemvs bestSolvedGen cost
1claude-exaclaude code · exa mcp · cheapest Claude arm60.0%reference18 / 3066.753.3$50.31
2claudeclaude code · builtin websearch · 1,632 builtin searches46.7%13.314 / 3046.746.7$120.06
3codexcodex cli · live web_search · 2.4M tokens46.7%13.314 / 3040.053.3~$1–6*
4codex-exacodex cli · exa mcp · 425 MCP calls40.0%20.012 / 3046.733.3~$1–5*
5claude-valyuclaude code · valyu mcp26.7%33.38 / 3026.726.7$117.93
6codex-valyucodex cli · valyu mcp · 858 MCP calls, 2.7M tokens20.0%40.06 / 3033.36.7~$1.5–7*

click a metric header to sort · hover a header for its definition

The same numbers as a 2 × 3

backend moves accuracy 27–33 points · agent at most 6.7

search ↓ · agent →Claude CodeCodex CLIbackend spread
builtin46.7%46.7%0.0 pts
exa60.0%40.0%20.0 pts
valyu26.7%20.0%6.7 pts
agent spread33.3 pts26.7 pts

Per-task verdicts are on the ; per-run search behavior from the raw traces is under and per task in the .

caveats

n=30, single run per cell, so one task is worth ~3.3 points; treat gaps under 10 points as noise (claude-exa’s lead and the valyu deficit clear that bar; builtin-vs-codex-exa does not). Codex ran at its CLI-default reasoning effort (“none” in this build). The judge is Gemini, not the paper’s GPT-OSS. The two funders’ backends finished first and last; disclosure below.

Funding

Exa provided $1,000 in API credits for this benchmark series (shared with RB-30 and DRB2-20), and Valyu provided $500 for this benchmark (disclosed post-freeze, pre-main-run; Amendment F). Both are compared search backends here; one funder’s backend won and the other’s lost, under configs frozen mechanically before the first run. Neither had input into task selection, configurations, prompts or judging.

30 tasks · 15 dated / 15 undated · 6 systems · 180 main runs, 0 failures · judge gemini-3.1-pro-preview, blinded · frozen & run 2026-08-13 · 8 amendments, none post-score · questions & answers © LiveBrowseComp authors · repository · RESULTS.md ·