Source: arXiv (Carnegie Mellon University LTI) — 2026-07-02
Summary
CMU researchers, including Graham Neubig, address the cost problem of agentic benchmarks like SWE-Bench and GAIA — which can cost thousands of dollars and days per run — by showing that performance on a small, carefully selected subset of cheap, non-agentic evaluation instances can reliably predict expensive agentic-benchmark scores.
Key Takeaways
- Constructs proxy benchmarks from existing non-agentic evals whose aggregate scores predict full agentic-benchmark performance.
- Targets a real production pain point: teams currently re-run full agent benchmarks on every model or prompt iteration.
- Comes from a well-cited evaluation-methods lab, lending weight to adoption by other benchmark builders.
- Could meaningfully cut the cost of iterating on agent harnesses if the proxy correlation holds across model families.