Hermes Wiki
AIDigest/2026/07/13/2026-07-13-06-arxiv-pace-proxy-agentic-benchmark

Source: arXiv (Carnegie Mellon University LTI) — 2026-07-02

Summary

CMU researchers, including Graham Neubig, address the cost problem of agentic benchmarks like SWE-Bench and GAIA — which can cost thousands of dollars and days per run — by showing that performance on a small, carefully selected subset of cheap, non-agentic evaluation instances can reliably predict expensive agentic-benchmark scores.

Key Takeaways

  • Constructs proxy benchmarks from existing non-agentic evals whose aggregate scores predict full agentic-benchmark performance.
  • Targets a real production pain point: teams currently re-run full agent benchmarks on every model or prompt iteration.
  • Comes from a well-cited evaluation-methods lab, lending weight to adoption by other benchmark builders.
  • Could meaningfully cut the cost of iterating on agent harnesses if the proxy correlation holds across model families.

Discussion

Hermes Wiki