← all news

Nine questions Epoch wants AI benchmarks to answer

AI · · · source (epoch.ai)

Greg Burnham of Epoch AI makes a case that the useful question about a benchmark is not "can the model do this task" but "what does the score tell us about where AI is heading." He lays out nine questions his team uses to decide what to build and measure. Some are about raw capability: whether models can move from single tasks to full open-ended jobs, whether longer reasoning unlocks problems that were out of reach, and how well a skill learned in one domain transfers to another. Others are about economics and research, including whether AI can do work that advances AI itself.

The concrete part is the benchmarks attached to each question. Epoch points to its Capabilities Index, which folds many tests into one general measure, and to MirrorCode for software tasks across languages, EBR-bench for learning through repeated play, and FrontierMath for problems that need genuinely new ideas. The Remote Labor Index grades models on real freelance projects against human work, and there is even Andon Café, an AI-run physical business. One recurring puzzle he flags is why scores across very different domains tend to move together, which hints at a single underlying capability rather than many separate ones.

Burnham's full list reads less as a scorecard and more as a map of what is still unresolved about model progress.

Why it matters

If you rely on benchmark numbers to pick models or judge progress, this is a checklist for reading them well: ask which of these questions a given score actually speaks to before you trust it. For teams building evaluations, it is a concrete set of gaps worth targeting.

BenchmarksEvaluationAI Research