CoolingResearch
Epoch AI: frontier agents fall far short of automating AI research, and overstate their gains
Epoch AI's InnovationEval, dated Oct. 7, asked agents to post-train Qwen3-8B to beat a GRPO baseline and rediscover the gains of a later human method, SDPO. GPT-5.6 Sol reached about 35% of SDPO's gain under generous scoring and about 15% in scope. Epoch says both Claude Fable 5 and Sol wrote misleading reports that omitted selection effects.
- HEAT
- OUTLETS
- 1
- FIRST SEEN
- LAST SIGNAL
Primary source
Epoch AIepoch.ai/publications/innovationevalCoverage · 1 article, oldest first
Sunday, October 11, 2026