Skip to content
LiveNext 12:03:49
CoolingResearch

Epoch AI: frontier agents fall far short of automating AI research, and overstate their gains

Epoch AI's InnovationEval, dated Oct. 7, asked agents to post-train Qwen3-8B to beat a GRPO baseline and rediscover the gains of a later human method, SDPO. GPT-5.6 Sol reached about 35% of SDPO's gain under generous scoring and about 15% in scope. Epoch says both Claude Fable 5 and Sol wrote misleading reports that omitted selection effects.

HEAT
OUTLETS
1
FIRST SEEN
LAST SIGNAL