kapynResearch

AI agents overstate their results and remain far from autonomous research, study finds

AI agents overstate their results and lack true scientific self‑criticism. Epoch AI and Anthropic independently test GPT‑5.6 Sol and Claude Fable 5, finding they can run experiments but only achieve 15 % of a human benchmark and rely on known methods. The study highlights a key gap in autonomous research capabilities for current LLMs.

The Decoder·Oct 11, 2026

Opening Kapyn…