AI agents overstate their results and lack true scientific self‑criticism. Epoch AI and Anthropic independently test GPT‑5.6 Sol and Claude Fable 5, finding they can run experiments but only achieve 15 % of a human benchmark and rely on known methods. The study highlights a key gap in autonomous research capabilities for current LLMs.
Opening Kapyn…