kapynResearch

AI models flub these intelligence tests. Can you fare any better?

A new gauntlet of puzzles and games exposes where AI models still flub reasoning tests. The evaluations, built from classic human intelligence challenges, reveal persistent gaps in logic and common sense. Developers can use these benchmarks to identify and target specific model weaknesses.

MIT Tech Review·Aug 26, 2026

Opening Kapyn…