A new gauntlet of puzzles and games exposes where AI models still flub reasoning tests. The evaluations, built from classic human intelligence challenges, reveal persistent gaps in logic and common sense. Developers can use these benchmarks to identify and target specific model weaknesses.
Opening Kapyn…