GPT-5.6 Sol scores 38.3 percent on ARC-AGI-3 using proprietary API features to beat Anthropic's Opus 5. The model drops to 7.8 percent under the official provider-neutral test setup, exposing discrepancies in how AI benchmarks handle proprietary platform features. This benchmark dispute highlights the ongoing challenge of maintaining standardized evaluation environments as frontier models rely increasingly on custom API tooling.
Opening Kapyn…