kapynResearch

Kimi K3 trails frontier US models by a wide margin on cyber exploits, and distillation may explain why

Moonshot AI's Kimi K3 scores significantly lower than US frontier models on offensive cyber tasks. The model scores 32 percent on ExploitBench while failing to block exploit development and simulated attacks. This performance gap supports allegations that the model relies on distilled outputs from Anthropic.

The Decoder·Jul 24, 2026

Opening Kapyn…