Deepseek V4.1‑Flash is a 552‑billion‑parameter multimodal model that cuts KV cache memory to a quarter of its predecessor. It narrowly outperforms Opus 5 and GPT‑5.6 Sol on the DeepSWE coding benchmark while using only 16 billion active parameters per token, and ships under an MIT license for cheaper AI agents.
Opening Kapyn…