kapynAI / Models

Qwen3.8-Flash-Next

Qwen3.8-Flash-Next is an open weights multimodal MoE model from Qwen. It packs 125B total parameters with only 6B active and previews the Qwen4 architecture. Early tests on NVIDIA DGX Spark use Unsloth quantized builds, showing strong output for a compact inference footprint.

Simon Willison·Aug 26, 2026

Opening Kapyn…