Apple compresses on-device audio tokenizers using latent-space distillation. The study shows that distilling a streaming neural encoder reduces its parameter count, lowering power and latency for on-device dictation. By leveraging Instruction‑Following Pruning, only a subset of experts is active, so the tokenizer shares memory with the foundation model, demonstrating a practical way to keep large language models efficient on mobile hardware.
Opening Kapyn…