kapynOpen Source

local-kimi — Optimized local serving engine for Kimi-Linear-48B: INT4 quantizer, fused decode kernels for a measured 3.18x, and an Op

Local-kimi is an optimized local serving engine designed for the Kimi-Linear-48B model. It features an INT4 quantizer, fused decode kernels delivering a 3.18x speedup, and an OpenAI-compatible server. The included k3 bridge automatically detects clients per request, ensuring seamless integration with coding tools like Claude Code and Aider without configuration changes.

GitHub·Jul 29, 2026

Opening Kapyn…