kapynOpen Source

v100-skinny — Hand-written NVFP4 W4A16 CUDA kernels and chain-MTP speculative serving — Qwen3.6-27B at up to 366 tok/s on four Tesla V

v100-skinny is an inference engine utilizing hand-written CUDA kernels for low-precision execution on older hardware. It delivers up to 366 tokens per second for the Qwen3.6-27B model on four Tesla V100 GPUs by implementing NVFP4 W4A16 kernels and chain-MTP speculative serving on hardware lacking native FP4 support. This project provides developers with a high-performance optimization pathway for running advanced large language models on legacy data center infrastructure.

GitHub·Aug 11, 2026

Opening Kapyn…