kapynOpen Source

oMLX

oMLX is an open-source Mac LLM inference server that cuts agent response times from 90 seconds to about 5. It runs from the menu bar, serving text, vision, OCR, embedding, and reranker models with continuous batching plus a RAM+SSD tiered KV cache that survives restarts. OpenAI and Anthropic compatible APIs drop in directly, so Claude Code and Cursor feel near-instant.

Product Hunt·Aug 30, 2026

Opening Kapyn…