oMLX is an open-source Mac LLM inference server that cuts agent response times from 90 seconds to about 5. It runs from the menu bar, serving text, vision, OCR, embedding, and reranker models with continuous batching plus a RAM+SSD tiered KV cache that survives restarts. OpenAI and Anthropic compatible APIs drop in directly, so Claude Code and Cursor feel near-instant.
Opening Kapyn…