Author: jundot
Stars: 366 stars today
Description: LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
LLM inference, optimized for your Mac
Continuous batching and tiered KV caching, managed directly from your menu bar.
junkim.dot@gmail.com · https://omlx.ai/me
Install · Quickstart · Features · Models · CLI Configuration · Benchmarks · oMLX.ai
Every LLM server I tried made me choose between convenience and control. I wanted to pin everyday models in memory, auto-swap heavier ones on demand, set context limits - and manage it all from a menu bar.
oMLX persists KV cache across a hot in-memory tier and cold SSD tier - even when context changes mid-conversation, all past context stays cached and reusable across requests, making local LLMs practical for real coding work with tools like Claude Code. That's why I built it.
Download the .dmg from Releases, drag to Applications, done. The app includes in-app auto-update, so future upgrades are just one click. The macOS app also installs a lightweight ~/.omlx/bin/omlx CLI shim so terminal commands and Apple Shortcuts can control the app-managed server.
```bash brew tap jundot/omlx https://github.com/jundot/omlx brew install jundot/omlx/omlx
brew update && brew upgrade omlx
omlx start
/opt/homebrew/opt/omlx/libexec/bin/pip install mcp ```
Optional GLM-5.2 / MiniMax M3 native custom kernels currently require a HEAD build:
bash
brew install jundot/omlx/omlx --HEAD --with-custom-kernel
```bash git clone https://github.com/jundot/omlx.git cd omlx pip install -e . # Core only pip install -e ".[mcp]" # With MCP (Model Context Protocol) support
OMLX_WITH_CUSTOM_KERNEL=1 pip install -e . ```
Requires macOS 15.0+ (Sequoia), Python 3.11–3.13, and Apple Silicon (M1/M2/M3/M4).
Note on native custom kernels: a plain
pip install -e .does NOT build them, and the affected model families then silently fall back to much slower generic paths -- for GLM-5.2 the fused DSA prefill is roughly 30x faster with the kernels (measured 845 vs ~29 tok/s on an M3 Ultra), and the fallback also uses more memory (#2137). Building them requires the Metal toolchain, which Command Line Tools alone do not provide (xcrun: error: unable to find utility "metal"): install full Xcode, or use the official DMG which ships the kernels precompiled. Homebrew can build them withbrew install jundot/omlx/omlx --HEAD --with-custom-kernel, but that build also needs full Xcode. To verify your install:
bash python -c "from omlx.custom_kernels import native_kernel_status; print(native_kernel_status())"
Launch oMLX from your Applications folder. The Welcome screen guides you through three steps - model directory, server start, and first model download. That's it. To connect OpenClaw, OpenCode, Codex, Hermes Agent, or Copilot, see Integrations.
```bash
omlx start omlx stop omlx restart
omlx serve --model-dir ~/models ```
The server discovers LLMs, VLMs, embedding models, and rerankers from subdirectories automatically. Any OpenAI-compatible client can connect to http://localhost:8000/v1. A built-in chat UI is also available at http://localhost:8000/admin/chat.
If you installed via Homebrew, you can run oMLX as a managed background service:
```bash omlx start # Start via brew services omlx stop # Stop omlx restart # Restart
brew services start omlx # Start (auto-restarts on crash) brew services stop omlx # Stop brew services restart omlx # Restart brew services info omlx # Check status ```
The service runs omlx serve with zero-config defaults (~/.omlx/models, port 8000). omlx start, omlx stop, and omlx restart are the portable lifecycle commands; Homebrew installs delegate them to brew services. To customize, either set environment variables (OMLX_MODEL_DIR, OMLX_PORT, etc.) or run omlx serve --model-dir /your/path once to persist settings to ~/.omlx/settings.json.
Logs are written to two locations:
- Service log: $(brew --prefix)/var/log/omlx.log (stdout/stderr)
- Server log: ~/.omlx/logs/server.log (structured application log)
Supports text LLMs, vision-language models (VLM), OCR models, embeddings, and rerankers on Apple Silicon.
Web UI at /admin for real-time monitoring, model management, chat, benchmark, and per-model settings. Supports English, Korean, Japanese, Chinese, French, Russian, Spanish, and Brazilian Portuguese. All CDN dependencies are vendored for fully offline operation.
Source builds can split one downloaded language model across unequal-memory Macs using MLX pipeline ranks over Ring or Thunderbolt RDMA/JACCL. The Cluster dashboard handles read-only peer discovery, strict SSH/runtime verification, byte-aware unequal shard planning, measured compute/link rebalancing, headroom-aware execution tuning, activation, and a live shard/performance map on both Macs. Interactive, balanced, and throughput profiles expose coalesced batching, prompt-cache affinity, rotating-KV limits, Ring connection tuning, and a capability-gated experimental token-only output path. See Distributed inference across Macs for setup, security boundaries, current limitations, and the physical-hardware validation checklist.
Run VLMs with the same continuous batching and tiered KV cache stack as text LLMs. Supports multi-image chat, base64/URL/file image inputs, and tool calling with vision context. OCR models (DeepSeek-OCR, DOTS-OCR, GLM-OCR) are auto-detected with optimized prompts.
Block-based KV cache management inspired by vLLM, with prefix sharing and Copy-on-Write. The cache operates across two tiers:
Handles concurrent requests through mlx-lm's BatchGenerator. Max concurrent requests is configurable via CLI or admin panel.
Context scaling support for running smaller context models with Claude Code. Scales reported token counts so that auto-compact triggers at the right timing, and SSE keep-alive prevents read timeouts during long prefill.
Load LLMs, VLMs, embedding models, and rerankers within the same server. Models are managed through a combination of automatic and manual controls:
Configure sampling parameters, chat template kwargs, TTL, model alias, model type override, and more per model directly from the admin panel. Changes apply immediately without server restart.
/v1/models returns the alias, and requests accept both the alias and directory name./v1/models then also lists <model>:<profile> (e.g. qwen3-8b:thinking), which serves on the same engine as the base model with the profile's settings overlaid per request — no extra memory, no reload. When the base model has an alias, the exposed ID is advertised as <alias>:<profile>; the directory-name form keeps working, just like for the base model.
Chat directly with any loaded model from the admin panel. Supports conversation history, model switching, dark mode, reasoning model output, and image upload for VLM/OCR models.
Search and download MLX models from HuggingFace directly in the admin dashboard. Browse model cards, check file sizes, and download with one click.
Set up OpenClaw, OpenCode, Codex, Hermes Agent, Copilot, and Pi directly from the admin dashboard with a single click. No manual config editing required.
One-click benchmarking from the admin panel. Measures prefill (PP) and text generation (TG) tokens per second, with partial prefix cache hit testing for realistic performance numbers.
Native Swift / SwiftUI menubar app (not Electron). Start, stop, and monitor the server without opening a terminal. Includes persistent serving stats (survives restarts), auto-restart on crash, and built-in auto-update.
Drop-in replacement for OpenAI and Anthropic APIs. Supports streaming usage stats (stream_options.include_usage), Anthropic adaptive thinking, and vision inputs (base64, URL).
| Endpoint | Description |
|----------|-------------|
| POST /v1/chat/completions | Chat completions (streaming) |
| POST /v1/completions | Text completions (streaming) |
| POST /v1/messages | Anthropic Messages API |
| POST /v1/embeddings | Text embeddings |
| POST /v1/rerank | Document reranking |
| GET /v1/models | List available models |
Supports all function calling formats available in mlx-lm, JSON schema validation, and MCP tool integration. Tool calling requires the model's chat template to support the tools parameter. The following model families are auto-detected via mlx-lm's built-in tool parsers:
| Model Family | Format |
|---|---|
| Llama, Qwen, DeepSeek, etc. | JSON <tool_call> |
| Qwen3.5 Series | XML <function=...> |
| Gemma | <start_function_call> |
| GLM (4.7, 5) | <arg_key>/<arg_value> XML |
| MiniMax | Namespaced <minimax:tool_call> |
| Mistral | [TOOL_CALLS] |
| Kimi K2 | <\|tool_calls_section_begin\|> |
| Longcat | <longcat_tool_call> |
Models not listed above may still work if their chat template accepts tools and their output uses a recognized <tool_call> XML format. For tool-enabled streaming, assistant text is emitted incrementally while known tool-call control markup is suppressed from visible content; structured tool calls are emitted after parsing the completed turn.
Point --model-dir at a directory containing MLX-format model subdirectories. Two-level organization folders (e.g., mlx-community/model-name/) are also supported.
~/models/
├── Step-3.5-Flash-8bit/
├── Qwen3-Coder-Next-8bit/
├── gpt-oss-120b-MXFP4-Q8/
├── Qwen3.5-122B-A10B-4bit/
└── bge-m3/
Models are auto-detected by type. You can also download models directly from the admin dashboard.
| Type | Models | |------|--------| | LLM | Any model supported by mlx-lm | | VLM | Qwen3.5 Series, GLM-4V, Pixtral, and other mlx-vlm models | | OCR | DeepSeek-OCR, DOTS-OCR, GLM-OCR | | Embedding | BERT, BGE-M3, ModernBERT | | Reranker | ModernBERT, XLM-RoBERTa |
```bash
omlx start omlx stop omlx restart
omlx serve --model-dir ~/models
omlx serve --model-dir ~/models --memory-guard safe
omlx serve --model-dir ~/models --memory-guard-gb 48
omlx serve --model-dir ~/models --paged-ssd-cache-dir ~/.omlx/cache
omlx serve --model-dir ~/models --hot-cache-max-size 20%
omlx serve --model-dir ~/models --max-concurrent-requests 16
omlx serve --model-dir ~/models --mcp-config mcp.json
omlx serve --model-dir ~/models --hf-endpoint https://hf-mirror.com
omlx serve --model-dir ~/models --api-key your-secret-key
```
All settings can also be configured from the web admin panel at /admin. Settings are persisted to ~/.omlx/settings.json, and CLI flags take precedence.
bash
git clone https://github.com/jundot/omlx.git
cd omlx
pip install -e ".[dev]"
pytest -m "not slow"
The native SwiftUI app lives at apps/omlx-mac/. Requires Xcode 26.5+ and Python 3.11+. venvstacks is declared as a dev dependency so pip install -e ".[dev]" (or uv sync --dev) brings the pinned version in. The build script also falls back to uvx venvstacks or pipx run venvstacks if you prefer a host-global tool runner.
```bash
apps/omlx-mac/Scripts/build.sh release
open apps/omlx-mac/build/Stage/oMLX.app
apps/omlx-mac/Scripts/build.sh release --rebuild-donor
apps/omlx-mac/Scripts/build.sh release --with-custom-kernel ```
First cold build takes 10–20 minutes (venvstacks Python layer assembly). Subsequent builds reuse the cached packaging/_export/ and finish in about 4 minutes. See packaging/README.md for the layer configuration and apps/omlx-mac/ for the Swift sources.
Contributions are welcome! See Contributing Guide for details.
Unable to fetch file structure.