{"id":44516111,"url":"https://github.com/jundot/omlx","last_synced_at":"2026-05-27T03:10:21.544Z","repository":{"id":338432502,"uuid":"1157171418","full_name":"jundot/omlx","owner":"jundot","description":"LLM inference server with continuous batching \u0026 SSD caching for Apple Silicon — managed from the macOS menu bar","archived":false,"fork":false,"pushed_at":"2026-02-27T11:51:13.000Z","size":8630,"stargazers_count":88,"open_issues_count":16,"forks_count":13,"subscribers_count":4,"default_branch":"main","last_synced_at":"2026-02-27T11:59:12.031Z","etag":null,"topics":["apple-silicon","inference-server","llm","macos","mlx","openai-api"],"latest_commit_sha":null,"homepage":"https://github.com/jundot/omlx","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/jundot.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":"docs/CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2026-02-13T14:13:27.000Z","updated_at":"2026-02-27T11:51:17.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/jundot/omlx","commit_stats":null,"previous_names":["jundot/omlx"],"tags_count":14,"template":false,"template_full_name":null,"purl":"pkg:github/jundot/omlx","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/jundot%2Fomlx","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/jundot%2Fomlx/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/jundot%2Fomlx/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/jundot%2Fomlx/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/jundot","download_url":"https://codeload.github.com/jundot/omlx/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/jundot%2Fomlx/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":30003813,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-03-02T12:19:43.414Z","status":"ssl_error","status_checked_at":"2026-03-02T12:19:02.215Z","response_time":60,"last_error":"SSL_connect returned=1 errno=0 peeraddr=140.82.121.6:443 state=error: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["apple-silicon","inference-server","llm","macos","mlx","openai-api"],"created_at":"2026-02-13T17:09:38.058Z","updated_at":"2026-03-02T13:15:02.597Z","avatar_url":"https://github.com/jundot.png","language":"Python","funding_links":["https://buymeacoffee.com/jundot"],"categories":["Model Serving \u0026 Inference","Inference engines","Repos","Python","AI开源项目","A01_文本生成_文本对话","\u003cimg src=\"./assets/cpu.svg\" width=\"16\" height=\"16\" style=\"vertical-align: middle;\"\u003e Backends","3. Inference Engines \u0026 Serving","LLM \u0026 Inference","*Ops for AI","🏆 القائمة الكاملة Top 200"],"sub_categories":["Model Serving Frameworks","AI 工具","大语言对话模型及数据","Model Serving \u0026 Inference"],"readme":"\u003cp align=\"center\"\u003e\n  \u003cpicture\u003e\n    \u003csource media=\"(prefers-color-scheme: dark)\" srcset=\"docs/images/icon-rounded-dark.svg\" width=\"140\"\u003e\n    \u003csource media=\"(prefers-color-scheme: light)\" srcset=\"docs/images/icon-rounded-light.svg\" width=\"140\"\u003e\n    \u003cimg alt=\"oMLX\" src=\"docs/images/icon-rounded-light.svg\" width=\"140\"\u003e\n  \u003c/picture\u003e\n\u003c/p\u003e\n\n\u003ch1 align=\"center\"\u003eoMLX\u003c/h1\u003e\n\u003cp align=\"center\"\u003e\u003cb\u003eLLM inference, optimized for your Mac\u003c/b\u003e\u003cbr\u003eContinuous batching and tiered KV caching, managed directly from your menu bar.\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"https://img.shields.io/badge/license-Apache%202.0-blue\" alt=\"License\"\u003e\n  \u003cimg src=\"https://img.shields.io/badge/python-3.10+-green\" alt=\"Python 3.10+\"\u003e\n  \u003cimg src=\"https://img.shields.io/badge/platform-Apple%20Silicon-black?logo=apple\" alt=\"Apple Silicon\"\u003e\n  \u003ca href=\"https://buymeacoffee.com/jundot\"\u003e\u003cimg src=\"https://img.shields.io/badge/Buy%20Me%20a%20Coffee-ffdd00?logo=buy-me-a-coffee\u0026logoColor=black\" alt=\"Buy Me a Coffee\"\u003e\u003c/a\u003e\n\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003ca href=\"#install\"\u003eInstall\u003c/a\u003e ·\n  \u003ca href=\"#quickstart\"\u003eQuickstart\u003c/a\u003e ·\n  \u003ca href=\"#features\"\u003eFeatures\u003c/a\u003e ·\n  \u003ca href=\"#models\"\u003eModels\u003c/a\u003e ·\n  \u003ca href=\"#cli-configuration\"\u003eCLI Configuration\u003c/a\u003e ·\n  \u003ca href=\"https://github.com/jundot/omlx\"\u003eGitHub\u003c/a\u003e\n\u003c/p\u003e\n\n---\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"docs/images/omlx_dashboard.png\" alt=\"oMLX Admin Dashboard\" width=\"800\"\u003e\n\u003c/p\u003e\n\n\u003e *Every LLM server I tried made me choose between convenience and control. I wanted to pin everyday models in memory, auto-swap heavier ones on demand, set context limits - and manage it all from a menu bar.*\n\u003e\n\u003e *oMLX persists KV cache across a hot in-memory tier and cold SSD tier - even when context changes mid-conversation, all past context stays cached and reusable across requests, making local LLMs practical for real coding work with tools like Claude Code. That's why I built it.*\n\n## Install\n\n### macOS App\n\nDownload the `.dmg` from [Releases](https://github.com/jundot/omlx/releases), drag to Applications, done. The app includes in-app auto-update, so future upgrades are just one click.\n\n### Homebrew\n\n```bash\nbrew tap jundot/omlx https://github.com/jundot/omlx\nbrew install omlx\n\n# Upgrade to the latest version\nbrew update \u0026\u0026 brew upgrade omlx\n\n# Run as a background service (auto-restarts on crash)\nbrew services start omlx\n```\n\n### From Source\n\n```bash\ngit clone https://github.com/jundot/omlx.git\ncd omlx\npip install -e .\n```\n\nRequires Python 3.10+ and Apple Silicon (M1/M2/M3/M4).\n\n## Quickstart\n\n### macOS App\n\nLaunch oMLX from your Applications folder. The Welcome screen guides you through three steps - model directory, server start, and first model download. That's it.\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"docs/images/Screenshot 2026-02-10 at 00.36.32.png\" alt=\"oMLX Welcome Screen\" width=\"360\"\u003e\n  \u003cimg src=\"docs/images/Screenshot 2026-02-10 at 00.34.30.png\" alt=\"oMLX Menubar\" width=\"240\"\u003e\n\u003c/p\u003e\n\n### CLI\n\n```bash\nomlx serve --model-dir ~/models\n```\n\nThe server discovers models from subdirectories automatically. Any OpenAI-compatible client can connect to `http://localhost:8000/v1`. A built-in chat UI is also available at `http://localhost:8000/admin/chat`.\n\n### Homebrew Service\n\nIf you installed via Homebrew, you can run oMLX as a managed background service:\n\n```bash\nbrew services start omlx    # Start (auto-restarts on crash)\nbrew services stop omlx     # Stop\nbrew services restart omlx  # Restart\nbrew services info omlx     # Check status\n```\n\nThe service runs `omlx serve` with zero-config defaults (`~/.omlx/models`, port 8000). To customize, either set environment variables (`OMLX_MODEL_DIR`, `OMLX_PORT`, etc.) or run `omlx serve --model-dir /your/path` once to persist settings to `~/.omlx/settings.json`.\n\nLogs are written to two locations:\n- **Service log**: `$(brew --prefix)/var/log/omlx.log` (stdout/stderr)\n- **Server log**: `~/.omlx/logs/server.log` (structured application log)\n\n## Features\n\noMLX is built on top of [vllm-mlx](https://github.com/waybarrios/vllm-mlx), extending it with tiered KV caching, multi-model serving, an admin dashboard, Claude Code optimization, and Anthropic API support. Currently supports text-based LLMs - VLM and OCR model support is planned for upcoming milestones.\n\n### Admin Dashboard\n\nWeb UI at `/admin` for real-time monitoring, model management, chat, benchmark, and per-model settings. All CDN dependencies are vendored for fully offline operation.\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"docs/images/Screenshot 2026-02-10 at 00.45.34.png\" alt=\"oMLX Admin Dashboard\" width=\"720\"\u003e\n\u003c/p\u003e\n\n### Tiered KV Cache (Hot + Cold)\n\nBlock-based KV cache management inspired by vLLM, with prefix sharing and Copy-on-Write. The cache operates across two tiers:\n\n- **Hot tier (RAM)**: Frequently accessed blocks stay in memory for fast access.\n- **Cold tier (SSD)**: When the hot cache fills up, blocks are offloaded to SSD in safetensors format. On the next request with a matching prefix, they're restored from disk instead of recomputed from scratch - even after a server restart.\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"docs/images/omlx_hot_cold_cache.png\" alt=\"oMLX Hot \u0026 Cold Cache\" width=\"720\"\u003e\n\u003c/p\u003e\n\n### Continuous Batching\n\nHandles concurrent requests through mlx-lm's BatchGenerator. Prefill and completion batch sizes are configurable.\n\n### Claude Code Optimization\n\nContext scaling support for running smaller context models with Claude Code. Scales reported token counts so that auto-compact triggers at the right timing, and SSE keep-alive prevents read timeouts during long prefill.\n\n### Multi-Model Serving\n\nLoad LLMs, embedding models, and rerankers within the same server. Models are managed through a combination of automatic and manual controls:\n\n- **LRU eviction**: Least-recently-used models are evicted automatically when memory runs low.\n- **Manual load/unload**: Interactive status badges in the admin panel let you load or unload models on demand.\n- **Model pinning**: Pin frequently used models to keep them always loaded.\n- **Per-model TTL**: Set an idle timeout per model to auto-unload after a period of inactivity.\n- **Process memory enforcement**: Total memory limit (default: system RAM - 8GB) prevents system-wide OOM.\n\n### Per-Model Settings\n\nConfigure sampling parameters, chat template kwargs, TTL, and more per model directly from the admin panel. Changes apply immediately without server restart.\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"docs/images/omlx_ChatTemplateKwargs.png\" alt=\"oMLX Chat Template Kwargs\" width=\"480\"\u003e\n\u003c/p\u003e\n\n### Built-in Chat\n\nChat directly with any loaded model from the admin panel. Supports conversation history, model switching, dark mode, and reasoning model output.\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"docs/images/Screenshot 2026-02-10 at 00.35.20.png\" alt=\"oMLX Chat\" width=\"720\"\u003e\n\u003c/p\u003e\n\n### Model Downloader\n\nSearch and download MLX models from HuggingFace directly in the admin dashboard. Browse model cards, check file sizes, and download with one click.\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"docs/images/downloader_omlx.png\" alt=\"oMLX Model Downloader\" width=\"720\"\u003e\n\u003c/p\u003e\n\n### Performance Benchmark\n\nOne-click benchmarking from the admin panel. Measures prefill (PP) and text generation (TG) tokens per second, with partial prefix cache hit testing for realistic performance numbers.\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"docs/images/benchmark_omlx.png\" alt=\"oMLX Benchmark Tool\" width=\"720\"\u003e\n\u003c/p\u003e\n\n### macOS Menubar App\n\nNative PyObjC menubar app (not Electron). Start, stop, and monitor the server without opening a terminal. Includes real-time serving stats, auto-restart on crash, and in-app auto-update.\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"docs/images/Screenshot 2026-02-10 at 00.51.54.png\" alt=\"oMLX Menubar Stats\" width=\"400\"\u003e\n\u003c/p\u003e\n\n### API Compatibility\n\nDrop-in replacement for OpenAI and Anthropic APIs. Supports streaming usage stats (`stream_options.include_usage`) and Anthropic adaptive thinking.\n\n| Endpoint | Description |\n|----------|-------------|\n| `POST /v1/chat/completions` | Chat completions (streaming) |\n| `POST /v1/completions` | Text completions (streaming) |\n| `POST /v1/messages` | Anthropic Messages API |\n| `POST /v1/embeddings` | Text embeddings |\n| `POST /v1/rerank` | Document reranking |\n| `GET /v1/models` | List available models |\n\n### Tool Calling \u0026 Structured Output\n\nSupports all function calling formats available in mlx-lm, JSON schema validation, and MCP tool integration. Tool calling requires the model's chat template to support the `tools` parameter. The following model families are auto-detected via mlx-lm's built-in tool parsers:\n\n| Model Family | Format |\n|---|---|\n| Llama, Qwen, DeepSeek, etc. | JSON `\u003ctool_call\u003e` |\n| Qwen3 Coder | XML `\u003cfunction=...\u003e` |\n| Gemma | `\u003cstart_function_call\u003e` |\n| GLM (4.7, 5) | `\u003carg_key\u003e/\u003carg_value\u003e` XML |\n| MiniMax | Namespaced `\u003cminimax:tool_call\u003e` |\n| Mistral | `[TOOL_CALLS]` |\n| Kimi K2 | `\u003c\\|tool_calls_section_begin\\|\u003e` |\n| Longcat | `\u003clongcat_tool_call\u003e` |\n\nModels not listed above may still work if their chat template accepts `tools` and their output uses a recognized `\u003ctool_call\u003e` XML format. Streaming requests with tool calls buffer all content and emit results at completion.\n\n## Models\n\nPoint `--model-dir` at a directory containing MLX-format model subdirectories. Two-level organization folders (e.g., `mlx-community/model-name/`) are also supported.\n\n```\n~/models/\n├── Step-3.5-Flash-8bit/\n├── Qwen3-Coder-Next-8bit/\n├── gpt-oss-120b-MXFP4-Q8/\n└── bge-m3/\n```\n\nModels are auto-detected by type. You can also download models directly from the admin dashboard.\n\n| Type | Models |\n|------|--------|\n| LLM | Any model supported by [mlx-lm](https://github.com/ml-explore/mlx-lm) |\n| Embedding | BERT, BGE-M3, ModernBERT |\n| Reranker | ModernBERT, XLM-RoBERTa |\n\n## CLI Configuration\n\n```bash\n# Memory limit for loaded models\nomlx serve --model-dir ~/models --max-model-memory 32GB\n\n# Process-level memory limit (default: auto = RAM - 8GB)\nomlx serve --model-dir ~/models --max-process-memory 80%\n\n# Enable SSD cache for KV blocks\nomlx serve --model-dir ~/models --paged-ssd-cache-dir ~/.omlx/cache\n\n# Set in-memory hot cache size\nomlx serve --model-dir ~/models --hot-cache-max-size 20%\n\n# Adjust batch sizes\nomlx serve --model-dir ~/models --prefill-batch-size 8 --completion-batch-size 32\n\n# With MCP tools\nomlx serve --model-dir ~/models --mcp-config mcp.json\n\n# API key authentication\nomlx serve --model-dir ~/models --api-key your-secret-key\n```\n\nAll settings can also be configured from the web admin panel at `/admin`. Settings are persisted to `~/.omlx/settings.json`, and CLI flags take precedence.\n\n\u003cdetails\u003e\n\u003csummary\u003eArchitecture\u003c/summary\u003e\n\n```\nFastAPI Server (OpenAI / Anthropic API)\n    │\n    ├── EnginePool (multi-model, LRU eviction, TTL, manual load/unload)\n    │   ├── BatchedEngine (LLMs, continuous batching)\n    │   ├── EmbeddingEngine\n    │   └── RerankerEngine\n    │\n    ├── ProcessMemoryEnforcer (total memory limit, TTL checks)\n    │\n    ├── Scheduler (FCFS, configurable batch sizes)\n    │   └── mlx-lm BatchGenerator\n    │\n    └── Cache Stack\n        ├── PagedCacheManager (GPU, block-based, CoW, prefix sharing)\n        ├── Hot Cache (in-memory tier, write-back)\n        └── PagedSSDCacheManager (SSD cold tier, safetensors format)\n```\n\n\u003c/details\u003e\n\n## Development\n\n### CLI Server\n\n```bash\ngit clone https://github.com/jundot/omlx.git\ncd omlx\npip install -e \".[dev]\"\npytest -m \"not slow\"\n```\n\n### macOS App\n\nRequires Python 3.11+ and [venvstacks](https://venvstacks.lmstudio.ai) (`pip install venvstacks`).\n\n```bash\ncd packaging\n\n# Full build (venvstacks + app bundle + DMG)\npython build.py\n\n# Skip venvstacks (code changes only)\npython build.py --skip-venv\n\n# DMG only\npython build.py --dmg-only\n```\n\nSee [packaging/README.md](packaging/README.md) for details on the app bundle structure and layer configuration.\n\n## Contributing\n\nContributions are welcome! See [Contributing Guide](docs/CONTRIBUTING.md) for details.\n\n- Bug fixes and improvements\n- Performance optimizations\n- Documentation improvements\n\n## License\n\n[Apache 2.0](LICENSE)\n\n## Acknowledgments\n\n- [MLX](https://github.com/ml-explore/mlx) and [mlx-lm](https://github.com/ml-explore/mlx-lm) by Apple\n- [vllm-mlx](https://github.com/waybarrios/vllm-mlx) - oMLX originated as a fork of vllm-mlx v0.1.0, since re-architected with multi-model serving, paged SSD caching, an admin panel, and a standalone macOS menu bar app\n- [venvstacks](https://venvstacks.lmstudio.ai) - Portable Python environment layering for the macOS app bundle\n- [mlx-embeddings](https://github.com/Blaizzy/mlx-embeddings) - Embedding model support for Apple Silicon\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fjundot%2Fomlx","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fjundot%2Fomlx","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fjundot%2Fomlx/lists"}