{"id":50329487,"url":"https://github.com/m4stanuj/mast-llm-router","last_synced_at":"2026-05-29T09:00:28.244Z","repository":{"id":361099884,"uuid":"1253075553","full_name":"m4stanuj/mast-llm-router","owner":"m4stanuj","description":"Task-aware MCP LLM fallback router: 13 provider routes, 10 chains, semantic cache, auto fallback, $0/month.","archived":false,"fork":false,"pushed_at":"2026-05-29T07:13:01.000Z","size":339,"stargazers_count":1,"open_issues_count":0,"forks_count":0,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-05-29T08:25:12.548Z","etag":null,"topics":["ai-agents","claude-code","codex","cursor","fallback-router","free-tier","gemini","groq","llm-router","local-first","m4st","mcp","openrouter","python","windsurf"],"latest_commit_sha":null,"homepage":"https://github.com/m4stanuj/mast-llm-router","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/m4stanuj.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":"AGENTS.md","dco":null,"cla":null}},"created_at":"2026-05-29T06:12:09.000Z","updated_at":"2026-05-29T07:12:19.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/m4stanuj/mast-llm-router","commit_stats":null,"previous_names":["m4stanuj/mast-llm-router"],"tags_count":2,"template":false,"template_full_name":null,"purl":"pkg:github/m4stanuj/mast-llm-router","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/m4stanuj%2Fmast-llm-router","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/m4stanuj%2Fmast-llm-router/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/m4stanuj%2Fmast-llm-router/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/m4stanuj%2Fmast-llm-router/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/m4stanuj","download_url":"https://codeload.github.com/m4stanuj/mast-llm-router/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/m4stanuj%2Fmast-llm-router/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":33644313,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-05-29T02:00:06.066Z","response_time":107,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["ai-agents","claude-code","codex","cursor","fallback-router","free-tier","gemini","groq","llm-router","local-first","m4st","mcp","openrouter","python","windsurf"],"created_at":"2026-05-29T09:00:25.885Z","updated_at":"2026-05-29T09:00:28.234Z","avatar_url":"https://github.com/m4stanuj.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# 🚀 MAST LLM Router — Intelligent LLM Request Distribution Engine\n\n[![CI](https://github.com/m4stanuj/mast-llm-router/actions/workflows/ci.yml/badge.svg)](https://github.com/m4stanuj/mast-llm-router/actions)\n[![Python](https://img.shields.io/badge/python-3.10%2B-blue)](https://www.python.org/)\n[![MCP](https://img.shields.io/badge/MCP-compatible-green)](https://modelcontextprotocol.io)\n[![License](https://img.shields.io/badge/license-MIT-brightgreen)](LICENSE)\n[![Cost](https://img.shields.io/badge/monthly%20cost-%240-success)](/#)\n[![Providers](https://img.shields.io/badge/providers-13-orange)](/#)\n[![PRESENTATION](https://img.shields.io/badge/view-Presentation-blueviolet)](PRESENTATION.md)\n[![SOCIAL](https://img.shields.io/badge/social-kit-ff69b4)](SOCIAL.md)\n[![Download](https://img.shields.io/badge/download-zip-success)](mast-llm-router.zip)\n[![Stars](https://img.shields.io/github/stars/m4stanuj/mast-llm-router?style=social)](https://github.com/m4stanuj/mast-llm-router)\n\n\u003e **🏆 Task-aware LLM fallback router — 13 provider routes · 10 chains · 6 fallbacks · $0/month**  \n\u003e Works with Claude Code, Cursor, Windsurf, Continue.dev, Codex CLI, and any MCP-compatible client.\n\n---\n\n```ascii\n╔══════════════════════════════════════════════════════════════╗\n║                                                              ║\n║   Every AI pipeline breaks when a provider hits a rate       ║\n║   limit. This one doesn't.                                   ║\n║                                                              ║\n║   ┌─────────┐   ┌──────────┐   ┌──────────┐   ┌──────────┐  ║\n║   │  Task   │──▶│  Chain   │──▶│ Fallback │──▶│ Response │  ║\n║   │ Detect  │   │ Select   │   │  Loop x6 │   │          │  ║\n║   └─────────┘   └──────────┘   └──────────┘   └──────────┘  ║\n║                                                              ║\n║   🔄 One fails → Next takes over → resilient by design      ║\n║   💰 13 provider routes · free-tier APIs · $0/month         ║\n║   🧠 Semantic caching · Auto key detection · MCP native     ║\n║                                                              ║\n╚══════════════════════════════════════════════════════════════╝\n```\n\n---\n\n## 📊 At a Glance\n\n| Metric | Value |\n|--------|-------|\n| **Provider Routes** | 13 (Groq, Cerebras, Gemini, DeepSeek, OpenRouter, SambaNova, Together, NVIDIA NIM, Mistral, xAI/Grok, HuggingFace, Kimi K2, Nemotron) |\n| **Task Chains** | 10 (speed, reason, code, vision, research, write, agent, pentest, hinglish, vision_reason) |\n| **Fallback Depth** | 6 models per chain — auto-failover on 429/503/empty response |\n| **Reliability Strategy** | Best-first chain routing with automatic fallback across provider routes |\n| **Cache Strategy** | Fuzzy semantic matching at 0.82 threshold with a 500-entry LRU cache |\n| **Monthly Cost** | **$0.00** (100% free-tier APIs) |\n| **Protocol** | MCP (Model Context Protocol) — stdio + HTTP |\n| **Clients** | Claude Code, Cursor, Windsurf, Continue.dev, Codex CLI, Antigravity |\n\n---\n\n## 🎯 What is this?\n\n`mast-llm-router` is a **task-aware intelligent fallback router** for LLM requests that:\n\n- **Detects** what kind of task you're doing from your prompt keywords\n- **Routes** to the optimal model chain for that specific task\n- **Auto-falls** to the next provider if one hits rate limits, errors out, or returns garbage\n- **Caches** semantically similar prompts to eliminate redundant API calls\n- **Detects** API keys automatically from their prefix — just paste and go\n- **Costs** exactly nothing — runs entirely on free-tier quotas\n\nBuilt as part of [M4ST](https://github.com/m4stanuj) — a personal AI OS running on an RTX 2060 Super in Bareilly, India.\n\n---\n\n## M4ST Ecosystem\n\n| Repo | Role |\n|------|------|\n| [MAST](https://github.com/m4stanuj/MAST) | Flagship AI operator stack |\n| [mast-llm-router](https://github.com/m4stanuj/mast-llm-router) | This repo: task-aware LLM fallback router |\n| [semantic-cache-engine](https://github.com/m4stanuj/semantic-cache-engine) | Standalone semantic cache module |\n| [openwork](https://github.com/m4stanuj/openwork) | Universal MCP workspace/config layer |\n| [m4stclaw-legacy-archive](https://github.com/m4stanuj/m4stclaw-legacy-archive) | Historical archive and lineage |\n\n---\n\n## 🎬 Quick Demo\n\n```\nUser: \"Write a Python script to scrape Hacker News\"\nRouter: Detected task → code\nChain:  kimi-k2 → qwen3-coder → mimo-pro → nvidia-deepseek → deepseek → sambanova\nResult: Response from kimi-k2 in 1.2s (cache miss)\n\nUser: \"yeh kya hai samjhao\"\nRouter: Detected task → hinglish\nChain:  sarvam → gemini-flash → groq-llama → cerebras → openrouter → mistral\nResult: Response in Hindi-English mix\n\nUser: \"Explain quantum computing in simple terms\"\nRouter: Detected task → reason\nChain:  deepseek-r1 → nemotron → gemini-pro → openrouter → together → mistral\nResult: Response from deepseek-r1 in 3.4s (cached from similar query)\n```\n\n---\n\n## 🧠 Algorithm: How It Works\n\n### Step 1: Task Detection\n```\nInput: \"Write a Python web scraper\"\n           │\n           ▼\n    ┌───────────────┐\n    │ Keyword Scan  │\n    │               │\n    │ \"Python\"  → 📝 code\n    │ \"scraper\" → 📝 code\n    │ \"write\"   → 📝 code\n    └───────┬───────┘\n            │\n    ┌───────▼───────┐\n    │ Chain: CODE   │\n    │ Confidence 94%│\n    └───────────────┘\n```\n\n### Step 2: Fallback Loop\n```\nChain: CODE\n        │\n    ┌───▼────────────┐\n    │ Provider 1     │── 429 Rate Limited ──┐\n    │ Kimi K2        │                      │\n    └────────────────┘                      │\n                                            ▼\n    ┌────────────────┐             ┌────────────────┐\n    │ Provider 2     │── 503 Error ─▶   Fallback     │\n    │ Qwen3 Coder    │── ──────────▶   Loop auto-    │\n    └────────────────┘               selects next    │\n                                            │\n    ┌────────────────┐                      │\n    │ Provider 3     │◀─────────────────────┘\n    │ Mimo Pro       │── ✅ Success 1.2s\n    └────────────────┘\n            │\n    ┌───────▼───────┐\n    │ Response sent │\n    │ to MCP Client │\n    └───────────────┘\n```\n\n### Step 3: Semantic Cache\n```\nPrompt ──▶ Embedding ──▶ Fuzzy Match (\u003e0.82) ──▶ Cache Hit? ──▶ Return cached\n                                                      │\n                                                   Miss ──▶ Call API ──▶ Store\n```\n\n---\n\n## ✨ Feature Highlights\n\n- Task-aware chain selection for code, reasoning, research, writing, agents, vision, pentest, and Hinglish flows\n- Automatic fallback when a provider rate-limits, errors, or returns an empty response\n- SMART_KEY detection so mixed API keys can be pasted without manual provider mapping\n- Semantic cache for repeated or similar prompts, tuned with a 0.82 fuzzy-match threshold\n\n---\n\n## Feature Matrix\n\n| Feature | Detail |\n|---|---|\n| **13 provider routes** | Groq, Cerebras, Gemini, OpenRouter, SambaNova, DeepSeek, Together, NVIDIA NIM, Mistral, xAI/Grok, HuggingFace, Kimi K2, Nemotron |\n| **10 task chains** | speed, reason, code, vision, research, write, agent, pentest, hinglish, vision_reason |\n| **6 models per chain** | Best-first, auto-falls to next on failure |\n| **SMART_KEY detection** | Paste any API key — provider auto-detected by prefix |\n| **Semantic cache** | Fuzzy match at 0.82 threshold, 500 entry LRU |\n| **Thread-safe cooldowns** | Per-key 429/auth cooldown, not per-provider |\n| **Both transports** | stdio (local) + HTTP (remote) |\n| **$0/month** | 100% free-tier APIs |\n\n---\n\n## Task Chains\n\n```\nspeed         →  groq → cerebras → gemini-flash → openrouter → sambanova → deepseek\nreason        →  deepseek-r1 → nemotron → gemini-pro → openrouter → together → mistral\ncode          →  kimi-k2 → qwen3-coder → mimo-pro → nvidia → deepseek → sambanova\nvision        →  gemini-vision → openrouter-vision → together-vision → ...\nresearch      →  perplexity → gemini-pro → deepseek-r1 → openrouter → ...\nwrite         →  gemini-pro → mistral → together → openrouter → groq → cerebras\nagent         →  deepseek-r1 → gemini-pro → openrouter → together → groq → ...\npentest       →  nvidia-deepseek → nemotron → deepseek-r1 → glm → mistral → ...\nhinglish      →  sarvam → gemini-flash → groq → cerebras → openrouter → mistral\nvision_reason →  gemini-vision → openrouter-vision → together-vision → ...\n```\n\n---\n\n## MCP Tools Exposed\n\n| Tool | Description |\n|---|---|\n| `llm_chat` | Single-turn prompt → best model |\n| `llm_chat_multi_turn` | Full conversation history support |\n| `llm_detect_task` | Preview which chain will handle your prompt |\n| `llm_router_status` | Provider health, key counts, cooldowns |\n| `llm_list_providers` | All providers + chains in JSON |\n| `llm_cache_control` | Cache stats or clear |\n\n---\n\n## Installation\n\n### 1. Clone\n\n```bash\ngit clone https://github.com/m4stanuj/mast-llm-router.git\ncd mast-llm-router\n```\n\n### 2. Install dependencies\n\n```bash\npip install -r requirements.txt\n```\n\n### 3. Configure keys\n\n```bash\ncp .env.example .env\n# Edit .env — paste your free-tier API keys\n```\n\n\u003e **SMART_KEY tip:** Just paste any key into `SMART_KEY_1`, `SMART_KEY_2`, etc.  \n\u003e The router detects the provider automatically from the key prefix.\n\n### 4. Test it\n\n```bash\npython src/server.py --help\n```\n\n---\n\n## Client Setup\n\n### Claude Code\n\nAdd to `~/.claude/claude_desktop_config.json` (or via `claude mcp add`):\n\n```json\n{\n  \"mcpServers\": {\n    \"mast-router\": {\n      \"command\": \"python\",\n      \"args\": [\"/absolute/path/to/mast-llm-router/src/server.py\"]\n    }\n  }\n}\n```\n\nOr one-liner:\n```bash\nclaude mcp add mast-router python /absolute/path/to/mast-llm-router/src/server.py\n```\n\n### Cursor / Windsurf\n\nSettings → MCP → Add Server:\n\n```json\n{\n  \"name\": \"mast-router\",\n  \"type\": \"stdio\",\n  \"command\": \"python\",\n  \"args\": [\"/absolute/path/to/mast-llm-router/src/server.py\"]\n}\n```\n\n### Continue.dev\n\nIn `.continue/config.json`:\n\n```json\n{\n  \"mcpServers\": [\n    {\n      \"name\": \"mast-router\",\n      \"command\": \"python\",\n      \"args\": [\"/absolute/path/to/mast-llm-router/src/server.py\"]\n    }\n  ]\n}\n```\n\n### Codex CLI\n\n```bash\ncodex --mcp-server \"python /absolute/path/to/mast-llm-router/src/server.py\"\n```\n\n### HTTP Mode (Antigravity, Magnus, remote clients)\n\n```bash\npython src/server.py --http --port 8000\n```\n\nThen point your client to: `http://localhost:8000/mcp`\n\n---\n\n## Environment Variables\n\n| Variable | Description |\n|---|---|\n| `SMART_KEY_1` … `SMART_KEY_30` | Auto-detected keys (recommended) |\n| `GROQ_API_KEY` | Groq (+ `_1` through `_20` for rotation) |\n| `CEREBRAS_API_KEY` | Cerebras |\n| `GEMINI_API_KEY` | Google Gemini |\n| `OPENROUTER_API_KEY` | OpenRouter |\n| `NVIDIA_API_KEY` | NVIDIA NIM |\n| `SAMBANOVA_API_KEY` | SambaNova |\n| `DEEPSEEK_API_KEY` | DeepSeek |\n| `TOGETHER_API_KEY` | Together AI |\n| `MISTRAL_API_KEY` | Mistral |\n| `GROKAI_API_KEY` | xAI / Grok |\n| `HUGGINGFACE_API_KEY` | HuggingFace |\n\n---\n\n## How Key Detection Works\n\n```\ngsk_...      → Groq\ncsk-...      → Cerebras\nAIza...      → Gemini\nsk-or-...    → OpenRouter\nnvapi-...    → NVIDIA NIM\nsk-...       → DeepSeek / Together / Mistral (length-based split)\nxai-...      → Grok\nhf_...       → HuggingFace\n```\n\n---\n\n## Architecture\n\n```\n┌─────────────────────────────────────────────────────┐\n│                  MCP Client                          │\n│  (Claude Code / Cursor / Codex / Antigravity / …)   │\n└────────────────────┬────────────────────────────────┘\n                     │  stdio / HTTP\n┌────────────────────▼────────────────────────────────┐\n│              server.py  (FastMCP)                    │\n│   llm_chat │ multi_turn │ detect_task │ status …     │\n└────────────────────┬────────────────────────────────┘\n                     │\n┌────────────────────▼────────────────────────────────┐\n│           llm_fallback.py  (Core Router)             │\n│                                                      │\n│  ┌─────────────┐  ┌──────────┐  ┌─────────────────┐ │\n│  │ Task Detect │  │  Cache   │  │  Key Manager    │ │\n│  │  (keyword)  │  │  (fuzzy) │  │  (cooldown)     │ │\n│  └──────┬──────┘  └──────────┘  └─────────────────┘ │\n│         │                                            │\n│  ┌──────▼──────────────────────────────────────┐    │\n│  │            Task Chain Selector               │    │\n│  │  speed/reason/code/vision/pentest/hinglish…  │    │\n│  └──────┬──────────────────────────────────────┘    │\n│         │                                            │\n│  ┌──────▼──────────────────────────────────────┐    │\n│  │         Fallback Loop (6 models)             │    │\n│  │  Provider 1 → fail → Provider 2 → …         │    │\n│  └─────────────────────────────────────────────┘    │\n└─────────────────────────────────────────────────────┘\n```\n\n---\n\n## Free Tier Limits (as of 2026)\n\n| Provider | Free RPM | Free TPD |\n|---|---|---|\n| Groq | 30 | 14,400 |\n| Cerebras | 30 | ~1M |\n| Gemini Flash | 15 | 1M |\n| NVIDIA NIM | 40 | — |\n| SambaNova | 10 | — |\n| OpenRouter (free models) | varies | varies |\n| DeepSeek | 50 | — |\n\n---\n\n## Project Structure\n\n```\nmast-llm-router/\n├── src/\n│   ├── server.py          # MCP server (FastMCP)\n│   └── llm_fallback.py    # Core router logic\n├── config/\n│   ├── claude_code.json   # Claude Code config\n│   └── cursor_windsurf.json\n├── .env.example           # Key template\n├── .gitignore\n├── requirements.txt\n└── README.md\n```\n\n---\n\n## Part of M4ST Ecosystem\n\n```\nM4ST OS\n├── llm_fallback.py     ← this repo\n├── mcp_servers/        86 MCP tools\n├── OpenWork            MCP-based AI workspace\n├── CAI                 Pentest agent layer\n└── voice / memory / browser automation\n```\n\n---\n\n## 📂 Project Resources\n\n| Resource | Description |\n|----------|-------------|\n| [📖 PRESENTATION.md](./PRESENTATION.md) | Full slide deck — algorithm walkthrough, benchmarks, use cases |\n| [📱 SOCIAL.md](./SOCIAL.md) | Social media kit — tweets, LinkedIn posts, hashtags, captions |\n| [🎬 DEMO_STORYBOARD.md](./DEMO_STORYBOARD.md) | GIF/video storyboard — task detection, fallback, cache hit |\n| [🤖 AGENTS.md](./AGENTS.md) | Guide for AI agents using this MCP server |\n| [📋 CHANGELOG.md](./CHANGELOG.md) | Version history and roadmap |\n| [📦 mast-llm-router.zip](./mast-llm-router.zip) | Downloadable ZIP archive |\n| [🔗 GitHub Release](https://github.com/m4stanuj/mast-llm-router/releases) | Latest release with assets |\n\n## 🏆 Why MAST LLM Router?\n\n```\n✅ 13 provider routes → Redundancy across free-tier LLM APIs\n✅ 10 task chains → Optimal model for every use case\n✅ 6 fallbacks → Best-first recovery when providers fail\n✅ $0/month → Free tiers only\n✅ SMART_KEY → Paste any key, auto-detected\n✅ Semantic cache → Repeated/similar prompts can return instantly\n✅ MCP native → Works with every major AI coding tool\n✅ Open source → MIT license, fork and build\n```\n\n## ⭐ Star History\n\n[![Star History Chart](https://api.star-history.com/svg?repos=m4stanuj/mast-llm-router\u0026type=Date)](https://star-history.com/#m4stanuj/mast-llm-router\u0026Date)\n\n## 📢 Share\n\n```markdown\n**Twitter/X:**\n🧵 I built a $0/month LLM router with 13 provider routes and auto-fallback.\n6 models per chain. Semantic cache. MCP native.\ngithub.com/m4stanuj/mast-llm-router\n#LLM #AI #OpenSource #MCP #Python\n\n**LinkedIn:**\n🏗️ MAST LLM Router — task-aware fallback router for 13 LLM provider routes.\n100% free-tier. Zero config. Full code on GitHub.\nhttps://github.com/m4stanuj/mast-llm-router\n```\n\n## License\n\nMIT — use it, fork it, build on it.\n\n---\n\n*Built by [@m4stanuj](https://github.com/m4stanuj) | [LinkedIn](https://linkedin.com/in/mast-anuj) | RTX 2060 Super | Bareilly, India*  \n*Zero VC money. Zero monthly cost. Full control.*\n\n## 🔖 Hashtags\n\n```\n#LLM #AI #OpenSource #MCP #Python #MachineLearning #DeveloperTools\n#AIAgents #LLMRouter #FreeAPI #ArtificialIntelligence #PythonDev\n#ModelContextProtocol #LLMFallback #MultiProvider #AIIndex\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fm4stanuj%2Fmast-llm-router","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fm4stanuj%2Fmast-llm-router","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fm4stanuj%2Fmast-llm-router/lists"}