{"id":49278109,"url":"https://github.com/usr-wwelsh/turbolab","last_synced_at":"2026-05-30T21:00:34.139Z","repository":{"id":352261045,"uuid":"1214490962","full_name":"usr-wwelsh/turbolab","owner":"usr-wwelsh","description":"Minimal local AI gateway with web UI, HuggingFace search, TurboQuant, and OpenAI API. Single binary, no cloud, no API keys.","archived":false,"fork":false,"pushed_at":"2026-05-26T16:36:23.000Z","size":393,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-05-26T18:23:05.438Z","etag":null,"topics":["ai-gateway","cpu-inference","gguf","go","huggingface","llama-cpp","local-ai","openai-api","safetensors","self-hosted","svelte"],"latest_commit_sha":null,"homepage":"","language":"Go","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/usr-wwelsh.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2026-04-18T16:44:30.000Z","updated_at":"2026-05-26T16:35:42.000Z","dependencies_parsed_at":"2026-05-14T19:04:37.544Z","dependency_job_id":null,"html_url":"https://github.com/usr-wwelsh/turbolab","commit_stats":null,"previous_names":["usr-wwelsh/turbolab"],"tags_count":17,"template":false,"template_full_name":null,"purl":"pkg:github/usr-wwelsh/turbolab","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/usr-wwelsh%2Fturbolab","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/usr-wwelsh%2Fturbolab/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/usr-wwelsh%2Fturbolab/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/usr-wwelsh%2Fturbolab/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/usr-wwelsh","download_url":"https://codeload.github.com/usr-wwelsh/turbolab/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/usr-wwelsh%2Fturbolab/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":33709269,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-05-30T02:00:06.278Z","response_time":92,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["ai-gateway","cpu-inference","gguf","go","huggingface","llama-cpp","local-ai","openai-api","safetensors","self-hosted","svelte"],"created_at":"2026-04-25T17:04:49.225Z","updated_at":"2026-05-30T21:00:34.124Z","avatar_url":"https://github.com/usr-wwelsh.png","language":"Go","funding_links":[],"categories":[],"sub_categories":[],"readme":"# turbolab\n[![turbolab status](https://vitals.wwel.sh/badge/proxmox/turbolab/status.svg)](https://github.com/usr-wwelsh/vitalSVG) [![turbolab cpu](https://vitals.wwel.sh/badge/proxmox/turbolab/cpu.svg)](https://github.com/usr-wwelsh/vitalSVG) [![turbolab ram](https://vitals.wwel.sh/badge/proxmox/turbolab/ram.svg)](https://github.com/usr-wwelsh/vitalSVG) [![turbolab uptime](https://vitals.wwel.sh/badge/proxmox/turbolab/uptime.svg)](https://github.com/usr-wwelsh/vitalSVG)\n[![turbolab cpu trend](https://vitals.wwel.sh/badge/proxmox/turbolab/sparkline.svg?metric=cpu)](https://github.com/usr-wwelsh/vitalSVG) [![turbolab ram trend](https://vitals.wwel.sh/badge/proxmox/turbolab/sparkline.svg?metric=ram)](https://github.com/usr-wwelsh/vitalSVG)\n\u003ctable\u003e\n  \u003ctr\u003e\n    \u003ctd\u003e\u003cimg src=\"demo1.webp\" width=\"400\"/\u003e\u003c/td\u003e\n    \u003ctd\u003e\u003cimg src=\"demo2.webp\" width=\"400\"/\u003e\u003c/td\u003e\n  \u003c/tr\u003e\n  \u003ctr\u003e\n    \u003ctd\u003e\u003cimg src=\"demo3.webp\" width=\"400\"/\u003e\u003c/td\u003e\n    \u003ctd\u003e\u003cimg src=\"demo4.webp\" width=\"400\"/\u003e\u003c/td\u003e\n  \u003c/tr\u003e\n\u003c/table\u003e\n\nSelf-hosted AI model server. Single binary, no cloud, no API keys. Run HuggingFace models locally with a web chat UI and an OpenAI-compatible API.\n\n## How it works\n\nturbolab manages two inference backends depending on the model format:\n\n- **turboquant** — handles HuggingFace SafeTensor models with configurable KV cache quantization (2/4/8-bit). Keeps memory footprint small without sacrificing too much quality.\n- **llama-server** (llama.cpp) — handles GGUF models. turbolab auto-selects the best quant variant for your available RAM: reads `/proc/meminfo`, filters out files that won't fit (with 1GB headroom), then ranks by quant type — Q4_0 first for raw CPU throughput on AVX2, Q4_K_M second for quality. Falls back to smallest file if everything is filtered.\n\nThread count propagates to OMP, MKL, and OpenBLAS simultaneously so you're not leaving cores on the table.\n\nIf the inference process crashes, turbolab restarts it automatically. Three consecutive fast crashes (under 5s uptime) and it gives up and logs the failure rather than looping forever.\n\n## Features\n\n- Web UI for browsing HuggingFace, loading models, and chatting\n- OpenAI-compatible `/v1/` API — works with any client that supports it\n- RAM-aware GGUF quant selection at load time\n- Dual backend: turboquant for SafeTensors, llama-server for GGUF\n- CPU-first by default, GPU opt-in via `--no-cpu-only`\n- Usage tracking per model (tokens, sessions) stored in SQLite\n- System monitor (CPU, RAM, disk) via `/api/status`\n- SSE log stream at `/api/events`\n- Self-update: `turbolab update`\n- Optional systemd service setup via `turbolab setup`\n\n### Memory\n\nturbolab ships a built-in memory store backed by SQLite, accessible from the web UI and via an [MCP](https://modelcontextprotocol.io) server at `/mcp`.\n\n- **Add memories** — plain text, tagged, with an optional source URL or file path. Drop a file (PDF, DOCX, HTML, RST, LaTeX) to auto-convert via pandoc.\n- **Search** — FTS5 full-text search with OR term matching and prefix expansion. Automatically upgrades to vector similarity search when a model is loaded.\n- **Semantic embeddings** — memories are embedded via the loaded model's `/v1/embeddings` endpoint and stored as float32 vectors. Hit \"embed ↺\" to index existing memories. Cosine similarity search finds conceptually related memories even when keywords don't overlap.\n- **Auto-relate** — on insert, similar memories (cosine ≥ 0.75) are automatically linked with a `similar` edge in the background.\n- **Relationships** — manually link any two memories with a typed edge (`uses`, `depends_on`, `related`, etc.)\n- **Traversal** — `get_related` walks connected memories via BFS to surface related context at configurable depth\n- **Graph view** — interactive force-directed canvas graph of all memories and their edges\n- **Auto-inject** — when enabled, relevant memories are prepended to every chat request as a system message. The chat UI shows a \"↑ N memories injected\" badge above each response; click to expand and see exactly what was pulled in.\n- **MCP server** — exposes `add_memory`, `search_memory`, `semantic_search_memory`, `get_related`, and `relate_memories` as MCP tools so any MCP-compatible AI client (Claude Code, Cursor, etc.) can read and write the same store\n\n## Install\n\nDownload the latest binary for your platform from [Releases](https://github.com/usr-wwelsh/turbolab/releases/latest):\n\n| Platform | Binary |\n|---|---|\n| Linux x86_64 | `turbolab_linux_amd64` |\n| Linux ARM64 | `turbolab_linux_arm64` |\n\n```bash\nchmod +x turbolab_linux_amd64\nsudo mv turbolab_linux_amd64 /usr/local/bin/turbolab\n```\n\n## Usage\n\n```bash\n# First-time setup — installs turboquant into a venv, downloads llama-server\n# Requires python3 on PATH\nturbolab setup\n\n# Start the server (default port 7860)\nturbolab serve\n\n# Open http://localhost:7860\n```\n\n```bash\nturbolab models search \u003cquery\u003e   # Search HuggingFace\nturbolab update                  # Self-update to latest release\nturbolab serve --port 8080       # Custom port\nturbolab serve --bits 8          # Higher precision (default 4)\nturbolab serve --no-cpu-only     # Enable GPU layers\n```\n\n## Requirements\n\n- Python 3.x on PATH (for turboquant backend — `turbolab setup` handles the rest)\n- llama-server auto-installed on linux/amd64; elsewhere install from [llama.cpp releases](https://github.com/ggerganov/llama.cpp/releases)\n\n## Use Cases\n\n- **Homelab AI gateway** — serve models to your local network\n- **Dev agent harnesses** — OpenAI-compatible API makes swapping models trivial\n- **Local alternative to cloud** — no API keys, no data leaving your machine\n\n**Not network-safe** — no authentication or security middleware. Designed for trusted networks (homelab/localhost) only.\n\n## Build from source\n\n```bash\ngit clone https://github.com/usr-wwelsh/turbolab\ncd turbolab\nmake build\n```\n\nRequires Go 1.23+ and Node 20+.\n\n## License\n\nMIT — [usr-wwelsh](https://github.com/usr-wwelsh)\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fusr-wwelsh%2Fturbolab","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fusr-wwelsh%2Fturbolab","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fusr-wwelsh%2Fturbolab/lists"}