{"id":50645433,"url":"https://github.com/somnusochi/vlm-autoyolo","last_synced_at":"2026-06-08T13:01:10.898Z","repository":{"id":361998907,"uuid":"1256747268","full_name":"Somnusochi/VLM-AutoYOLO","owner":"Somnusochi","description":"AI Auto Annotation \u0026 YOLO Training Pipeline, End-to-end object detection auto-labeling and YOLO training platform. VLM-powered annotation with NVIDIA LocateAnything-3B, manual refinement, one-click YOLO training, video keyframe extraction, and model validation. Supports image and video.","archived":false,"fork":false,"pushed_at":"2026-06-07T12:12:51.000Z","size":58843,"stargazers_count":77,"open_issues_count":0,"forks_count":7,"subscribers_count":0,"default_branch":"master","last_synced_at":"2026-06-07T12:23:12.691Z","etag":null,"topics":["auto-labeling","computer-vision","data-annotation","deep-learning","fastapi","locate-anything","machine-learning","nvidia","object-detection","pytorch","react","ultralytics","video-annotation","vlm","yolo","yolo-training"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"agpl-3.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/Somnusochi.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2026-06-02T03:55:53.000Z","updated_at":"2026-06-07T12:12:54.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/Somnusochi/VLM-AutoYOLO","commit_stats":null,"previous_names":["somnusochi/autolabeling","somnusochi/locateanything","somnusochi/vlm-autoyolo"],"tags_count":28,"template":false,"template_full_name":null,"purl":"pkg:github/Somnusochi/VLM-AutoYOLO","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Somnusochi%2FVLM-AutoYOLO","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Somnusochi%2FVLM-AutoYOLO/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Somnusochi%2FVLM-AutoYOLO/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Somnusochi%2FVLM-AutoYOLO/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/Somnusochi","download_url":"https://codeload.github.com/Somnusochi/VLM-AutoYOLO/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Somnusochi%2FVLM-AutoYOLO/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":34063159,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-06-08T02:00:07.615Z","response_time":111,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["auto-labeling","computer-vision","data-annotation","deep-learning","fastapi","locate-anything","machine-learning","nvidia","object-detection","pytorch","react","ultralytics","video-annotation","vlm","yolo","yolo-training"],"created_at":"2026-06-07T12:02:13.537Z","updated_at":"2026-06-08T13:01:10.822Z","avatar_url":"https://github.com/Somnusochi.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# VLM-AutoYOLO\n\n[简体中文](README_ZH.md) | English\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"https://img.shields.io/badge/License-AGPL%20v3-blue.svg\" alt=\"License\"\u003e\n  \u003cimg src=\"https://img.shields.io/badge/Python-3.12+-blue\" alt=\"Python\"\u003e\n  \u003cimg src=\"https://img.shields.io/badge/Node.js-22+-green\" alt=\"Node.js\"\u003e\n  \u003cimg src=\"https://img.shields.io/badge/Platform-macOS%20%7C%20Windows%20%7C%20Linux-lightgrey\" alt=\"Platform\"\u003e\n  \u003cimg src=\"https://img.shields.io/badge/GPU-MPS%20%7C%20CUDA-orange\" alt=\"GPU\"\u003e\n  \u003ca href=\"mailto:somnusochi@gmail.com\"\u003e\u003cimg src=\"https://img.shields.io/badge/Open_to_Work-🤝-brightgreen?style=flat\" alt=\"Open to Work\"\u003e\u003c/a\u003e\n  \u003cimg src=\"https://img.shields.io/github/stars/Somnusochi/VLM-AutoYOLO?style=social\" alt=\"Stars\"\u003e\n\u003c/p\u003e\n\n```\n🖼️ image/video → 🔍 VLM / SAM3 detection → 🎯 SAM2/SAM3 mask → ✏️ refine → 📦 export → 🚀 YOLO → ✅ model\n```\n\n**Images or videos in → YOLO model out**, with VLM auto-labeling (LocateAnything-3B), SAM2.1 / SAM3 mask refinement, and human-in-the-loop correction. Multi-format export, one-click YOLO training (detect \u0026 segment), video keyframe extraction, and model validation — all GPU-accelerated on macOS MPS and Windows/Linux CUDA.\n\n## Key Features\n- 🤖 **VLM auto-labeling**: Open-vocabulary object detection with LocateAnything-3B\n- 🎯 **SAM2 / SAM3 segmentation**: Bbox → pixel-precise mask with SAM 2.1 or SAM3 text-driven detection+segmentation in one pass, BBox/Mask toggle on canvas\n- 🎥 **Video annotation**: Intelligent keyframe extraction (scene / motion / interval), SSIM dedup\n- ✏️ **Manual refinement**: Canvas draw mode, NMS filtering, hide/show individual boxes\n- 📦 **Multi-format export**: YOLO, YOLO-Seg, COCO JSON, Pascal VOC XML, CreateML JSON\n- 🚀 **One-click training**: YOLOv8 / v11 / v26, detect \u0026 segment, real-time SSE progress\n- ✅ **Model validation**: Batch image / video testing, MJPEG live stream, SSE video inference\n- 💾 **Smart model management**: Lazy loading, idle auto-unload, MPS/CUDA strategy pattern cleanup\n- 🌐 **i18n**: English / 简体中文 / 日本語 · 🎨 **Theme**: Light / dark mode\n\n## Documentation\n\n📚 **[User Guide (English)](docs/guide/en/README.md)** | 📚 **[用户指南 (中文)](docs/guide/README.md)**\n\nComprehensive guides: quick start, annotation best practices, training parameter tuning, model deployment.\n\n## Screenshots\n\n| VLM Pre-annotation \u0026 Refinement | YOLO Training |\n|--------------------------------|---------------|\n| ![VLM pre-annotation and refinement](docs/1.png) | ![YOLO training](docs/2.png) |\n\n| Video Keyframe Entry | Model Validation |\n|---------------------|-----------------|\n| ![Video keyframe entry](docs/4.png) | ![Model validation](docs/3.png) |\n\n## Tech Stack\n\n| Layer | Technology |\n|-------|-----------|\n| Visual Grounding | NVIDIA LocateAnything-3B (Qwen2.5-3B + MoonViT) |\n| Segmentation | SAM 2.1 / SAM3 — Segment Anything Model 2 / 3 |\n| Object Detection | YOLOv8 / v11 / v26 — Detect \u0026 Segment (Ultralytics) |\n| Backend | Python FastAPI + PostgreSQL + SSE |\n| Frontend | React + TypeScript + Vite + Tailwind CSS + antd |\n| GPU Memory | Strategy Pattern (`gpu_memory.py`) — CUDA expandable segments / MPS synchronize + empty_cache |\n| State | Zustand + TanStack Query + ahooks |\n| i18n | i18next (English / 简体中文 / 日本語) |\n| Video | ffmpeg (scene / motion / interval extraction) |\n| Tooling | pnpm, ESLint, Prettier, Husky, commitlint, Playwright |\n\n## Quick Start\n\n### Docker Deployment\n\n\u003e **Requirements:** Linux or Windows (WSL2) with NVIDIA GPU + [NVIDIA Container Toolkit](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html).\n\u003e **macOS is not supported** — Docker on Mac has no GPU passthrough. Use [Manual Setup](#manual-setup) instead.\n\n**Quick start with pre-built images:**\n\n```bash\ncurl -O https://raw.githubusercontent.com/Somnusochi/VLM-AutoYOLO/master/docker-compose.yml\ndocker compose up -d\nopen http://localhost        # Frontend\nopen http://localhost:8000/docs  # API docs\n```\n\n**Build from source:**\n\n```bash\ngit clone https://github.com/Somnusochi/VLM-AutoYOLO.git\ncd VLM-AutoYOLO\ndocker compose up -d --build\n```\n\n**Services:**\n\n| Service | Port | Description |\n|---------|------|-------------|\n| Frontend | 80 | React web UI (Nginx) |\n| Backend | 8000 | FastAPI server |\n| SAM3 | 8002 | SAM3 standalone inference service |\n| Database | 5432 | PostgreSQL |\n\n**GPU Support** — add to `docker-compose.yml`:\n\n```yaml\nbackend:\n  deploy:\n    resources:\n      reservations:\n        devices:\n          - driver: nvidia\n            count: 1\n            capabilities: [gpu]\n  environment:\n    DEVICE: cuda\n```\n\n**Persistent Storage (Docker volumes):**\n- `pgdata` — Database · `model-cache` — VLM, SAM2 \u0026 SAM3 models · `uploads` — User images/videos · `training-data` — YOLO training outputs\n\n**Backup / Restore:**\n\n```bash\ndocker compose exec db pg_dump -U postgres autolabeling \u003e backup.sql\ncat backup.sql | docker compose exec -T db psql -U postgres autolabeling\n```\n\n### Manual Setup\n\n**Requirements:**\n\n| Resource | Minimum | Recommended |\n|----------|---------|-------------|\n| Python | 3.12+ | 3.12+ |\n| Node.js | 22+ | 22+ |\n| PostgreSQL | 16+ | 16+ |\n| ffmpeg | Any | — |\n| macOS | Apple Silicon 16GB | 24GB+ |\n| NVIDIA GPU | 12GB VRAM | 16GB+ |\n\n**Setup:**\n\n```bash\ngit clone https://github.com/Somnusochi/VLM-AutoYOLO.git\ncd VLM-AutoYOLO\n\n# Backend\ncd backend\npython3 -m venv .venv\nsource .venv/bin/activate  # Windows: .venv\\Scripts\\activate\npip install -r requirements.txt\ncd ..\n\n# Frontend\ncd frontend\npnpm install\ncd ..\n\n# Database (PostgreSQL recommended, but SQLite is supported out of the box)\n# If using PostgreSQL:\n# psql -d postgres -c \"CREATE DATABASE autolabeling;\"\n# cp backend/.env.example backend/.env\n# If you prefer a zero-setup SQLite database, just skip the two steps above. The system will auto-generate autolabeling.db\n\n# Migrations\ncd backend\nPYTHONPATH=. alembic upgrade head\n```\n\n**Pre-download models (optional):**\n\n```bash\nhuggingface-cli download nvidia/LocateAnything-3B --local-dir backend/model\n```\n\n**Launch:**\n\n```bash\n./start.sh   # macOS / Linux\nstart.bat    # Windows\n```\n\n| Service | URL |\n|---------|-----|\n| Frontend | http://localhost:5173 |\n| Backend | http://localhost:8000 |\n| API Docs | http://localhost:8000/docs |\n\n## Project Structure\n\nFull directory tree: **[docs/STRUCTURE.md](docs/STRUCTURE.md)**\n\n## Features\n\n### VLM Pre-annotation\n\nUpload images or video keyframes with open-vocabulary descriptions (e.g. `fire, smoke`, `red car`). LocateAnything-3B automatically detects and draws bounding boxes.\n\n- Open-vocabulary natural language descriptions\n- Auto-resize by long-side cap (VRAM-based: 800–1333px)\n- Batch upload folders or video keyframes, streaming results\n\n### SAM2 Segmentation\n\nEnable SAM2 (Segment Anything Model 2) to refine VLM bounding boxes into pixel-precise masks.\n\n- Check \"Enable SAM2 Segmentation\" before detection — runs automatically after VLM\n- SAM 2.1 model (base+), lazy-loaded with idle auto-unload\n- Score threshold slider for mask quality filtering\n- Masks rendered as semi-transparent overlays on canvas\n- BBox and Mask independently toggled on both main canvas and hover preview\n- Result table shows polygon vertex count per box\n\n### SAM3 Detection + Segmentation\n\nSwitch to SAM3 mode for text-driven detection and segmentation in a single pass — no VLM required.\n\n- Toggle between VLM+SAM2 and SAM3 via the model selector in the sidebar\n- Enter open-vocabulary text prompts (e.g. `cat`, `red car`) — SAM3 detects and segments all matching instances\n- **Confidence threshold** slider (0.0–1.0, default 0.5) controls detection sensitivity\n- **Mask threshold** slider (0.0–1.0, default 0.5) controls mask tightness\n- Enable/disable segmentation independently — bbox-only mode skips mask extraction for faster results\n- SAM3 runs as a standalone HTTP service on port 8002 with its own venv (`backend/sam3-venv/`)\n- **Requires `HF_TOKEN`** — set this env var before starting the backend. Two steps:\n  1. Open [huggingface.co/facebook/sam3](https://huggingface.co/facebook/sam3) in browser, click **\"Agree and access repository\"**\n  2. Create a **Read** token at [huggingface.co/settings/tokens](https://huggingface.co/settings/tokens) (no need for Fine-grained — a plain Read token inherits your account's permissions)\n  Model cached in `~/.cache/huggingface/hub/` after first download\n- Auto-starts on first use, idle auto-unload after 10 min\n- Real-time loading status via SSE (`starting` → `loading` → `loaded`)\n- Manual unload button to free GPU memory\n- Backend auto-switches: using SAM3 unloads VLM/SAM2, and vice versa\n- Detection records tagged with `model_type` (VLM / VLM+SAM2 / SAM3) for traceability\n\n### Video Annotation\n\nUpload a video, extract keyframes, select and batch-annotate.\n\n- **Three extraction modes**: scene change, motion detection (optical flow), fixed interval\n- **SSIM deduplication**: auto-removes near-duplicate frames\n- **Timeline preview**: horizontal scrollable strip, click for full-size view\n- **Multi-select**: check frames, select/cancel all, load to annotation queue\n\n### Manual Annotation\n\nCanvas-based annotation with View / Draw modes.\n\n- Category quick-fill from history\n- VLM pre-annotation baseline → delete mistakes → draw missing boxes\n- All / Best / NMS filter modes, settings saved per detection\n- Hide individual boxes while inspecting dense results\n- Per-frame re-detection\n\n### History Management\n\n- Thumbnail + category tag previews, tag-based multi-select filtering\n- Click to view details, re-detect with updated labels, frontend pagination\n- Single / batch export in **5 formats**: YOLO, YOLO-Seg, COCO JSON, Pascal VOC XML, CreateML JSON\n- Format selection via dropdown menu, one-click zip download\n\n### YOLO Training\n\n- **Series**: YOLOv8 / v11 / v26 (n/s/m/l/x)\n- **Task types**: Object Detection (Detect), Instance Segmentation (Segment)\n- Segmentation training auto-uses SAM2 polygon labels; falls back to bbox when unavailable\n- Tag filter + thumbnail preview for precise data selection\n- Dataset split presets (70/20/10, 80/20, 90/10, 60/20/20)\n- Real-time SSE progress: Epoch / Loss / mAP50\n- Auto ONNX export; download PT / ONNX / dataset zip\n\n### Model Validation\n\n- **Dual source**: trained models or externally uploaded `.pt` files\n- **Conf / IoU sliders** for real-time threshold tuning\n- **Batch image validation** with bounding boxes and confidence scores\n- **Video validation** (three modes):\n  - MJPEG live stream with interactive play/pause\n  - SSE prediction stream with per-frame JSON events\n  - Sync batch prediction — all frames at once\n- Temporary results; export predictions as YOLO `.txt` files\n\n### Model Management\n\n- **Lazy loading**: VLM, SAM2, and SAM3 load on first use, unload after idle (default 10 min)\n- **Idle watchdog**: all three models auto-unload after `MODEL_IDLE_TIMEOUT_SECONDS` of inactivity\n- **Unified SSE status**: `GET /api/v1/model/events` streams VLM, SAM2, SAM3 status in one connection\n- **Manual unload**: each model has its own unload button and API endpoint\n- **GPU memory**: Strategy Pattern (`gpu_memory.py`) — CUDA `expandable_segments` / MPS `synchronize`+`empty_cache`+`gc`\n\n## API Reference\n\nFull API documentation with request/response examples: **[docs/API.md](docs/API.md)**\n\n## Cross-Platform\n\n| Platform | Inference | Training |\n|----------|-----------|----------|\n| macOS (Apple Silicon) | MPS | MPS |\n| Linux / Windows (NVIDIA) | CUDA | CUDA |\n\nAuto-detection: CUDA → MPS. Override via `DEVICE` env. **CPU not supported.**\n\n## Inference Benchmarks\n\nTested locally on an **Apple MacBook Pro (M4 Pro, 24GB Unified Memory)** using Apple MPS hardware acceleration.\n\n| Image Resolution (Max Side) | Inference Latency | Actual Memory Footprint |\n| :--- | :--- | :--- |\n| **Thumbnail (256px)** | `~0.68s` | Stable around `~11.8GB` |\n| **High-Res (1024px)** | `~4.35s` | Stable around `~11.8GB` |\n\nFull detailed benchmarks across different hardware configurations: **[docs/BENCHMARKS.md](docs/BENCHMARKS.md)**\n\n## Highlights\n\n- **MPS / CUDA full-pipeline GPU acceleration** — VLM, SAM2, and YOLO training all GPU-accelerated\n- **Strategy Pattern GPU memory** — `gpu_memory.py` centralizes CUDA / MPS cleanup; `expandable_segments:True`\n- **SAM2 / SAM3 mask refinement** — SAM2 refines VLM bboxes; SAM3 does text-driven detection+segmentation in one pass\n- **5 export formats** — YOLO, YOLO-Seg, COCO, Pascal VOC, CreateML\n- **Detect \u0026 Segment training** — polygon labels auto-used when SAM2 masks are available\n- **Cross-platform** — macOS MPS, Windows / Linux CUDA, unified codebase\n- **Unified SSE model status** — single EventSource for VLM, SAM2, SAM3 states; no polling\n\n## Development\n\n```bash\n# Frontend\ncd frontend \u0026\u0026 pnpm install \u0026\u0026 pnpm run lint \u0026\u0026 pnpm run build\n\n# Backend\ncd backend \u0026\u0026 source .venv/bin/activate\nPYTHONPATH=. alembic upgrade head\npython -m compileall app alembic\n```\n\n## Stargazers\n\n[![Star History Chart](https://api.star-history.com/svg?repos=Somnusochi/VLM-AutoYOLO\u0026type=Date)](https://star-history.com/#Somnusochi/VLM-AutoYOLO\u0026Date)\n\n## License\n\nCode: [AGPL-3.0](LICENSE).\n\nThird-party dependencies:\n- LocateAnything-3B model — [NVIDIA License](https://huggingface.co/nvidia/LocateAnything-3B/blob/main/LICENSE) (non-commercial use only)\n- SAM3 model — [Facebook Research License](https://huggingface.co/facebook/sam3) (gated repository, requires HuggingFace access token)\n- Ultralytics YOLO — [AGPL-3.0](https://github.com/ultralytics/ultralytics/blob/main/LICENSE) (copyleft; training/deployment may trigger obligations)\n\n---\n\nIf this project helps you, please ⭐ [star it on GitHub](https://github.com/Somnusochi/VLM-AutoYOLO). I'm open to new opportunities — reach out: somnusochi@gmail.com\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fsomnusochi%2Fvlm-autoyolo","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fsomnusochi%2Fvlm-autoyolo","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fsomnusochi%2Fvlm-autoyolo/lists"}