awesome-vision-ai-stack
A curated, builder-first list of Vision Language Models (VLMs), local runtimes, document AI tools, UI agents, robotics vision stacks, datasets, benchmarks, and production resources.
https://github.com/edujbarrios/awesome-vision-ai-stack
Last synced: 4 days ago
JSON representation
-
Agents, Grounding, and Robotics
-
Grounding and perception
- CLIP & OpenCLIP - Foundation models for zero-shot detection, retrieval, and visual embeddings; critical for flexible grounding and cross-modal search.
- YOLO-World - Open-vocabulary detection combining YOLO efficiency with open-set capabilities.
- Segment Anything - Foundational segmentation model.
- OWL-ViT - Open-vocabulary detection.
- Detic - Detection with image-level supervision and open-vocabulary flavor.
- Florence-2 - Dense prediction and structured understanding for documents, captions, and visual grounding.
-
Robotics-oriented projects
- OpenVLA - Vision-language-action direction for robotics.
- LeRobot - Hugging Face robotics stack.
- Open X-Embodiment - Cross-robot learning ecosystem.
- RT-1 - Robotics transformer reference.
- RT-2 - Vision-language-action robotics reference.
-
-
Applications and Demos
-
Useful open projects
- ComfyUI - Important ecosystem for visual model workflows.
- InvokeAI - Image generation stack, useful alongside multimodal systems.
- Label Studio - Data labeling for multimodal training loops.
- FiftyOne - Dataset inspection and vision evaluation.
- Unstructured - Document preprocessing for multimodal pipelines.
- Haystack - RAG framework adaptable to multimodal retrieval.
- LlamaIndex - Retrieval and agents, with multimodal extensions.
- LangChain - Orchestration stack with multimodal integrations.
-
-
Benchmarks and Evaluation
-
General VLM evaluation
- MMMU - Massive multitask multimodal reasoning benchmark.
- MMBench - Broad multimodal benchmark.
- MM-Vet - Challenging evaluation for integrated multimodal capabilities.
- POPE - Evaluates object hallucination.
- HallusionBench - Hallucination-focused benchmark.
- ScienceQA - Science reasoning benchmark with multimodal settings.
- TextVQA - Text-centric visual QA.
- VQAv2 - Classic visual question answering benchmark.
- DocVQA - Document visual question answering benchmark.
-
-
Datasets
-
Document and chart data
- InfographicVQA
- PubLayNet - Document layout annotations.
- FUNSD - Form understanding benchmark.
-
Grounding and detection data
- RefCOCO / RefCOCO+ / RefCOCOg - Referring expression benchmarks.
- Objects365 - Detection dataset.
- Open Images - Large detection and localization dataset.
-
Image-text pretraining
- LAION-5B - Large-scale image-text resource.
- COCO Captions - Standard image captioning dataset.
- Visual Genome - Dense visual annotations and relationships.
- DataComp - Dataset curation benchmark ecosystem.
-
Instruction and conversational multimodal data
- LLaVA-Instruct-150K - Foundational multimodal instruction dataset.
- ShareGPT4V - Large multimodal instruction-style dataset.
-
-
Document AI, OCR, and Chart Understanding
-
Models and tools
- PaddleOCR - Strong OCR toolkit and a common baseline for document pipelines.
- docTR - OCR for document text detection and recognition.
- Nougat - OCR-style document understanding for scientific PDFs.
- Donut - OCR-free document understanding model.
- Pix2Struct - Document understanding without OCR; excellent for structured layouts, charts, and complex documents.
- LayoutLM - Important family for document layout understanding.
- DocLayout-YOLO - Modern layout detection and segmentation for complex documents.
- MinerU - Open document parsing and PDF extraction tooling.
- Marker - PDF-to-markdown/document extraction workflow.
- Surya - OCR and layout toolkit.
- ChartOCR - Chart understanding reference.
- ChartQA - Dataset and benchmark for chart reasoning.
-
-
Foundation VLMs
-
General-purpose open models
- LLaVA - One of the most influential open visual instruction-tuned models.
- Qwen-VL - Earlier Qwen multimodal line with broad ecosystem support.
- InternVL - Strong family of open large vision-language models.
- CogVLM - Open visual language model family from THUDM.
- MiniGPT-4 - Early and influential image-chat system.
- InstructBLIP - Instruction-tuned extension of BLIP-style architectures.
- BLIP-2 - Efficient VLM architecture connecting frozen vision and language models.
- IDEFICS - Hugging Face open multimodal family.
- DeepSeek-VL - Open multimodal reasoning models from DeepSeek.
- Molmo - Open multimodal assistant from Ai2 with strong grounding focus.
- Phi-3 Vision - Compact multimodal model useful for practical deployments.
- Fuyu - Multimodal autoregressive model with a distinct design.
- Gemma Vision - Google's open multimodal model; efficient vision-language understanding with strong performance on document and image reasoning.
- Moondream - Lightweight open-source VLM optimized for efficiency and local deployment.
- MedGEMMA - Medical-focused vision language model from Google for healthcare applications.
-
Research landmarks
- Flamingo - Landmark few-shot visual language model.
- Kosmos-1 - Early multimodal reasoning and grounding work.
- PaLI - Scalable multilingual vision-language model.
- PaLI-X - Larger multimodal extension of PaLI.
- Kosmos-2 - Grounded multimodal large language model.
- SEED-Bench ecosystem - Useful benchmark family around multimodal reasoning.
-
-
Learning Resources
-
Repositories and hubs
- Hugging Face Multimodal Tasks
- OpenCompass - Evaluation ecosystem.
-
-
Local Inference and Serving
-
APIs and interfaces
- Open WebUI - Popular self-hosted chat UI for local models.
- Lobe Chat - Polished interface for model backends.
- LibreChat - Open chat UI with multi-backend support.
- Flowise - Visual builder for LLM and multimodal pipelines.
-
Run in browser (zero-setup inference)
- WebLLM - Run VLMs natively in the browser without backend server.
- Transformers.js - Hugging Face models (vision and audio) in JavaScript; enables client-side multimodal inference.
- ONNX Runtime Web - Cross-platform model inference in browser; standardized format.
- TensorFlow.js - TensorFlow models in browser and Node.js; useful for vision tasks and edge deployment.
-
Run locally
- Ollama - Local model runtime with growing multimodal support.
- LM Studio - Desktop app for running local models with a friendly UI.
- Jan - Open local AI runtime and desktop app.
- llama.cpp - Core local inference stack; important for lightweight experimentation.
- MLC-LLM - Compile and deploy models on edge and mobile devices.
- OpenVINO - Useful for Intel-optimized deployments.
-
Serve at scale
- vLLM - High-throughput inference engine increasingly relevant for multimodal serving.
- SGLang - Fast serving and structured generation framework.
- TensorRT-LLM - NVIDIA-optimized inference stack.
- TGI - Hugging Face serving stack.
- BentoML - Production model serving and packaging.
- Ray Serve - Scalable service orchestration for model workloads.
-
-
Training and Fine-Tuning
-
Libraries
- Transformers - The default ecosystem for many multimodal models.
- TRL - Preference optimization and instruction-tuning workflows.
- PEFT - Parameter-efficient fine-tuning.
- Axolotl - Popular fine-tuning framework.
- LLaMA-Factory - Large fine-tuning platform with multimodal support.
- DeepSpeed - Distributed training and optimization.
- PyTorch Lightning - Structured model training workflows.
- OpenFlamingo - Open framework for Flamingo-style multimodal modeling.
- LAVIS - Vision-language research and training toolkit.
-
-
UI Understanding and Computer Use
-
Projects and references
- SeeAct - Visual web agent framework.
- OpenHands - Software agent platform; relevant for multimodal and browser-use workflows.
- Browser Use - Browser automation with model control.
- Stagehand - Browser automation framework aimed at AI-native workflows.
- OmniParser - Screen parsing for GUI grounding and action planning.
- UI-TARS - UI-centric agent/model direction.
- GroundingDINO - Key building block for screen and visual grounding.
- SAM 2 - Segmentation backbone useful for visual agents and annotation loops.
-
-
Video and Long-Context Multimodality
-
Models and systems
- Video-LLaVA - Video extension of LLaVA-style instruction tuning.
- VideoChat2 - Video multimodal conversation direction.
- LLaVA-NeXT-Video - Video-capable branch of the LLaVA family.
- LongVU - Long video understanding direction.
- VideoMAE - Self-supervised masked autoencoder for video; strong backbone for video understanding and retrieval.
-
Programming Languages
Categories
Foundation VLMs
21
Local Inference and Serving
20
Datasets
12
Document AI, OCR, and Chart Understanding
12
Agents, Grounding, and Robotics
11
Training and Fine-Tuning
9
Benchmarks and Evaluation
9
Applications and Demos
8
UI Understanding and Computer Use
8
Video and Long-Context Multimodality
5
Learning Resources
2
Sub Categories
General-purpose open models
15
Models and tools
12
Libraries
9
General VLM evaluation
9
Projects and references
8
Useful open projects
8
Run locally
6
Serve at scale
6
Grounding and perception
6
Research landmarks
6
Models and systems
5
Robotics-oriented projects
5
APIs and interfaces
4
Run in browser (zero-setup inference)
4
Image-text pretraining
4
Grounding and detection data
3
Document and chart data
3
Repositories and hubs
2
Instruction and conversational multimodal data
2
Keywords
llm
19
deep-learning
16
pytorch
13
python
10
machine-learning
10
chatgpt
10
ai
9
openai
7
gpt
7
rag
6
large-language-models
6
language-model
6
nlp
6
deepseek
5
transformer
5
artificial-intelligence
5
inference
5
ocr
5
llama
5
qwen
5
computer-vision
5
generative-ai
4
agent
4
langchain
4
transformers
4
agents
4
vision-language-model
4
multi-modal
3
llm-inference
3
llms
3
llama3
3
llm-serving
3
gemini
3
ollama
3
data-science
3
glm
3
llama2
3
stable-diffusion
3
mlops
3
instruction-tuning
3
natural-language-processing
3
image-classification
3
foundation-models
3
vlm
2
vision-language-transformer
2
speech-recognition
2
mcp
2
vision-language-pretraining
2
video-understanding
2
object-detection
2