An open API service indexing awesome lists of open source software.

awesome-vision-ai-stack

A curated, builder-first list of Vision Language Models (VLMs), local runtimes, document AI tools, UI agents, robotics vision stacks, datasets, benchmarks, and production resources.
https://github.com/edujbarrios/awesome-vision-ai-stack

Last synced: 4 days ago
JSON representation

  • Agents, Grounding, and Robotics

    • Grounding and perception

      • CLIP & OpenCLIP - Foundation models for zero-shot detection, retrieval, and visual embeddings; critical for flexible grounding and cross-modal search.
      • YOLO-World - Open-vocabulary detection combining YOLO efficiency with open-set capabilities.
      • Segment Anything - Foundational segmentation model.
      • OWL-ViT - Open-vocabulary detection.
      • Detic - Detection with image-level supervision and open-vocabulary flavor.
      • Florence-2 - Dense prediction and structured understanding for documents, captions, and visual grounding.
    • Robotics-oriented projects

      • OpenVLA - Vision-language-action direction for robotics.
      • LeRobot - Hugging Face robotics stack.
      • Open X-Embodiment - Cross-robot learning ecosystem.
      • RT-1 - Robotics transformer reference.
      • RT-2 - Vision-language-action robotics reference.
  • Applications and Demos

    • Useful open projects

      • ComfyUI - Important ecosystem for visual model workflows.
      • InvokeAI - Image generation stack, useful alongside multimodal systems.
      • Label Studio - Data labeling for multimodal training loops.
      • FiftyOne - Dataset inspection and vision evaluation.
      • Unstructured - Document preprocessing for multimodal pipelines.
      • Haystack - RAG framework adaptable to multimodal retrieval.
      • LlamaIndex - Retrieval and agents, with multimodal extensions.
      • LangChain - Orchestration stack with multimodal integrations.
  • Benchmarks and Evaluation

    • General VLM evaluation

      • MMMU - Massive multitask multimodal reasoning benchmark.
      • MMBench - Broad multimodal benchmark.
      • MM-Vet - Challenging evaluation for integrated multimodal capabilities.
      • POPE - Evaluates object hallucination.
      • HallusionBench - Hallucination-focused benchmark.
      • ScienceQA - Science reasoning benchmark with multimodal settings.
      • TextVQA - Text-centric visual QA.
      • VQAv2 - Classic visual question answering benchmark.
      • DocVQA - Document visual question answering benchmark.
  • Datasets

  • Document AI, OCR, and Chart Understanding

    • Models and tools

      • PaddleOCR - Strong OCR toolkit and a common baseline for document pipelines.
      • docTR - OCR for document text detection and recognition.
      • Nougat - OCR-style document understanding for scientific PDFs.
      • Donut - OCR-free document understanding model.
      • Pix2Struct - Document understanding without OCR; excellent for structured layouts, charts, and complex documents.
      • LayoutLM - Important family for document layout understanding.
      • DocLayout-YOLO - Modern layout detection and segmentation for complex documents.
      • MinerU - Open document parsing and PDF extraction tooling.
      • Marker - PDF-to-markdown/document extraction workflow.
      • Surya - OCR and layout toolkit.
      • ChartOCR - Chart understanding reference.
      • ChartQA - Dataset and benchmark for chart reasoning.
  • Foundation VLMs

    • General-purpose open models

      • LLaVA - One of the most influential open visual instruction-tuned models.
      • Qwen-VL - Earlier Qwen multimodal line with broad ecosystem support.
      • InternVL - Strong family of open large vision-language models.
      • CogVLM - Open visual language model family from THUDM.
      • MiniGPT-4 - Early and influential image-chat system.
      • InstructBLIP - Instruction-tuned extension of BLIP-style architectures.
      • BLIP-2 - Efficient VLM architecture connecting frozen vision and language models.
      • IDEFICS - Hugging Face open multimodal family.
      • DeepSeek-VL - Open multimodal reasoning models from DeepSeek.
      • Molmo - Open multimodal assistant from Ai2 with strong grounding focus.
      • Phi-3 Vision - Compact multimodal model useful for practical deployments.
      • Fuyu - Multimodal autoregressive model with a distinct design.
      • Gemma Vision - Google's open multimodal model; efficient vision-language understanding with strong performance on document and image reasoning.
      • Moondream - Lightweight open-source VLM optimized for efficiency and local deployment.
      • MedGEMMA - Medical-focused vision language model from Google for healthcare applications.
    • Research landmarks

      • Flamingo - Landmark few-shot visual language model.
      • Kosmos-1 - Early multimodal reasoning and grounding work.
      • PaLI - Scalable multilingual vision-language model.
      • PaLI-X - Larger multimodal extension of PaLI.
      • Kosmos-2 - Grounded multimodal large language model.
      • SEED-Bench ecosystem - Useful benchmark family around multimodal reasoning.
  • Learning Resources

  • Local Inference and Serving

    • APIs and interfaces

      • Open WebUI - Popular self-hosted chat UI for local models.
      • Lobe Chat - Polished interface for model backends.
      • LibreChat - Open chat UI with multi-backend support.
      • Flowise - Visual builder for LLM and multimodal pipelines.
    • Run in browser (zero-setup inference)

      • WebLLM - Run VLMs natively in the browser without backend server.
      • Transformers.js - Hugging Face models (vision and audio) in JavaScript; enables client-side multimodal inference.
      • ONNX Runtime Web - Cross-platform model inference in browser; standardized format.
      • TensorFlow.js - TensorFlow models in browser and Node.js; useful for vision tasks and edge deployment.
    • Run locally

      • Ollama - Local model runtime with growing multimodal support.
      • LM Studio - Desktop app for running local models with a friendly UI.
      • Jan - Open local AI runtime and desktop app.
      • llama.cpp - Core local inference stack; important for lightweight experimentation.
      • MLC-LLM - Compile and deploy models on edge and mobile devices.
      • OpenVINO - Useful for Intel-optimized deployments.
    • Serve at scale

      • vLLM - High-throughput inference engine increasingly relevant for multimodal serving.
      • SGLang - Fast serving and structured generation framework.
      • TensorRT-LLM - NVIDIA-optimized inference stack.
      • TGI - Hugging Face serving stack.
      • BentoML - Production model serving and packaging.
      • Ray Serve - Scalable service orchestration for model workloads.
  • Training and Fine-Tuning

    • Libraries

      • Transformers - The default ecosystem for many multimodal models.
      • TRL - Preference optimization and instruction-tuning workflows.
      • PEFT - Parameter-efficient fine-tuning.
      • Axolotl - Popular fine-tuning framework.
      • LLaMA-Factory - Large fine-tuning platform with multimodal support.
      • DeepSpeed - Distributed training and optimization.
      • PyTorch Lightning - Structured model training workflows.
      • OpenFlamingo - Open framework for Flamingo-style multimodal modeling.
      • LAVIS - Vision-language research and training toolkit.
  • UI Understanding and Computer Use

    • Projects and references

      • SeeAct - Visual web agent framework.
      • OpenHands - Software agent platform; relevant for multimodal and browser-use workflows.
      • Browser Use - Browser automation with model control.
      • Stagehand - Browser automation framework aimed at AI-native workflows.
      • OmniParser - Screen parsing for GUI grounding and action planning.
      • UI-TARS - UI-centric agent/model direction.
      • GroundingDINO - Key building block for screen and visual grounding.
      • SAM 2 - Segmentation backbone useful for visual agents and annotation loops.
  • Video and Long-Context Multimodality

    • Models and systems

      • Video-LLaVA - Video extension of LLaVA-style instruction tuning.
      • VideoChat2 - Video multimodal conversation direction.
      • LLaVA-NeXT-Video - Video-capable branch of the LLaVA family.
      • LongVU - Long video understanding direction.
      • VideoMAE - Self-supervised masked autoencoder for video; strong backbone for video understanding and retrieval.