An open API service indexing awesome lists of open source software.

Projects in Awesome Lists tagged with mllm

A curated list of projects in awesome lists tagged with mllm .

https://github.com/microsoft/unilm

Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities

beit beit-3 bitnet deepnet document-ai foundation-models kosmos kosmos-1 layoutlm layoutxlm llm minilm mllm multimodal nlp pre-trained-model textdiffuser trocr unilm xlm-e

Last synced: 13 May 2025

https://github.com/robbyant-research/MagicQuill

[CVPR'25] Official Implementations for Paper - MagicQuill: An Intelligent Interactive Image Editing System

aigc gradio image-editing mllm

Last synced: 23 Sep 2026

https://github.com/ant-research/MagicQuill

[CVPR'25] Official Implementations for Paper - MagicQuill: An Intelligent Interactive Image Editing System

aigc gradio image-editing mllm

Last synced: 25 Sep 2025

https://github.com/next-gpt/next-gpt

Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model

chatgpt foundation-models gpt-4 instruction-tuning large-language-models llm mllm multi-modal-chatgpt multimodal visual-language-learning

Last synced: 14 May 2025

https://github.com/ant-research/magicquill

[CVPR'25] Official Implementations for Paper - MagicQuill: An Intelligent Interactive Image Editing System

aigc gradio image-editing mllm

Last synced: 14 May 2025

https://nvlabs.github.io/Eagle/

Eagle: Frontier Vision-Language Models with Data-Centric Strategies

demo eagle gpt4 huggingface large-language-models llama llama3 llava llm lmm lvlm mllm nvdia

Last synced: 16 Jul 2026

https://github.com/x-plug/mplug-docowl

mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

chart-understanding document-understanding mllm multimodal multimodal-large-language-models table-understanding

Last synced: 14 May 2025

https://github.com/X-PLUG/mPLUG-DocOwl

mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

chart-understanding document-understanding mllm multimodal multimodal-large-language-models table-understanding

Last synced: 11 May 2025

https://github.com/EvolvingLMMs-Lab/LLaVA-OneVision-2

Fully Open Framework for Democratized Multimodal Training

llava llava-onevision llm mllm qwen3 vision-language-model

Last synced: 23 Jul 2026

https://github.com/BAAI-DCAI/Bunny

A family of lightweight multimodal models.

chatgpt chinese english gpt-4 mllm multimodal-large-language-models vlm

Last synced: 01 May 2025

https://github.com/SkyworkAI/Skywork-R1V

Pioneering Multimodal Reasoning with CoT

deepseek-r1 llm mllm

Last synced: 01 Apr 2025

https://github.com/baai-dcai/bunny

A family of lightweight multimodal models.

chatgpt chinese english gpt-4 mllm multimodal-large-language-models vlm

Last synced: 21 Apr 2025

https://github.com/ChocoWu/Awesome-Scene-Graph-Generation

This is a repository for listing papers on scene graph generation and application.

mllm mllm-for-sg scene-graph scene-graph-generation scene-graph-to-image

Last synced: 12 Sep 2026

https://github.com/nvlabs/eagle

Eagle Family: Exploring Model Designs, Data Recipes and Training Strategies for Frontier-Class Multimodal LLMs

demo eagle gpt4 huggingface large-language-models llama llama3 llava llm lmm lvlm mllm nvdia

Last synced: 15 May 2025

https://github.com/vita-mllm/woodpecker

✨✨Woodpecker: Hallucination Correction for Multimodal Large Language Models

hallucination hallucinations large-language-models llm mllm multimodal-large-language-models multimodality

Last synced: 15 May 2025

https://github.com/bradyfu/woodpecker

✨✨Woodpecker: Hallucination Correction for Multimodal Large Language Models

hallucination hallucinations large-language-models llm mllm multimodal-large-language-models multimodality

Last synced: 31 Mar 2025

https://github.com/VITA-MLLM/Woodpecker

✨✨Woodpecker: Hallucination Correction for Multimodal Large Language Models. The first work to correct hallucinations in MLLMs.

hallucination hallucinations large-language-models llm mllm multimodal-large-language-models multimodality

Last synced: 09 May 2025

https://github.com/NVlabs/EAGLE

EAGLE: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

demo eagle gpt4 huggingface large-language-models llama llama3 llava llm lmm lvlm mllm nvdia

Last synced: 27 Sep 2025

https://github.com/foundationvision/groma

[ECCV2024] Grounded Multimodal Large Language Model with Localized Visual Tokenization

foundation-models grounding large-language-models llama llama2 llm mllm multimodal vision-language-model

Last synced: 04 Apr 2025

https://github.com/skyworkai/vitron

NeurIPS 2024 Paper: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing

mllm multimodal-large-language-models segmentation

Last synced: 16 May 2025

https://github.com/Coobiw/MPP-LLaVA

Personal Project: MPP-Qwen14B & MPP-Qwen-Next(Multimodal Pipeline Parallel based on Qwen-LM). Support [video/image/multi-image] {sft/conversations}. Don't let the poverty limit your imagination! Train your own 8B/14B LLaVA-training-like MLLM on RTX3090/4090 24GB.

deepspeed fine-tuning mllm model-parallel multimodal-large-language-models pipeline-parallelism pretraining qwen video-language-model video-large-language-models

Last synced: 14 Aug 2026

https://github.com/evolvinglmms-lab/llava-onevision-1.5

Fully Open Framework for Democratized Multimodal Training

llava llm mllm qwen3 vision-language-model

Last synced: 26 Dec 2025

https://github.com/ZJU-REAL/Awesome-GUI-Agents

A curated collection of resources, tools, and frameworks for developing GUI Agents.

agents guiagents llm mllm

Last synced: 08 Jul 2026

https://github.com/dvlab-research/llmga

This project is the official implementation of 'LLMGA: Multimodal Large Language Model based Generation Assistant', ECCV2024 Oral

aigc image-design-assistant image-editing image-generation large-language-model llm mllm multi-modal

Last synced: 03 Jul 2025

https://github.com/x-plug/youku-mplug

Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Pre-training Dataset and Benchmarks

benchmark chinese dataset mllm multimodal multimodal-large-language-models multimodal-pretraining video video-question-answering video-retrieval youku

Last synced: 26 Jun 2025

https://github.com/X-PLUG/Youku-mPLUG

Youku-mPLUG: A 10 Million Large-scale Chinese Video-Language Pre-training Dataset and Benchmarks

benchmark chinese dataset mllm multimodal multimodal-large-language-models multimodal-pretraining video video-question-answering video-retrieval youku

Last synced: 20 Apr 2025

https://github.com/VITA-MLLM/Long-VITA

✨✨Long-VITA: Scaling Large Multi-modal Models to 1 Million Tokens with Leading Short-Context Accuracy

long-context mllm vision-language-model

Last synced: 31 Mar 2025

https://github.com/gokayfem/comfyui_vlm_nodes

Custom ComfyUI nodes for Vision Language Models, Large Language Models, Image to Music, Text to Music, Consistent and Random Creative Prompt Generation

comfyui custom-nodes image-captioning img2sfx img2text joytag llava llm mllm nodes phi15 siglip vlm

Last synced: 07 Apr 2025

https://github.com/gokayfem/ComfyUI_VLM_nodes

Custom ComfyUI nodes for Vision Language Models, Large Language Models, Image to Music, Text to Music, Consistent and Random Creative Prompt Generation

comfyui custom-nodes image-captioning img2sfx img2text joytag llava llm mllm nodes phi15 siglip vlm

Last synced: 14 Jul 2025

https://github.com/CircleRadon/TokenPacker

The code for "TokenPacker: Efficient Visual Projector for Multimodal LLM".

connector lmm mllm token-reduction tokenpacker visual-projector

Last synced: 23 Sep 2025

https://github.com/MIV-XJTU/JanusVLN

Official implementation for "JanusVLN: Decoupling Semantics and Spatiality with Dual Implicit Memory for Vision-Language Navigation"

llm mllm vla vln

Last synced: 23 Dec 2025

https://github.com/x-plug/mplug-2

mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video (ICML 2023)

foundation-models image-retrieval mllm mplug multimodal multimodal-pretraining video video-question-answering video-retrieval vqa

Last synced: 09 Sep 2025

https://github.com/tiger-ai-lab/mantis

Official code for Paper "Mantis: Multi-Image Instruction Tuning" [TMLR2024]

fuyu language llava-llama3 lmm mantis mllm multi-image-understanding multimodal video vision vlm

Last synced: 13 Jun 2025

https://tiger-ai-lab.github.io/Mantis/

Official code for Paper "Mantis: Multi-Image Instruction Tuning" [TMLR2024]

fuyu language llava-llama3 lmm mantis mllm multi-image-understanding multimodal video vision vlm

Last synced: 12 Apr 2025

https://github.com/damo-nlp-sg/videorefer

[CVPR 2025] The code for "VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM"

mllm pixel-understanding sam2 video-understanding

Last synced: 08 May 2025

https://github.com/idea-research/chatrex

Code for ChatRex: Taming Multimodal LLM for Joint Perception and Understanding

detection llm mllm multimodal-deep-learning vlm

Last synced: 28 Jun 2025

https://github.com/waybarrios/vllm-mlx

OpenAI-compatible server for Apple Silicon. Run LLMs and vision-language models (Llama, Qwen-VL, LLaVA) with continuous batching, MCP tool calling, and multimodal support. Native MLX backend, 400+ tok/s.

apple-silicon audio-processing computer-vision image-understanding inference llm machine-learning macos mllm mlx multimodal-ai speech-to-text stt text-to-speech tts video-understanding vision-language-model vllm

Last synced: 23 Jan 2026

https://github.com/tidedra/vl-rlhf

A RLHF Infrastructure for Vision-Language Models

dpo llm lmm mllm rlhf vlm

Last synced: 20 Jun 2025

https://github.com/foundationvision/generateu

[CVPR2024] Generative Region-Language Pretraining for Open-Ended Object Detection

mllm multimodality object-detection open-vocabulary open-vocabulary-detection open-world

Last synced: 05 Apr 2025

https://github.com/JackYFL/awesome-VLLMs

This repository collects papers on VLLM applications. We will update new papers irregularly.

application embodied llm mllm reasoning-agent survey vllm vlm vqa

Last synced: 06 Nov 2025

https://github.com/zjrwtx/sft-data-builder

利用免费的大模型api来结合你的私域数据来生成sft训练数据(妥妥白嫖)支持llamafactory等工具的训练数据格式synthetic data

agents alpaca cot datagene gpt40 llm mllm multiagents o1 python react sharegpt slm synthetic-data tailwindcss visionlanguagemodel

Last synced: 07 Mar 2026

https://github.com/thu-ml/mmtrusteval

A toolbox for benchmarking trustworthiness of multimodal large language models (MultiTrust, NeurIPS 2024 Track Datasets and Benchmarks)

benchmark claude fairness gpt-4 mllm multi-modal privacy robustness safety toolbox trustworthy-ai truthfulness

Last synced: 05 Apr 2025

https://github.com/baaivision/densefusion

DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

image-descriptions mllm multimodal-large-language-models vision-language-models visual-perception vlm

Last synced: 05 Apr 2025

https://github.com/microsoft/eureka-ml-insights

A framework for standardizing evaluations of large foundation models, beyond single-score reporting and rankings.

ai artificial-intelligence evaluation-framework llm machine-learning mllm

Last synced: 05 Apr 2025

https://github.com/thu-ml/MMTrustEval

A toolbox for benchmarking trustworthiness of multimodal large language models (MultiTrust, NeurIPS 2024 Track Datasets and Benchmarks)

benchmark claude fairness gpt-4 mllm multi-modal privacy robustness safety toolbox trustworthy-ai truthfulness

Last synced: 27 Jul 2025

https://github.com/bz-lab/AUITestAgent

AUITestAgent is the first automatic, natural language-driven GUI testing tool for mobile apps, capable of fully automating the entire process of GUI interaction and function verification.

agent automation gpt-4o gui llm mllm mobile-app multi-agent multimodal multimodal-agent testing

Last synced: 14 Apr 2025

https://github.com/niutrans/vision-llm-alignment

This repository contains the code for SFT, RLHF, and DPO, designed for vision-based LLMs, including the LLaVA models and the LLaMA-3.2-vision models.

alignment dpo llama3-vision llava llm mllm multi-model ppo reward rlhf sft vision

Last synced: 06 Apr 2025

https://github.com/NiuTrans/Vision-LLM-Alignment

This repository contains the code for SFT, RLHF, and DPO, designed for vision-based LLMs, including the LLaVA models and the LLaMA-3.2-vision models.

alignment dpo llama3-vision llava llm mllm multi-model ppo reward rlhf sft vision

Last synced: 07 May 2025

https://github.com/x-plug/mplug-halowl

mPLUG-HalOwl: Multimodal Hallucination Evaluation and Mitigating

benchmark contrastive-learning hallucinations mllm multimodal-hallucination multimodal-large-language-models

Last synced: 31 Aug 2025

https://github.com/buaadreamer/chinese-llava-med

中文医学多模态大模型 Large Chinese Language-and-Vision Assistant for BioMedicine

ai chinese gpt4v huggingface-datasets llama-factory llava medical minigpt4 mllm multimodal qwen1-5 transformers

Last synced: 11 Sep 2025

https://github.com/xingchensong/touchnet

A native-PyTorch library for large scale M-LLM (text/audio) training with tp/cp/dp/pp.

audio large-scale mllm pytorch text

Last synced: 27 Oct 2025

https://github.com/kwaivgi/uniaa

Unified Multi-modal IAA Baseline and Benchmark

benchmark dataset image-aesthetic-assessment llava mllm

Last synced: 13 Apr 2025

https://github.com/orvillex/machinelearning

本项目以应用为主出发,结合了从基础的机器学习、深度学习到目标检测以及目前最新的大模型,采用目前成熟的 第三方库、开源预训练模型以及相关论文的最新技术,目的是记录学习的过程同时也进行分享以供更多人可以直接进行使用。

knn llm machine-learning mllm numpy scipy siglip sklearn spark-mllib svm tensorflow

Last synced: 09 Apr 2025

https://github.com/parsee-ai/parsee-datasets

Datasets, case studies and benchmarks for extracting structured information from PDFs, HTML files or images, created by the Parsee.ai team. Datasets also on Hugging Face: https://huggingface.co/parsee-ai

datasets llm mllm rag

Last synced: 11 Sep 2025

https://github.com/hewei2001/reachqa

Code & Dataset for Paper: "Distill Visual Chart Reasoning Ability from LLMs to MLLMs"

data-synthesis llm mllm

Last synced: 30 Jul 2025

https://github.com/manycore-research/SpatialLM

SpatialLM: Large Language Model for Spatial Understanding

mllm point-clouds scene-understanding spatial-intelligence

Last synced: 24 Mar 2025

https://github.com/aimagelab/reflectiva

[CVPR 2025] Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering

knowledge-base mllm multimodal vlm vqa

Last synced: 26 Aug 2025

https://github.com/buaadreamer/mllm-finetuning-demo

使用LLaMA-Factory微调多模态大语言模型的示例代码 Demo of Finetuning Multimodal LLM with LLaMA-Factory

finetune-llm huggingface-datasets llama-factory llava lora mllm paligemma pretraining supervised-finetuning transformers yi-vl

Last synced: 11 Apr 2025

https://github.com/waltonfuture/Diff-eRank

Code for https://arxiv.org/abs/2401.17139 (NeurIPS 2024)

evaluation-metrics llm llm-inference machine-learning mllm neurips-2024

Last synced: 19 Jul 2025

https://wanghao9610.github.io/X2SAM/

X2SAM: Any Segmentation in Images and Videos

any-segmentation mllm sam

Last synced: 23 May 2026

https://github.com/wanghao9610/X2SAM

X2SAM: Any Segmentation in Images and Videos

any-segmentation mllm sam

Last synced: 23 May 2026

https://github.com/aidc-ai/wings

The code repository for "Wings: Learning Multimodal LLMs without Text-only Forgetting" [NeurIPS 2024]

deep-learning mllm multimodal-large-language-models multimodal-llm text-only-forgetting

Last synced: 17 Nov 2025

https://github.com/DCDmllm/InstructSAM

The code for "InstructSAM: Segment Any Instance with Any Instructions"

instruction-driven-segmentation mllm multi-instance-segmentation

Last synced: 25 Jun 2026

https://github.com/showlab/visincontext

Official implementation of Leveraging Visual Tokens for Extended Text Contexts in Multi-Modal Learning

efficient in-context-learning llm mllm

Last synced: 22 Apr 2025

https://github.com/freedomintelligence/trim

We introduce new approach, Token Reduction using CLIP Metric (TRIM), aimed at improving the efficiency of MLLMs without sacrificing their performance.

llm mllm multimodal vision-and-language vision-language-model vlm

Last synced: 30 Apr 2025

https://github.com/tychenjiajun/exif-ai

A Node.js CLI and library that uses OpenAI, Ollama, ZhipuAI, Google Gemini or Coze to write AI-generated image descriptions and/or tags to EXIF metadata by its content.

ai cli cli-tool coze exif gemini image jpeg jpg llm metadata mllm ollama openai openai-api photo zhipu

Last synced: 26 Oct 2025

https://github.com/FreedomIntelligence/TRIM

We introduce new approach, Token Reduction using CLIP Metric (TRIM), aimed at improving the efficiency of MLLMs without sacrificing their performance.

llm mllm multimodal vision-and-language vision-language-model vlm

Last synced: 23 Sep 2025

https://github.com/cilabuniba/i-dream-my-painting

[WACV 2025] I Dream My Painting: Connecting MLLMs and Diffusion Models via Prompt Generation for Text-Guided Multi-Mask Inpainting

computer-vision diffusion inpainting llava mllm multimodal prompt-generation stable-diffusion text-to-image

Last synced: 06 Jul 2025

https://github.com/ming-zch/cii-bench

Can MLLMs Understand the Deep Implication Behind Chinese Images?

benchmark datasets image-analysis llm mllm vision-and-language vlm

Last synced: 03 Mar 2025

https://github.com/xirui-li/MOSSBench

An implementation for MLLM oversensitivity evaluatio

alignment attack mllm oversensitivity vlm

Last synced: 27 Jul 2025

https://github.com/buaadreamer/qwen2-vl-history

Qwen2-VL在文旅领域的LLaMA-Factory微调案例 The case for fine-tuning Qwen2-VL in the field of historical literature and museums

beauty history llama-factory mllm multimodal-large-language-models museum qwen2-vl supervised-finetuning

Last synced: 02 Feb 2026

https://github.com/black-yt/IrisGUI

A Lightweight Desktop GUI Agent via Dynamic Focus Vision and Hierarchical Memory. Your AI-powered hands and eyes for desktop automation.

agent ai ai-agents automomous-agents computer-use computer-use-agent computer-use-agents desktop gui guiagent interaction lightweight linux llm macos mllm vllm windows

Last synced: 08 Jul 2026

https://nicolafan.github.io/tamart/

Understanding How MLLMs Describe Artworks Using Token Activation Maps

art computer-vision deep-learning explainability interpretability mllm multimodal

Last synced: 02 Aug 2026

https://github.com/km1994/awesomemultimodel

【AIGC 实战入门笔记 —— AIGC 摩天大楼】分享 大语言模型(LLMs),大模型高效微调(SFT),检索增强生成(RAG),智能体(Agent),PPT自动生成, 角色扮演,文生图(Stable Diffusion) ,图像文字识别(OCR),语音识别(ASR),语音合成(TTS),人像分割(SA),多模态(VLM),Ai 换脸(Face Swapping), 文生视频(VD),图生视频(SVD),Ai 动作迁移,Ai 虚拟试衣,数字人,全模态理解(Omni),Ai音乐生成 干货学习 等 实战与经验。

agent animate asr face-recognition llm llms mllm ocr omni peft-fine-tuning-llm ppt rag sft stable-diffusion svd text-to-music text-to-sql video-diffusion-model virtual-try-on vlm

Last synced: 14 May 2025

https://github.com/kenyony/flexllm

High-Performance LLM Client for Production Batch processing with checkpoint recovery, response caching, load balancing, and cost tracking

claude claude-code gemini llm mllm openai python rate-limit

Last synced: 24 Feb 2026

https://github.com/neemiasbsilva/mllm-persona-evaluation

Official implementation of Stable Behavior, Limited Variation: Persona Validity in LLM Agents for Urban Sentiment Perception

llm-agents mllm ollama perceptsent persona-evaluation qwen-vl sentiment-analysis urban-sentiment vision-language-model vllm

Last synced: 28 May 2026

https://github.com/willoscar/awesome-hci-llm

Awesome-HCI (Ubiquitous, LLM, MLLM, Agent, RAG, Embodied-AI)

agent awesome awesome-list embodied-ai hci human-computer-interaction imu llm mllm rag sensor survey ubiquitous

Last synced: 04 Mar 2025

https://github.com/nidhiyashwanth/spatiallm

Trying out SpatialLM (SpatialLM: Large Language Model for Spatial Understanding). Impressed with results 💖

mllm point-clouds scene-understanding spatial-intelligence

Last synced: 28 Mar 2025

https://github.com/hemangjoshi37a/factoryaioptimize

AI-Powered Multi-Camera Vision LLM System for Factory Optimization

ai automation factory gidc llm machines manufacturing mllm vision vlm

Last synced: 02 Apr 2026

https://github.com/kaicheng001/awesome-r1

A curated list of research papers, models, and resources related to R1-style reasoning models following DeepSeek-R1's breakthrough in January 2025.

awesome deepseek-r1 llm lmm mllm r1 reasoning-models reward-model thinking vlm

Last synced: 22 Jun 2025