{"id":21361220,"url":"https://github.com/thisisiron/llava-pool","last_synced_at":"2025-07-08T07:36:15.779Z","repository":{"id":262801482,"uuid":"887592736","full_name":"thisisiron/LLaVA-Pool","owner":"thisisiron","description":"🌋 A flexible framework for training and configuring Vision-Language Models","archived":false,"fork":false,"pushed_at":"2025-03-14T15:01:09.000Z","size":1586,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-03-14T16:22:11.087Z","etag":null,"topics":["llava","multimodal-large-language-models","vision-language-model","vlm"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/thisisiron.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2024-11-13T00:07:43.000Z","updated_at":"2025-03-14T15:01:09.000Z","dependencies_parsed_at":"2024-12-29T16:23:22.292Z","dependency_job_id":"498fb1c0-0cf3-4bdb-a805-bc7e8294fa1c","html_url":"https://github.com/thisisiron/LLaVA-Pool","commit_stats":null,"previous_names":["thisisiron/vlm-finetuning","thisisiron/llava-pool"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/thisisiron%2FLLaVA-Pool","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/thisisiron%2FLLaVA-Pool/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/thisisiron%2FLLaVA-Pool/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/thisisiron%2FLLaVA-Pool/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/thisisiron","download_url":"https://codeload.github.com/thisisiron/LLaVA-Pool/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":243835937,"owners_count":20355611,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["llava","multimodal-large-language-models","vision-language-model","vlm"],"created_at":"2024-11-22T06:09:01.858Z","updated_at":"2025-07-08T07:36:15.773Z","avatar_url":"https://github.com/thisisiron.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# LLaVA-Pool\n\n\u003cp align=\"center\"\u003e\n    \u003cimg src=\"assets/llavapool_nobg.png\" width=300\u003e\n\u003c/p\u003e\n\n\u003cdiv align=\"center\"\u003e\n\n**A Comprehensive Framework for Training and Fine-tuning Vision-Language Models**\n\u003c/div\u003e\n\n## 📖 Overview\n\nLLaVA-Pool is a powerful and flexible framework designed for training and fine-tuning Vision-Language Models (VLMs). It provides a unified interface for working with various state-of-the-art VLMs, supporting both pre-training and supervised fine-tuning workflows. With LLaVA-Pool, you can easily customize and optimize VLMs for your specific multimodal tasks.\n\n## ✨ Key Features\n\n- **Multiple Model Support**: Compatible with leading VLMs including Qwen2-VL, Qwen2.5-VL, LLama 3.2 Vision, Pixtral, and InternVL2.5\n- **Flexible Training Methods**: Support for pre-training, supervised fine-tuning (SFT), and direct preference optimization (DPO)\n- **Efficient Data Processing**: Streamlined data loading and processing pipelines for multimodal datasets\n- **Distributed Training**: Built-in support for multi-GPU and multi-node training\n- **Customizable Configuration**: YAML-based configuration system for easy experiment management\n- **Inference Tools**: Ready-to-use inference scripts for model evaluation and deployment\n\n## 🛠️ Installation\n\nTo install LLaVA-Pool, follow these commands:\n\n```bash\ngit clone https://github.com/thisisiron/LLaVA-Pool.git\ncd LLaVA-Pool\npip install -r requirements.txt\n```\n\nFor optimal performance with GPU acceleration, we recommend installing flash-attention:\n\n```bash\npip install flash-attn --no-build-isolation\n```\n\n## 📊 Supported Models\n\n| Model | Converter | Description |\n| --- | --- | --- |\n| Qwen2-VL | qwen2_vl | Qwen2's vision-language model |\n| Qwen2.5-VL | qwen2_vl | Qwen2.5's vision-language model |\n| Llama 3.2 Vision | llama3.2_vision | Meta's Llama 3.2 with vision capabilities |\n| Pixtral | pixtral | Pixtral vision-language model |\n| InternVL2.5 | internvl2_5 | InternVL's 2.5 version |\n\n## 📚 Data Preparation\n\nLLaVA-Pool supports various data formats for training and fine-tuning. The data directory structure should be organized as follows:\n\n```\ndata/\n├── dataset_config.json  # Configuration for datasets\n├── demo.json            # Example data format\n└── demo_data/           # Example images\n    ├── image1.jpg\n    ├── image2.jpg\n    └── ...\n```\n\nThe `dataset_config.json` file defines the datasets to be used for training. Each dataset should follow the format specified in the demo.json file, which includes conversations and image references.\n\n## 🚀 Training\n\n### Pre-training\n\nFor pre-training a vision-language model, you can use the provided configuration files in the `examples` directory:\n\n```bash\nexport PYTHONPATH=src:$PYTHONPATH\ntorchrun --nnodes 1 --nproc_per_node 4 --master_port 20001 src/llavapool/run.py examples/pretrain_config.yaml\n```\n\n### Supervised Fine-tuning (SFT)\n\nFor fine-tuning a pre-trained model on your specific task:\n\n```bash\nexport PYTHONPATH=src:$PYTHONPATH\ntorchrun --nnodes 1 --nproc_per_node 4 --master_port 20001 src/llavapool/run.py examples/qwen2vl_full_sft.yaml\n```\n\n### Direct Preference Optimization (DPO)\n\nTo align your model with human preferences using DPO:\n\n```bash\nexport PYTHONPATH=src:$PYTHONPATH\ntorchrun --nnodes 1 --nproc_per_node 4 --master_port 20001 src/llavapool/run.py examples/qwen2vl_dpo.yaml\n```\n\nDPO fine-tuning requires a dataset of preferred and rejected responses for given prompts. This alignment technique helps improve the model's response quality by learning from human preferences.\n\nYou can customize the training parameters by modifying the YAML configuration files in the `examples` directory.\n\n## 🔍 Inference\n\nTo run inference with a trained model:\n\n```bash\npython infer.py --model_path /path/to/your/model\n```\n\nThis will launch a Gradio interface for interactive testing of your model.\n\n## 📋 Configuration\n\nLLaVA-Pool uses YAML configuration files to define training parameters. Here's an example configuration for fine-tuning Qwen2-VL:\n\n```yaml\n### model\nmodel_name_or_path: \"Qwen/Qwen2-VL-7B-Instruct\"\n### method\nstage: sft\ndo_train: true\nfinetuning_type: full\nfreeze_vision_tower: false\ntrain_mm_proj_only: false\ndeepspeed: scripts/deepspeed/zero3.json\n### dataset\ndataset: aihub_table,aihub_chart,aihub_math,aihub_ocr,docvqa,arxivqa,ocr-vqa-200k,figureqa\ntemplate: qwen2_vl\ncutoff_len: 32768\noverwrite_cache: true\npreprocessing_num_workers: 80\n### output\noutput_dir: output/qwen2_vl-7b/full/ko-en_docai_qwen2vl_train-vit_cl8192_wodocvqa25k\nlogging_steps: 100\nsave_steps: 20000\noverwrite_output_dir: true\n### train\nper_device_train_batch_size: 1\ngradient_accumulation_steps: 8\nlearning_rate: 1.0e-5\nnum_train_epochs: 1\nlr_scheduler_type: cosine\nwarmup_ratio: 0.1\nbf16: true\n```\n\n## 🤝 Contributing\n\nContributions to LLaVA-Pool are welcome! Please feel free to submit a Pull Request.\n\n## 📄 License\n\nThis project is licensed under the Apache License 2.0. See the [LICENSE](LICENSE) file for more details.\n\n## 🔗 References\n\nThis repository was built with inspiration from:\n\n- [LLaMA-Factory](https://github.com/hiyouga/LLaMA-Factory)\n- [LLaVA-NeXT](https://github.com/haotian-liu/LLaVA)\n- [InternVL](https://github.com/OpenGVLab/InternVL)\n\n## 📧 Contact\n\nFor questions or feedback, please open an issue on GitHub.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fthisisiron%2Fllava-pool","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fthisisiron%2Fllava-pool","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fthisisiron%2Fllava-pool/lists"}