{"id":29294869,"url":"https://github.com/aidc-ai/ovis-u1","last_synced_at":"2026-03-16T09:36:56.294Z","repository":{"id":301745567,"uuid":"992512658","full_name":"AIDC-AI/Ovis-U1","owner":"AIDC-AI","description":"An unified model that seamlessly integrates multimodal understanding, text-to-image generation, and image editing within a single powerful framework.","archived":false,"fork":false,"pushed_at":"2025-08-04T06:26:01.000Z","size":10501,"stargazers_count":394,"open_issues_count":2,"forks_count":10,"subscribers_count":2,"default_branch":"main","last_synced_at":"2025-08-14T06:07:00.425Z","etag":null,"topics":["image-editing","multimodal-large-language-models","text-to-image"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/AIDC-AI.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2025-05-29T09:21:44.000Z","updated_at":"2025-08-13T09:13:48.000Z","dependencies_parsed_at":"2025-07-31T12:36:45.898Z","dependency_job_id":null,"html_url":"https://github.com/AIDC-AI/Ovis-U1","commit_stats":null,"previous_names":["aidc-ai/ovis-u1"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/AIDC-AI/Ovis-U1","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/AIDC-AI%2FOvis-U1","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/AIDC-AI%2FOvis-U1/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/AIDC-AI%2FOvis-U1/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/AIDC-AI%2FOvis-U1/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/AIDC-AI","download_url":"https://codeload.github.com/AIDC-AI/Ovis-U1/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/AIDC-AI%2FOvis-U1/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":284902571,"owners_count":27081910,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-11-17T02:00:06.431Z","response_time":55,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["image-editing","multimodal-large-language-models","text-to-image"],"created_at":"2025-07-06T13:40:31.029Z","updated_at":"2026-03-16T09:36:56.263Z","avatar_url":"https://github.com/AIDC-AI.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cp align=\"center\"\u003e\n\u003cimg src=https://cdn-uploads.huggingface.co/production/uploads/637aebed7ce76c3b834cea37/3IK823BZ8w-mz_QfeYkDn.png width=\"30%\"/\u003e\u003c/p\u003e\n\n\u003ch1 align=\"center\"\u003e\nOvis-U1: Unified Understanding, Generation, and Editing\n\u003c/h1\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003ca href=\"https://arxiv.org/abs/2506.23044\"\u003e\u003cimg src=\"https://img.shields.io/badge/arXiv_paper-2506.23044-b31b1b.svg\" alt=\"arxiv\"\u003e\u003c/a\u003e\n\u003c!--   \u003ca href=\"https://github.com/AIDC-AI/Ovis-U1/blob/main/docs/Ovis_U1_Report.pdf\"\u003e\u003cimg src=\"https://img.shields.io/badge/Paper-Tech_Report-b31b1b\" alt=\"paper\"\u003e\u003c/a\u003e --\u003e\n  \u003ca href=\"https://github.com/AIDC-AI/Ovis\"\u003e\u003cimg src=\"https://img.shields.io/badge/GitHub-AIDC--AI/Ovis--U1-blue?style=flat\u0026logo=github\" alt=\"code\"\u003e\u003c/a\u003e\n  \u003ca href=\"https://huggingface.co/spaces/AIDC-AI/Ovis-U1-3B\"\u003e\u003cimg src=\"https://img.shields.io/badge/🎨_HF_Spaces-AIDC--AI/Ovis--U1--3B-lightblack\" alt=\"demo\"\u003e\u003c/a\u003e\n  \u003ca href=\"https://huggingface.co/AIDC-AI/Ovis-U1-3B\"\u003e\u003cimg src=\"https://img.shields.io/badge/🤗_Model-AIDC--AI/Ovis--U1--3B-yellow\" alt=\"model\"\u003e\u003c/a\u003e\n\u003c/p\u003e\n\n\u003cp align=\"left\"\u003e\nBuilding on the foundation of the Ovis series, Ovis-U1 is a 3-billion-parameter unified model that  seamlessly integrates \u003cb\u003emultimodal understanding\u003c/b\u003e, \u003cb\u003etext-to-image generation\u003c/b\u003e, and \u003cb\u003eimage editing\u003c/b\u003e within a single powerful framework. \n\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"docs/imgs/Ovis-U1.jpg\" width=\"95%\"\u003e\n  \u003cbr\u003e\n  \u003cem\u003eThe overall architecture of Ovis-U1 (cf. Fig.2 in our report).\u003c/em\u003e\n\u003c/p\u003e\n\n## 🏆 Highlights\n\n*   **Unified Capabilities**: A single model excels at three core tasks: understanding complex scenes, generating images from text, and performing precise edits based on instructions.\n*   **Advanced Architecture**: Ovis-U1 features a powerful diffusion-based visual decoder (MMDiT) and a bidirectional token refiner, enabling high-fidelity image synthesis and enhanced interaction between text and vision.\n*   **Synergistic Unified Training**: Unlike models trained on single tasks, Ovis-U1 is trained on a diverse mix of understanding, generation, and editing data simultaneously. Our findings show that this approach achieves improved generalization, seamlessly handling real-world multimodal challenges with high accuracy.\n*   **State-of-the-Art Performance**: Ovis-U1 achieves leading scores on multiple academic benchmarks, surpassing strong contemporary models in multimodal understanding (69.6 on OpenCompass), generation (83.72 on DPG-Bench), and editing (4.00 on ImgEdit-Bench).\n\n\n## ✨ Showcase\n\nHere are some examples demonstrating the capabilities of Ovis-U1.\n\n\u003cfigure\u003e\n  \u003cimg src=\"docs/imgs/examples.png\" alt=\"Ovis-U1 examples\"\u003e\n  \u003cfigcaption style=\"text-align: center;\"\u003e\u003c/figcaption\u003e\n\u003c/figure\u003e\n\n\n## 🚀 News\n\n- [2025/11/29] 🔥 Announcing Ovis-Image ([GitHub](https://github.com/AIDC-AI/Ovis-Image), [Model](https://huggingface.co/AIDC-AI/Ovis-Image-7B), [Demo](https://huggingface.co/spaces/AIDC-AI/Ovis-Image-7B))!\n- [2025/6/28] Announcing Ovis-U1-3B ([Model](https://huggingface.co/AIDC-AI/Ovis-U1-3B), [Demo](https://huggingface.co/spaces/AIDC-AI/Ovis-U1-3B))!\n\n\n## 📦 Installation\n\nOvis-U1 has been tested with Python 3.10, Torch 2.4.0, Transformers 4.51.3, and DeepSpeed 0.15.4. For a full list of package dependencies, please see `requirements.txt`.\n\n```bash\ngit clone git@github.com:AIDC-AI/Ovis-U1.git\nconda create -n ovis-u1 python=3.10 -y\nconda activate ovis-u1\ncd Ovis-U1\npip install -r requirements.txt\npip install -e .\n```\n\n## 🛠️ Inference\n\nWe provide simple scripts to test the different capabilities of Ovis-U1.\n\nFor single image understanding, please run\n\n```bash\npython test_img_to_txt.py\n```\n\nFor multi-image understanding, please run\n\n```bash\npython test_multi_img_to_txt.py\n```\n\nFor text-to-image, please run\n```bash\npython test_txt_to_img.py \\\n    --height 1024 \\\n    --width 1024  \\\n    --steps 50 \\\n    --seed 42 \\\n    --txt_cfg 5  \n```\n\nFor image editing, please run\n```bash\npython test_img_edit.py \\\n    --steps 50 \\\n    --img_cfg 1.5 \\\n    --txt_cfg 6  \n```\n\nAlternatively, you can try Ovis-U1 directly in your browser on [![Hugging Face Space](https://img.shields.io/badge/🎨_HF_Spaces-AIDC--AI/Ovis--U1--3B-lightblack)](https://huggingface.co/spaces/AIDC-AI/Ovis-U1-3B)\n\n\n## 📊 Performance\n\n#### OpenCompass Multi-modal Academic Benchmarks\n\n| Model | Avg | MMB | MMS | MMMU | MathVista | Hallusion | AI2D | OCRBench | MMVet | \n|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|\n| GPT-4o | **75.4** | **86**  |**70.2** | **72.9** | **71.6** | **57** | **86.3** | 82.2 | **76.9** | \n| InternVL2.5-2B | 59.9 | 70.9 | 54.3 | 43.2 | 51.1 | 42.3 | 74.9 | 80.2 | 62.6 | \n| SAIL-VL-2B | 61 | 73.7 |56.5 | 44.1 | 62.8 | 45.9 | 77.4 | 83.1 | 44.2 | \n| InternVL3-2B | 61.1 | 78 |61.1 | 48.7 | 57.6 | 41.9 | 78.6 | 83.1 | \u003cins\u003e67\u003c/ins\u003e | \n| Qwen2.5-VL-3B | 64.5 | 76.8 | 56.3 | 51.2 | 61.2 | 46.6 | 81.4 | 82.8 | 60 | \n| Ovis2-2B | 65.2 | 76.9 | 56.7 | 45.6 | 64.1 | 50.2 | 82.7 | 87.3 | 58.3 | \n| SAIL-VL-1.5-2B | 67  | 78.5 | 62.6 | 46.4 | 67 | 50 | 83.7 | **89.1** | 58.8 | \n| Ristretto-3B | 67.7 | \u003cins\u003e80.2\u003c/ins\u003e | \u003cins\u003e62.8\u003c/ins\u003e | \u003cins\u003e51.3\u003c/ins\u003e | 67.6 | 50.2 | 84.2 | 84.7 | 60.7 | \n| Ovis-U1 |  \u003cins\u003e69.6\u003c/ins\u003e  | 77.8 |61.3 | 51.1 | \u003cins\u003e69.4\u003c/ins\u003e | \u003cins\u003e56.3\u003c/ins\u003e | \u003cins\u003e85.6\u003c/ins\u003e |  \u003cins\u003e88.3\u003c/ins\u003e | 66.7 | \n\n#### GenEval\n\n| Model | Overall |Single object | Two object | Counting | Colors | Position | Attribute binding | \n|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|\n| GPT-4o | 0.84 | \u003cins\u003e0.99\u003c/ins\u003e | 0.92 | \u003cins\u003e0.85\u003c/ins\u003e | 0.92 | 0.75 | 0.61 | \n| BAGEL | 0.82  | \u003cins\u003e0.99\u003c/ins\u003e | 0.94 | 0.81 | 0.88 | 0.64 | 0.63 | \n| BAGEL 📝 | \u003cins\u003e0.88\u003c/ins\u003e | 0.98 | 0.95 | 0.84 | \u003cins\u003e0.95\u003c/ins\u003e | \u003cins\u003e0.78\u003c/ins\u003e | **0.77** |\n| UniWorld-V1 | 0.80 | \u003cins\u003e0.99\u003c/ins\u003e | 0.93 | 0.79 | 0.89 | 0.49 | 0.70 |\n| UniWorld-V1 📝 | 0.84 | 0.98 | 0.93 | 0.81 | 0.89 | 0.74 | 0.71 | \n| OmniGen | 0.68 |  0.98 | 0.84 | 0.66 | 0.74 | 0.40 | 0.43 | \n| OmniGen2 |0.80 |  **1** | 0.95 | 0.64 | 0.88 | 0.55 | \u003cins\u003e0.76\u003c/ins\u003e | \n| OmniGen2 📝 | 0.86 | \u003cins\u003e0.99\u003c/ins\u003e | \u003cins\u003e0.96\u003c/ins\u003e | 0.74 | **0.98** | 0.71 | 0.75 | \n| Ovis-U1 |**0.89** |  0.98 | **0.98** | **0.90** | 0.92 | **0.79** | 0.75 | \n\n*📝 denotes using the rewritten prompts*\n\n#### DPG-Bench\n\n| Model | Overall | Global | Entity | Attribute | Relation | Other | \n|:---:|:---:|:---:|:---:|:---:|:---:|:---:|\n| BAGEL | **85.07** | **88.94** | **90.37** | **91.29** | \u003cins\u003e90.82\u003c/ins\u003e | \u003cins\u003e88.67\u003c/ins\u003e | \n| UniWorld-V1 |81.38 |  83.64 | 88.39 | 88.44 | 89.27 | 87.22 | \n| OmniGen |81.16 | 87.90 | 88.97 | 88.47 | 87.95 | 83.56 | \n| OmniGen2 |83.57 | \u003cins\u003e88.81\u003c/ins\u003e | 88.83 | \u003cins\u003e90.18\u003c/ins\u003e | 89.37 | **90.27** | \n| Ovis-U1 | \u003cins\u003e83.72\u003c/ins\u003e | 82.37 | \u003cins\u003e90.08\u003c/ins\u003e | 88.68 | **93.35** | 85.20 |\n\n#### ImgEdit-Bench\n\n| Model | Overall |Add | Adjust | Extract | Replace | Remove | Background | Style | Hybrid | Action | \n|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|\n| GPT-4o | **4.2** | **4.61** | **4.33** | \u003cins\u003e2.9\u003c/ins\u003e | \u003cins\u003e4.35\u003c/ins\u003e | \u003cins\u003e3.66\u003c/ins\u003e | **4.57** | **4.93** | **3.96** | **4.89** | \n| MagicBrush | 1.90 | 2.84 | 1.58 | 1.51 | 1.97 | 1.58 | 1.75 | 2.38 | 1.62 | 1.22 | \n| Instruct-P2P | 1.88 | 2.45 | 1.83 | 1.44 | 2.01 | 1.50 | 1.44 | 3.55 | 1.2 | 1.46 | \n| AnyEdit | 2.45 | 3.18 | 2.95 | 1.88 | 2.47 | 2.23 | 2.24 | 2.85 | 1.56 | 2.65 | \n| UltraEdit |2.7 | 3.44 | 2.81 | 2.13 | 2.96 | 1.45 | 2.83 | 3.76 | 1.91 | 2.98 | \n| OmniGen |  2.96 | 3.47 | 3.04 | 1.71 | 2.94 | 2.43 | 3.21 | 4.19 | 2.24 | 3.38 |\n| Step1X-Edit |3.06 |  3.88 | 3.14 | 1.76 | 3.40 | 2.41 | 3.16 | 4.63 | 2.64 | 2.52 | \n| ICEdit |3.05 | 3.58 | 3.39 | 1.73 | 3.15 | 2.93 | 3.08 | 3.84 | 2.04 | 3.68 | \n| BAGEL |3.2 | 3.56 | 3.31 | 1.7 | 3.3 | 2.62 | 3.24 | 4.49 | 2.38 | 4.17 | \n| UniWorld-V1 |3.26 | 3.82 | 3.64 | 2.27 | 3.47 | 3.24 | 2.99 | 4.21 | 2.96 | 2.74 | \n| OmniGen2 | 3.44 | 3.57 | 3.06 | 1.77 | 3.74 | 3.2 | 3.57 | \u003cins\u003e4.81\u003c/ins\u003e | 2.52 | \u003cins\u003e4.68\u003c/ins\u003e |\n| Ovis-U1 |\u003cins\u003e4.00\u003c/ins\u003e | \u003cins\u003e4.13\u003c/ins\u003e | \u003cins\u003e3.62\u003c/ins\u003e | **2.98** | **4.45** | **4.06** | \u003cins\u003e4.22\u003c/ins\u003e | 4.69 | \u003cins\u003e3.45\u003c/ins\u003e | 4.61 | \n\n\n#### GEdit-Bench-EN\n\n|  Model | Avg | Background Change | Color Alteration   | Material Modification  | Motion Change | Portrait Beautification  | Style Transfer  | Subject Addition  | Subject Removal  | Subject Replacement  | Text Modification  | Tone Transformation  | \n|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|:---:|\n| GPT-4o |**7.534** | 7.205 |\t6.491 |\t**6.607** | **8.096** |\t**7.768** |\t\u003cins\u003e6.961\u003c/ins\u003e |\t7.622 |\t**8.331** |\t**8.067** |\t**7.427** |\t**8.301** |\t\n| AnyEdit | 3.212 | 4.663\t| 4.260 |\t2.537 |\t2.024 |\t3.479\t| 2.032 |\t3.995 |\t3.089 |\t3.180 |\t0.922 |\t5.151 |\t\n| Instruct-Pix2Pix | \t3.684 | 3.825 |\t5.182 |\t3.688 |\t3.509 |\t4.339 |\t4.560 |\t3.461 |\t2.031 |\t4.237 |\t0.955 |\t4.733 |\n| MagicBrush |4.518 |\t5.637 |\t5.136 |\t5.078 |\t4.513 |\t4.487 |\t4.439 |\t5.252 |\t3.704 |\t4.941 |\t1.384 |\t5.130 |\t\n| OmniGen | 5.062 | 5.281 |\t6.003 |\t5.308 |\t2.916 |\t3.087 |\t4.903 |\t6.628 |\t6.352 |\t5.616 |\t4.519 |\t5.064 |\t\n| Gemini |6.315 | \t6.781 |\t6.369 |\t6.040 |\t6.938 |\t5.591 |\t4.676 |\t7.501 |\t6.447 |\t7.003 |\t5.765 |\t6.350 |\t\n| Step1X-Edit |\t6.701 | 6.547 |\t6.545 |\t6.204 |\t6.483 |\t6.787 |\t**7.221** |\t6.975 |\t6.512 |\t7.068 |\t\u003cins\u003e6.921\u003c/ins\u003e |\t6.448 |\t\n| Doubao |\u003cins\u003e6.754\u003c/ins\u003e | \t\u003cins\u003e7.430\u003c/ins\u003e |\t**7.095** |\t6.339 |\t\u003cins\u003e6.973\u003c/ins\u003e |\t\u003cins\u003e6.972\u003c/ins\u003e |\t6.767 |\t\u003cins\u003e7.674\u003c/ins\u003e |\t6.748 |\t\u003cins\u003e7.447\u003c/ins\u003e |\t3.471 |\t\u003cins\u003e7.383\u003c/ins\u003e |\t\n| BAGEL | 6.519 | 7.324 |\t\u003cins\u003e6.909\u003c/ins\u003e |\t\u003cins\u003e6.381\u003c/ins\u003e |\t4.753 |\t4.573 |\t6.150 |\t**7.896** |\t7.164 |\t7.021 |\t7.320 |\t6.218 |\t\n| Ovis-U1 |6.420 | **7.486** |\t6.879 |\t6.208 |\t4.790 |\t5.981 |\t6.463 |\t7.491 |\t\u003cins\u003e7.254\u003c/ins\u003e |\t7.266 |\t4.482 |\t6.314 |\t\n\n* Note that the leaderboard has been updated by this [commit](https://github.com/step1x-edit/step1x-edit.github.io/commit/b45f822d64a1b5b3239509fb7905efb2afad0300). The results shown here are from an earlier version.\n\n## 📚 Citation\n\nIf you find Ovis-U1 useful for your research or applications, please cite our technical report:\n\n```bibtex\n@article{wang2025ovisu1,\n  title={Ovis-U1 Technical Report}, \n  author={Wang, Guo-Hua and Zhao, Shanshan and Zhang, Xinjie and Cao, Liangfu and Zhan, Pengxin and Duan, Lunhao and Lu, Shiyin and Fu, Minghao and Zhao, Jianshan and Li, Yang and Chen, Qing-Guo},\n  journal={arXiv preprint arXiv:2506.23044},\n  year={2025}\n}\n```\n\n## 🙏 Acknowledgments\n\nThe code is built upon [Ovis](https://github.com/AIDC-AI/Ovis) and [FLUX](https://github.com/black-forest-labs/flux). We thank their authors for open-sourcing their great work.\n\n## 📄 License\n\nThis project is released under Apache License 2.0 (http://www.apache.org/licenses/LICENSE-2.0, SPDX-License-identifier: Apache-2.0).\n\n## 🚨 Disclaimer\n\nWe used compliance checking algorithms during the training process, to ensure the compliance of the trained model to the best of our ability. Due to complex data and the diversity of language model usage scenarios, we cannot guarantee that the model is completely free of copyright issues or improper content. If you believe anything infringes on your rights or generates improper content, please contact us, and we will promptly address the matter.\n\n\n## 🔥 We are hiring!\n\nWe are looking for both interns and full-time researchers to join our team, focusing on multimodal understanding, generation, reasoning, AI agents, and unified multimodal models. If you are interested in exploring these exciting areas, please reach out to us at qingguo.cqg@alibaba-inc.com.\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Faidc-ai%2Fovis-u1","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Faidc-ai%2Fovis-u1","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Faidc-ai%2Fovis-u1/lists"}