{"id":31921782,"url":"https://github.com/alpha-vllm/lumina-dimoo","last_synced_at":"2025-10-13T22:54:00.763Z","repository":{"id":314123092,"uuid":"1053643576","full_name":"Alpha-VLLM/Lumina-DiMOO","owner":"Alpha-VLLM","description":"Lumina-DiMOO - An Open-Sourced Multi-Modal Large Diffusion Language Model","archived":false,"fork":false,"pushed_at":"2025-10-05T20:34:50.000Z","size":55223,"stargazers_count":689,"open_issues_count":3,"forks_count":41,"subscribers_count":20,"default_branch":"main","last_synced_at":"2025-10-05T21:14:26.430Z","etag":null,"topics":["diffusion-large-language-model","discrete-diffusion-models","unified-multimodal-understanding-and-generation"],"latest_commit_sha":null,"homepage":"https://synbol.github.io/Lumina-DiMOO/","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/Alpha-VLLM.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2025-09-09T18:19:24.000Z","updated_at":"2025-10-05T20:34:53.000Z","dependencies_parsed_at":"2025-10-05T21:16:46.511Z","dependency_job_id":null,"html_url":"https://github.com/Alpha-VLLM/Lumina-DiMOO","commit_stats":null,"previous_names":["alpha-vllm/lumina-dimoo"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/Alpha-VLLM/Lumina-DiMOO","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Alpha-VLLM%2FLumina-DiMOO","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Alpha-VLLM%2FLumina-DiMOO/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Alpha-VLLM%2FLumina-DiMOO/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Alpha-VLLM%2FLumina-DiMOO/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/Alpha-VLLM","download_url":"https://codeload.github.com/Alpha-VLLM/Lumina-DiMOO/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Alpha-VLLM%2FLumina-DiMOO/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":279017088,"owners_count":26085984,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-10-13T02:00:06.723Z","response_time":61,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["diffusion-large-language-model","discrete-diffusion-models","unified-multimodal-understanding-and-generation"],"created_at":"2025-10-13T22:53:59.199Z","updated_at":"2025-10-13T22:54:00.756Z","avatar_url":"https://github.com/Alpha-VLLM.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cp align=\"center\"\u003e\n \u003cimg src=\"./assets/Lumina-DiMOO.png\" width=\"20%\"/\u003e\n\u003c/p\u003e\n\n\u003cdiv align=\"center\"\u003e\n \u003ch1\u003e Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding \u003c/h1\u003e\n\n  [[📑 Technical Report ](http://arxiv.org/abs/2510.06308)] \u0026emsp; [[🌐 Project Page (Demo \u0026 Benchmark)](https://synbol.github.io/Lumina-DiMOO/)] \u0026emsp; [[🤗 Model ](https://huggingface.co/Alpha-VLLM/Lumina-DiMOO)]\n \n \u003cb\u003e¹Shanghai AI Laboratory, ²Shanghai Innovation Institute, ³Shanghai Jiao Tong University, ⁴Nanjing University \u003c/b\u003e\n \n \u003cb\u003e⁵The University of Sydney, ⁶The Chinese University of Hong Kong, ⁷Tsinghua University\u003c/b\u003e\n\n \u003cimg src=\"./assets/teaser.png\" width=\"95%\"/\u003e\n\u003c/div\u003e\n\n## 📚 Introduction \nWe introduce Lumina-DiMOO, an omni foundational model for seamless multimodal generation and understanding. Lumina-DiMOO is distinguished by four key innovations:\n\n - **Unified Discrete Diffusion Architecture:** Lumina-DiMOO sets itself apart from prior unified models by utilizing a fully discrete diffusion modeling to handle inputs and outputs across various modalities.\n - **Versatile Multimodal Capabilities:** Lumina-DiMOO supports a broad spectrum of multimodal tasks, including text-to-image generation (allowing for arbitrary and high-resolution), image-to-image generation (e.g., image editing, subject-driven generation, and image inpainting, etc.), alongside advanced image understanding.\n\n - **Higher Sampling Efficiency:** Compared to previous AR or hybrid AR-diffusion paradigms, Lumina-DiMOO demonstrates remarkable sampling efficiency. Additionally, we design a bespoke caching method to further speed up the sampling speed by 2x.\n\n - **Superior Performance:** Lumina-DiMOO achieves state-of-the-art performance on multiple benchmarks, surpassing existing open-source unified multimodal models, setting a new standard in the field.\n\n\n   \n \u003cimg src=\"./assets/architecture.png\" width=\"100%\"/\u003e\n\n\n## 🔥 News\n- **[2025-10-06]** Training code is released.\n- **[2025-09-25]** We have released the Technical Report.\n- **[2025-09-20]** 🎉 In the latest [UniGenBench Leaderboard](https://huggingface.co/spaces/CodeGoat24/UniGenBench_Leaderboard)(maintained by Tencent Hunyuan Team), Lumina-DiMOO's generation evaluation ranks 1st 🥇 among all open-source unified models. \n- **[2025-09-12]** We have open-sourced Image Inpainting \u0026 Extrapolation code.\n- **[2025-09-11]** We have open-sourced the Max Logit-based Cache solution, offering a 2x speed improvement for sampling.\n- **[2025-09-10]** 🎉 We release the initial version of **Lumina-DiMOO**, including:\n  - 🎯 Model Checkpoints on [HuggingFace](https://huggingface.co/Alpha-VLLM/Lumina-DiMOO)!\n  - 🎯 Text-to-Image \u0026 Image-to-Image Generation Inference code!\n  - 🎯 Image Understanding Inference Code!\n  - 🎯 Website \u0026 Demo on [Project Page](https://synbol.github.io/Lumina-DiMOO/)!\n\n## 📝 Open-Source Plan\n - [x] Image Inpainting \u0026 Extrapolation Code\n - [x] Fast Sampling with Max Logit-based Cache\n - [ ] Gradio Demo\n - [ ] Bechmark Evaluation Code\n - [x] Fine-Tuning Code\n - [ ] Self-GRPO Training Code\n - [x] Technical Report\n\n## 📽️ Qualitative Results\nHere we present some comparative generation results with other models. **For additional visualization results, please see our [Project Page](https://synbol.github.io/Lumina-DiMOO/).**\n\u003cdetails open\u003e\n  \u003csummary\u003eText-to-Image Comparison\u003c/summary\u003e\n  \u003cimg src=\"./assets/demo_t2i.png\" width=\"100%\"/\u003e\n\u003c!--   \u003cdetails open\u003e\n  \u003csummary\u003eEffects of Max Logit-Based Cache (A800 GPU, 1536x768 resolution)\u003c/summary\u003e\n  Without Cache: Latency: 58.2 s; Peak GPU Memory: 38.9 GiB\n  \u003cimg src=\"./assets/nocache.png\" width=\"80%\"/\u003e\n\n\n  With Cache: Latency: 32.2 s; Peak GPU Memory: 45.9 GiB\n  \u003cimg src=\"./assets/cache.png\" width=\"80%\"/\u003e\n\u003c/details\u003e --\u003e\n\u003c/details\u003e\n\n\u003cdetails close\u003e\n  \u003csummary\u003eImage Editing Comparison\u003c/summary\u003e\n  \u003cimg src=\"./assets/demo_editing.png\" width=\"100%\"/\u003e\n\u003c/details\u003e\n\n\u003cdetails close\u003e\n  \u003csummary\u003eControllable \u0026 Subject-Driven Generation Comparison\u003c/summary\u003e\n  \u003cimg src=\"./assets/qualitative_control_subject.png\" width=\"100%\"/\u003e\n\u003c/details\u003e\n\n\u003cdetails close\u003e\n  \u003csummary\u003eImage Inpainting \u0026 Extrapolation\u003c/summary\u003e\n  \u003cimg src=\"./assets/demo_inpainting.jpg\" width=\"100%\"/\u003e\n\u003c/details\u003e\n\n\n## 📊 Quantitative Performance\n\u003cdetails open\u003e\n  \u003csummary\u003eGenEval Benchmark\u003c/summary\u003e\n  \u003cimg src=\"./assets/GenEval_benchmark.png\" width=\"100%\"/\u003e\n\u003c/details\u003e\n\n\n\u003cdetails close\u003e\n  \u003csummary\u003eDPG Benchmark\u003c/summary\u003e\n  \u003cimg src=\"./assets/DPG_benchmark.png\" width=\"100%\"/\u003e\n\u003c/details\u003e\n\n\u003cdetails close\u003e\n  \u003csummary\u003eOneIG-EN Benchmark\u003c/summary\u003e\n  \u003cimg src=\"./assets/OneIG-EN_benchmark.png\" width=\"100%\"/\u003e\n\u003c/details\u003e\n\n\n\u003cdetails close\u003e\n  \u003csummary\u003eTIIF Benchmark\u003c/summary\u003e\n  \u003cimg src=\"./assets/TIIF_benchmark.png\" width=\"100%\"/\u003e\n\u003c/details\u003e\n\n\u003cdetails close\u003e\n  \u003csummary\u003eImage-to-Image Benchmark\u003c/summary\u003e\n  \u003cimg src=\"./assets/i2i_benchmark.png\" width=\"100%\"/\u003e\n\u003c/details\u003e\n\n\u003cdetails close\u003e\n  \u003csummary\u003eImage Understanding Benchmark\u003c/summary\u003e\n  \u003cimg src=\"./assets/understanding_benchmark.png\" width=\"100%\"/\u003e\n\u003c/details\u003e\n\n## 🚀 Sampling Speed Analysis\n- Since text generation is performed in a block-wise manner, unlike image generation which uses a single global decoding step, its speed is influenced by both the number of blocks and the number of steps. Therefore, the speed improvement of image understanding is not as significant as that of image generation.\n\n- **Lumina-DiMOO Settings**: For image generation, we sample 64 steps. For image understanding, we set the block length to 256 and the number of sampling steps to 128.\n\n\u003cimg src=\"./assets/speed_comparison.png\" width=\"100%\"/\u003e\n\n\n## 📌 Quick Start\n### ⚙️ Installation\n#### 1. Create a conda environment\n```\ngit clone https://github.com/Alpha-VLLM/Lumina-DiMOO.git \u0026\u0026 cd Lumina-DiMOO\nconda create -n lumina_dimoo python=3.10 -y\nconda activate lumina_dimoo\n```\n#### 2. Install  dependencies\n```\npip install -r requirements.txt\n```\n\n### 🧨 How to Fine-Tuning Lumina-DiMOO\n#### Step 1: Pre-extract discrete codes of training images.\nThe final format after specific processing can refer to the sample json file ``assets/mmu_sample.json`` and ``assets/t2i_sample.json``.\n```\nbash pre_tokenizer/run_pre_token.sh\n```\n#### Step 2: Train Lumina-DiMOO model.\n```\nbash train/train.sh\n```\n\n### 🚗 Text-to-Image Generation Inference\n#### 1. Normal Sampling\n```\npython inference/inference_t2i.py\\\n    --checkpoint Alpha-VLLM/Lumina-DiMOO \\\n    --prompt \"A striking photograph of a glass of orange juice on a wooden kitchen table, capturing a playful moment. The orange juice splashes out of the glass and forms the word \\\"Smile\\\" in a whimsical, swirling script just above the glass. The background is softly blurred, revealing a cozy, homely kitchen with warm lighting and a sense of comfort.\" \\\n    --height 768 \\\n    --width 1536 \\\n    --timesteps 64 \\\n    --cfg_scale 4.0 \\\n    --seed 65513 \\\n    --vae_ckpt Alpha-VLLM/Lumina-DiMOO \\\n    --output_dir output/results_text_to_image\n```\n#### 2. DDP Sampling\nTo support large-scale sampling/testing, we provide additional ddp sampling scripts that support multi-GPU parallel sampling.\n```\ntorchrun --nproc_per_node=8 inference/inference_t2i_ddp.py \\\n    --checkpoint Alpha-VLLM/Lumina-DiMOO \\\n    --prompt_path /path/to/prompts.jsonl \\\n    --height 1024 \\\n    --width 1024 \\\n    --timesteps 64 \\\n    --cfg_scale 4.0 \\\n    --vae_ckpt Alpha-VLLM/Lumina-DiMOO \\\n    --output_dir output/results_image_to_image_ddp \\\n    --output_json output/results_image_to_image_ddp/results.json\n```\n#### 3. Faster Sampling with Cache\n- Add `--use-cache` to accelerate sampling through max logit-based cache (ML-Cache). The efficiency-quality tradeoff can be tuned by `cache_ratio` (in `(0,1)`; the higher the faster), `warmup_ratio` (in `[0,1)`; the lower the faster), and `refresh_interval` (in `(1, timesteps-int(warmup_ratio*timesteps)-1]`; the higher the faster). \n```\npython inference/inference_t2i.py\\\n    --checkpoint Alpha-VLLM/Lumina-DiMOO \\\n    --prompt \"A striking photograph of a glass of orange juice on a wooden kitchen table, capturing a playful moment. The orange juice splashes out of the glass and forms the word \\\"Smile\\\" in a whimsical, swirling script just above the glass. The background is softly blurred, revealing a cozy, homely kitchen with warm lighting and a sense of comfort.\" \\\n    --height 768 \\\n    --width 1536 \\\n    --timesteps 64 \\\n    --cfg_scale 4.0 \\\n    --seed 65513 \\\n    --vae_ckpt Alpha-VLLM/Lumina-DiMOO \\\n    --output_dir output/results_text_to_image_usecache \\\n    --use-cache \\\n    --cache_ratio 0.9 \\\n    --warmup_ratio 0.3 \\\n    --refresh_interval 5\n```\n\n- We provide the inference time and GPU memory on one A800 as a reference:\n\n| Method               | Inference Time | Inference GPU Memory |\n|----------------------|--------|----------|\n| Lumina-DiMOO      | 58.2s     | 38.9 GB  |\n| + ML-Cache        | 32.2s     | 45.9 GB  |\n\n### 🌟 Image-to-Image Inference\n \n#### 1. Controllable Generation: \"hed_control\", \"depth_control\", \"openpose_control\", \"subject_driven\".\n\n```\npython inference/inference_i2i.py \\\n    --checkpoint Alpha-VLLM/Lumina-DiMOO \\\n    --prompt \"A functional wooden printer stand.Nestled next to a brick wall in a bustling city street, it stands firm as pedestrians hustle by, illuminated by the warm glow of vintage street lamps.\" \\\n    --image_path examples/example_2.jpg \\\n    --edit_type depth_control \\\n    --timesteps 64 \\\n    --cfg_scale 2.5 \\\n    --cfg_img 4.0 \\\n    --vae_ckpt Alpha-VLLM/Lumina-DiMOO \\\n    --output_dir output/results_image_to_image\n```\n\n#### 2. Subject-Driven Generation.\n```\npython inference/inference_i2i.py \\\n    --checkpoint Alpha-VLLM/Lumina-DiMOO \\\n    --prompt \"A creamy, rich-flavored dark beverage.Captured in a bustling urban street at twilight, this item is placed on an outdoor café table, as city lights begin to twinkle and passersby create a lively atmosphere.\" \\\n    --image_path examples/example_3.jpg \\\n    --edit_type subject_driven \\\n    --timesteps 64 \\\n    --cfg_scale 2.5 \\\n    --cfg_img 4.0 \\\n    --vae_ckpt Alpha-VLLM/Lumina-DiMOO \\\n    --output_dir output/results_image_to_image\n```\n\n#### 3. Image Editing: \"edit_add\", \"edit_remove\", \"edit_replace\", \"edit_background\", \"edit_text_transfer\".\n```\npython inference/inference_i2i.py \\\n    --checkpoint Alpha-VLLM/Lumina-DiMOO \\\n    --prompt \"Add a beige shed with brown trim and double doors with a diamond pattern in the center-right, occupying more than a third of the image.\" \\\n    --image_path examples/example_4.png \\\n    --edit_type edit_add \\\n    --timesteps 64 \\\n    --cfg_scale 2.5 \\\n    --cfg_img 4.0 \\\n    --vae_ckpt Alpha-VLLM/Lumina-DiMOO \\\n    --output_dir output/results_image_to_image\n```\n\n#### 4. Style Transfer (An Image as Style Reference)\n```\npython inference/inference_i2i.py \\\n    --checkpoint Alpha-VLLM/Lumina-DiMOO \\\n    --prompt \"Transform the current image into the style of the provided image.\" \\\n    --image_path examples/example_5.png \\\n    --ref_image_path examples/example_5_style.png \\\n    --edit_type image_ref_transfer \\\n    --timesteps 64 \\\n    --cfg_scale 2.5 \\\n    --cfg_img 4.0 \\\n    --vae_ckpt Alpha-VLLM/Lumina-DiMOO \\\n    --output_dir output/results_image_to_image\n```\n\n#### 5. Dense Prediction: \"canny_pred\", \"hed_pred\", \"depth_pred\", \"openpose_pred\", \"canny_control\".\n```\npython inference/inference_i2i.py \\\n    --checkpoint Alpha-VLLM/Lumina-DiMOO \\\n    --prompt \"Generate a canny edge map accroding to the image.\" \\\n    --image_path examples/example_1.png \\\n    --edit_type canny_pred \\\n    --timesteps 64 \\\n    --cfg_scale 2.5 \\\n    --cfg_img 4.0 \\\n    --vae_ckpt Alpha-VLLM/Lumina-DiMOO \\\n    --output_dir output/results_image_to_image\n```\n\n### 🏃 Image Inpainting \u0026 Extrapolation Inference\n\n#### 1. Image Inpainting\n```\npython inference/inference_t2i.py\\\n    --checkpoint Alpha-VLLM/Lumina-DiMOO \\\n    --prompt \"Porsche showroom. Make there be a Porsche logo on the back wall behind the car.\" \\\n    --painting_mode inpainting \\\n    --painting_image examples/example_8.png \\\n    --mask_h_ratio 0.5 \\\n    --mask_w_ratio 0.5 \\\n    --timesteps 64 \\\n    --cfg_scale 4.0 \\\n    --seed 65513 \\\n    --vae_ckpt Alpha-VLLM/Lumina-DiMOO \\\n    --output_dir output/results_text_to_image\n```\n\n#### 2. Image Extrapolation\n```\npython inference/inference_t2i.py\\\n    --checkpoint Alpha-VLLM/Lumina-DiMOO \\\n    --prompt \"A photograph showcasing a pale gold moon, partially veiled by wispy cirrus clouds, dominating a dramatic twilight sky. The moon's soft glow reflects on the tranquil surface of a lake below, creating a shimmering mirror effect, while a small wooden rowboat gently bobs on the water's edge. Dark silhouettes of tall, ancient pine trees encircle the lake, their branches reaching towards the sky like skeletal fingers, as a gentle mist hangs low, diffusing the moonlight and adding a sense of serene mystery. The scene is bathed in soft, cool lighting, creating an ethereal and captivating atmosphere.\" \\\n    --painting_mode outpainting \\\n    --painting_image examples/example_7.png \\\n    --mask_h_ratio 1 \\\n    --mask_w_ratio 0.2 \\\n    --timesteps 64 \\\n    --cfg_scale 4.0 \\\n    --seed 65513 \\\n    --vae_ckpt Alpha-VLLM/Lumina-DiMOO \\\n    --output_dir output/results_text_to_image\n```\n\n\n### ⚡️ Image Understanding Inference\n```\npython inference/inference_mmu.py \\\n    --checkpoint Alpha-VLLM/Lumina-DiMOO \\\n    --prompt \"Please describe this image.\" \\\n    --image_path examples/example_6.jpg \\\n    --steps 128 \\\n    --gen_length 128 \\\n    --block_length 32 \\\n    --vae_ckpt Alpha-VLLM/Lumina-DiMOO \\\n    --output_dir output/outputs_text_understanding\n```\n\n\n## 💬 Discussion\nYou can reach us with this WeChat QR code!\n\u003cp align=\"left\"\u003e\n \u003cimg src=\"./assets/wechat.jpeg\" width=\"35%\"/\u003e\n \u003cbr\u003e\n\u003c/p\u003e\n\n## 📜 Acknowledgements\nThis work was also supported and implemented by [MindSpeed MM](https://gitee.com/ascend/MindSpeed-MM), an open-source training framework for large-scale multimodal models designed for distributed training, developed and maintained by Huawei's Computing Product Line. Specifically Optimized for Huawei‘s Ascend AI chips, MindSpeed MM offers comprehensive support for distributed training and is tailored for a wide range of multimodal tasks.\n\n## 🌟 Star History\n\n[![Star History Chart](https://api.star-history.com/svg?repos=Alpha-VLLM/Lumina-DiMOO\u0026type=Date)](https://www.star-history.com/#Alpha-VLLM/Lumina-DiMOO\u0026Date)\n\n## 📖 BibTeX\n```\n@article{xin2025lumina,\n  title={Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding},\n  author={Xin, Yi and Qin, Qi and Luo, Siqi and Zhu, Kaiwen and Yan, Juncheng and Tai, Yan and Lei, Jiayi and Cao, Yuewen and Wang, Keqi and Wang, Yibin and others},\n  journal={arXiv preprint arXiv:2510.06308},\n  year={2025}\n}\n```\n\n\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Falpha-vllm%2Flumina-dimoo","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Falpha-vllm%2Flumina-dimoo","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Falpha-vllm%2Flumina-dimoo/lists"}