{"id":56458,"url":"https://github.com/LMD0311/Awesome-World-Model","name":"Awesome-World-Model","description":"Collect some World Models for Autonomous Driving (and Robotic, etc.) papers. ","projects_count":955,"last_synced_at":"2026-09-25T16:00:28.233Z","repository":{"id":215072193,"uuid":"738046322","full_name":"LMD0311/Awesome-World-Model","owner":"LMD0311","description":"Collect some World Models for Autonomous Driving (and Robotic, etc.) papers. ","archived":false,"fork":false,"pushed_at":"2026-08-10T16:14:08.000Z","size":278,"stargazers_count":2208,"open_issues_count":0,"forks_count":87,"subscribers_count":60,"default_branch":"main","last_synced_at":"2026-08-17T00:10:02.053Z","etag":null,"topics":["artificial-intelligence","artificial-intelligence-algorithms","autonomous-driving","autonomous-vehicles","awesome","computer-vision","deep-learning","future-predict","robotics","world-model"],"latest_commit_sha":null,"homepage":"https://arxiv.org/abs/2502.10498","language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/LMD0311.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2024-01-02T09:38:21.000Z","updated_at":"2026-08-16T11:35:14.000Z","dependencies_parsed_at":"2026-06-04T20:00:23.040Z","dependency_job_id":null,"html_url":"https://github.com/LMD0311/Awesome-World-Model","commit_stats":null,"previous_names":["lmd0311/awesome-world-model"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/LMD0311/Awesome-World-Model","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/LMD0311%2FAwesome-World-Model","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/LMD0311%2FAwesome-World-Model/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/LMD0311%2FAwesome-World-Model/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/LMD0311%2FAwesome-World-Model/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/LMD0311","download_url":"https://codeload.github.com/LMD0311/Awesome-World-Model/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/LMD0311%2FAwesome-World-Model/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":341189360,"owners_count":37034043,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-08-22T15:14:58.755Z","status":"online","status_checked_at":"2026-09-05T02:00:06.064Z","response_time":108,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"created_at":"2024-02-26T20:13:12.532Z","updated_at":"2026-09-25T16:00:28.234Z","primary_language":null,"list_of_lists":false,"displayable":true,"categories":["Papers","Other World Model Paper","Workshop \u0026 Challenge","2024"],"sub_categories":["2024","2023","Technical blog or video","Survey","World model original paper","2022","2021","2020","2018","2025","2026"],"readme":"# Awesome World Models for Autonomous Driving\n\n[![Awesome](https://cdn.rawgit.com/sindresorhus/awesome/d7305f38d29fed78fa85652e3a63e154dd8e8829/media/badge.svg)](https://github.com/sindresorhus/awesome) [![arXiv](https://img.shields.io/badge/Arxiv-2502.10498-b31b1b.svg?logo=arXiv)](https://arxiv.org/abs/2502.10498)\n\nThis repo is used for recording, tracking, and benchmarking several recent World Models (for Autonomous Driving or Robotic) methods, as a supplement to our [**survey**](https://arxiv.org/abs/2502.10498).\n\nIf you find some ignored papers, **feel free to [*create pull requests*](https://github.com/LMD0311/Awesome-World-Model/blob/main/ContributionGuidelines.md), or [*open issues*](https://github.com/LMD0311/Awesome-World-Model/issues/new)**. Contributions in any form to make this list more comprehensive are welcome. 📣📣📣\n\nIf you find this repository useful, please consider  **giving us a star** 🌟 and a [**cite**](https://github.com/LMD0311/Awesome-World-Model#citation).\n\n## 📚 Citation\nIf you find this repository useful in your research, please kindly consider giving a star ⭐ and a citation:\n```bibtex\n@article{tu2025drivingworldmodel,\n  title={The Role of World Models in Shaping Autonomous Driving: A Comprehensive Survey}, \n  author={Tu, Sifan and Zhou, Xin and Liang, Dingkang and Jiang, Xingyu and Zhang, Yumeng and Li, Xiaofan and Bai, Xiang},\n  journal={Frontiers of Computer Science},\n  year={2026}\n}\n\n@article{zhao2026simwam,\n  title={SimWAM: A Simple World Action Model for End-to-End Autonomous Driving},\n  author={Zongchuang Zhao and Xin Zhou and Tianyang Xu and Zhengyang Sun and Kaixuan Zhou and Honglin Li and Dingkang Liang and Xiang Bai},\n  journal={arXiv preprint arXiv:2608.07468},\n  year={2026}\n}\n\n@inproceedings{zhou2025hermes,\n  title={HERMES: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation},\n  author={Zhou, Xin and Liang, Dingkang and Tu, Sifan and Chen, Xiwu and Ding, Yikang and Zhang, Dingyuan and Tan, Feiyang and Zhao, Hengshuang and Bai, Xiang},\n  booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision},\n  year={2025}\n}\n\n@inproceedings{liang2025UniFuture,\n  title={UniFuture: A 4D Driving World Model for Future Generation and Perception},\n  author={Liang, Dingkang and Zhang, Dingyuan and Zhou, Xin and Tu, Sifan and Feng, Tianrui and Li, Xiaofan and Zhang, Yumeng and Du, Mingyang and Tan, Xiao and Bai, Xiang},\n  booktitle={Proceedings of the IEEE International Conference on Robotics Automation},\n  year={2026}\n}\n\n@article{chen2026out,\n  title={Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models},\n  author={Chen, Kaijin and Liang, Dingkang and Zhou, Xin and Ding, Yikang and Liu, Xiaoqiang and Wan, Pengfei and Bai, Xiang},\n  journal={arXiv preprint arXiv:2603.25716},\n  year={2026}\n}\n\n@article{zhou2026hermespp,\n  title={HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation},\n  author={Zhou, Xin and Liang, Dingkang and Chen, Xiwu and Tan, Feiyang and Zhang, Dingyuan and Zhao, Hengshuang and Bai, Xiang},\n  journal={arXiv preprint arXiv:2604.28196},\n  year={2026}\n}\n\n@inproceedings{xiao2026divide,\n  title={Divide and Conquer: Decoupled Representation Alignment for Multimodal World Models},\n  author={Xiao, Junyuan and Liang, Dingkang and Zhou, Xin and Ye, Yixuan and Su, Tongtong and Yi, Guangmo and Xia, Bin and Lyu, Qiang and Shi, Shurui and Huang, Jun and Si, Jianlou and Yang, Wenming},\n  booktitle={European Conference on Computer Vision},\n  year={2026}\n}\n```\n\n## Workshop \u0026 Challenge\n\n- [`CVPR 25 Workshop \u0026 Challenge | OpenDriveLab`](https://opendrivelab.com/challenge25/#1x-wm) Track: World Model.\n\u003e A world model is a computer program that can imagine how the world evolves in response to an agent's behavior. It has the potential to solve general-purpose simulation and evaluation, enabling robots that are safe, reliable, and intelligent in a wide variety of scenarios.\n- [`World Model Bench @ CVPR'25`](https://worldmodelbench.github.io/) WorldModelBench: The 1st Workshop on Benchmarking World Models\n\u003e World models refer to predictive models of physical phenomena in the world surrounding us. These models are fundamental for Physical AI agents, enabling crucial capabilities such as decision-making, planning, and counterfactual analysis. Effective world models must integrate several key components, including perception, instruction following, controllability, physical plausibility, and future prediction.\n- [`CVPR 24 Workshop \u0026 Challenge | OpenDriveLab`](https://opendrivelab.com/challenge24/#predictive_world_model) Track #4: Predictive World Model.\n- [`CVPR 23 Workshop on Autonomous Driving`](https://cvpr23.wad.vision/) CHALLENGE 3: ARGOVERSE CHALLENGES, [3D Occupancy Forecasting](https://eval.ai/web/challenges/challenge-page/1977/overview) using the [Argoverse 2 Sensor Dataset](https://www.argoverse.org/av2.html#sensor-link). Predict the spacetime occupancy of the world for the next 3 seconds.\n\n## Papers\n\n### World model original paper\n\n- Using Occupancy Grids for Mobile Robot Perception and Navigation [[paper](http://www.sci.brooklyn.cuny.edu/~parsons/courses/3415-fall-2011/papers/elfes.pdf)]\n\n### Technical blog or video\n\n- **`Yann LeCun`**: A Path Towards Autonomous Machine Intelligence [[paper](https://openreview.net/pdf?id=BZ5a1r-kVsf)] [[Video](https://www.youtube.com/watch?v=OKkEdTchsiE)]\n- **`ICCV'25 workshop`** Keynote - Ashok Elluswamy, Tesla [[Video](https://www.bilibili.com/video/BV1oasHzTEe3/?vd_source=9ef518a6c349809d9fa8ab9427bd8b2c)]\n- **`CVPR'23 workshop`** Keynote - Ashok Elluswamy, Tesla [[Video](https://www.youtube.com/watch?v=6x-Xb_uT7ts)]\n- **`Wayve`** Introducing GAIA-1: A Cutting-Edge Generative AI Model for Autonomy [[blog](https://wayve.ai/thinking/introducing-gaia1/)] \n  \u003e World models are the basis for the ability to predict what might happen next, which is fundamentally important for autonomous driving. They can act as a learned simulator, or a mental “what if” thought experiment for model-based reinforcement learning (RL) or planning. By incorporating world models into our driving models, we can enable them to understand human decisions better and ultimately generalise to more real-world situations.\n  \n\n### Survey\n- The Role of World Models in Shaping Autonomous Driving: A Comprehensive Survey. **`FCS 26`** [[Paper](https://arxiv.org/abs/2502.10498)] [[Journal](https://journal.hep.com.cn/fcs/EN/home)]\n- Toward Unified Robot Learning: Bridging Representation, Vision-Language-Action, and World Models. **`TMLR 26`** [[Paper](https://arxiv.org/abs/2609.03927)]\n- A survey of world models for physical AI with uncertainty representation and control. **`Discover Artificial Intelligence 26`** [[Paper](https://doi.org/10.1007/s44163-026-02122-1)]\n- Rethinking World Models for Safety-Critical Embodied Systems. **`arXiv 26.9`** [[Paper](https://arxiv.org/abs/2609.03774)]\n- Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.28226)]\n- Medical world models in healthcare: foundations, applications, and challenges for trustworthy clinical translation. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.25242)]\n- From World Models to World Action Models: A Concise Tutorial for Robotics. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.00836)]\n- Multi-Agent Embodied Autonomous Driving: From V2X Information Exchange to Shared World Models. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.13840)]\n- A Tutorial on World Models and Physical AI. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.12783)]\n- Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses. **`arXiv 26.05`** [[Paper](https://arxiv.org/abs/2605.02900)] [[Code](https://github.com/x-zheng16/Awesome-Embodied-AI-Safety)]\n- Aligning Cyber Space with Physical World: A Comprehensive Survey on Embodied AI. **`TMECH 25`** [[Paper](https://arxiv.org/abs/2407.06886)] [[Code](https://github.com/HCPLab-SYSU/Embodied_AI_Paper_List)]\n- A Survey on Future Physical World Generation for Autonomous Driving. **`MMAsia 25`** [[Paper](https://dl.acm.org/doi/full/10.1145/3769748.3773345)]\n- A survey on multimodal large language models for autonomous driving. **`WACVW 24`** [[Paper](https://arxiv.org/abs/2311.12320)] [[Code](https://github.com/IrohXu/Awesome-Multimodal-LLM-Autonomous-Driving)]\n- World Models: The Safety Perspective. **`ISSREW`** [[Paper](https://arxiv.org/abs/2411.07690)]\n- Progressive Robustness-Aware World Models in Autonomous Driving: A Review and Outlook. **`techrXiv 25.11`** [[Paper](https://doi.org/10.36227/techrxiv.176523308.84756413/v1)] [[Project](https://github.com/MoyangSensei/AwesomeRobustDWM)]\n- A Survey of Unified Multimodal Understanding and Generation: Advances and Challenges. **`techrXiv 25.11`** [[Paper](https://www.techrxiv.org/doi/full/10.36227/techrxiv.176289261.16802577)]\n- Simulating the Visual World with Artificial Intelligence: A Roadmap. **`arXiv 25.11`** [[Paper](https://arxiv.org/abs/2511.08585)] [[Project](https://world-model-roadmap.github.io/)]\n- A Step Toward World Models: A Survey on Robotic Manipulation. **`arXiv 25.11`** [[Paper](https://arxiv.org/abs/2511.02097)]\n- A Comprehensive Survey on World Models for Embodied AI. **`arXiv 25.10`** [[Paper](https://arxiv.org/abs/2510.16732)] [[Project](https://github.com/Li-Zn-H/AwesomeWorldModels)]\n- The Safety Challenge of World Models for Embodied AI Agents: A Review. **`arXiv 25.10`** [[Paper](https://arxiv.org/abs/2510.05865)]\n- A Survey on World Models Grounded in Acoustic Physical Information. **`arXiv 25.09`** [[Paper](https://arxiv.org/abs/2506.13833)]\n- 3D and 4D World Modeling: A Survey. **`arXiv 25.09`** [[Paper](https://arxiv.org/abs/2509.07996)] [[Code](https://github.com/worldbench/survey)]\n- A Survey of Embodied World Models. **`25.09`** [[Paper](https://www.researchgate.net/publication/395713824_A_Survey_of_Embodied_World_Models)]\n- One Flight Over the Gap: A Survey from Perspective to Panoramic Vision. **`arXiv 25.09`** [[Paper](https://arxiv.org/abs/2509.04444)] [[Page](https://insta360-research-team.github.io/Survey-of-Panorama/)]\n- Edge General Intelligence Through World Models and Agentic AI: Fundamentals, Solutions, and Challenges. **`arXiv 25.08`** [[Paper](https://arxiv.org/abs/2508.09561)]\n- A Survey: Learning Embodied Intelligence from Physical Simulators and World Models. **`arXiv 25.07`** [[Paper](https://arxiv.org/abs/2507.00917)]\n- From 2D to 3D Cognition: A Brief Survey of General World Models. **`arXiv 25.06`** [[Paper](https://arxiv.org/abs/2506.20134)]\n- World Models for Cognitive Agents: Transforming Edge Intelligence in Future Networks. **`arXiv 25.05`** [[Paper](https://arxiv.org/abs/2506.00417)]\n- Exploring the Evolution of Physics Cognition in Video Generation: A Survey. **`arXiv 25.03`** [[Paper](https://arxiv.org/abs/2503.21765)] [[Code](https://github.com/minnie-lin/Awesome-Physics-Cognition-based-Video-Generation)]\n- A Survey of World Models for Autonomous Driving. **`arXiv 25.01`** [[Paper](https://arxiv.org/abs/2501.11260)]\n- Generative Physical AI in Vision: A Survey. **`arXiv 25.01`** [[Paper](https://arxiv.org/abs/2501.10928)] [[Code](https://github.com/BestJunYu/Awesome-Physics-aware-Generation)]\n- Understanding World or Predicting Future? A Comprehensive Survey of World Models. **`arXiv 24.11`** [[Paper](https://arxiv.org/abs/2411.14499)]\n- Exploring the Interplay Between Video Generation and World Models in Autonomous Driving: A Survey. **`arXiv 24.11`** [[Paper](https://arxiv.org/abs/2411.02914)]\n- Is Sora a World Simulator? A Comprehensive Survey on General World Models and Beyond. **`arXiv 24.5`** [[Paper](https://arxiv.org/abs/2405.03520)] [[Code](https://github.com/GigaAI-research/General-World-Models-Survey)]\n- World Models for Autonomous Driving: An Initial Survey. **`arXiv 24.3`** [[Paper](https://arxiv.org/abs/2403.02622)]\n\n### 2026\n- **SimWAM**: A Simple World Action Model for End-to-End Autonomous Driving. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.07468)] [[Code](https://github.com/H-EmbodVis/SimWAM)]\n- [**UniFuture**] UniFuture: A 4D Driving World Model for Future Generation and Perception. **`ICRA 26`** [[Paper](https://arxiv.org/abs/2503.13587)] [[Code](https://github.com/dk-liang/UniFuture)] [[Project](https://dk-liang.github.io/UniFuture/)]\n- **HERMES++**: Toward a Unified Driving World Model for 3D Scene Understanding and Generation. **`arXiv 26.5`** [[Paper](https://arxiv.org/abs/2604.28196)] [[Code](https://github.com/H-EmbodVis/HERMESV2)] [[Project](https://h-embodvis.github.io/HERMESV2/)]\n- **RAYNOVA**: Scale-Temporal Autoregressive World Modeling in Ray Space. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2602.20685)] [[Project](https://raynova-ai.github.io/)]\n- **WAM-Flow**: Parallel Coarse-to-Fine Motion Planning via Discrete Flow Matching for Autonomous Driving. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2512.06112)] [[Code](https://github.com/fudan-generative-vision/WAM-Flow)]\n- See Tomorrow, Act Today: Foresight-Driven Autonomous Driving. **`CVPR 26 Findings`** [[Paper](https://arxiv.org/abs/2605.07195)]\n- **ResWorld**: Temporal Residual World Model for End-to-End Autonomous Driving. **`ICLR 26`** [[Paper](https://arxiv.org/abs/2602.10884)] [[Code](https://github.com/mengtan00/ResWorld.git)]\n- **WorldRFT**: Latent World Model Planning with Reinforcement Fine-Tuning for Autonomous Driving.  **`AAAI 26`** [[Paper](https://arxiv.org/abs/2512.19133)]\n- **GaussianDWM**: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2512.23180)] [[Code](https://github.com/dtc111111/GaussianDWM)]\n- **DriveLaW**: Unifying Planning and Video Generation in a Latent Driving World. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2512.23421)]\n- Latent Chain-of-Thought World Modeling for End-to-End Autonomous Driving. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2512.10226)]\n- **GenieDrive**: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2512.12751)] [[Project](https://huster-yzy.github.io/geniedrive_project_page/)]\n- **WorldLens**: Full-Spectrum Evaluations of Driving World Models in Real World. **`CVPR 26 Oral`** [[Paper](https://arxiv.org/abs/2512.10958)] [[Project](https://worldbench.github.io/worldlens)]\n- **U4D**: Uncertainty-Aware 4D World Modeling from LiDAR Sequences. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2512.02982)]\n- **Think Before You Drive**: World Model-Inspired Multimodal Grounding for Autonomous Vehicles. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2512.03454)]\n- **SparseWorld-TC**: Trajectory-Conditioned Sparse Occupancy World Model. **`CVPR 26 Oral`** [[Paper](https://arxiv.org/abs/2511.22039)]\n- **MoVieDrive**: Urban Scene Synthesis with Multi-Modal Multi-View Video Diffusion Transformer. **`CVPR 26 Findings`** [[Paper](https://arxiv.org/abs/2508.14327)]\n- **LaGen**: Towards Autoregressive LiDAR Scene Generation. **`ECCV 26`** [[Paper](https://arxiv.org/abs/2511.21256)]\n- **OmniNWM**: Omniscient Driving Navigation World Models. **`ECCV 26`** [[Paper](https://arxiv.org/abs/2510.18313)] [[Project](https://arlo0o.github.io/OmniNWM/)]\n- Long-term Traffic Simulation via Structured Autoregressive Modeling. **`ECCV 26`** [[Paper](https://arxiv.org/abs/2606.31209)]\n- **Driver-WM**: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics Rollout. **`ECCV 26`** [[Paper](https://arxiv.org/abs/2605.05092)]\n- **GEM**: Generating LiDAR World Model via Deformable Mamba. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2605.07326)]\n- **UniDrive-WM**: Unified Understanding, Planning and Generation World Model For Autonomous Driving. **`ECCV 26`** [[Paper](https://arxiv.org/abs/2601.04453)] [[Project](https://unidrive-wm.github.io/UniDrive-WM)]\n- **MAD**: Motion Appearance Decoupling for efficient Driving World Models. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2601.09452)] [[Project](https://vita-epfl.github.io/MAD-World-Model/)]\n- **Conductor**: Scalable Edge-assisted Fusion and Path Prediction for Connected Autonomous Vehicles. **`arXiv 26.9`** [[Paper](https://arxiv.org/abs/2609.04364)]\n- **SV-WAM**: An Efficient Surround-View World-Action Model for End-to-End Autonomous Driving. **`arXiv 26.9`** [[Paper](https://arxiv.org/abs/2609.03602)]\n- **Drive-HWM**: Hierarchical World Models for Dynamic-Latent Guided Autonomous Driving. **`arXiv 26.9`** [[Paper](https://arxiv.org/abs/2609.03572)]\n- **StyleDrive**: Long-Horizon Consistent and Interaction-Aware World Models for Multi-Style End-to-End Driving. **`arXiv 26.9`** [[Paper](https://arxiv.org/abs/2609.03225)]\n- **PDT**: From Proxy Learning to Driving Decisions: A Transfer-Based Framework for Evaluating Future-Aware Autonomous Driving Planners. **`arXiv 26.9`** [[Paper](https://arxiv.org/abs/2609.02688)]\n- **SAGE**: Self-Aware Active Learning Enables Continual Improvement in Autonomous Driving. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.29772)]\n- **GeoWAM**: Visual Geometry World Action Models for Autonomous Driving. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.23486)] [[Project](https://yiren-lu.com/project_pages/geowam/)]\n- **WA-JEPA**: Rethinking the Video JEPA Paradigm for World-Action Modeling in Autonomous Driving. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.20974)]\n- **DA-WAM**: Decision-Aligned Future Latents for Driving World Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.19085)]\n- **DriveCache**: Action-Aware Caching for Driving World Model Inference. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.16354)]\n- **GaussianDWM++**: Language-Grounded 3D Gaussian Driving World Model for Unified Scene Understanding, Editing, and Multi-Modal Generation. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.16234)]\n- **BrainWAM**: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.12854)]\n- How Can Driving World Models Do Counterfactual Prediction? **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.11601)]\n- Toward the Cognitive--Physical Limits of Embodied Intelligence through a World-Model-Centric Autonomous Racing Agent. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.10618)]\n- **Dreamer-SAC**: Off-Policy Learning in Latent World Models for Sample-Efficient Autonomous Driving. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.10386)]\n- **Adaptive-WAM**: Quality-Guided Early-Exit Planning from Intermediate Video-Diffusion Features. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.06008)]\n- **muSync-GS**: Physics-Synchronized Driving Video Synthesis for Weather and Geometric Road Hazards. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.04412)]\n- **Auto-JEPA**: A Latent World Model of Continuous Intent for End-to-End Autonomous Driving. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.29031)]\n- **HyWorldVLA**: A Vision-Language-Action Model with Hybrid World Modeling for Autonomous Driving. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.20988)]\n- **GeoWorldAD**: Geometry World Action Model for Autonomous Driving. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.17521)]\n- **Orbis 2**: A Hierarchical World Model for Driving. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.15898)]\n- **M4World**: A Multi-view Multimodal Driving World Model for Interactive Object Manipulation and Minute-long Streaming. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.14005)]\n- Ego-Dynamics-Augmented World Model for Autonomous Driving with Zero-Shot Cross-Chassis Adaptation. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.13410)]\n- **LIDAR-AD**: A Decoder-Free Latent-Interaction Dreamer with Action-Residual Chains for Autonomous Driving. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.11964)]\n- Is Energy Guidance All You Need? Training-Free Norm Injection for Driving World Models. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.10781)]\n- World Models as Adversaries: Multi-Agent Self-Play Fine-Tuning for Robust Motion Planning. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.10630)]\n- **WCog-VLA**: A Dual-Level World-Cognitive Vision-Language-Action Model for End-to-End Autonomous Driving. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.08375)]\n- **CRISP**: A Spatiotemporal Camera-Radar Backbone for Driving via Forecasting-Based World-Model Pretraining. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.04541)]\n- Geographic Diversity Beats Data Volume for Cross-Domain Generalization in Zero-Label JEPA Driving World Models. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.04500)]\n- **ForgeDrive**: Bidirectional Cross-Conditioning for Unified Visual-Action Generation in Autonomous Driving. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.31226)]\n- **OWMDrive**: Causality-Aware End-to-End Autonomous Driving via 4D Occupancy World Model. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.30421)]\n- **LWDrive**: Layer-Wise World-Model-Guided Vision-Language Model Planning for Autonomous Driving. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.29879)]\n- **X-Mind**: Efficient Visual Chain-of-Thought via Predictive World Model for End-to-End Driving. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.28758)]\n- A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.28757)]\n- **CascadeOcc**: Rethinking 3D Occupancy World Models with Cascaded VQ Representations. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.27644)]\n- **ReWorld**: Learning Better Representations for World Action Models. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.27504)]\n- **BadDreamer**: Transferable Backdoor Attacks against Video World Models for Autonomous Driving. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.21172)]\n- **OmniDrive**: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.17536)]\n- **GraphWorld**: Long-Horizon Planning with World Models for End-to-End Autonomous Driving. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.16274)]\n- **CausalDrive**: Real-time Causal World Models for Autonomous Driving. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.15341)]\n- **ReactSim-Bench**: Benchmarking Reactive Behavior World Model Simulation in Autonomous Driving. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.14058)]\n- **VISA**: VLM-Guided Instance Semantic Auditing for 3D Occupancy World Models. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.13460)]\n- Diffusion Transformer World-Action Model for AV Scene Prediction. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.12987)]\n- **PLAN-S**: Bridging Planning with Latent Style Dynamics for Autonomous Driving World Models. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.06014)]\n- **Discrete-WAM**: Unified Discrete Vision-Action Token Editing for World-Policy Learning. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.05645)]\n- **NVIDIA OmniDreams**: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.03159)]\n- **Unified Driving Tokens**: Representation- and Geometry-Guided Discrete Tokenizer for Driving World Models and Planning. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.01935)]\n- **Xiaomi EV World Model**: A Joint World Model Integrating Reconstruction and Generation for Autonomous Driving. **`arXiv 26.5`** [[Paper](https://arxiv.org/abs/2605.18137)]\n- **DriveCtrl**: Conditioned Sim-to-Real Driving Video Generation. **`arXiv 26.5`** [[Paper](https://arxiv.org/abs/2605.15116)]\n- **CoWorld-VLA**: Thinking in a Multi-Expert World Model for Autonomous Driving. **`arXiv 26.5`** [[Paper](https://arxiv.org/abs/2605.10426)]\n- **HEAT**: Heterogeneous End-to-End Autonomous Driving via Trajectory-Guided World Models. **`arXiv 26.5`** [[Paper](https://arxiv.org/abs/2605.19631)]\n- **X-World**: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.19979)]\n- **Vega**: Learning to Drive with Natural Language Instructions. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.25741)] [[Code](https://github.com/zuosc19/Vega)]\n- **DCARL**: A Divide-and-Conquer Framework for Autoregressive Long-Trajectory Video Generation. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.24835)] [[Project](https://junyiouy.github.io/projects/dcarl)]\n- **DreamerAD**: Efficient Reinforcement Learning via Latent World Model for Autonomous Driving. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.24587)]\n- **Latent-WAM**: Latent World Action Modeling for End-to-End Autonomous Driving. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.24581)]\n- Toward Physically Consistent Driving Video World Models under Challenging Trajectories. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.24506)] [[Project](https://wm-research.github.io/PhyGenesis/)]\n- **FAR-Drive**: Frame-AutoRegressive Video Generation in Closed-Loop Autonomous Driving. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.14938)]\n- **WorldVLM**: Combining World Model Forecasting and Vision-Language Reasoning. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.14497)]\n- [**WorldDrive**] Bridging Scene Generation and Planning: Driving with World Model via Unifying Vision and Motion Representation. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.14948)] [[Code](https://github.com/TabGuigui/WorldDrive)]\n- **DynVLA**: Learning World Dynamics for Action Reasoning in Autonomous Driving. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.11041)]\n- Latent World Models for Automated Driving: A Unified Taxonomy, Evaluation Framework, and Open Challenges. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.09086)]\n- **SAMoE-VLA**: A Scene Adaptive Mixture-of-Experts Vision-Language-Action Model for Autonomous Driving. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.08113)]\n- Kinematics-Aware Latent World Models for Data-Efficient Autonomous Driving. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.07264)]\n- **ShareVerse**: Multi-Agent Consistent Video Generation for Shared World Modeling. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.02697)]\n- Risk-Aware World Model Predictive Control for Generalizable End-to-End Autonomous Driving. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.23259)]\n- A Mechanistic View on Video Generation as World Models: State and Dynamics. **`arXiv 26.1`** [[Paper](https://arxiv.org/abs/2601.17067)]\n- **Drive-JEPA**: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving. **`arXiv 26.1`** [[Paper](https://arxiv.org/abs/2601.22032)]\n- **DrivingGen**: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving. **`arXiv 26.1`** [[Paper](https://arxiv.org/abs/2601.01528)] [[Project](https://drivinggen-bench.github.io/)]\n\n### 2025\n- **HERMES**: A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation.  **`ICCV 25`** [[Paper](https://arxiv.org/abs/2501.14729)] [[Code](https://github.com/LMD0311/HERMES)] [[Project](https://lmd0311.github.io/HERMES/)]\n- [**FSDrive**] FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving. **`NeurIPS 25`** [[Paper](https://arxiv.org/abs/2505.17685)] [[Code](https://github.com/MIV-XJTU/FSDrive)]\n- **DINO-Foresight**: Looking into the Future with DINO. **`NeurIPS 25`** [[Paper](https://arxiv.org/abs/2412.11673)] [[Code](https://github.com/Sta8is/DINO-Foresight)]\n- **From Forecasting to Planning**: Policy World Model for Collaborative State-Action Prediction. **`NeurIPS 25`** [[Paper](https://arxiv.org/abs/2510.19654)] [[Code](https://github.com/6550Zhao/Policy-World-Model)]\n- **InfiniCube**: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models.  **`ICCV 25`** [[Paper](https://arxiv.org/abs/2412.03934)] [[Project](https://research.nvidia.com/labs/toronto-ai/infinicube/)]\n- **DiST-4D**: Disentangled Spatiotemporal Diffusion with Metric Depth for 4D Driving Scene Generation.  **`ICCV 25`** [[Paper](https://arxiv.org/abs/2503.15208)] [[Project](https://royalmelon0505.github.io/DiST-4D/)]\n- **Epona**: Autoregressive Diffusion World Model for Autonomous Driving.  **`ICCV 25`** [[Paper](https://arxiv.org/abs/2506.24113)] [[Code](https://github.com/Kevin-thu/Epona/)]\n- **UniOcc**: A Unified Benchmark for Occupancy Forecasting and Prediction in Autonomous Driving. **`ICCV 25`** [[Paper](https://arxiv.org/abs/2503.24381)] [[Code](https://uniocc.github.io/)]\n- **DriVerse**: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion Alignment.  **`ACM MM 25`** [[Paper](https://arxiv.org/abs/2504.19614)] [[Code](https://github.com/shalfun/DriVerse)]\n- **OmniGen**: Unified Multimodal Sensor Generation for Autonomous Driving. **`ACM MM 25`** [[Paper](https://arxiv.org/abs/2512.14225)]\n- **World4Drive**: End-to-End Autonomous Driving via Intention-aware Physical Latent World Model.  **`ICCV 25`** [[Paper](https://arxiv.org/abs/2507.00603)]\n- [**PIWM**] Dream to Drive with Predictive Individual World Model.  **`TIV 25`** [[Paper](https://arxiv.org/abs/2501.16733)]  [[Code](https://github.com/gaoyinfeng/PIWM)]\n- **DriveDreamer4D**: World Models Are Effective Data Machines for 4D Driving Scene Representation. **`CVPR 25`** [[Paper](https://arxiv.org/abs/2410.13571)] [[Project Page](https://drivedreamer4d.github.io/)]\n- **GaussianWorld**: Gaussian World Model for Streaming 3D Occupancy Prediction. **`CVPR 25`** [[Paper](https://arxiv.org/abs/2412.10373)] [[Code](https://github.com/zuosc19/GaussianWorld)]\n- **ReconDreamer**: Crafting World Models for Driving Scene Reconstruction via Online Restoration. **`CVPR 25`** [[Paper](https://arxiv.org/abs/2411.19548)] [[Code](https://github.com/GigaAI-research/ReconDreamer)]\n- **FUTURIST**: Advancing Semantic Future Prediction through Multimodal Visual Sequence Transformers. **`CVPR 25`** [[Paper](https://arxiv.org/abs/2501.08303)] [[Code](https://github.com/Sta8is/FUTURIST)]\n- **MaskGWM**: A Generalizable Driving World Model with Video Mask Reconstruction.  **`CVPR 25`** [[Paper](https://arxiv.org/abs/2502.11663)] [[Code](https://github.com/SenseTime-FVG/OpenDWM)]\n- **UniScene**: Unified Occupancy-centric Driving Scene Generation. **`CVPR 25`** [[Paper](https://arxiv.org/abs/2412.05435)] [[Project](https://arlo0o.github.io/uniscene/)]\n- **DrivingGPT**: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers. **`CVPR 25`** [[Paper](https://arxiv.org/abs/2412.18607)] [[Project](https://rogerchern.github.io/DrivingGPT/)]\n- **GEM**: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control. **`CVPR 25`** [[Paper](https://arxiv.org/abs/2412.11198)] [[Project](https://vita-epfl.github.io/GEM.github.io/)]\n- [**UMGen**] Generating Multimodal Driving Scenes via Next-Scene Prediction. **`CVPR 25`** [[Paper](https://arxiv.org/abs/2503.14945)] [[Project](https://yanhaowu.github.io/UMGen/)] [[Code](https://github.com/YanhaoWu/UMGen/)]\n- **DIO**: Decomposable Implicit 4D Occupancy-Flow World Model. **`CVPR 25`** [[Paper](https://openaccess.thecvf.com/content/CVPR2025/html/Diehl_DIO_Decomposable_Implicit_4D_Occupancy-Flow_World_Model_CVPR_2025_paper.html)]\n- **SceneDiffuser++**: City-Scale Traffic Simulation via a Generative World Model. **`CVPR 25`** [[Paper](https://openaccess.thecvf.com/content/CVPR2025/html/Tan_SceneDiffuser_City-Scale_Traffic_Simulation_via_a_Generative_World_Model_CVPR_2025_paper.html)]\n- **DynamicCity**: Large-Scale LiDAR Generation from Dynamic Scenes  **`ICLR 25`** [[Paper](https://arxiv.org/abs/2410.18084)] [[Code](https://github.com/3DTopia/DynamicCity)]\n- **AdaWM**: Adaptive World Model based Planning for Autonomous Driving.  **`ICLR 25`** [[Paper](https://arxiv.org/abs/2501.13072)]\n- **OccProphet**: Pushing Efficiency Frontier of Camera-Only 4D Occupancy Forecasting with Observer-Forecaster-Refiner Framework.  **`ICLR 25`** [[Paper](https://arxiv.org/abs/2502.15180)] [[Code](https://github.com/JLChen-C/OccProphet)]\n- [**PreWorld**] Semi-Supervised Vision-Centric 3D Occupancy World Model for Autonomous Driving.  **`ICLR 25`** [[Paper](https://arxiv.org/abs/2502.07309)] [[Code](https://github.com/getterupper/PreWorld)]\n- [**SSR**] Does End-to-End Autonomous Driving Really Need Perception Tasks? **`ICLR 25`** [[Paper](https://arxiv.org/abs/2409.18341)] [[Code](https://github.com/PeidongLi/SSR)]\n- **Occ-LLM**: Enhancing Autonomous Driving with Occupancy-Based Large Language Models.  **`ICRA 25`** [[Paper](https://arxiv.org/abs/2502.06419)]\n- **STAGE**: A Stream-Centric Generative World Model for Long-Horizon Driving-Scene Simulation.  **`IROS 25`** [[Paper](https://arxiv.org/abs/2506.13138)] [[Project](https://4dvlab.github.io/STAGE/)]\n- **Drive\u0026Gen**: Co-Evaluating End-to-End Driving and Video Generation Models.  **`IROS 25`** [[Paper](https://arxiv.org/abs/2510.06209)]\n- Learning to Generate 4D LiDAR Sequences. **`ICCVW 25`** [[Paper](https://arxiv.org/abs/2509.11959)]\n- World model-based end-to-end scene generation for accident anticipation in autonomous driving. **`Communications Engineering 25`** [[Paper](https://www.nature.com/articles/s44172-025-00474-7)]\n- World Models for Autonomous Navigation of Terrestrial Robots from LIDAR Observations. **`JIFS 25`** [[Paper](https://arxiv.org/abs/2512.03429)]\n- **InDRiVE**: Reward-Free World-Model Pretraining for Autonomous Driving via Latent Disagreement. **`arXiv 25.12`** [[Paper](https://arxiv.org/abs/2512.18850)]\n- **UniUGP**: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving. **`arXiv 25.12`** [[Paper](https://arxiv.org/abs/2512.09864)] [[Project](https://seed-uniugp.github.io/)]\n- **MindDrive**: An All-in-One Framework Bridging World Models and Vision-Language Model for End-to-End Autonomous Driving. **`arXiv 25.12`** [[Paper](https://arxiv.org/abs/2512.04441)]\n- **RadarGen**: Automotive Radar Point Cloud Generation from Cameras. **`arXiv 25.12`** [[Paper](https://arxiv.org/abs/2512.17897)] [[Project](https://radargen.github.io/)]\n- Vehicle Dynamics Embedded World Models for Autonomous Driving. **`arXiv 25.12`** [[Paper](https://arxiv.org/abs/2512.02417)]\n- **LiSTAR**: Ray-Centric World Models for 4D LiDAR Sequences in Autonomous Driving. **`arXiv 25.11`** [[Paper](https://arxiv.org/abs/2511.16049)] [[Project](https://ocean-luna.github.io/LiSTAR.github.io/)]\n- **OpenTwinMap**: An Open-Source Digital Twin Generator for Urban Autonomous Driving. **`arXiv 25.11`** [[Paper](https://arxiv.org/abs/2511.21925)]\n- **AD-R1**: Closed-Loop Reinforcement Learning for End-to-End Autonomous Driving with Impartial World Models. **`arXiv 25.11`** [[Paper](https://arxiv.org/abs/2511.20325)]\n- **CorrectAD**: A Self-Correcting Agentic System to Improve End-to-end Planning in Autonomous Driving. **`arXiv 25.11`** [[Paper](https://arxiv.org/abs/2511.13297)]\n- [**UniScenev2**] Scaling Up Occupancy-centric Driving Scene Generation: Dataset and Method. **`arXiv 25.10`** [[Paper](https://arxiv.org/abs/2510.22973)]\n- Vision-Centric 4D Occupancy Forecasting and Planning via Implicit Residual World Models. **`arXiv 25.10`** [[Paper](https://arxiv.org/abs/2510.16729)]\n- **SparseWorld**: A Flexible, Adaptive, and Efficient 4D Occupancy World Model Powered by Sparse and Dynamic Queries. **`arXiv 25.10`** [[Paper](https://arxiv.org/abs/2510.17482)] [[Code](https://github.com/MSunDYY/SparseWorld)]\n- [**ORAD-3D**] Advancing Off-Road Autonomous Driving: The Large-Scale ORAD-3D Dataset and Comprehensive Benchmarks. **`arXiv 25.10`** [[Paper](https://arxiv.org/abs/2510.16500)] [[Code](https://github.com/chaytonmin/ORAD-3D)]\n- [**Dream4Drive**] Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks. **`arXiv 25.10`** [[Paper](https://arxiv.org/abs/2510.19195)] [[Project](https://wm-research.github.io/Dream4Drive/)]\n- **DriveVLA-W0**: World Models Amplify Data Scaling Law in Autonomous Driving. **`arXiv 25.10`** [[Paper](https://arxiv.org/abs/2510.12796)]\n- **CoIRL-AD**: Collaborative-Competitive Imitation-Reinforcement Learning in Latent World Models for Autonomous Driving. **`arXiv 25.10`** [[Paper](https://arxiv.org/abs/2510.12560)]\n- **CVD-STORM**: Cross-View Video Diffusion with Spatial-Temporal Reconstruction Model for Autonomous Driving. **`arXiv 25.10`** [[Paper](https://arxiv.org/abs/2510.07944)]\n- [**PhiGensis**] 4D Driving Scene Generation With Stereo Forcing. **`arXiv 25.9`** [[Paper](https://arxiv.org/abs/2509.20251)] [[Project](https://jiangxb98.github.io/PhiGensis/)]\n- **TeraSim-World**: Worldwide Safety-Critical Data Synthesis for End-to-End Autonomous Driving. **`arXiv 25.9`** [[Paper](https://arxiv.org/abs/2509.13164)]\n- **OccTENS**: 3D Occupancy World Model via Temporal Next-Scale Prediction. **`arXiv 25.9`** [[Paper](https://arxiv.org/abs/2509.03887)]\n- [**G^2Editor**] Realistic and Controllable 3D Gaussian-Guided Object Editing for Driving Video Generation. **`arXiv 25.8`** [[Paper](https://arxiv.org/abs/2508.20471)]\n- **LSD-3D**: Large-Scale 3D Driving Scene Generation with Geometry Grounding. **`arXiv 25.8`** [[Paper](https://arxiv.org/abs/2508.19204)] [[Project](https://princeton-computational-imaging.github.io/LSD-3D/)]\n- Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation. **`arXiv 25.8`** [[Paper](https://arxiv.org/abs/2508.16512)]\n- **ImagiDrive**: A Unified Imagination-and-Planning Framework for Autonomous Driving. **`arXiv 25.8`** [[Paper](https://arxiv.org/abs/2508.11428)] [[Code](https://github.com/fudan-zvg/ImagiDrive)]\n- **LiDARCrafter**: Dynamic 4D World Modeling from LiDAR Sequences. **`arXiv 25.8`** [[Paper](https://arxiv.org/abs/2508.03692)] [[Project](https://lidarcrafter.github.io/)]\n- **FASTopoWM**: Fast-Slow Lane Segment Topology Reasoning with Latent World Models. **`arXiv 25.7`** [[Paper](https://arxiv.org/abs/2507.23325)]\n- World Model-Based End-to-End Scene Generation for Accident Anticipation in Autonomous Driving. **`arXiv 25.7`** [[Paper](https://arxiv.org/abs/2507.12762)]\n- **Orbis**: Overcoming Challenges of Long-Horizon Prediction in Driving World Models.  **`arXiv 25.7`** [[Paper](https://arxiv.org/abs/2507.13162)] [[Code](https://lmb-freiburg.github.io/orbis.github.io/)]\n- **I2 -World**: Intra-Inter Tokenization for Efficient Dynamic 4D Scene Forecasting.  **`arXiv 25.7`** [[Paper](https://arxiv.org/abs/2507.09144)] [[Code](https://github.com/lzzzzzm/II-World)]\n- **NRSeg**: Noise-Resilient Learning for BEV Semantic Segmentation via Driving World Models.  **`arXiv 25.7`** [[Paper](https://arxiv.org/abs/2507.04002)] [[Code](https://github.com/lynn-yu/NRSeg)]\n- Towards foundational LiDAR world models with efficient latent flow matching.  **`arXiv 25.6`** [[Paper](https://arxiv.org/abs/2506.23434)]\n- **ReSim**: Reliable World Simulation for Autonomous Driving.  **`arXiv 25.6`** [[Paper](https://arxiv.org/abs/2506.09981)] [[Project](https://opendrivelab.com/ReSim)]\n- **Cosmos-Drive-Dreams**: Scalable Synthetic Driving Data Generation with World Foundation Models.  **`arXiv 25.6`** **`NVIDIA`** [[Paper](https://arxiv.org/abs/2506.09042)] [[Project](https://research.nvidia.com/labs/toronto-ai/cosmos_drive_dreams/)]\n- **Dreamland**: Controllable World Creation with Simulator and Generative Models.  **`arXiv 25.6`** [[Paper](https://arxiv.org/abs/2506.08006)] [[Project](https://metadriverse.github.io/dreamland/)]\n- **LongDWM**: Cross-Granularity Distillation for Building a Long-Term Driving World Model.  **`arXiv 25.6`** [[Paper](https://arxiv.org/abs/2506.01546)] [[Code](https://wang-xiaodong1899.github.io/longdwm/)]\n- **ProphetDWM**: ProphetDWM: A Driving World Model for Rolling Out Future Actions and Videos. **`arXiv 25.5`** [[Paper](https://arxiv.org/abs/2505.18650)]\n- **GeoDrive**: 3D Geometry-Informed Driving World Model with Precise Action Control.  **`arXiv 25.5`** [[Paper](https://arxiv.org/abs/2505.22421)] [[Code](https://github.com/antonioo-c/GeoDrive)]\n- **DriveX**: Omni Scene Modeling for Learning Generalizable World Knowledge in\nAutonomous Driving.  **`arXiv 25.5`** [[Paper](https://arxiv.org/abs/2505.19239)]\n- **VL-SAFE**: Vision-Language Guided Safety-Aware Reinforcement Learning with World Models for Autonomous Driving.  **`arXiv 25.5`** [[Paper](https://arxiv.org/abs/2505.16377)] [[Project](https://ys-qu.github.io/vlsafe-website/)]\n- **Raw2Drive**: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2).  **`arXiv 25.5`** [[Paper](https://arxiv.org/abs/2505.16394)]\n- [**RAMBLE**] From Imitation to Exploration: End-to-end Autonomous Driving based on World Model.  **`arXiv 25.4`** [[Paper](https://arxiv.org/abs/2410.02253)] [[Code](https://github.com/SCP-CN-001/ramble)]\n- **DiVE**: Efficient Multi-View Driving Scenes Generation Based on Video Diffusion Transformer.  **`arXiv 25.4`** [[Paper](https://arxiv.org/abs/2504.18576)]\n- [**WoTE**] End-to-End Driving with Online Trajectory Evaluation via BEV World Model.  **`ICCV 25`** [[Paper](https://arxiv.org/abs/2504.01941)] [[Code](https://github.com/liyingyanUCAS/WoTE)]\n- **MagicDrive-V2**: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control. **`arXiv 25.3`** [[Paper](https://arxiv.org/abs/2411.13807)] [[Project](https://gaoruiyuan.com/magicdrive-v2/)]\n- **CoGen**: 3D Consistent Video Generation via Adaptive Conditioning for Autonomous Driving.  **`arXiv 25.3`** [[Paper](https://arxiv.org/abs/2503.22231)] \n- **GAIA-2**: A Controllable Multi-View Generative World Model for Autonomous Driving.  **`arXiv 25.3`** [[Paper](https://arxiv.org/abs/2503.20523)] \n- **Semi-SD**: Semi-Supervised Metric Depth Estimation via Surrounding Cameras for Autonomous Driving.  **`arXiv 25.3`** [[Paper](https://arxiv.org/abs/2503.19713)] [[Code](https://github.com/xieyuser/Semi-SD)]\n- **MiLA**: Multi-view Intensive-fidelity Long-term Video Generation World Model for Autonomous Driving.  **`arXiv 25.3`** [[Paper](https://arxiv.org/abs/2503.15875)] [[Project](https://xiaomi-mlab.github.io/mila.github.io/)]\n- **SimWorld**: A Unified Benchmark for Simulator-Conditioned Scene Generation via World Model **`arXiv 25.3`** [[Paper](https://arxiv.org/abs/2503.13952)] [[Code](https://github.com/Li-Zn-H/SimWorld)]\n- [**EOT-WM**] Other Vehicle Trajectories Are Also Needed: A Driving World Model Unifies Ego-Other Vehicle Trajectories in Video Latant Space. **`arXiv 25.3`** [[Paper](https://arxiv.org/abs/2503.09215)]\n- [**T^3Former**] Temporal Triplane Transformers as Occupancy World Models. **`arXiv 25.3`** [[Paper](https://arxiv.org/abs/2503.07338)]\n- **AVD2**: Accident Video Diffusion for Accident Video Description. **`arXiv 25.3`** [[Paper](https://arxiv.org/abs/2502.14801)] [[Project](https://an-answer-tree.github.io/)]\n- **VaViM and VaVAM**: Autonomous Driving through Video Generative Modeling.  **`arXiv 25.2`** [[Paper](https://arxiv.org/abs/2502.15672)] [[Code](https://github.com/valeoai/VideoActionModel)]\n- **Dream to Drive**: Model-Based Vehicle Control Using Analytic World Models.  **`arXiv 25.2`** [[Paper](https://arxiv.org/abs/2502.10012)]\n- **AD-L-JEPA**: Self-Supervised Spatial World Models with Joint Embedding Predictive Architecture for Autonomous Driving with LiDAR Data.  **`arXiv 25.1`** [[Paper](https://arxiv.org/abs/2501.04969)] [[Code](https://github.com/HaoranZhuExplorer/AD-L-JEPA-Release)]\n\n### 2024\n- [**SEM2**] Enhance Sample Efficiency and Robustness of End-to-end Urban Autonomous Driving via Semantic Masked World Model. **`TITS`** [[Paper](https://ieeexplore.ieee.org/abstract/document/10538211/)]\n- **Vista**: A Generalizable Driving World Model with High Fidelity and Versatile Controllability. **`NeurIPS 24`** [[Paper](https://arxiv.org/abs/2405.17398)] [[Code](https://github.com/OpenDriveLab/Vista)]\n- **SceneDiffuser**: Efficient and Controllable Driving Simulation Initialization and Rollout. **`NeurIPS 24`** [[Paper](https://arxiv.org/abs/2412.12129)]\n- **DrivingDojo Dataset**: Advancing Interactive and Knowledge-Enriched Driving World Model. **`NeurIPS 24`** [[Paper](https://arxiv.org/abs/2410.10738)] [[Project](https://drivingdojo.github.io/)]\n- **Think2Drive**: Efficient Reinforcement Learning by Thinking in Latent World Model for Quasi-Realistic Autonomous Driving. **`ECCV 24`** [[Paper](https://arxiv.org/abs/2402.16720)]\n- [**MARL-CCE**] Modelling Competitive Behaviors in Autonomous Driving Under Generative World Model. **`ECCV 24`** [[Paper](https://www.ecva.net/papers/eccv_24/papers_ECCV/papers/05085.pdf)] [[Code](https://github.com/qiaoguanren/MARL-CCE)]\n- **DriveDreamer**: Towards Real-world-driven World Models for Autonomous Driving. **`ECCV 24`** [[Paper](https://arxiv.org/abs/2309.09777)] [[Code](https://github.com/JeffWang987/DriveDreamer)]\n- **OccWorld**: Learning a 3D Occupancy World Model for Autonomous Driving. **`ECCV 24`** [[Paper](https://arxiv.org/abs/2311.16038)] [[Code](https://github.com/wzzheng/OccWorld)]\n- [**NeMo**] Neural Volumetric World Models for Autonomous Driving. **`ECCV 24`** [[Paper](https://www.ecva.net/papers/eccv_24/papers_ECCV/papers/02571.pdf)]\n- **CarFormer**: Self-Driving with Learned Object-Centric Representations. **`ECCV 24`** [[Paper](https://arxiv.org/abs/2407.15843)] [[Code](https://kuis-ai.github.io/CarFormer/)]\n- [**MARL-CCE**] Modelling-Competitive-Behaviors-in-Autonomous-Driving-Under-Generative-World-Model. **`ECCV 24`** [[Code](https://github.com/qiaoguanren/MARL-CCE)]\n- [**GUMP**] Solving Motion Planning Tasks with a Scalable Generative Model. **`ECCV 24`** [[Paper](https://arxiv.org/abs/2407.02797)] [[Code](https://github.com/HorizonRobotics/GUMP/)]\n- **WoVoGen**: World Volume-aware Diffusion for Controllable Multi-camera Driving Scene Generation. **`ECCV 24`** [[Paper](https://arxiv.org/abs/2312.02934)] [[Code](https://github.com/fudan-zvg/WoVoGen)]\n- **DrivingDiffusion**: Layout-Guided multi-view driving scene video generation with latent diffusion model. **`ECCV 24`** [[Paper](https://arxiv.org/abs/2310.07771)] [[Code](https://github.com/shalfun/DrivingDiffusion)]\n- **3D-VLA**: A 3D Vision-Language-Action Generative World Model.  **`ICML 24`** [[Paper](https://arxiv.org/abs/2403.09631)]\n- [**ViDAR**] Visual Point Cloud Forecasting enables Scalable Autonomous Driving. **`CVPR 24`** [[Paper](https://arxiv.org/abs/2312.17655)] [[Code](https://github.com/OpenDriveLab/ViDAR)]\n- [**GenAD**] Generalized Predictive Model for Autonomous Driving. **`CVPR 24`** [[Paper](https://arxiv.org/abs/2403.09630)] [[Data](https://github.com/OpenDriveLab/DriveAGI?tab=readme-ov-file#genad-dataset-opendv-youtube)]\n- **Cam4DOCC**: Benchmark for Camera-Only 4D Occupancy Forecasting in Autonomous Driving Applications. **`CVPR 24`** [[Paper](https://arxiv.org/abs/2311.17663)] [[Code](https://github.com/haomo-ai/Cam4DOcc)]\n- [**Drive-WM**] Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving. **`CVPR 24`** [[Paper](https://arxiv.org/abs/2311.17918)] [[Code](https://github.com/BraveGroup/Drive-WM)]\n- **DriveWorld**: 4D Pre-trained Scene Understanding via World Models for Autonomous Driving. **`CVPR 24`** [[Paper](https://arxiv.org/abs/2405.04390)]\n- **Panacea**: Panoramic and Controllable Video Generation for Autonomous Driving. **`CVPR 24`** [[Paper](https://arxiv.org/abs/2311.16813)] [[Code](https://panacea-ad.github.io/)]\n- **UnO**: Unsupervised Occupancy Fields for Perception and Forecasting. **`CVPR 24`** [[Paper](https://arxiv.org/abs/2406.08691)] [[Code](https://waabi.ai/research/uno)]\n- **MagicDrive**: Street View Generation with Diverse 3D Geometry Control. **`ICLR 24`** [[Paper](https://arxiv.org/abs/2310.02601)] [[Code](https://github.com/cure-lab/MagicDrive)]\n- **Copilot4D**: Learning Unsupervised World Models for Autonomous Driving via Discrete Diffusion. **`ICLR 24`** [[Paper](https://arxiv.org/abs/2311.01017)]\n- **SafeDreamer**: Safe Reinforcement Learning with World Models. **`ICLR 24`** [[Paper](https://openreview.net/forum?id=tsE5HLYtYg)] [[Code](https://github.com/PKU-Alignment/SafeDreamer)]\n- **DrivingWorld**: Constructing World Model for Autonomous Driving via Video GPT. **`arXiv 24.12`** [[Paper](https://arxiv.org/abs/2412.19505)] [[Code](https://github.com/YvanYin/DrivingWorld)]\n- An Efficient Occupancy World Model via Decoupled Dynamic Flow and Image-assisted Training. **`arXiv 24.12`** [[Paper](https://arxiv.org/abs/2412.13772)]\n- **Doe-1**: Closed-Loop Autonomous Driving with Large World Model. **`arXiv 24.12`** [[Paper](https://arxiv.org/abs/2412.09627)] [[Code](https://github.com/wzzheng/Doe)]\n- [**DrivePhysica**] Physical Informed Driving World Model. **`arXiv 24.12`** [[Paper](https://arxiv.org/abs/2412.08410)] [[Code](https://metadrivescape.github.io/papers_project/DrivePhysica/page.html)]\n- **Terra** **ACT-Bench**: Towards Action Controllable World Models for Autonomous Driving. **`arXiv 24.12`** [[Paper](https://arxiv.org/abs/2412.05337)] [[Code](https://github.com/turingmotors/ACT-Bench)] [[Project](https://turingmotors.github.io/actbench/)] [[Hugging Face](https://huggingface.co/turing-motors/Terra)] \n- **UniMLVG**: Unified Framework for Multi-view Long Video Generation with Comprehensive Control Capabilities for Autonomous Driving. **`arXiv 24.12`** [[Paper](https://arxiv.org/abs/2412.04842)] [[Project](https://sensetime-fvg.github.io/UniMLVG/)] [[Code](https://github.com/SenseTime-FVG/OpenDWM)]\n- **HoloDrive**: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving. **`arXiv 24.12`** [[Paper](https://arxiv.org/abs/2412.01407)]\n- **InfinityDrive**: Breaking Time Limits in Driving World Models. **`arXiv 24.12`** [[Paper](https://arxiv.org/abs/2412.01522)] [[Project Page](https://metadrivescape.github.io/papers_project/InfinityDrive/page.html)]\n- Generating Out-Of-Distribution Scenarios Using Language Models. **`arXiv 24.11`** [[Paper](https://arxiv.org/abs/2411.16554)]\n- **Imagine-2-Drive**: High-Fidelity World Modeling in CARLA for Autonomous Vehicles. **`arXiv 24.11`** [[Paper](https://arxiv.org/abs/2411.10171)] [[Project Page](https://anantagrg.github.io/Imagine-2-Drive.github.io/)]\n- **WorldSimBench**: Towards Video Generation Models as World Simulator. **`arXiv 24.10`** [[Paper](https://arxiv.org/abs/2410.18072)] [[Project Page](https://iranqin.github.io/WorldSimBench.github.io/)]\n- **DOME**: Taming Diffusion Model into High-Fidelity Controllable Occupancy World Model. **`arXiv 24.10`** [[Paper](https://arxiv.org/abs/2410.10429)] [[Project Page](https://gusongen.github.io/DOME)]\n- **OCCVAR**: Scalable 4D Occupancy Prediction via Next-Scale Prediction. **`OpenReview`** [[Paper](https://openreview.net/forum?id=X2HnTFsFm8)]\n- Mitigating Covariate Shift in Imitation Learning for Autonomous Vehicles Using Latent Space Generative World Models. **`arXiv 24.9`** [[Paper](https://arxiv.org/abs/2409.16663)]\n- [**LatentDriver**] Learning Multiple Probabilistic Decisions from Latent World Model in Autonomous Driving. **`arXiv 24.9`** [[Paper](https://arxiv.org/abs/2409.15730)] [[Code](https://github.com/Sephirex-X/LatentDriver)]\n- **RenderWorld**: World Model with Self-Supervised 3D Label. **`arXiv 24.9`** [[Paper](https://arxiv.org/abs/2409.11356)]\n- **OccLLaMA**: An Occupancy-Language-Action Generative World Model for Autonomous Driving. **`arXiv 24.9`** [[Paper](https://arxiv.org/abs/2409.03272)]\n- **DriveGenVLM**: Real-world Video Generation for Vision Language Model based Autonomous Driving. **`arXiv 24.8`** [[Paper](https://arxiv.org/abs/2408.16647)]\n- [**Drive-OccWorld**] Driving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous Driving. **`arXiv 24.8`** [[Paper](https://arxiv.org/abs/2408.14197)]\n- **BEVWorld**: A Multimodal World Model for Autonomous Driving via Unified BEV Latent Space. **`arXiv 24.7`** [[Paper](https://arxiv.org/abs/2407.05679)] [[Code](https://github.com/zympsyche/BevWorld)]\n- [**TOKEN**] Tokenize the World into Object-level Knowledge to Address Long-tail Events in Autonomous Driving. **`arXiv 24.7`** [[Paper](https://arxiv.org/abs/2407.00959)]\n- **UMAD**: Unsupervised Mask-Level Anomaly Detection for Autonomous Driving. **`arXiv 24.6`** [[Paper](https://arxiv.org/abs/2406.06370)]\n- **SimGen**: Simulator-conditioned Driving Scene Generation. **`arXiv 24.6`** [[Paper](https://arxiv.org/abs/2406.09386)] [[Code](https://metadriverse.github.io/simgen/)]\n- [**AdaptiveDriver**] Planning with Adaptive World Models for Autonomous Driving. **`arXiv 24.6`** [[Paper](https://arxiv.org/abs/2406.10714)] [[Code](https://arunbalajeev.github.io/world_models_planning/world_model_paper.html)]\n- [**LAW**] Enhancing End-to-End Autonomous Driving with Latent World Model. **`arXiv 24.6`** [[Paper](https://arxiv.org/abs/2406.08481)] [[Code](https://github.com/BraveGroup/LAW)]\n- [**Delphi**] Unleashing Generalization of End-to-End Autonomous Driving with Controllable Long Video Generation. **`arXiv 24.6`** [[Paper](https://arxiv.org/abs/2406.01349)] [[Code](https://github.com/westlake-autolab/Delphi)]\n- **OccSora**: 4D Occupancy Generation Models as World Simulators for Autonomous Driving. **`arXiv 24.5`** [[Paper](https://arxiv.org/abs/2405.20337)] [[Code](https://github.com/wzzheng/OccSora)]\n- **MagicDrive3D**: Controllable 3D Generation for Any-View Rendering in Street Scenes. **`arXiv 24.5`** [[Paper](https://arxiv.org/abs/2405.14475)] [[Code](https://gaoruiyuan.com/magicdrive3d/)]\n- **CarDreamer**: Open-Source Learning Platform for World Model based Autonomous Driving. **`arXiv 24.5`** [[Paper](https://arxiv.org/abs/2405.09111)] [[Code](https://github.com/ucd-dare/CarDreamer)]\n- [**DriveSim**] Probing Multimodal LLMs as World Models for Driving. **`arXiv 24.5`** [[Paper](https://arxiv.org/abs/2405.05956)] [[Code](https://github.com/sreeramsa/DriveSim)]\n- **LidarDM**: Generative LiDAR Simulation in a Generated World. **`arXiv 24.4`** [[Paper](https://arxiv.org/abs/2404.02903)] [[Code](https://github.com/vzyrianov/lidardm)]\n- **SubjectDrive**: Scaling Generative Data in Autonomous Driving via Subject Control. **`arXiv 24.3`** [[Paper](https://arxiv.org/abs/2403.19438)] [[Project](https://subjectdrive.github.io/)]\n- **DriveDreamer-2**: LLM-Enhanced World Models for Diverse Driving Video Generation. **`arXiv 24.3`** [[Paper](https://arxiv.org/abs/2403.06845)] [[Code](https://drivedreamer2.github.io/)]\n\n### 2023\n\n- **TrafficBots**: Towards World Models for Autonomous Driving Simulation and Motion Prediction. **`ICRA 23`** [[Paper](https://arxiv.org/abs/2303.04116)] [[Code](https://github.com/zhejz/TrafficBots)]\n- [**CTT**] Categorical Traffic Transformer: Interpretable and Diverse Behavior Prediction with Tokenized Latent. **`arXiv 23.11`** [[Paper](https://arxiv.org/abs/2311.18307)]\n- **MUVO**: A Multimodal Generative World Model for Autonomous Driving with Geometric Representations. **`arXiv 23.11`** [[Paper](https://arxiv.org/abs/2311.11762)]\n- **GAIA-1**: A Generative World Model for Autonomous Driving. **`arXiv 23.9`** [[Paper](https://arxiv.org/abs/2309.17080)]\n- **ADriver-I**: A General World Model for Autonomous Driving. **`arXiv 23.9`** [[Paper](https://arxiv.org/abs/2311.13549)]\n- **UniWorld**: Autonomous Driving Pre-training via World Models. **`arXiv 23.8`** [[Paper](https://arxiv.org/abs/2308.07234)] [[Code](https://github.com/chaytonmin/UniWorld)]\n\n### 2022\n\n- [**MILE**] Model-Based Imitation Learning for Urban Driving. **`NeurIPS 22`** [[Paper](https://proceedings.neurips.cc/paper_files/paper/2022/hash/827cb489449ea216e4a257c47e407d18-Abstract-Conference.html)] [[Code](https://github.com/wayveai/mile)]\n- **Symphony**: Learning Realistic and Diverse Agents for Autonomous Driving Simulation. **`ICRA 22`** [[Paper](https://arxiv.org/abs/2205.03195)] \n- Hierarchical Model-Based Imitation Learning for Planning in Autonomous Driving. **`IROS 22`** [[Paper](https://arxiv.org/abs/2210.09539)]\n\n## Other World Model Paper\n\u003e Due to the large number of relevant papers, we will no longer update the list of related papers (related to general world models and robotics), but we still welcome contributions from the community. If you have a paper you’d like to add, please feel free to submit a pull request.\n\n### 2026\n- [**HyDRA**] Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.25716)] [[Code](https://github.com/H-EmbodVis/HyDRA)] [[Project](https://kj-chen666.github.io/Hybrid-Memory-in-Video-World-Models/)]\n- [**VEGA-3D**] Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding. **`ECCV 26`** [[Paper](https://arxiv.org/abs/2603.19235)] [[Code](https://github.com/H-EmbodVis/VEGA-3D)]\n- **Divide and Conquer**: Decoupled Representation Alignment for Multimodal World Models. **`ECCV 26`** [[Paper](https://arxiv.org/abs/2605.01896)]\n- Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2603.05438)]\n- **GeoWorld**: Geometric World Models. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2602.23058)] [[Project](https://steve-zeyu-zhang.github.io/GeoWorld)]\n- [**EAWM**] From Observations to Events: Event-Aware World Model for Reinforcement Learning. **`ICLR 26`** [[Paper](https://arxiv.org/abs/2601.19336)] [[Code](https://github.com/MarquisDarwin/EAWM)]\n- **R2-Dreamer**: Redundancy-Reduced World Models without Decoders or Augmentation. **`ICLR 26`** [[Paper](https://arxiv.org/abs/2603.18202)] [[Code](https://github.com/NM512/r2dreamer)]\n- [**SeqWM**] Empowering Multi-Robot Cooperation via Sequential World Models. **`ICLR 26`** [[Paper](https://arxiv.org/abs/2509.13095)] [[Code](https://github.com/zhaozijie2022/seqwm)]\n- **NeuroHex**: Highly-Efficient Hex Coordinate System for Creating World Models to Enable Adaptive AI. **`NICE 26`** [[Paper](https://arxiv.org/abs/2603.00376)]\n- Foundation World Models for Agents that Learn, Verify, and Adapt Reliably Beyond Static Environments. **`AAMAS 26`** [[Paper](https://arxiv.org/abs/2602.23997)]\n- Probabilistic Dreaming for World Models. **`ICLRW 26`** [[Paper](https://arxiv.org/abs/2603.04715)]\n- From Part to Whole: 3D Generative World Model with an Adaptive Structural Hierarchy. **`ICME 26`** [[Paper](https://arxiv.org/abs/2603.21557)]\n- Value-guided action planning with JEPA world models. **`World Modeling Workshop 26`** [[Paper](https://arxiv.org/abs/2601.00844)]\n- Self-Supervised Multi-Modal World Model with 4D Space-Time Embedding. **`World Modeling Workshop 26`** [[Paper](https://arxiv.org/abs/2603.07039)] [[Project](https://github.com/legel/deepearth)]\n- Explicit World Models for Reliable Human-Robot Collaboration. **`AAAIW 26`** [[Paper](https://arxiv.org/abs/2601.01705)]\n- **Yume-1.5**: A Text-Controlled Interactive World Generation Model. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2512.22096)]\n- [**ORCA**] Active Intelligence in Video Avatars via Closed-loop World Modeling. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2512.20615)] [[Project](https://xuanhuahe.github.io/ORCA/)]\n- Dexterous World Models. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2512.17907)] [[Project](http://snuvclab.github.io/dwm)]\n- **Motus**: A Unified Latent Action World Model. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2512.13030)]\n- **CLARITY**: Medical World Model for Guiding Treatment Decisions by Modeling Context-Aware Disease Trajectories in Latent Space. **`ECCV 26`** [[Paper](https://arxiv.org/abs/2512.08029)]\n- **ModularAgent**: A Task-Aware Modular Framework for Joint Optimization of Multimodal Large Language Models and World Models. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2512.04513)]\n- **NavForesee**: A Unified Vision-Language World Model for Hierarchical Planning and Dual-Horizon Navigation Prediction. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2512.01550)]\n- **TraceGen**: World Modeling in 3D Trace Space Enables Learning from Cross-Embodiment Videos. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2511.21690)]\n- **4DWorldBench**: A Comprehensive Evaluation Framework for 3D/4D World Generation Models. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2511.19836)]\n- **Thinking Ahead**: Foresight Intelligence in MLLMs and World Models. **`ECCV 26`** [[Paper](https://arxiv.org/abs/2511.18735)]\n- **X-WIN**: Building Chest Radiograph World Model via Predictive Sensing. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2511.14918)]\n- **IPR-1**: Interactive Physical Reasoner. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2511.15407)]\n- Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World Models. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2510.26782)]\n- **ORV**: 4D Occupancy-centric Robot Video Generation. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2506.03079)] [[Project](https://orangesodahub.github.io/ORV/)]\n- **KineBench**: Benchmarking Embodied World Models via IDM-Free Kinematic Grounding. **`ECCV 26`** [[Paper](https://arxiv.org/abs/2607.19876)]\n- **Stereo World Model**: Camera-Guided Stereo Video Generation. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2603.17375)] [[Project](https://sunyangtian.github.io/StereoWorld-web/)]\n- **DreamSAC**: Learning Hamiltonian World Models via Symmetry Exploration. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2603.07545)]\n- Inference-time Physics Alignment of Video Generative Models with Latent World Models. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2601.10553)] [[Code](https://github.com/facebookresearch/WMReward)]\n- **PointWorld**: Scaling 3D World Models for In-The-Wild Robotic Manipulation. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2601.03782)] [[Project](https://point-world.github.io/)]\n- **VerseCrafter**: Dynamic Realistic Video World Model with 4D Geometric Control. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2601.05138)] [[Project](https://sixiaozheng.github.io/VerseCrafter_page/)]\n- **NeoVerse**: Enhancing 4D World Model with in-the-wild Monocular Videos. **`CVPR 26`** [[Paper](https://arxiv.org/abs/2601.00393)] [[Project](https://neoverse-4d.github.io/)]\n- **UniJEPA**: A Unified Joint-Embedding Predictive Architecture for Task-Agnostic Visual World Modeling. **`ICML 26`** [[Paper](https://arxiv.org/abs/2608.07409)]\n- **DreamMimic**: Learning Visuomotor Whole-Body Loco-Manipulation via World Model. **`IROS 26`** [[Paper](https://arxiv.org/abs/2608.22278)]\n- **Orbit-Planner**: Towards Latent World Models for On-Orbit Obstacle Avoidance of Satellite Agents. **`AP-GARSS 26`** [[Paper](https://arxiv.org/abs/2608.16651)] [[Project](https://zhijianli2003.github.io/Orbit_Planner/)]\n- **OMR**: Off-Manifold Refinement: Guiding Video Generators with a Frozen World Model. **`BMVC 26`** [[Paper](https://arxiv.org/abs/2608.29904)] [[Project](https://itruonghai.github.io/omr)]\n- How do World Models and Policies Compose in LLM Agents? A Joint Spectral and Behavioral Account. **`EMNLP 26`** [[Paper](https://arxiv.org/abs/2608.30067)]\n- **WebWorld**: The Browser as a World Model for Self-Improving Web Code. **`EMNLP 26`** [[Paper](https://arxiv.org/abs/2608.30530)]\n- **FactoSR**: Unfold The World: Factorize 4D Properties in Reinforcing Spatial Reasoning. **`ECCV 26`** [[Paper](https://arxiv.org/abs/2609.03729)]\n- **SPAR3S**: Sparse Auto-Regressive Modeling for Scene Generation from Multi-View Images. **`ECCV 26`** [[Paper](https://arxiv.org/abs/2609.03931)]\n- Toward Physically Grounded JEPA World Models for Goal-Conditioned Robotic Planning. **`IROS 26 Workshop`** [[Paper](https://arxiv.org/abs/2609.03565)]\n- **ActSafeGuard**: Differentiable and Training-Aligned Constraint Enforcement for Flow-Matching Policies. **`arXiv 26.9`** [[Paper](https://arxiv.org/abs/2609.11697)]\n- **FolDeX**: A Physical-World Benchmark for Long-Horizon Robotic Manipulation of Deformable Objects. **`arXiv 26.9`** [[Paper](https://arxiv.org/abs/2609.10243)] [[Project](https://ai.midea.com/#/fold-challenge)]\n- **TourPhysics**: Bringing Physics to World Models for Exploration and Manipulation from a Single Image. **`arXiv 26.9`** [[Paper](https://arxiv.org/abs/2609.04911)]\n- Coupled Control and Wireless World Models for Resilient Remote Robotic Control. **`arXiv 26.9`** [[Paper](https://arxiv.org/abs/2609.04851)]\n- Spectral-Target Physical Latent Structuring for JEPA-Style World Models. **`arXiv 26.9`** [[Paper](https://arxiv.org/abs/2609.04264)]\n- **Principia**: Relational Physics Tests for Video Models. **`arXiv 26.9`** [[Paper](https://arxiv.org/abs/2609.04200)] [[Project](https://principiabench.github.io/)]\n- **Puffin-World**: Scaling a Unified Multimodal Model with Native 3D World States. **`arXiv 26.9`** [[Paper](https://arxiv.org/abs/2609.04196)] [[Project](https://kangliao929.github.io/projects/puffin-world/)]\n- **GIFT**: Guided Intermediate Feature Training via Action-Oriented Structural Supervision for Robotic Manipulation. **`arXiv 26.9`** [[Paper](https://arxiv.org/abs/2609.04193)] [[Project](https://openphoenix-team.github.io/GIFT-pages)]\n- **WorldReward**: Reward Modeling for Camera-Conditioned World Models. **`arXiv 26.9`** [[Paper](https://arxiv.org/abs/2609.03952)] [[Project](https://codegoat24.github.io/WorldReward)]\n- Semantic Bayesian World Models. **`arXiv 26.9`** [[Paper](https://arxiv.org/abs/2609.03834)]\n- **WISE**: World-Model-Guided Imagination Scheduling for Efficient Post-Training of Vision-Language-Action Models. **`arXiv 26.9`** [[Paper](https://arxiv.org/abs/2609.03681)]\n- **StateAgent**: Do Video Generators Track the World Across Segments? A Benchmark and Method for World-State Reasoning in Video Continuation. **`arXiv 26.9`** [[Paper](https://arxiv.org/abs/2609.03673)] [[Code](https://github.com/AMAP-ML/StateAgent)]\n- Building Pretraining Data for World Models: An Unreal Engine-Based Pipeline for Action-Conditioned Video Generation. **`arXiv 26.9`** [[Paper](https://arxiv.org/abs/2609.03557)]\n- **World-Coherent Decoding**: Self-Verifying Test-Time Planning for World Action Models. **`arXiv 26.9`** [[Paper](https://arxiv.org/abs/2609.02159)]\n- Towards a Belief-Based World Model for LLM Agents. **`arXiv 26.9`** [[Paper](https://arxiv.org/abs/2609.00455)] [[Code](https://github.com/skumar-ml/belief-world-models)]\n- **CAER**: Causal Action Effect Reweighting for World Model Training. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.30897)] [[Project](https://manifoldai-research.github.io/CAER/)]\n- **Motus2**: A Self-Evolving General World Model for Dexterous Manipulation. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.30237)]\n- **The Intervention Gap in Latent World Models**. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.29998)]\n- **AcrossWAM1.0**: A Modular Latent World-Action Stack for Compact Robot Policies. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.29937)]\n- **CLAP**: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.27406)] [[Project](https://omni-clap.github.io/)]\n- Successive Capacity Growth: Task-Complexity-Driven Width and Depth Expansion for Vision Transformer Encoders in JEPA World Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.27367)]\n- **PAWBench**: How Far Are We from Probabilistically Aligned World Modeling? **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.27345)]\n- **R2M-Bench**: Evaluating Revisit Memory via Relative Consistency in Interactive Video World Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.27328)] [[Code](https://github.com/AMAP-ML/R2MBench)]\n- Making Latent Evolution Explicit: Operator-Structured Transitions for World Action Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.27259)]\n- **SpatialCrafter**: Single Image World Modeling with Generative 3D Proxies. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.27073)]\n- **Riemann-1.0**: An Embodied World Action Model for Physical AI. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.27033)]\n- **Zero-WAM**: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.26103)] [[Project](https://robbyant-research.github.io/Zero-WAM/)]\n- **WALL-SS**: Scaling Long-horizon World Models via Next-Scale Autoregression. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.26239)]\n- **4DGS-WAM**: Bridging Past and Future with an Object-Centric World Action Model based on 4D Gaussian Splatting. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.25956)]\n- **Code World Model**: Coding Agent as World Brain. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.25927)] [[Project](https://buaacyw.github.io/cwm/)]\n- **GaussianDream++**: Efficient 3D Gaussian World Modeling for Robotic Manipulation. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.25659)]\n- **ConfAL-WM**: Confidence-Guided Active Learning for Action-Conditioned World Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.25572)] [[Project](https://confal-wm.github.io/)]\n- **Rollout-Decoded Reconstruction**: Rollout-Decoded Reconstruction for Long-Horizon Prediction in Latent World Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.25017)]\n- Do Robotic World Models Really Follow Actions? Diagnosing and Aligning Action-Conditioned Generation for Policy Learning. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.24885)]\n- Latent Action as Intention Enables Efficient Future Imagination for World Action Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.24882)]\n- **LeFlow**: Generative Latent Flow Planning for World Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.24855)]\n- **GameWAM**: A World Action Model for Video Games. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.26200)]\n- **GaussianWAM**: Distilling Geometry and Semantics from 3D Gaussian Fields into World-Action Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.24714)]\n- **GlanceWAM**: Sparse Test-Time Imagination for World-Action Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.23927)]\n- **ReWorld**: An Interactive World Model with Long-Horizon Memory. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.23565)] [[Project](https://zhifeichen097.github.io/ReWorld/)]\n- Correcting a Learned Physical Invariant Improves World-Model Rollouts. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.23526)] [[Code](https://github.com/Zarand3r/world-model-invariants)]\n- **EchoWM**: Open and Enterable Omnimodal World Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.23189)]\n- From Generation to Simulation: How Far Are World Models from Being True Simulators? **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.23070)] [[Code](https://github.com/AtongWang/world-model-simulators)]\n- **LpWM**: A Case for Sparse Representations in World Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.22764)]\n- **MOSH-WM**: Mask-Grounded Soft-Hamiltonian Dynamics for Object-Centric World Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.22750)]\n- Where World Models Break: Natural-Input Failure Discovery. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.22421)]\n- **LD4WAM**: Learning Latent Dynamics from Human Videos for World Action Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.22403)]\n- **WAM-OPD**: On-Policy Distillation for World Action Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.22364)]\n- Beyond Instance Slots: Semantically Rich World Models for Physical Interaction Planning. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.22294)]\n- **DELE-w0.5**: Inferring Action from Future Latent State for Robotic Manipulation. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.22067)]\n- **ForeTime-VLA**: Causal Future-Token Distillation from a World Action Model for Conveyor-Belt Manipulation. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.20735)]\n- **DECOWAM**: Decoupled Whole-Body World-Action Model for Legged Mobile Manipulation. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.20114)]\n- **Orthogonal JEPA**: Factorized Predictive States for Latent World Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.20065)]\n- **RISE**: Adaptive Imagination for World Action Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.20430)]\n- **GigaBrain-WBC-0.5**: A Behavior World Model for Robust Whole-Body Control with Environment Interaction. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.18234)] [[Project](https://shepherd1226.github.io/gigabrain-wbc-0.5/)]\n- **Hydra-0**: Action Flow for Generalist World Modeling and Control. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.18077)] [[Project](https://nvidia-isaac.github.io/video_to_data/hydra-0/)]\n- **No Gaussian Required**: Contrastive Inverse Dynamics for JEPA World Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.17542)] [[Code](https://github.com/jackboyla/action-contrastive-jepa)]\n- **SCALE**: State-Calibrated Latent Embeddings for JEPA Planning in the Right Geometry. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.16287)]\n- **Marionette**: Predicting World States, Rendering Geometry, Painting Appearance. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.14530)] [[Project](https://alayalab.github.io/Marionette/)]\n- **Twin**: Playing an Unknown Game with a Test-Time Digital Twin. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.14490)] [[Project](https://arc-agi-3-twin.vercel.app/)] [[Code](https://github.com/Alexyskoutnev/TWIN-ARC-AGI-3)]\n- **Traj-LeWM**: Path-Aware World-Model Planning via Latent Trajectory Cost. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.14125)]\n- **ForgeWM**: Progressive Causal Training for Few-Step Action-Conditioned Video World Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.14022)]\n- **hint$^2$**: Hierarchical World Models for Inference-Time Temporal Logic Guidance. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.13678)] [[Project](https://anonymous-hint2.github.io/)]\n- **PlayWorld**: Benchmarking World Models with Agent Players over Long-Horizon Objectives. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.13552)] [[Project](https://kxding.github.io/project/PlayWorld/)]\n- **Alaya-EVOKE**: From Linear-Scaling Supervision to Endless World. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.13546)]\n- **AlayaWorld**: Interactive Long-Horizon World Modeling - Full Technical Report. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.13492)]\n- **DreamX-Phi 1.0**: Action-Conditioned Video World Model for Robotic Manipulation. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.13489)] [[Code](https://github.com/AMAP-ML/DreamX-Phi)]\n- A Unifying Perspective on Causal World Models: From Observations to Representations to Structure. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.13456)]\n- **ContactGuard**: Pre-Contact Execution Monitoring with Action-Conditioned Latent World Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.13438)]\n- **S2-HWM**: Sparse Event-Structured Hierarchical World Model for Long-Horizon Surgical Robot Manipulation. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.13103)]\n- **H2R-Bench**: Benchmarking Human-to-Robot Manipulation Video Generation in World Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.13049)]\n- Diagnosing JEPA World Models with Action-Conditioned Predictive Consistency. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.12939)]\n- **Foresight Without Seeing**: Latent Futures for World Action Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.11605)]\n- **RIFT**: Keep the Future, Drop the Rollout for World Action Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.11521)]\n- **Surgical WAM**: A World-Action Model for Data-Efficient Surgical Robot Learning. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.11204)]\n- **Flex-$\\pi$**: A Multi-Stream World-Action Model with Compute Flexibility. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.10860)] [[Project](https://flex-pi.github.io/)]\n- **StageWAM**: Joint-Embedding Stage Prediction for World-Action Models in Robot Manipulation. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.10780)]\n- **PBD-AG**: Persistent Baseline-Delta Active Graphs with Uncertainty-Aware Inspection for Long-Horizon Service Robots. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.10449)] [[Project](https://shuobao214.github.io/PBD-AG/)]\n- **Stream Forcing**: Constructing Unified Training Trajectory for Robust Streaming Video Generation. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.10439)]\n- **Beyond Myopic World Models**: Long-Horizon End-to-End Training for Direct Future Prediction. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.07420)]\n- **Addressable Memory for Video World Models**. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.07408)]\n- **WNM-3D**: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.07267)]\n- **MemWM**: Memory-Augmented Text-Based World Model. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.07107)]\n- Transformers Struggle to Use Their Emergent World Models: Revisiting the Tower of Hanoi, and the Illusion of Thinking. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.07077)]\n- **PILOT**: Decoupling Intention from Trajectory: A Representational Deduction Framework for World Action Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.06994)]\n- **PSG-JEPA**: Is Forward Prediction Enough? Physical State Grounding for JEPA World Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.06799)]\n- **Surg-UniWorld**: A Unified Surgical World Model with Multimodal Control Experts. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.06770)]\n- **Dueling World Models**: Advantage-Style Action Channels for Common-Mode Distractor Rejection. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.06706)]\n- **TaskSense**: Focusing on What Matters in World Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.06544)]\n- **$\\omega$-0**: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.06375)]\n- **GeniWorld**: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.06332)]\n- **MASS**: Multiplayer World Models with Authoritative Shared State. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.06257)]\n- **EnvACE**: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.06197)]\n- **GAUGE**: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.05948)]\n- **Robust-WAM**: Bridging Generative Pretraining and Semantic Foresight in World-Action Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.05903)]\n- **AppDeltaWorld**: Transition-Grounded Delta Code World Model for Mobile GUI Agents. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.05891)]\n- **XEWorld**: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments? **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.05799)]\n- **PhyLatent**: Learning Dynamics-Relevant Representations for JEPA World Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.05720)]\n- **LAWM-3D**: Learning 3D-Aware Latent Actions from Human Videos for Generalizable Robot World Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.05706)]\n- **DreamGuard**: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.05695)]\n- **Uncertainty-Aware World Model for Aerial Image-Goal Navigation**. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.05597)]\n- **HERA**: Historical Evidence Routing Adapter for Physical Prediction in Latent World Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.05523)]\n- **HelloWorld**: Enabling Socially Interactive Characters in Video World Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.05070)]\n- **DreamWAM**: Beyond RGB Future Prediction for World Action Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.04996)]\n- **WorldCycle**: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.04964)]\n- **MobileWAM**: Bridging World Action Models to Mobile Manipulation with Chain-of-Foresight. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.04657)]\n- Overcoming Statistical Bias in Action-Controllable World Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.04653)]\n- **Faster-WAM**: Efficient Inference-Time Future Conditioning for Robust World Action Models. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.04404)]\n- **LiLa-WAM**: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.03701)]\n- **UniNav**: A Unified World-Action Diffusion Model for Visual Navigation. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.03244)]\n- **CrossScope**: A Role-Asymmetric World Model for Joint Dual-Scope Surgical Video Prediction. **`arXiv 26.8`** [[Paper](https://arxiv.org/abs/2608.03211)]\n- **WCM**: A World Critic Model for Vision-Language-Action Reinforcement Learning. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.29613)]\n- **AquaJEPA**: Action-Conditioned Multimodal Predictive Representations for Underwater Robot Dynamics. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.29393)]\n- **BWM**: A Low-Cost High-Fidelity World Simulator for Robot Learning. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.29302)]\n- **FBFM**: A Training-Free Asynchronous Feedback Mechanism for Flow-Matching in World-Action Models Execution. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.29235)]\n- **ST-WAM**: Semantic-Temporal World Action Model for Robust Manipulation under Visual Distribution Shifts. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.28993)]\n- **PhiZero**: A World Model Built Around Physical Language. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.28624)]\n- **QuantWAMs**: Calibrating at the Right Granularity for World Action Models. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.28405)]\n- **TacWAM**: Anchor-Guided World Action Model with Mechanics-Aware Tactile Prediction. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.28391)]\n- **ShadowDancer**: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.28362)]\n- **EgoGenesis**: Egocentric World-Action Modeling with Online Anchored Projective Memory and Action-3D RoPE. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.28243)]\n- **ODEWorld**: A Continuous Predictive Architecture via Physical-Time Flow. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.27924)]\n- **World Action Planner**: Generalizable Decision-Making with Action-Conditioned World Models. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.27599)]\n- Failure Detection for Surgical Robot Imitation Policies via Flow-Matching World Modeling. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.27511)]\n- **What Can Latent World Models Know?** Physical Parameter Identifiability in Multimodal Predictive Representations. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.27017)]\n- **CheckVLA**: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.26789)]\n- **StatePlay**: State-Aware Game World Models for Mechanics-Consistent Generation. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.26754)]\n- **ActSWM**: Action-Sensitive World Models for Long-Horizon Planning in Open-World Games. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.26712)]\n- **ContactFlow**: A video action conditioning that transfers across embodiments. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.26579)]\n- **CG-World**: A Large-Scale World-State Dataset and Protocol for World Models. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.26452)]\n- Learning Implicit Causal World Models from Multi-Agent Demonstrations. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.26336)]\n- **INTACT**: Isomorphic Intent-to-Action Learning for Search-Free World Models. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.26056)]\n- **Reinformed Dreamer**: An Asymmetric World Model Efficiently Trained through Latent Guidance. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.26040)]\n- **Wonder**: Video World Model Done Better. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.26037)]\n- **DC-WAM**: Dynamic-Centric Visual Supervision and Reasoning for World-Action Models. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.25918)]\n- **Temporal-Distance JEPA**: Plan-Aware Representation Learning for Latent World Model Predictive Control. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.25337)]\n- **VisualPatchWorld**: Code World Models as Latent Structured Representations for Planning. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.25236)]\n- **FeelWorld**: Visuo-Tactile World Model for Hierarchical Contact Prediction and Planning. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.24267)]\n- **LeapBot-WA**: World-Anchor Action Models via Predictive Latent Alignments. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.23969)]\n- **WorldDiT**: A Unified Diffusion Architecture for World and Action Modeling. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.23909)]\n- **ViTacWorld**: Scaling Visuo-Tactile World Models for Contact-Rich Robot Manipulation. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.22530)] [[Project](https://vitacworld.github.io/)]\n- Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.21594)]\n- Masked Visual Actions for Unified World Modeling. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.19343)]\n- **ABot-World-0**: Infinite Interactive World Rollout on a Single Desktop GPU. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.19191)]\n- **DriftWorld**: Fast World Modeling through Drifting. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.15065)]\n- **AeroAct**: Action-Centered World-Action Models for Language-Conditioned Quadrotor Flight. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.14997)]\n- **MECo-WAM**: Learning 4D Geometric Priors for Inference-Efficient World Action Models. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.05468)]\n- **RoboWorld**: Fast and Reliable Neural Simulators for Generalist Robot Policy Evaluation. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.01060)]\n- 3D Point World Models: Point Completion Enables More Accurate Dynamics Learning. **`arXiv 26.7`** [[Paper](https://arxiv.org/abs/2607.00148)]\n- **DVG-WM**: Disentangled Video Generation Enables Efficient Embodied World Model for Robotic Manipulation. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.32028)]\n- **ViPSim**: Collaborating Visual and Parameter Spaces for Consistent Long-Horizon Embodied World Models. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.28804)]\n- **Qwen-RobotWorld**: Unifying Embodied World Modeling through Language-Conditioned Video Generation. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.17030)]\n- **WorldOlympiad**: Can Your World Model Survive a Triathlon? **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.11129)]\n- **MotionWAM**: Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.09215)]\n- **WorldFly**: A World-Model-Based Vision-Language-Action Model for UAV Navigation. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.06147)]\n- World-Language-Action Model for Unified World Modeling, Language Reasoning, and Action Synthesis. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.05979)]\n- **PiL-World**: A Chunk-Wise World Model for VLA Policy-in-the-Loop Evaluation. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.05773)]\n- **OSCAR**: Omni-Embodiment Action-Conditioned World Model for Robotics. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.04463)]\n- **AirDreamer**: Generalist Drone Navigation with World Models. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.03252)]\n- **Cosmos 3**: Omnimodal World Models for Physical AI. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.02800)]\n- **RoboDream**: Compositional World Models for Scalable Robot Data Synthesis. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.02577)]\n- **RoboTrustBench**: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation. **`arXiv 26.6`** [[Paper](https://arxiv.org/abs/2606.01600)]\n- [**Helios**] Real Real-Time Long Video Generation Model. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.04379)] [[Code](https://github.com/PKU-YuanGroup/Helios)] [[Project](https://pku-yuangroup.github.io/Helios-Page/)]\n- Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.25685)]\n- **MMaDA-VLA**: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.25406)]\n- **ABot-PhysWorld**: Interactive World Foundation Model for Robotic Manipulation with Physics Alignment. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.23376)]\n- **Describe-Then-Act**: Proactive Agent Steering via Distilled Language-Action World Models. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.23149)]\n- Model Predictive Control with Differentiable World Models for Offline Reinforcement Learning. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.22430)]\n- **WorldCache**: Content-Aware Caching for Accelerated Video World Models. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.22286)] [[Code](https://umair1221.github.io/World-Cache/)]\n- **ThinkJEPA**: Empowering Latent World Models with Large Vision-Language Reasoning Model. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.22281)]\n- **Omni-WorldBench**: Towards a Comprehensive Interaction-Centric Evaluation for World Models. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.22212)]\n- Do World Action Models Generalize Better than VLAs? A Robustness Study. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.22078)]\n- **InSpatio-WorldFM**: An Open-Source Real-Time Generative Frame Model. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.11911)] [[Project](https://inspatio.github.io/worldfm/)]\n- **AcceRL**: A Distributed Asynchronous Reinforcement Learning and World Model Framework for Vision-Language-Action Models. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.18464)]\n- **EVA**: Aligning Video World Models with Executable Robot Actions via Inverse Dynamics Rewards. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.17808)] [[Project](https://eva-project-page.github.io/)]\n- **GigaWorld-Policy**: An Efficient Action-Centered World--Action Model. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.17240)]\n- **MosaicMem**: Hybrid Spatial Memory for Controllable Video World Models. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.17117)] [[Project](https://mosaicmem.github.io/mosaicmem/)]\n- **DreamPlan**: Efficient Reinforcement Fine-Tuning of Vision-Language Planners via Video World Models. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.16860)] [[Project](https://psi-lab.ai/DreamPlan/)]\n- **Simulation Distillation**: Pretraining World Models in Simulation for Rapid Real-World Adaptation. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.15759)] [[Project](https://sim-dist.github.io/)]\n- **ResWM**: Residual-Action World Model for Visual RL. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.11110)]\n- **World2Act**: Latent Action Post-Training via Skill-Compositional World Models. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.10422)] [[Project](https://wm2act.github.io/)]\n- **RAE-NWM**: Navigation World Model in Dense Visual Representation Space. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.09241)] [[Code](https://github.com/20robo/raenwm)]\n- **MWM**: Mobile World Models for Action-Conditioned Consistent Prediction. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.07799)]\n- **LiveWorld**: Simulating Out-of-Sight Dynamics in Generative Video World Models. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.07145)]\n- **WorldCache**: Accelerating World Models for Free via Heterogeneous Token Caching. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.06331)] [[Project](https://github.com/FofGofx/WorldCache)]\n- World Properties without World Models: Recovering Spatial and Temporal Structure from Co-occurrence Statistics in Static Word Embeddings. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.04317)]\n- Beyond Pixel Histories: World Models with Persistent 3D State. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.03482)]\n- **DreamWorld**: Unified World Modeling in Video Generation. **`arXiv 26.3`** [[Paper](https://arxiv.org/abs/2603.00466)]\n- **MetaOthello**: A Controlled Study of Multiple World Models in Transformers. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.23164)]\n- The Trinity of Consistency as a Defining Principle for General World Models. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.23152)]\n- **UCM**: Unifying Camera Control and Memory with Time-aware Positional Encoding Warping for World Models. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.22960)] [[Project](https://humanaigc.github.io/ucm-webpage/)]\n- **CWM**: Contrastive World Models for Action Feasibility Learning in Embodied Agent Pipelines. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.22452)]\n- **Solaris**: Building a Multiplayer Video World Model in Minecraft. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.22208)] [[Project](https://solaris-wm.github.io/)]\n- When World Models Dream Wrong: Physical-Conditioned Adversarial Attacks against World Models. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.18739)]\n- Learning Invariant Visual Representations for Planning with Joint-Embedding Predictive World Models. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.18639)]\n- Factored Latent Action World Models. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.16229)]\n- [**DreamZero**] World Action Models are Zero-shot Policies. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.15922)] [[Project](https://dreamzero0.github.io/)]\n- **VLM-DEWM**: Dynamic External World Model for Verifiable and Resilient Vision-Language Planning in Manufacturing. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.15549)]\n- Self-Supervised JEPA-based World Models for LiDAR Occupancy Completion and Forecasting. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.12540)]\n- GigaBrain-0.5M: a VLA That Learns From World Model-Based Reinforcement Learning. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.12099)] [[Project](https://gigabrain05m.github.io/)]\n- **VLAW**: Iterative Co-Improvement of Vision-Language-Action Policy and World Model. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.12063)] [[Project](https://sites.google.com/view/vlaw-arxiv)]\n- Scaling World Model for Hierarchical Manipulation Policies. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.10983)] [[Project](https://vista-wm.github.io/)]\n- Say, Dream, and Act: Learning Video World Models for Instruction-Driven Robot Manipulation. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.10717)]\n- **Olaf-World**: Orienting Latent Actions for Video World Modeling. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.10104)] [[Project](https://showlab.github.io/Olaf-World/)]\n- **VLA-JEPA**: Enhancing Vision-Language-Action Model with Latent World Model. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.10098)]\n- **Agent World Model**: Infinity Synthetic Environments for Agentic Reinforcement Learning. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.10090)] [[Code](https://github.com/Snowflake-Labs/agent-world-model)]\n- **MVISTA-4D**: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.09878)]\n- **Hand2World**: Autoregressive Egocentric Interaction Generation via Free-Space Hand Gestures. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.09600)] [[Project](https://hand2world.github.io/)]\n- **WorldArena**: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.08971)]\n- **MIND**: Benchmarking Memory Consistency and Action Control in World Models. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.08025)] [[Code](https://github.com/CSU-JPG/MIND)]\n- Cross-View World Models. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.07277)]\n- Interpreting Physics in Video World Models. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.07050)]\n- **DreamDojo**: A Generalist Robot World Model from Large-Scale Human Videos. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.06949)] [[Project](https://dreamdojo-world.github.io/)]\n- **World-VLA-Loop**: Closed-Loop Learning of Video World Model and VLA Policy. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.06508)] [[Project](https://showlab.github.io/World-VLA-Loop/)]\n- Self-Improving World Modelling with Latent Actions. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.06130)]\n- **BridgeV2W**: Bridging Video Generation Models to Embodied World Models via Embodiment Masks. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.03793)] [[Project](https://bridgev2w.github.io/)]\n- **LIVE**: Long-horizon Interactive Video World Modeling. **`arXiv 26.2`** [[Paper](https://arxiv.org/abs/2602.03747)]\n- [**Lingbot-World**] Advancing Open-source World Models. **`arXiv 26.1`** [[Paper](https://arxiv.org/abs/2601.20540)] [[Code](https://github.com/robbyant/lingbot-world)]\n- [**Lingbot-VA**] Causal World Modeling for Robot Control. **`arXiv 26.1`** [[Paper](https://arxiv.org/abs/2601.21998)] [[Code](https://github.com/robbyant/lingbot-va)]\n- **PathWise**: Planning through World Model for Automated Heuristic Design via Self-Evolving LLMs. **`arXiv 26.1`** [[Paper](https://arxiv.org/abs/2601.20539)]\n- **WorldBench**: Disambiguating Physics for Diagnostic Evaluation of World Modelsl. **`arXiv 26.1`** [[Paper](https://arxiv.org/abs/2601.21282)] [[Project](https://world-bench.github.io/)]\n- Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models. **`arXiv 26.1`** [[Paper](https://arxiv.org/abs/2601.19834)] [[Project](https://thuml.github.io/Reasoning-Visual-World)]\n- **PhysicsMind**: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World Models. **`arXiv 26.1`** [[Paper](https://arxiv.org/abs/2601.16007)]\n- **Boltzmann-GPT**: Bridging Energy-Based World Models and Language Generation. **`arXiv 26.1`** [[Paper](https://arxiv.org/abs/2601.17094)]\n- **MetaWorld**: Skill Transfer and Composition in ","projects_url":"https://awesome.ecosyste.ms/api/v1/lists/lmd0311%2Fawesome-world-model/projects"}