Awesome-Video-Diffusion
A curated list of recent diffusion models for video generation, editing, and various other applications.
https://github.com/showlab/Awesome-Video-Diffusion
Last synced: 12 days ago
JSON representation
-
Table of Contents <!-- omit in toc -->
-
3D
- Text2NeRF: Text-Driven 3D Scene Generation with Neural Radiance Fields
- RoomDreamer: Text-Driven 3D Indoor Scene Synthesis with Coherent Geometry and Texture
- NeuralField-LDM: Scene Generation with Hierarchical Latent Diffusion Models
- Single-Stage Diffusion NeRF: A Unified Approach to 3D Generation and Reconstruction
- Instruct-NeRF2NeRF: Editing 3D Scenes with Instructions
- DiffusioNeRF: Regularizing Neural Radiance Fields with Denoising Diffusion Models
- NerfDiff: Single-image View Synthesis with NeRF-guided Distillation from 3D-aware Diffusion
- DiffRF: Rendering-guided 3D Radiance Field Diffusion
- Vivid-ZOO: Multi-View Video Generation with Diffusion Model
- Director3D: Real-world Camera Trajectory and 3D Scene Generation from Text
- YouDream: Generating Anatomically Controllable Consistent Text-to-3D Animals
- MultiDiff: Consistent Novel View Synthesis from a Single Image
- SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View Consistency
- Shape of Motion: 4D Reconstruction from a Single Video
- WonderWorld: Interactive 3D Scene Generation from a Single Image
- WonderJourney: Going from Anywhere to Everywhere
- ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model
- Hi3D: Pursuing High-Resolution Image-to-3D Generation with Video Diffusion Models
- ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model
- L3DG: Latent 3D Gaussian Diffusion
- GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation
- MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
- Wonderland: Navigating 3D Scenes from a Single Image
- Difix3D+: Improving 3D Reconstructions with Single-Step Diffusion Models
- Voyager: Long-Range and World-Consistent Video Diffusion for Explorable 3D Scene Generation
- WorldExplorer: Towards Generating Fully Navigable 3D Scenes
- ViewDiff: 3D-Consistent Image Generation with Text-to-Image Models
- ![Star
- Text2Room: Extracting Textured 3D Meshes from 2D Text-to-Image Models
- ![Star
- MultiDiff: Consistent Novel View Synthesis from a Single Image
- ![Star - World/Voyager)
- ![Star
- ![Star
- ![Star - fdu/Hi3D-Official)
- ![Star - of-motion/)
- ![Star
- ![Star
- ![Star - zhengcheng/vividzoo)
- ![Star
- ![Star
- ![Star
- ![Star
- ![Star - nerf2nerf)
- Monocular Normal Estimation via Shading Sequence Estimation
-
3D / NeRF
-
4D
- DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion
- ![Star
- CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models
- 4Real-Video: Learning Generalizable Photo-Realistic 4D Video Diffusion
- PaintScene4D: Consistent 4D Scene Generation from Text Prompts
- ![Star
- Stereo4D Learning How Things Move in 3D from Internet Stereo Videos
- DreamDrive: Generative 4D Scene Modeling from Street View Images
- Not All Frame Features Are Equal: Video-to-4D Generation via Decoupling Dynamic-Static Features
- AvatarArtist: Open-Domain 4D Avatarization
- ![Star - research/AvatarArtist)
- Vivid4D: Improving 4D Reconstruction from Monocular Video by Video Inpainting
- Taming Video Diffusion Models for Panoramic 4D Scene Generation
- ![Star
- In-2-4D: Inbetweening from Two Single-View Images to 4D Generation
- ![Star - 2-4D)
- ![Star - research/AvatarArtist)
- Diffuman4D: 4D Consistent Human View Synthesis from Sparse-View Videos with Spatio-Temporal Diffusion Models
- ![Star
- ![Star
- ![Star
- ![Star - 2-4D)
- ![Star - research/AvatarArtist)
- ![Star
- ![Star
- Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models
- ![Star
- ![Video - pvCLd0)
- ![Star
-
Acceleration for Video Generation
-
AI Safety
-
Audio Synthesis for Video
- Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model Adaptation
- FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
- Network Bending of Diffusion Models for Audio-Visual Generation
- ![Star
- Masked Generative Video-to-Audio Transformers with Enhanced Synchronicity
- Video-to-Audio Generation with Hidden Alignment
- ![Star - ldm)
- Read, Watch and Scream! Sound Generation from Text and Video
- ![Star - ai/rewas)
- Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
- ![Star - omni/mini-omni)
- Speech To Speech: an effort for an open-sourced and modular GPT4-o
- ![Star - to-speech)
- Draw an Audio: Leveraging Multi-Instruction for Video-to-Audio Synthesis
- ![Star - an-Audio-Code)
- VMAs: Video-to-Music Generation via Semantic Alignment in Web Music Videos
- STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment
- MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization
- Video-Guided Foley Sound Generation with Multimodal Controls
- VinTAGe: Joint Video and Text Conditioning for Holistic Audio Generation
- YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls
- Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
- ![Star
- Stable-V2A: Synthesis of Synchronized Audio Effects with Temporal and Semantic Controls
- ![Star - V2A)
- AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation
- ![Star - research/AVLink)
- Video-Foley: Two-Stage Video-To-Sound Generation via Temporal Event Condition For Foley Sound
- ![Star - foley)
- XMusic: Towards a Generalized and Controllable Symbolic Music Generation Framework
- AGAV-Rater: Enhancing LMM for AI-Generated Audio-Visual Quality Assessment
- AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation
- UniForm: A Unified Diffusion Transformer for Audio-Video Generation
- ![Star
- ![Star - V2A)
- ![Star - ren16/STAV2A)
- ![Star - omni/mini-omni)
- ![Star - to-speech)
- ![Star - foley)
- ![Star - ldm)
- ![Star - ai/rewas)
- ![Star
- ![Star - research/AVLink)
- ![Star
-
Character Customization
- PersonalVideo: High ID-Fidelity Video Customization without Dynamic and Semantic Degradation
- Magic Mirror: ID-Preserved Video Generation in Video Diffusion Transformers
- ![Star - research/MagicMirror/)
- ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning
- DreamVideo-2: Zero-Shot Subject-Driven Video Customization with Precise Motion Control
- Multi-subject Open-set Personalization in Video Generation
- Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance
- Phantom: Subject-consistent video generation via cross-modal alignment
- ![Star - video/Phantom)
- Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts
- FantasyID: Face Knowledge Enhanced ID-Preserving Video Generation
- Dynamic Concepts Personalization from Single Videos
- VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models
- ![Star - CS/VideoMaker)
- CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities
- ![Star - CS/CustomCrafter)
- CINEMA: Coherent Multi-Subject Video Generation via MLLM-Based Guidance
- MagicID: Hybrid Preference Optimization for ID-Consistent and Dynamic-Preserved Video Customization
- ![Star
- Concat-ID: Towards Universal Identity-Preserving Video Synthesis
- ![Star - GSAI/Concat-ID)
- VideoMage: Multi-Subject and Motion Customization of Text-to-Video Diffusion Models
- HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation
- ![Star
- ![Star
- ![Star
- ![Star - video/Phantom)
- ![Star - research/MagicMirror/)
- ![Star - CS/VideoMaker)
- ![Star - CS/CustomCrafter)
- ![Star - GSAI/Concat-ID)
-
Code-rendered Video Generation
- ![Website
- ![arXiv
- Paper2Video: Automatic Video Generation from Scientific Papers
- Code2Video: A Code-centric Paradigm for Educational Video Generation
- ![Star
- ![Star
- ![Star
- ![Star
- AgentMarket - B2A marketplace for AI agents. 189 APIs, 28M+ data.
-
Commercial Product
-
Controllable Video Generation
- Moonshot: Towards Controllable Video Generation and Editing with Multimodal Conditions
- TrailBlazer: Trajectory Control for Diffusion-Based Video Generation
- Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation
- SparseCtrl: Adding Sparse Controls to Text-to-Video Diffusion Models
- DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory
- Control-A-Video: Controllable Text-to-Video Generation with Diffusion Models
- ControlVideo: Training-free Controllable Text-to-Video Generation
- Motion-Conditioned Diffusion Model for Controllable Video Synthesis
- Motion-Zero: Zero-Shot Moving Object Control Framework for Diffusion-Based Video Generation
- Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance
- CameraCtrl: Enabling Camera Control for Video Diffusion Models
- MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model
- ![Star - qiu/FreeTraj)
- Training-free Camera Control for Video Generation
- MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance
-
Programming Languages
Categories
Sub Categories
Video Generation
367
Video Editing
181
Controllable Video Generation
150
Long Video / Film Generation
110
Open-source Toolboxes and Foundation Models
84
Motion Customization
70
Human or Subject Motion
69
Evaluation Benchmarks and Metrics
50
3D
45
Audio Synthesis for Video
44
Talking Head Generation
42
Video Generation with 3D/Physical Prior
37
Open-World Model
35
Character Customization
31
4D
29
Video Understanding
25
Policy Learning
18
Efficient Video Generation
16
Commercial Product
14
Reinforcement Learning for Video Generation
11
3D / NeRF
11
Video Enhancement and Restoration
10
Healthcare and Biology
10
Video Generation with Physical Prior / 3D
10
Code-rendered Video Generation
9
Policy Learning with Video Generation
8
Other Applications
6
Rendering with Virtual Engine
5
Virtual Try-On
4
Long-form Video Generation and Completion
3
New Video Generation Benchmark and Metrics
3
Other Application for Video Gen
2
World Model
2
Game Generation
1
Acceleration for Video Generation
1
Human Feedback for Video Generation
1
Human/AI Feedback for Video Generation
1
Try On with Video Generation
1
Unified Model for Generation and Understanding
1
Video Generation with Physical Prior/3D
1
Efficiency for Video Generation
1
AI Safety
1
Keywords
video-generation
5
text-to-video
5
diffusion-models
4
ai
3
image-to-video
2
text-to-video-generation
2
machine-learning
2
image-generation
2
txt2video
1
video-editing
1
segment-anything
1
restyle
1
remover
1
remove-background
1
voice-clone
1
public-api
1
photo-editing
1
wunjo
1
lip-sync
1
img2video
1
free
1
face-swap
1
face-animation
1
deepfake
1
controlnet
1
text-to-gif
1
sdxl
1
hotshot-xl
1
hotshot
1
t2v
1
colaboratory
1
colab-notebook
1
colab
1
text-to-image-generation
1
education
1
multi-agent
1
computer-vision
1
deep-learning
1
flow-matching
1
generative-ai
1
huggingface
1
open-source
1
pytorch
1
transformer
1
video-diffusion
1
agentic-ai
1
ai-agents
1
llm-agents
1
multimodal
1
prompt-engineering
1