awesome-video-generation
A curated list of AI video generation APIs, SDKs, and tools including text-to-video, video editing, multimodal generation, diffusion models, and generative AI platforms. Covers commercial services, open source models with APIs, and production-ready infrastructure for developers building video applications.
https://github.com/backblaze-labs/awesome-video-generation
Last synced: about 4 hours ago
JSON representation
-
Avatar and Talking Head APIs
- Anam AI - time interactive avatar API. CARA-3 model, 180ms median latency at 25fps/720x480. Connects to any LLM. JavaScript and Python SDKs. [Docs](https://anam.ai/api) | SDK: [JavaScript](https://github.com/anam-org/javascript-sdk), [Python](https://github.com/anam-org/python-sdk)
- Captions / Mirage - head videos from script + image + actor ID. Natural gestures, eye contact, synchronized audio. [Docs](https://help.mirage.app/api-reference/api)
- Colossyan
- D-ID - time WebRTC streaming. [Docs](https://docs.d-id.com/) | SDK: [Python](https://pypi.org/project/did-api/)
- DeepBrain AI (AI Studios)
- Elai.io - learning. Turns documents and scripts into avatar-presented videos. [Docs](https://elai.readme.io)
- Hedra - 3 omnimodal talking-avatar API (image + text + audio in one pass). Long-form up to 10 min; LiveKit plugin for realtime; Node SDK, REST, Make.com integration. Omnia Fast Alpha adds full scene control. [Docs](https://docs.hedra.com) | SDK: Node
- HeyGen - time streaming avatars via WebRTC. Template-based workflows. [Docs](https://docs.heygen.com/) | SDK: [JS/TS](https://github.com/HeyGen-Official/StreamingAvatarSDK)
- Hour One
- Simli - time speech-to-video avatar API using 3D Gaussian splatting for full-face animation (not just lip-sync). WebRTC streaming, <200ms latency. Python and JS SDKs. [Docs](https://docs.simli.com) | SDK: [Python](https://docs.simli.com/api-reference/python), [JavaScript](https://github.com/simliai/simli-client)
- Sync.so - grade lipsync and visual dubbing API. sync-3 flagship at native 4K with obstruction detection; lipsync-2-pro for diffusion super-res; react-1 for expressive emotion. Batch up to 500 videos per job. From $0.02/sec. [Docs](https://sync.so/docs/introduction) | SDK: [Python](https://pypi.org/project/syncsdk/), [TypeScript](https://www.npmjs.com/package/@sync.so/sdk)
- Synthesia - based video creation from scripts. 140+ languages, custom avatars, template workflows. API in beta. [Docs](https://docs.synthesia.io/)
- SynthLife - scheduling and unlimited content generation.
- Tavus - 4 model does real-time gaussian-diffusion facial synthesis at ~600ms latency. Replica API clones face + voice. Integrates with Pipecat and LiveKit. [Docs](https://www.tavus.io/ai-video-api)
- VEED Fabric 1.0 - video API. Diffusion Transformer drives lip-sync, head, and body gestures from any image (photo, illustration, mascot). Up to 5 min; 480p/720p MP4 output. REST via fal.ai; Python and JS SDKs. [Docs](https://fal.ai/models/veed/fabric-1.0) | SDK: [Python](https://docs.fal.ai), [JavaScript](https://docs.fal.ai)
-
Evaluation and Observability
- VBench / VBench-2.0 - grained dimensions including subject consistency, motion smoothness, temporal flickering. VBench-2.0 adds Physics and Commonsense evaluation. [Docs](https://huggingface.co/spaces/Vchitect/VBench_Leaderboard)
- Artificial Analysis Video Arena - based blind-comparison leaderboard for text-to-video and image-to-video models, with separate tracks for audio-enabled output. Embeddable leaderboard widgets and HuggingFace Space. [Docs](https://artificialanalysis.ai/video/leaderboard/text-to-video)
- Video-Bench - aligned video generation benchmark. Uses MLLMs to evaluate nine quality dimensions (imaging quality, aesthetic quality, temporal consistency, motion effects, text alignment). pip install videobench. Apache-2.0. [Docs](https://video-bench.github.io/)
-
Infrastructure and Deployment
- HandBrake - source video transcoder wrapping FFmpeg. GUI and CLI. [Docs](https://github.com/HandBrake/HandBrake)
- hls.js
- Shaka Player - source DASH + HLS player.
- Backblaze B2 - compatible object storage at low cost. Free egress via Cloudflare. [Docs](https://www.backblaze.com/docs/cloud-storage-s3-compatible-api) | [B2 integration](https://www.backblaze.com/docs/cloud-storage-s3-compatible-api)
- Cloudflare Stream
- CoreWeave - native AI cloud with enterprise-scale GPU infrastructure.
- fal.ai
- FFmpeg - standard multimedia processing. Encode, decode, transcode, stream, filter. [Docs](https://github.com/FFmpeg/FFmpeg)
- Lambda Labs - demand H100/B200 GPUs. SSH and JupyterLab access with REST API for instance management.
- Modal - first GPU platform. Container spin-up in ~1 second. [Docs](https://modal.com/docs/guide)
- Mux - first video infrastructure. Upload, encode, stream (VOD + live), analytics. SDKs for Node, Python, Ruby, Go, and more. [Docs](https://docs.mux.com)
- Pollo AI
- Replicate - source video models via REST API. [Docs](https://replicate.com/docs)
- RunPod
- SiliconFlow - source video models including Wan2.1/2.2 T2V and I2V, and HunyuanVideo-HD. OpenAI-compatible REST API at api.siliconflow.cn/v1. English docs available. [Docs](https://docs.siliconflow.cn/en/userguide/capabilities/video)
- Together AI - service GPU clusters.
- WaveSpeedAI
-
Open Source Models
- Open-Sora - like generation. 2s–15s at 144p–720p. T2V, I2V, V2V.
- Wan 2.1 (Alibaba) - 14B model. Also supports I2V, editing, T2I, and V2A. T2V-1.3B runs on consumer GPUs. Also on Replicate and fal.ai. [Docs](https://huggingface.co/Wan-AI/Wan2.1-T2V-14B)
- Wan 2.2 (Alibaba) - source MoE video diffusion model. +65.6% more image training data and +83.2% more video data vs 2.1. 5B and 14B variants. [Docs](https://huggingface.co/Wan-AI/Wan2.2-T2V-A14B)
- CogVideoX (Zhipu AI / Z.AI) - 5B flagship; supports 10s videos. Commercial product "Ying" available via API. [Docs](https://huggingface.co/zai-org/CogVideoX-5b)
- AnimateDiff - and-play animation module for Stable Diffusion models. Merged into HuggingFace diffusers. [Docs](https://huggingface.co/docs/diffusers/api/pipelines/animatediff)
- HunyuanVideo (Tencent) - Hunyuan/HunyuanVideo-1.5)
- LTX-Video / LTX-2 (Lightricks) - based real-time video gen model. LTX-2 adds native 4K at 50fps with synchronized audio. ComfyUI nodes available. [Docs](https://github.com/Lightricks/LTX-2)
- SkyReels (SkyworkAI) - centric video fine-tuned on HunyuanVideo. V2: infinite-length video via Autoregressive Diffusion-Forcing. V3: multimodal reaching closed-source SOTA levels. [Docs](https://github.com/SkyworkAI/SkyReels-V3)
- MAGI-1 (Sand AI) - by-chunk (24 frames/chunk). T2V, I2V, V2V with streaming generation. Outperforms Wan 2.1 and HunyuanVideo on benchmarks. [Docs](https://huggingface.co/sand-ai/MAGI-1)
- Step-Video-T2V (StepFun) - to-video model, up to 204 frames, bilingual (EN/ZH). [Docs](https://huggingface.co/stepfun-ai/stepvideo-t2v)
- Pyramid Flow - flow-sd3)
- LongCat-Video (Meituan) - to-video, image-to-video, and video continuation, tuned for efficient long-form 720p generation. FlashAttention-2 acceleration. Streamlit demo included. Avatar variant released Dec 2025. [Docs](https://huggingface.co/meituan-longcat)
- OmniAvatar - driven full-body avatar video generation with adaptive body animation. Pixel-wise multi-hierarchical audio embedding for diverse scenes. 1.3B and 14B variants (LoRA on Wan 2.1). Local inference via CLI. [Docs](https://omni-avatar.github.io)
- Allegro (Rhymes AI) - ai/Allegro)
- MOVA (OpenMOSS) - tower architecture with cross-attention fusion. Multilingual lip-sync and environment-aware SFX. 360p and 720p weights. Diffusers integration planned. [Docs](https://huggingface.co/OpenMOSS-Team/MOVA-720p)
- daVinci-MagiHuman - source audio-video generation model (Sand.ai + GAIR). Unified transformer jointly processes text, video, and audio for lip-synced talking-head generation. T2V and TI2V modes; 7 languages; 2s for 5s/256p clip on H100. Apache-2.0. [Docs](https://huggingface.co/GAIR/daVinci-MagiHuman)
- Duix Avatar - Avatar/wiki)
- EchoMimicV2 (Ant Group) - driven semi-body human animation from a single image + audio. Supports English and Mandarin. Accelerated inference (~50s for 120 frames on A100). Gradio UI included. [Docs](https://huggingface.co/BadToBest/EchoMimicV2)
- FramePack (lllyasviel) - frame prediction I2V model that packs input context to constant length so generation cost is independent of video length. Generates 60s/30fps video with a 13B model on 6 GB VRAM. Gradio GUI; Windows one-click package and Linux pip install. [Docs](https://lllyasviel.github.io/frame_pack_gitpage/)
- LivePortrait (KwaiVGI)
- Meta Movie Gen
- Mochi 1 (Genmo)
- MuseTalk (Tencent Music) - time audio-driven lip-sync in latent space. 30fps+ on a V100. v1.5 improves identity consistency. Supports Chinese, English, and Japanese. [Docs](https://huggingface.co/TMElyralab/MuseTalk)
- NVIDIA Cosmos - Predict2.5 generates physics-based video simulations from text/image/video/sensor inputs. [Docs](https://docs.nvidia.com/cosmos/latest/introduction.html)
- OmniHuman-1 (ByteDance) - body, any aspect ratio. On Replicate.
- Open-Sora-Plan (PKU Yuan Group) - style generation. v1.5.0 (Jun 2025) uses an 8B-scale sparse DiT (SUV) with WFVAE achieving HunyuanVideo-level quality. Supports T2V, I2V, and image generation. HuggingFace and ModelScope weights. [Docs](https://huggingface.co/PKU-Yuan-Lab)
- SANA-WM (NVIDIA) - parameter image-to-video world model. Generates 720p minute-scale video with 6-DoF camera control on a single GPU. Three inference variants (bidirectional, chunk-causal, distilled). Weights at Efficient-Large-Model/SANA-WM_bidirectional on HuggingFace; diffusers-compatible. Apache-2.0. [Docs](https://nvlabs.github.io/Sana/WM/)
- SkyReels-A2 (SkyworkAI) - to-video (E2V) controllable generation framework. Compose arbitrary visual elements (characters, objects, backgrounds) from reference images into coherent video via text prompts. 14B Wan2.1-based weights on HuggingFace. ComfyUI-compatible; Gradio demo included. [Docs](https://huggingface.co/Skywork/SkyReels-A2)
- UniVidX - to-any conditioning. Built on Wan2.1-T2V-14B; Apache-2.0 weights on HuggingFace. [Docs](https://huggingface.co/houyuanchen111/UniVidX)
- Vid2World - conditioned interactive world models for robotics, game simulation, and navigation. Non-causal backbone is converted to autoregressive causal architecture with frame-level action conditioning. Pretrained checkpoints on HuggingFace. [Docs](https://knightnemo.github.io/vid2world/)
- Wan-Alpha - to-video model that outputs video with transparent alpha channels via a Shiftable RGB-A Distribution Learner. Enables compositing with arbitrary backgrounds. v1.0 and v2.0 weights on HuggingFace; ComfyUI nodes available. [Docs](https://huggingface.co/htdong/Wan-Alpha)
-
Real-Time and Interactive Video
- Krea Realtime 14B - weight 14B autoregressive video model distilled from Wan 2.1 via Self-Forcing. ~11fps on a single B200, ~1s time-to-first-frame. WebSocket streaming server for mid-generation prompt edits. Research/non-commercial license. [Docs](https://huggingface.co/krea/krea-realtime-video)
- Causal Forcing - time interactive video generation on a single RTX 4090. Frame-wise and chunk-wise inference; builds on Wan 2.1. HuggingFace weights at zhuhz22/Causal-Forcing. [Docs](https://thu-ml.github.io/CausalForcing.github.io/)
- CausVid - step autoregressive generator. Enables streaming video generation at 9.4 FPS on a single GPU via KV caching. Supports T2V, I2V, and video-to-video translation. [Docs](https://huggingface.co/tianweiy/CausVid)
- Decart (Lucy 2) - time video transformation at 30fps 1080p with near-zero latency. Live-stream style transfer, character swaps, environment transformation, product placement. ~$3/hour. [Docs](https://docs.platform.decart.ai/models/video/video-generation)
- Helios (PKU Yuan Group) - scale video at 19.5 FPS on a single H100 without KV-cache or causal masking. Three weight variants (Base, Mid, Distilled) on HuggingFace. Diffusers-compatible; group offloading reduces VRAM to ~6 GB. [Docs](https://pku-yuangroup.github.io/Helios-Page/)
- HY-WorldPlay (Tencent) - time interactive world generation at 24fps. Accepts image or text prompt; responds to keyboard/mouse camera inputs. WorldPlay-5B and 8B weights open. Built on HunyuanVideo 1.5. [Docs](https://huggingface.co/tencent/HY-WorldPlay)
- LongLive (NVIDIA Labs) - time interactive long video generation at 20.7 FPS on a single H100. Accepts sequential user prompts; generates up to 240s with visual consistency via KV-recache and frame-level attention sink. LongLive-1.3B weights on HuggingFace. [Docs](https://nvlabs.github.io/LongLive/)
-
SDKs and Developer Tooling
- HuggingFace Diffusers
- MiniMax MCP Server - AI/MiniMax-MCP-JS)
- Replicate SDK - tuning. [Docs](https://replicate.com/docs/get-started/python) | SDK: Python (pip install replicate)
- fal.ai SDK - client), Node (npm install @fal-ai/client)
- Luma AI SDK
- Creatomate - driven videos from JSON templates. REST API with Node.js, PHP, and Python SDKs. Scales automatically; free trial available. [Docs](https://creatomate.com/docs/api/introduction) | SDK: [Node](https://github.com/Creatomate/creatomate-node), [PHP](https://packagist.org/packages/creatomate/creatomate), [JavaScript](https://github.com/Creatomate/creatomate-preview)
- FastVideo - training and inference framework for accelerating video diffusion models. pip install fastvideo. Supports Wan2.1/2.2 distillation; 3x speedup via SageAttention and Teacache. Scales to 64 GPUs. [Docs](https://hao-ai-lab.github.io/FastVideo/) | SDK: [Python](https://pypi.org/project/fastvideo/)
- HeyGen Streaming Avatar SDK - time WebRTC interactive avatar sessions. SDK: TypeScript (npm install @heygen/streaming-avatar)
- LTX Desktop - source nonlinear video editor with on-device AI generation (text-to-video, image-to-video, audio-to-video, retake). Runs locally on Windows/Linux with NVIDIA GPUs; API mode for macOS. Built on LTX-2.3. v1.0.5 released Apr 2026. [Docs](https://docs.ltx.video/open-source-model/getting-started/quick-start)
- rCM (NVIDIA Labs) - based continuous-time consistency distillation for 10B+ video diffusion models (Wan2.1, Cosmos-Predict2). Generates high-quality videos in 2–4 steps, 15–50x faster than full diffusion. FlashAttention-2 JVP kernel open-sourced.
- Remotion - side via CLI or Lambda (AWS). TypeScript-first; supports any CSS, Canvas, SVG, WebGL. [Docs](https://www.remotion.dev/docs) | SDK: [TypeScript](https://www.npmjs.com/package/remotion)
- Runway SDK - in polling. SDK: Python (pip install runwayml), Node (npm install @runwayml/sdk)
- Shotstack - driven and AI-generated videos at scale via JSON templates. Node, Python, PHP, and Ruby SDKs. [Docs](https://shotstack.io/docs/guide/getting-started/core-concepts/) | SDK: [Node](https://github.com/shotstack/shotstack-sdk-node), [Python](https://github.com/shotstack/shotstack-sdk-python), [PHP](https://github.com/shotstack/shotstack-sdk-php), [Ruby](https://github.com/shotstack/shotstack-sdk-ruby)
- TalkingHead (met4citizen) - time 3D avatar lip-sync. Full-body RPM/VRM avatars driven by TTS audio. Integrates ElevenLabs, OpenAI, Google TTS, and Azure. Installable via npm; CDN import also supported. MIT license; v1.7.0 released Dec 2025. [Docs](https://met4citizen.github.io/TalkingHead/) | SDK: [JavaScript](https://www.npmjs.com/package/@met4citizen/talkinghead)
- TurboDiffusion - Linear Attention, and rCM timestep distillation for 100–200x speedup. Supports TurboWan 2.1/2.2 T2V and I2V at 480p/720p. pip install + CLI inference.
- Wan2GP - VRAM inference GUI and server for Wan 2.1/2.2, HunyuanVideo, LTX-Video, and Flux. Runs on 6 GB VRAM via int8/fp8/GGUF/NV FP4 quantization. Supports AMD RDNA and older Nvidia GPUs. Gradio web UI with mask editor and motion designer. Custom community license; free for non-commercial use.
-
Start building with Genblaze
- Genblaze - source Python SDK for AI-generated video, audio, and images. It orchestrates multi-provider generation pipelines with built-in, tamper-evident provenance and native Backblaze B2 storage.
-
Templates and Example Projects
- fal.ai Next.js Video Generator - click Vercel deploy. [Docs](https://vercel.com/templates/next.js/fal-video-generator)
- B2 Video Object Detection with Transformers - b2-samples/b2-transformers-video-object-detection)
- Google Gemini Streamlit + Cloud Run
- HeyGen Streaming Avatar Demo - time WebRTC avatar sessions.
- OpenMontage - to-end. Works with Claude Code, Cursor, Copilot; pip install; free local-only mode via Piper TTS and open stock archives. AGPL-3.0.
- Pixelle-Video - automated short video pipeline. Input a topic; the tool writes the script, generates AI images via ComfyUI, synthesizes voice (Edge-TTS/Index-TTS), adds music, and outputs a finished video. Supports GPT, Qwen, DeepSeek, and Ollama. Apache-2.0; v0.1.15 released Jan 2026. [Docs](https://aidc-ai.github.io/Pixelle-Video/)
- Stability AI SVD Streamlit Demo
- ViMax - agent agentic video generation framework (Director, Screenwriter, Producer all-in-one). Orchestrates 12 specialized agents for end-to-end script-to-video with character and scene consistency. uv install; supports Gemini, MiniMax, and Veo 2 backends. MIT license.
-
Text-to-Video APIs
- Stability AI (SVD) - to-video via Stable Video Diffusion. Hosted API deprecated July 2025; open weights available for self-hosting. [Docs](https://github.com/Stability-AI/generative-models)
- Fliki - to-video and text-to-speech platform. Enterprise API with 2,500+ voices in 80+ languages. [Docs](https://developer.fliki.ai)
- HappyHorse-1.0 - to-video and image-to-video with integrated audio and lip-sync in one pass. Task-based REST with Bearer auth. [Docs](https://happyhorse.app/docs)
- Higgsfield - sync. 15M+ users.
- InVideo AI
- Kling AI - to-video and image-to-video from Kuaishou. Up to 30s clips at 1080p/30fps. Async task-based API. Also on fal.ai. [Docs](https://app.klingai.com/global/dev/document-api/quickStart/productIntroduction/overview)
- Krea - 4.5, Ray 2, Seedance Pro). Job-based with webhooks. OpenAPI spec, Python/Node/Go examples. [Docs](https://docs.krea.ai) | SDK: Python, Node, Go
- Leonardo.AI (Motion API) - to-video and image-to-video REST API. Supports Motion 2.0, Veo 3, Veo 3 Fast, Kling 2.1/2.5 models. 480p–1080p, 4–8s clips, frame interpolation, style control. [Docs](https://docs.leonardo.ai/reference/createtexttovideogeneration)
- Luma Dream Machine - quality text-to-video with character reference and style reference inputs. Ray 3 is the latest model. [Docs](https://docs.lumalabs.ai/docs/api) | SDK: [Python](https://github.com/lumalabs/lumaai-python), [JS](https://www.npmjs.com/package/lumaai)
- Magic Hour - modal AI video generation API. Text-to-video, image-to-video, style transfer, 4K upscaling. Scales to zero when idle. [Docs](https://docs.magichour.ai)
- MiniMax / Hailuo - to-video and image-to-video up to 1080p, 10s clips. [Docs](https://platform.minimax.io/docs/guides/video-generation) | SDK: [Python](https://pypi.org/project/minimax/), [Node](https://www.npmjs.com/package/minimax)
- Morph Studio - code AI video studio aggregating Wan 2.6, Kling 2.6 Pro, Seedance, Sora 2, Veo 3 into a single canvas with storyboarding and style transfer.
- OpenAI Sora - to-video and image-to-video via the v1/videos endpoint. Sora 2 supports up to 90s at 4K with spatial audio. [Docs](https://platform.openai.com/docs/api-reference/videos) | SDK: Python, Node
- Pika (v2.2) - to-video and image-to-video with Pikaframes multi-keyframe interpolation. API powered by fal.ai. [Docs](https://fal.ai/models/fal-ai/pika/v2.2/text-to-video/api) | SDK: fal Python, fal JS
- Runway (Gen-4) - to-video and image-to-video with Gen-4 Turbo. Async task-based REST API with polling helpers. [Docs](https://docs.dev.runwayml.com) | SDK: [Python](https://pypi.org/project/runwayml/), [Node](https://github.com/runwayml/sdk-node)
- Seedance 2.0 (ByteDance) - Branch Diffusion Transformer for simultaneous video + audio generation. Up to 15s at 2K resolution. Available via Dreamina. [Docs](https://dreamina.capcut.com)
- Vidu (Shengshu Technology) - form AI video model with native audio-video generation in a single output. Ranked
- Wan 2.7 (Alibaba) - weight Wan 2.2. Native 1080p, 2–15s clips, first-and-last-frame control, up to 5 reference inputs, instruction-based video editing. Via DashScope and fal.ai at $0.10/sec. [Docs](https://help.aliyun.com/zh/model-studio/wanx-video-generation)
- xAI Aurora / Grok Imagine - to-video and image-to-video using xAI's Aurora autoregressive MoE model. 6–15s clips at 720p with synchronized audio. [Docs](https://x.ai/news/grok-imagine-api)
-
Uncategorized
-
Video Enhancement and Understanding
- FlashVSR - step diffusion framework for streaming 4x video super-resolution at ~17 FPS for 768x1408 on a single A100. Locality-constrained sparse attention and tiny conditional decoder. Apache-2.0; v1.1 weights on HuggingFace. [Docs](https://huggingface.co/JunhaoZhuang/FlashVSR-v1.1)
- REAL Video Enhancer - platform GUI/CLI for frame interpolation (RIFE), upscaling (Real-ESRGAN, Waifu2x), denoising, and decompression. TensorRT and NCNN backends. Windows/macOS/Linux; v2.4.1 released Jan 2026.
- Topaz Video AI
- Twelve Labs
- Video2X - resolution and frame interpolation. Supports Real-ESRGAN, Real-CUGAN, and RIFE via Vulkan. GUI and CLI. Docker images on GHCR. [Docs](https://github.com/k4yt3x/video2x/wiki)
- WebSR - time AI video and image upscaling in the browser. Multiple pre-trained network sizes (animation, real-life, 3D content), worker thread support, and a custom training pipeline. Installable via npm. SDK: [JavaScript](https://www.npmjs.com/package/@websr/websr)
-
Video Style Transfer and Motion
- DomoAI - to-video style remixer. 50+ styles (anime, Ghibli, cinematic). v2.4.1 supports text-to-video, image-to-video, talking avatars, and animation.
- Viggle AI - transfer video tool that animates static characters to match a motion video or live webcam input. Mix Mode, Live Mode, and VTubing support. 40M+ users.
Programming Languages
Categories
Open Source Models
31
Text-to-Video APIs
19
Infrastructure and Deployment
17
SDKs and Developer Tooling
16
Avatar and Talking Head APIs
15
Templates and Example Projects
8
Real-Time and Interactive Video
7
Uncategorized
6
Video Enhancement and Understanding
6
Evaluation and Observability
3
Video Style Transfer and Motion
2
Start building with Genblaze
1
Sub Categories
Keywords
video-generation
10
text-to-video
6
image-generation
4
diffusion
4
stable-diffusion
3
text-to-speech
3
image-to-video
3
diffusion-models
3
talking-head
2
pytorch
2
flux
2
hls
2
distillation
2
few-step-generation
2
generative-ai
2
wan-video
2
text-to-video-generation
2
world-models
2
dit
2
rife
2
javascript
2
aigc
2
video-streaming
2
playback
2
lip-sync
2
ai
2
video
2
benchmark
1
dataset
1
video-editing
1
image-animation
1
face-animation
1
video2video
1
evaluation-kit
1
text2video
1
gen-ai
1
dash
1
text2image
1
virtualhumans
1
drm
1
anime4k
1
frame-interpolation
1
machine-learning
1
vulkan
1
upscale-video
1
neural-networks
1
realcugan
1
realesrgan
1
super-resoluion
1
live
1