{"id":30577032,"url":"https://github.com/FreedomIntelligence/TalkVid","last_synced_at":"2025-08-29T02:03:15.416Z","repository":{"id":310761101,"uuid":"1033056764","full_name":"FreedomIntelligence/TalkVid","owner":"FreedomIntelligence","description":"TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis","archived":false,"fork":false,"pushed_at":"2025-08-20T02:35:24.000Z","size":61738,"stargazers_count":2,"open_issues_count":0,"forks_count":0,"subscribers_count":0,"default_branch":"main","last_synced_at":"2025-08-20T04:42:14.752Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":"https://arxiv.org/abs/2508.13618","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/FreedomIntelligence.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2025-08-06T08:32:24.000Z","updated_at":"2025-08-20T03:18:19.000Z","dependencies_parsed_at":"2025-08-20T04:52:25.050Z","dependency_job_id":null,"html_url":"https://github.com/FreedomIntelligence/TalkVid","commit_stats":null,"previous_names":["freedomintelligence/talkvid"],"tags_count":null,"template":false,"template_full_name":null,"purl":"pkg:github/FreedomIntelligence/TalkVid","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/FreedomIntelligence%2FTalkVid","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/FreedomIntelligence%2FTalkVid/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/FreedomIntelligence%2FTalkVid/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/FreedomIntelligence%2FTalkVid/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/FreedomIntelligence","download_url":"https://codeload.github.com/FreedomIntelligence/TalkVid/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/FreedomIntelligence%2FTalkVid/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":272607539,"owners_count":24963759,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-08-29T02:00:10.610Z","response_time":87,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2025-08-29T02:01:20.390Z","updated_at":"2025-08-29T02:03:15.382Z","avatar_url":"https://github.com/FreedomIntelligence.png","language":"Python","funding_links":[],"categories":["Datasets","App"],"sub_categories":[],"readme":"![teaser](assets/teaser.png)\n\n# TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis\n\n🚀🚀🚀 Official implementation of **TalkVid**: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis\n\n![diversity](assets/diversity.webp)\n\n* **Authors**: [Shunian Chen*](https://github.com/Shunian-Chen), \n[Hejin Huang*](https://orcid.org/0009-0003-6700-8840), \n[Yexin Liu*](https://scholar.google.com/citations?user=Y8zBpcoAAAAJ), \n[Zihan Ye](), \n[Pengcheng Chen](https://github.com/cppppppc), \n[Chenghao Zhu](), [Michael Guan](), \n[Rongsheng Wang](https://scholar.google.com/citations?user=SSaBaioAAAAJ), \n[Junying Chen](https://scholar.google.com/citations?user=I0raPTYAAAAJ), \n[Guanbin Li](https://scholar.google.com/citations?user=2A2Bx2UAAAAJ), \n[Ser-Nam Lim†](https://scholar.google.com/citations?user=HX0BfLYAAAAJ), \n[Harry Yang†](https://scholar.google.com/citations?user=jpIFgToAAAAJ), \n[Benyou Wang†](https://scholar.google.com/citations?user=Jk4vJU8AAAAJ)\n\n* **Institutions**: The Chinese University of Hong Kong, Shenzhen; Sun Yat-sen University; The Hong Kong University of Science and Technology\n* **Resources**: [📄Paper](https://arxiv.org/abs/2508.13618)  [🤗Dataset](https://huggingface.co/datasets/FreedomIntelligence/TalkVid)  [🌐Project Page](https://freedomintelligence.github.io/talk-vid/)\n\n## 💡 Highlights\n\n* 🔥 **Large-scale high-quality** talking head dataset **TalkVid** with over 1,244 hours of HD/4K footage\n* 🔥 **Multimodal diversified content** covering 15 languages and wide age ranges (0–60+ years)\n* 🔥 **Advanced data pipeline** with comprehensive quality filtering and motion analysis\n* 🔥 **Full-body presence** including upper-body visual context unlike previous datasets\n* 🔥 **Rich annotations** with high-quality captions and comprehensive metadata\n\n## 📜 News\n**\\[2025/08/19\\]** 🚀 Our paper [TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis](https://arxiv.org/abs/2508.13618) is available!\n\n**\\[2025/08/19\\]** 🚀 Released TalkVid [dataset](https://huggingface.co/datasets/FreedomIntelligence/TalkVid) and training/inference code!\n\n**\\[2025/08/19\\]** 🚀 Released comprehensive data processing pipeline including quality filtering and motion analysis tools!\n\n## 📊 Dataset\n\n### TalkVid Dataset Overview\n\n**TalkVid** is a large-scale and diversified open-source dataset for audio-driven talking head synthesis, featuring:\n\n- **Scale**: 7,729 unique speakers with over 1,244 hours of HD/4K footage\n- **Diversity**: Covers 15 languages and wide age range (0–60+ years)\n- **Quality**: High-resolution videos (1080p \u0026 2160p) with comprehensive quality filtering\n- **Rich Context**: Full upper-body presence unlike head-only datasets\n- **Annotations**: High-quality captions and comprehensive metadata\n\n**Download Link**: [🤗 Hugging Face](https://huggingface.co/datasets/FreedomIntelligence/TalkVid)\n\n**More example videos** can be found in our [🌐 Project Page](https://freedomintelligence.github.io/talk-vid).\n\n#### Data Format\n\n```json\n{\n    \"id\": \"videovideoTr6MMsoWAog-scene1-scene1\",\n    \"height\": 1080,\n    \"width\": 1920,\n    \"fps\": 24.0,\n    \"start-time\": 0.1,\n    \"start-frame\": 0,\n    \"end-time\": 5.141666666666667,\n    \"end-frame\": 121,\n    \"durations\": \"5.042s\",\n    \"info\": {\n        \"Person ID\": \"597\",\n        \"Ethnicity\": \"White\",\n        \"Age Group\": \"60+\",\n        \"Gender\": \"Male\",\n        \"Video Link\": \"https://www.youtube.com/watch?v=Tr6MMsoWAog\",\n        \"Language\": \"English\",\n        \"Video Category\": \"Personal Experience\"\n    },\n    \"description\": \"The provided image sequence shows an older man in a suit, likely being interviewed or participating in a recorded conversation. He is seated and maintains a consistent, upright posture. Across the frames, his head rotates incrementally towards the camera's right, suggesting he is addressing someone off-screen in that direction. His facial expressions also show subtle shifts, likely related to speaking or reacting. No significant movements of the hands, arms, or torso are observed.  Because these are still images, any dynamic motion analysis is limited to inferring likely movements from the subtle positional changes between frames.\",\n    \"dover_scores\": 8.9,\n    \"cotracker_ratio\": 0.9271857142448425,\n    \"head_detail\": {\n        \"scores\": {\n            \"avg_movement\": 97.92236052453518,\n            \"min_movement\": 89.4061028957367,\n            \"avg_rotation\": 93.79223716779671,\n            \"min_rotation\": 70.42514759667668,\n            \"avg_completeness\": 100.0,\n            \"min_completeness\": 100.0,\n            \"avg_resolution\": 383.14267156972596,\n            \"min_resolution\": 349.6849455656829,\n            \"avg_orientation\": 80.29047955896623,\n            \"min_orientation\": 73.27433271185937\n        }\n    }\n}\n```\n\n### Data Statistics\n\n![statistics](assets/data_distribution.png)\n\nThe dataset exhibits excellent diversity across multiple dimensions:\n\n- **Languages**: English, Chinese, Arabic, Polish, German, Russian, French, Korean, Portuguese, Japanese, Thai, Spanish, Italian, Hindi\n- **Age Groups**: 0–19, 19–30, 31–45, 46–60, 60+\n- **Video Quality**: HD (1080p) and 4K (2160p) resolution with Dover score (mean ≈ 8.55), Cotracker ratio (mean ≈ 0.92), and head-detail scores concentrated in the 90–100 range\n- **Duration Distribution**: Balanced segments from 3-30 seconds for optimal training\n\n## ⚖️ Comparison with Other Datasets\n\n![compare](assets/compare_table.png)\n\nTalkVid stands as the **largest and most diverse** open-source dataset for audio-driven talking-head generation to date.\n\n| 🔍 Aspect                          | Description                                                         |\n| ---------------------------- | ---------------------------------------------------------------------- |\n| 📈 **Scale**                 | 7,729 speakers, over 1,244 hours of HD/4K footage                      |\n| 🌍 **Diversity**             | Covers **15 languages** and a wide age range (0–60+ years)             |\n| 🧍‍♀️ **Upper-body presence** | Unlike many prior datasets, TalkVid includes upper-body visual context |\n| 📝 **Rich Annotations**      | Comes with **high-quality captions** for every sample                  |\n| 🏞️ **In-the-wild quality**  | Entirely collected in real-world, unconstrained environments           |\n| 🎯 **Quality Assurance**     | Multi-stage filtering with DOVER, CoTracker, and head quality assessment |\n\nCompared to existing benchmarks such as GRID, VoxCeleb, MEAD, or MultiTalk, **TalkVid is the first dataset** to combine:\n\n* **Large-scale multilinguality** across 15+ languages\n* **Wild setting with upper-body inclusion** for more natural synthesis\n* **High-resolution (1080p \u0026 2160p) video** for detailed facial features\n* **Comprehensive metadata** including age, language, quality scores, and captions\n\n\u003e 🧪 Want to push the boundaries of talking-head generation, personalization, or cross-lingual synthesis? TalkVid is your new go-to dataset.\n\n## 🏗️ Data Filtering Pipeline\n\nOur comprehensive data filtering pipeline ensures high-quality dataset construction:\n\n### 1. Video Rough Segmentation\n```bash\ncd data_pipeline/1_video_rough_segmentation\nconda env create -f datapipe.yaml\nconda activate video-py310\nbash rough_segementation.sh\n```\n\n### 2. Video Quality \u0026 Motion Filtering\n```bash\ncd data_pipeline/2_video_quality_motion_filtering\n\n# Quality assessment using DOVER\nbash video_quality_dover.sh\n\n# Motion analysis using CoTracker  \nbash video_motion_cotracker.sh\n```\n\n### 3. Head Detail Filtering\n```bash\ncd data_pipeline/3_head_detail_filtering\nconda env create -f env_head.yml\nconda activate env_head\nbash head_filter.sh\n```\n\n![filter](assets/data_filter.png)\n\n## 🚀 Quick Start\n\n### Environment Setup\n\n```bash\n# Create conda environment\nconda create -n talkvid python=3.10 -y\nconda activate talkvid\n\n# Install dependencies\npip install -r requirements.txt\n\n# Install additional dependencies for video processing\nconda install -c conda-forge 'ffmpeg\u003c7' -y\nconda install torchaudio==2.4.0 pytorch-cuda=12.1 -c pytorch -c nvidia -y\n```\n\n### Model Downloads\n\nBefore running inference, download the required model checkpoints:\n\n```bash\n# Download the model checkpoints\nhuggingface-cli download tk93/V-Express --local-dir V-Express\nmv V-Express/model_ckpts model_ckpts\nmv V-Express/*.bin model_ckpts/v-express\nrm -rf V-Express/\n```\n\n### Quick Inference\n\nWe provide an easy-to-use inference script for generating talking head videos.\n\n#### Command Line Usage\n\n```bash\n# Single sample inference\nbash scripts/inference.sh\n\n# Or run directly with Python\ncd src\npython src/inference.py \\\n    --reference_image_path \"./test_samples/short_case/tys/ref.jpg\" \\\n    --audio_path \"./test_samples/short_case/tys/aud.mp3\" \\\n    --kps_path \"./test_samples/short_case/tys/kps.pth\" \\\n    --output_path \"./output.mp4\" \\\n    --retarget_strategy \"naive_retarget\" \\\n    --num_inference_steps 25 \\\n    --guidance_scale 3.5 \\\n    --context_frames 24\n```\n\n#### Key Parameters\n\n- `--reference_image_path`: Path to the reference portrait image\n- `--audio_path`: Path to the driving audio file\n- `--kps_path`: Path to keypoints file (can be generated automatically)\n- `--retarget_strategy`: Keypoint retargeting strategy (`fix_face`, `naive_retarget`, etc.)\n- `--num_inference_steps`: Number of denoising steps (trade-off between quality and speed)\n- `--context_frames`: Number of context frames for temporal consistency\n\n## 🏋️ Training\n\n### Data Preprocessing\n\nBefore training, preprocess your data:\n\n```bash\ncd src/data_preprocess\nbash env.sh  # Setup preprocessing environment\n# Follow data preprocessing instructions in data_preprocess/readme.md\n```\n\n### Multi-Stage Training\n\nOur model uses a progressive 3-stage training strategy:\n\n```bash\n# Stage 1: Basic motion learning\nexport STAGE=1 TRAIN=\"TalkVid-Core\" GPU=\"0,1\"\nbash scripts/train.sh\n\n# Stage 2: Audio-visual alignment  \nexport STAGE=2 TRAIN=\"TalkVid-Core\" GPU=\"0,1\"\nbash scripts/train.sh\n\n# Stage 3: Temporal consistency and refinement\nexport STAGE=3 TRAIN=\"TalkVid-Core\" GPU=\"0,1\"\nbash scripts/train.sh\n```\n\n### Training Configuration\n\nKey configuration files:\n- `src/configs/stage_1.yaml`: Basic motion and reference net training\n- `src/configs/stage_2.yaml`: Audio projection and alignment training  \n- `src/configs/stage_3.yaml`: Full model with motion module training\n\nTraining supports:\n- **Multi-GPU training** with DeepSpeed ZeRO-2\n- **Mixed precision** (fp16/bf16) for memory efficiency\n- **Gradient checkpointing** to reduce memory usage\n- **Flexible data loading** with configurable batch sizes and augmentations\n\n\u003c!-- ## 🛠️ Model Downloads\n\n| Model Component | Purpose | Download Link | Size |\n|---------|------|----------|------|\n| **TalkVid-Core** | Main talking head model | [🤗 HuggingFace](https://huggingface.co/FreedomIntelligence/TalkVid-Core) | ~2.3GB |\n| **Stable Diffusion VAE** | Video autoencoder | [🤗 HuggingFace](https://huggingface.co/stabilityai/sd-vae-ft-mse) | ~334MB |\n| **Wav2Vec2 Audio Encoder** | Audio feature extraction | [🤗 HuggingFace](https://huggingface.co/facebook/wav2vec2-base-960h) | ~378MB |\n| **InsightFace Models** | Face analysis and landmarks | [Official Site](https://github.com/deepinsight/insightface) | ~1.7GB |\n\n### Pre-trained Checkpoints\n\nWe provide checkpoints for different training stages:\n\n- **Stage 1**: Basic motion and reference learning\n- **Stage 2**: Audio-visual alignment and projection\n- **Stage 3**: Full model with temporal consistency (recommended) --\u003e\n\n## 📊 Evaluation \u0026 Benchmarks\n\n### Evaluation Metrics\n\nWe evaluate our model on multiple aspects:\n\n- **Lip Synchronization**: Sync-C, Sync-D,\n- **Perceptual Quality**: FID, FVD\n\n### TalkVid-Bench\n\nTalkVid-Bench comprises 500 carefully sampled and stratified video clips along four critical demographic and language dimensions: age, gender, ethnicity, and language. This stratified design enables granular analysis of model performance across diverse subgroups, mitigating biases hidden in traditional aggregate evaluations. Each dimension is divided into balanced categories:\n\n- **Age**: 0–19, 19–30, 31–45, 46–60, 60+, with a total of 105 samples.\n- **Gender**: Male, Female, with a total of 100 samples.\n- **Ethnicity**: Black, White, Asian, with a total of 100 samples.\n- **Language**: English, Chinese, Arabic, Polish, German, Russian, French, Korean, Portuguese, Japanese, Thai, Spanish, Italian, Hindi, and Other languages, with a total of 195 samples.\n\n\n### Benchmark Results\n\n![results](assets/benchmark_results.png)\n\nComparison with other baseline training datasets, including HDTF and Hallo3 on **TalkVid-bench** across four dimensions in general.\n\n## 🤝 Contributing\n\nWe welcome contributions to improve TalkVid! Here's how you can help:\n\n### How to Contribute\n\n1. **Fork the repository** and create your feature branch\n2. **Follow our coding standards** and add appropriate tests\n3. **Update documentation** for any new features\n4. **Submit a pull request** with detailed description\n\n### Areas for Contribution\n\n- 🎨 **Model improvements**: New architectures, loss functions, training strategies\n- 🔧 **Data processing**: Enhanced filtering, augmentation techniques\n- 📊 **Evaluation metrics**: New benchmarks and evaluation protocols\n- 🌐 **Multi-language support**: Extend to more languages and cultures\n- ⚡ **Optimization**: Speed and memory improvements\n\n## ❤️ Acknowledgments\n\nWe gratefully acknowledge the following projects and datasets that made TalkVid possible:\n\n* **[V-Express](https://github.com/tencent-ailab/V-Express)**: Foundation architecture and training framework\n* **[Stable Diffusion](https://github.com/Stability-AI/stablediffusion)**: Diffusion model backbone\n* **[InsightFace](https://github.com/deepinsight/insightface)**: Face detection and analysis tools\n* **[DOVER](https://github.com/QualityAssessment/DOVER)**: Video quality assessment\n* **[CoTracker](https://github.com/facebookresearch/co-tracker)**: Motion tracking and analysis\n* **[Wav2Vec2](https://github.com/pytorch/fairseq/tree/main/examples/wav2vec)**: Audio feature extraction\n* **Open source community**: All contributors and researchers advancing talking head synthesis\n\nSpecial thanks to the **V-Express team** for providing excellent open-source infrastructure that enabled this work.\n\n## 📚 Citation\n\nIf our work is helpful for your research, please consider giving a star ⭐ and citing our paper 📝\n\n```bibtex\n@misc{chen2025talkvidlargescalediversifieddataset,\n      title={TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis}, \n      author={Shunian Chen and Hejin Huang and Yexin Liu and Zihan Ye and Pengcheng Chen and Chenghao Zhu and Michael Guan and Rongsheng Wang and Junying Chen and Guanbin Li and Ser-Nam Lim and Harry Yang and Benyou Wang},\n      year={2025},\n      eprint={2508.13618},\n      archivePrefix={arXiv},\n      primaryClass={cs.CV},\n      url={https://arxiv.org/abs/2508.13618}, \n}\n```\n\n## 📄 License\n\n### Dataset License\nThe **TalkVid dataset** is released under [Creative Commons Attribution-NonCommercial 4.0 International License (CC BY-NC 4.0)](https://creativecommons.org/licenses/by-nc/4.0/), allowing only non-commercial research use.\n\n### Code License\nThe **source code** is released under [Apache License 2.0](LICENSE), allowing both academic and commercial use with proper attribution.\n\n## 🌟 Star History\n\n[![Star History Chart](https://api.star-history.com/svg?repos=FreedomIntelligence/TalkVid\u0026type=Date)](https://star-history.com/#FreedomIntelligence/TalkVid\u0026Date)\n\n---\n\n\u003cdiv align=\"center\"\u003e\n\n**🌟 If this project helps you, please give us a Star! 🌟**\n\n[![GitHub stars](https://img.shields.io/github/stars/FreedomIntelligence/TalkVid.svg?style=social\u0026label=Star)](https://github.com/FreedomIntelligence/TalkVid)\n[![GitHub forks](https://img.shields.io/github/forks/FreedomIntelligence/TalkVid.svg?style=social\u0026label=Fork)](https://github.com/FreedomIntelligence/TalkVid)\n\n[🏠 Homepage](https://freedomintelligence.github.io/talk-vid/) | [📄 Paper](https://arxiv.org/abs/2508.13618) | [🤗 Dataset](https://huggingface.co/datasets/FreedomIntelligence/TalkVid) | [💬 Discord](https://discord.gg/talkvid)\n\n\u003c/div\u003e\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FFreedomIntelligence%2FTalkVid","html_url":"https://awesome.ecosyste.ms/projects/github.com%2FFreedomIntelligence%2FTalkVid","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FFreedomIntelligence%2FTalkVid/lists"}