{"id":31034637,"url":"https://github.com/freedomintelligence/medgen","last_synced_at":"2026-02-15T15:01:35.688Z","repository":{"id":303718657,"uuid":"968710420","full_name":"FreedomIntelligence/MedGen","owner":"FreedomIntelligence","description":"MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos.","archived":false,"fork":false,"pushed_at":"2025-07-09T03:25:39.000Z","size":236,"stargazers_count":22,"open_issues_count":1,"forks_count":1,"subscribers_count":12,"default_branch":"main","last_synced_at":"2025-09-14T02:50:04.122Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":"","language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/FreedomIntelligence.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2025-04-18T15:31:35.000Z","updated_at":"2025-09-13T09:21:16.000Z","dependencies_parsed_at":"2025-07-09T04:29:23.995Z","dependency_job_id":"4e7300e8-b4c0-4d7f-aee2-6b8456528917","html_url":"https://github.com/FreedomIntelligence/MedGen","commit_stats":null,"previous_names":["freedomintelligence/medgen"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/FreedomIntelligence/MedGen","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/FreedomIntelligence%2FMedGen","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/FreedomIntelligence%2FMedGen/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/FreedomIntelligence%2FMedGen/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/FreedomIntelligence%2FMedGen/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/FreedomIntelligence","download_url":"https://codeload.github.com/FreedomIntelligence/MedGen/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/FreedomIntelligence%2FMedGen/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":29481924,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-02-15T11:35:25.641Z","status":"ssl_error","status_checked_at":"2026-02-15T11:34:57.128Z","response_time":118,"last_error":"SSL_connect returned=1 errno=0 peeraddr=140.82.121.5:443 state=error: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2025-09-14T02:46:35.751Z","updated_at":"2026-02-15T15:01:35.654Z","avatar_url":"https://github.com/FreedomIntelligence.png","language":null,"funding_links":[],"categories":[],"sub_categories":[],"readme":"# Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos\n![](https://i.imgur.com/waxVImv.png)\n\n\u003cp align=\"center\"\u003e\n[📃 \u003ca href=\"https://arxiv.org/abs/2507.05675\" target=\"_blank\"\u003ePaper\u003c/a\u003e] ｜ [🤗 \u003ca href=\"https://huggingface.co/datasets/FreedomIntelligence/MedVideoCap-55K\" target=\"_blank\"\u003eDataset\u003c/a\u003e] ｜ [🤗 \u003ca href=\"https://huggingface.co/FreedomIntelligence/MedGen\" target=\"_blank\"\u003eModel (coming)\u003c/a\u003e] ｜ [🚀 \u003ca href=\"https://huggingface.co/blog/wangrongsheng/medvideocap-55k\" target=\"_blank\"\u003eBlog\u003c/a\u003e]\n\u003c/p\u003e\n\n## 🛎️ News\n\n* **`Jul 8, 2025`:** We released our `paper`, `data` and `project`. Models are coming soon. Please stay tuned!\n\n## ⚡ Introduction\n\nRecent advances in video generation have shown remarkable progress in open-domain settings, yet medical video generation remains largely underexplored. Medical videos are critical for applications such as clinical training, education, and simulation, requiring not only high visual fidelity but also strict medical accuracy. However, current models often produce unrealistic or erroneous content when applied to medical prompts, largely due to the lack of large-scale, high-quality datasets tailored to the medical domain. To address this gap, we introduce **MedVideoCap-55K**, the first large-scale, diverse, and caption-rich dataset for medical video generation. It comprises over 55,000 curated clips spanning real-world medical scenarios, providing a strong foundation for training generalist medical video generation models. Built upon this dataset, we develop **MedGen**, which achieves leading performance among open-source models and rivals commercial systems across multiple benchmarks in both visual quality and medical accuracy.\nWe hope our dataset and model can serve as a valuable resource and help catalyze further research in medical video generation.\n\n\u003e [!NOTE]\n\u003e We open-sourced our models, data, and code here.\n\n## 📚 Data\n\n\u003cdiv align=\"center\"\u003e\n\u003cimg src=\"./assets/data.png\" alt=\"MedVideoCap-55K\"\u003e\n\u003c/div\u003e\n\nYou can [⬇️download our full MedVideoCap-55K](https://huggingface.co/datasets/FreedomIntelligence/MedVideoCap-55K) from HuggingFace. Our dataset has several features:\n\n1. **Superior in quantity**. Our dataset comprising 55k medical videos. Supporting video generation across various medical scenarios, it includes medical education, medical imaging, clinical practice, and more. \n2. **Superior in visual quality**. Our dataset is strictly selected from the aspects of aesthetics, temporal consistency, motion smoothness, and clarity assessment. \n3. **Expressive in caption**. Previously proposed medical video datasets typically use category labels as captions. In contrast, our dataset provides expressive and coherent video descriptions with the help of MLLMs.\n\n## 🤩 Acknowledgement\n\nOur works are inspired by the following works.\n\n- [FastVideo](https://github.com/hao-ai-lab/FastVideo): a lightweight framework for accelerating large video diffusion models.\n- [VBench](https://github.com/Vchitect/VBench): a comprehensive benchmark suite for video generative models.\n- [VideoScore](https://github.com/TIGER-AI-Lab/VideoScore): a automatic metrics to simulate fine-grained human feedback for video generation.\n\n## 📖 Citation\n```\n@misc{wang2025medgenunlockingmedicalvideo,\n      title={MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos}, \n      author={Rongsheng Wang and Junying Chen and Ke Ji and Zhenyang Cai and Shunian Chen and Yunjin Yang and Benyou Wang},\n      year={2025},\n      eprint={2507.05675},\n      archivePrefix={arXiv},\n      primaryClass={cs.CV},\n      url={https://arxiv.org/abs/2507.05675}, \n}\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Ffreedomintelligence%2Fmedgen","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Ffreedomintelligence%2Fmedgen","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Ffreedomintelligence%2Fmedgen/lists"}