{"id":13545849,"url":"https://github.com/wangxiao5791509/MultiModal_BigModels_Survey","last_synced_at":"2025-04-02T17:31:52.890Z","repository":{"id":43282611,"uuid":"437808232","full_name":"wangxiao5791509/MultiModal_BigModels_Survey","owner":"wangxiao5791509","description":"[MIR-2023-Survey] A continuously updated paper list for multi-modal pre-trained big models","archived":false,"fork":false,"pushed_at":"2025-02-15T23:36:29.000Z","size":13785,"stargazers_count":286,"open_issues_count":0,"forks_count":17,"subscribers_count":9,"default_branch":"main","last_synced_at":"2025-02-16T00:23:14.625Z","etag":null,"topics":["anhui-university","audio","big-models","depth","event-camera","multi-modal","natural-language","pengchenglab","point-cloud","pre-training","radar","review","rgb-text-audio","self-attention","survey","thermal-infrared","transformers"],"latest_commit_sha":null,"homepage":"","language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/wangxiao5791509.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2021-12-13T09:19:44.000Z","updated_at":"2025-02-15T23:36:33.000Z","dependencies_parsed_at":"2024-03-08T11:24:35.397Z","dependency_job_id":"d8fed4d1-fb36-4713-919e-a86e02c15448","html_url":"https://github.com/wangxiao5791509/MultiModal_BigModels_Survey","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/wangxiao5791509%2FMultiModal_BigModels_Survey","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/wangxiao5791509%2FMultiModal_BigModels_Survey/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/wangxiao5791509%2FMultiModal_BigModels_Survey/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/wangxiao5791509%2FMultiModal_BigModels_Survey/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/wangxiao5791509","download_url":"https://codeload.github.com/wangxiao5791509/MultiModal_BigModels_Survey/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":246860110,"owners_count":20845601,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["anhui-university","audio","big-models","depth","event-camera","multi-modal","natural-language","pengchenglab","point-cloud","pre-training","radar","review","rgb-text-audio","self-attention","survey","thermal-infrared","transformers"],"created_at":"2024-08-01T12:00:23.421Z","updated_at":"2025-04-02T17:31:52.883Z","avatar_url":"https://github.com/wangxiao5791509.png","language":null,"funding_links":[],"categories":["🌟 Topics","Others"],"sub_categories":["Multimodal"],"readme":"\n\u003cdiv align=\"center\"\u003e\n\u003cimg src=\"https://github.com/wangxiao5791509/MultiModal_BigModels/blob/main/figures/MM_PTMs.png\" width=\"1000px\"\u003e\n\u003c/div\u003e\n\n\n\n\n## This github will be continuously updated for the survey paper: \n\n\u003cdiv align=\"center\"\u003e  \n\n**Large-scale Multi-Modal Pre-trained Models: A Comprehensive Survey**, [Xiao Wang](https://wangxiao5791509.github.io/), [Guangyao Chen](https://icgy96.github.io/), Guangwu Qian, Pengcheng Gao, [Xiao-Yong Wei](https://scholar.google.com/citations?user=8kxWTokAAAAJ\u0026hl=zh-CN\u0026oi=ao), [Yaowei Wang](https://scholar.google.com/citations?user=o_DllmIAAAAJ\u0026hl=zh-CN\u0026oi=ao), [Yonghong Tian](https://scholar.google.com/citations?user=fn6hJx0AAAAJ\u0026hl=zh-CN\u0026oi=ao), [Wen Gao](https://scholar.google.com/citations?user=b0vWahYAAAAJ\u0026hl=zh-CN\u0026oi=ao). \n[[arXiv](https://arxiv.org/abs/2302.10035)] \n[[MIR](https://www.mi-research.net/article/doi/10.1007/s11633-022-1410-8)]\n[[极市平台公众号](https://mp.weixin.qq.com/s/5eELXfACI67yZT7WUtMFMA)]\n[[机器智能研究MIR(MIR编辑部)](https://mp.weixin.qq.com/s/yX1DdDCA-nMluzOB6Qz3sw)]\n[[Machine Intelligence Research (Youtube)](https://youtu.be/zQxV-SUz6zU?si=e27cyVjMUdU-XEwd)]\n\n\n------\n  \n\u003c/div\u003e\n\n\n\u003cimg src=\"https://github.com/wangxiao5791509/MultiModal_BigModels_Survey/blob/main/MIRtop3_2025.01.14.png\" width=\"1000px\"\u003e\n\n\n\n## News \n* [2025.01.14] [MIR 下载量 TOP10 好文 (我们的综述下载量：10K次)] [[MIR编辑部-机器智能研究MIR](https://mp.weixin.qq.com/s/UawMKDBEkuPrB4AnlYlatg)]\n* [2024.06.20] [MIR 下载量 TOP10 好文 (我们的综述下载量：6618次)] [[MIR编辑部-机器智能研究MIR](https://mp.weixin.qq.com/s/R9uZqe2ZByYHziTp0nZIIA)]\n\n\n## Framework of this survey\n\u003cimg src=\"https://github.com/wangxiao5791509/MultiModal_BigModels/blob/main/figures/framework.png\" width=\"1000px\"\u003e\n\u003cimg src=\"https://github.com/wangxiao5791509/MultiModal_BigModels/blob/main/figures/milestone.jpg\" width=\"1000px\"\u003e\n\n\n## Review and Surveys\nPlease check this file [[Surveys.md](https://github.com/wangxiao5791509/MultiModal_BigModels/blob/main/Surveys.md)]\n\n\n## Datasets \nPlease check this file [[Datasets.md](https://github.com/wangxiao5791509/MultiModal_BigModels/blob/main/Datasets.md)]\n\n\n## Publications \nPlease check this file [[paperList.md](https://github.com/wangxiao5791509/MultiModal_BigModels/blob/main/paperList.md)]\n\n\n\n\n\n## Experimental Analysis \n\u003cimg src=\"https://github.com/wangxiao5791509/MultiModal_BigModels/blob/main/figures/experimentResults.png\" width=\"1000px\"\u003e\n\u003cimg src=\"https://github.com/wangxiao5791509/MultiModal_BigModels/blob/main/figures/modelsGPUsParmas.png\" width=\"1000px\"\u003e\n\n\n\n## Other Useful Materials \n* [Awesome-Multimodal-Large-Language-Models](https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models)\n\n\n\n## :page_with_curl: BibTex: \nIf you find this survey useful for your research, please cite the following papers: \n\n```bibtex\n@article{wang2022MMPTMSurvey,\n  title={Large-scale Multi-Modal Pre-trained Models: A Comprehensive Survey},\n  author={Wang, Xiao and Chen, Guangyao and Qian, Guangwu and Gao, Pengcheng and Wei, Xiao-Yong and Wang, Yaowei and Tian, Yonghong and Gao, Wen},\n  url={https://github.com/wangxiao5791509/MultiModal_BigModels_Survey},\n  year={2022}\n}\n\n```\n\nIf you have any questions about this survey, please email me via: xiaowang@ahu.edu.cn or wangxiaocvpr@foxmail.com \n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fwangxiao5791509%2FMultiModal_BigModels_Survey","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fwangxiao5791509%2FMultiModal_BigModels_Survey","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fwangxiao5791509%2FMultiModal_BigModels_Survey/lists"}