{"id":13958457,"url":"https://github.com/yuxie11/R2D2","last_synced_at":"2025-07-21T00:30:54.779Z","repository":{"id":38271343,"uuid":"496928310","full_name":"yuxie11/R2D2","owner":"yuxie11","description":null,"archived":false,"fork":false,"pushed_at":"2023-11-09T02:45:05.000Z","size":963,"stargazers_count":155,"open_issues_count":13,"forks_count":22,"subscribers_count":2,"default_branch":"main","last_synced_at":"2024-08-08T13:13:04.037Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/yuxie11.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null}},"created_at":"2022-05-27T08:59:52.000Z","updated_at":"2024-08-07T12:42:32.000Z","dependencies_parsed_at":"2022-07-14T22:30:30.074Z","dependency_job_id":"576be9e0-cbc4-41c7-bfee-401988b9c219","html_url":"https://github.com/yuxie11/R2D2","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/yuxie11%2FR2D2","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/yuxie11%2FR2D2/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/yuxie11%2FR2D2/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/yuxie11%2FR2D2/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/yuxie11","download_url":"https://codeload.github.com/yuxie11/R2D2/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":226849989,"owners_count":17691894,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-08-08T13:01:36.491Z","updated_at":"2024-11-28T02:30:45.866Z","avatar_url":"https://github.com/yuxie11.png","language":"Python","funding_links":[],"categories":["其他_机器视觉"],"sub_categories":["网络服务_其他"],"readme":"# CCMB and R2D2: A Large-scale Chinese Cross-modal Benchmark and A  Vision-Language Framework\n\n\n\n🔥🔥🔥 **CCMB: A Large-scale Chinese Cross-modal Benchmark (ACM MM 2023)**\n\n\n\nThis repo is the official implementation of CCMB and R2D2\u003c/a\u003e. \n\n\u003c!-- \u0026#x2705;\u003ca href=\"http://zero.so.com\"\u003eZero benchmark\u003c/a\u003e is available. The detailed introduction and download URL are in \u003cfont size=4\u003e**http://zero.so.com**\u003c/font\u003e. The 250M data is in --\u003e\n\nCCMB is available. It include pre-train dataset (Zero) and 5 downstream datasets. The detailed introduction and download URL are in \u003cfont size=3\u003e**http://zero.so.com**\u003c/font\u003e. The 250M data is in \u003cfont size=3\u003e**https://pan.baidu.com/s/1gnNbjOdCQdqZ4bRNN1S-Vw?pwd=iau8**\u003c/font\u003e.\n\nR2D2 is a vision-language framework. We release the following code and models:\n\n\u0026#x2705;Pre-trained checkpoints.\n\n\u0026#x2705;Inference demo.\n\n\u0026#x2705;Fine-tuning code and checkpoints for Image-Text Retrieval and Image-Text Matching tasks.\n\n\u003c!-- \u0026#x274C;Pre-training code (coming soon). --\u003e\n\n\u003cimg src=\"image/framework.png\"\u003e\n\n## Performance\nWe show the performance of R2D2\u003csub\u003e\u003cfont size=1.5\u003eViT-L\u003c/font\u003e\u003c/sub\u003e fine-tuned on Flickr30k-CNA dataset. The output of R2D2 is a similarity score between 0 and 1.\n中文 (English) | 乔丹投篮 (Jordan shot) | 乔丹运球 (Jordan dribble)|詹姆斯投篮 (James shot)\n--- | :---: | :---:|--\nSimilarity score|0.99033021|0.91078649|0.61231128\n\n\u003cimg src=\"image/jordan.jpg\"\u003e\n\n## Requirements\n\u003cpre/\u003epip install -r requirements.txt\u003c/pre\u003e \n\n\n\n## Pre-trained checkpoints\nPre-trained image-text pairs | R2D2\u003csub\u003e\u003cfont size=1.5\u003eViT-L\u003c/font\u003e\u003c/sub\u003e | PRD2\u003csub\u003e\u003cfont size=1.5\u003eViT-L\u003c/font\u003e\u003c/sub\u003e\n--- | :---: | :---:\n250M | \u003ca href=\"https://drive.google.com/file/d/18Fd3vGvj0Dz8rPlxROxugjZaF8Z4jf7g/view?usp=sharing\"\u003eDownload\u003c/a\u003e | \u003ca href=\"https://drive.google.com/file/d/15zDdam7_-YT0suA3Wc226vvxcyBxWZ_O/view?usp=sharing\"\u003eDownload\n23M |  \u003ca href=\"https://drive.google.com/file/d/1vvvMv3mTRFGAUojbSJoZiTuqYPJqIquh/view?usp=sharing\"\u003eDownload\u003c/a\u003e | -\n\u003c!-- 2.3M | - | \u003ca href=\"https://drive.google.com/file/d/1SKH-d1Vd-1wn3qUt6YKnep7VsTXfbTK0/view?usp=sharing\"\u003eDownload\u003c/a\u003e | - --\u003e\n\n## Fine-tuned checkpoints\nDataset | R2D2\u003csub\u003e\u003cfont size=1.5\u003eViT-B\u003c/font\u003e\u003c/sub\u003e(23M) | \n--- | :---: \nFlickr-CNA | \u003ca href=\"https://drive.google.com/file/d/1qgbDIqSUBqGz6rGCGKtW14wzPTIcnLLg/view?usp=sharing\"\u003eDownload\u003c/a\u003e \nIQR | \u003ca href=\"https://drive.google.com/file/d/1lQ6rqMXukzul6XQJ8uZe_BQh-tuL1KNm/view?usp=sharing\"\u003eDownload\u003c/a\u003e \nICR | \u003ca href=\"https://drive.google.com/file/d/15Zsr8n49AEjOi2MkOfp1ZtUAKGss_Xbz/view?usp=sharing\"\u003eDownload\u003c/a\u003e \nIQM | \u003ca href=\"https://drive.google.com/file/d/1JxLL6mlhDz_pjoUuyeeRVTHw0q8gW5et/view?usp=sharing\"\u003eDownload\u003c/a\u003e \nICM | \u003ca href=\"https://drive.google.com/file/d/1FI9RzJT-0j30ftcfkx0zDF2v3T7iXZtG/view?usp=sharing\"\u003eDownload\u003c/a\u003e \n\n## Inference demo\n- To evaluate the pretrained R2D2 model on image-text pairs, run:\n    \u003cpre\u003epython r2d2_inference_demo.py\u003c/pre\u003e \n- To evaluate the pretrained PRD2 model on image-text pairs, run:\n    \u003cpre\u003epython prd2_inference_demo.py\u003c/pre\u003e \n\n## Downstream Tasks\n1. Download datasets and pretrained models.\n    for ICR, IQR, ICM, IQM tasks, after downloading you should see the following folder structure:\n    ```\n    ├── IQR_IQM_ICR_ICM_images\n    │   \n    ├── IQR\n    │   ├── train\n    │   └── val\n    ├── ICR\n    │   ├── train\n    │   └── val\n    ├── IQM\n    │   ├── train\n    │   └── val\n    │── ICM\n    │   ├── train\n    │   └── val\n    for Flickr30k-CNA, after downloading you should see the following folder structure:\n    ```\n    ├── Flickr30k-images\n    │   \n    ├── train\n    │   \n    ├── val\n    │  \n    └── test\n    ```\n  2. In config/retrieval_*.yaml, set the paths for the dataset and pretrain model paths.\n  3. Run fine-tuning for the Image-Text Retrieval task.\n      ```\n      sh train_r2d2_retrieval.sh\n      ```\n  4. Run fine-tuning for the Image-Text Matching task.\n      ```\n      sh train_r2d2_matching.sh\n      ```\n    \n### Citation\nIf you find this dataset and code useful for your research, please consider citing.\n\u003cpre\u003e\n@inproceedings{xie2023ccmb,\n  title={CCMB: A Large-scale Chinese Cross-modal Benchmark},\n  author={Xie, Chunyu and Cai, Heng and Li, Jincheng and Kong, Fanjing and Wu, Xiaoyu and Song, Jianfei and Morimitsu, Henrique and Yao, Lin and Wang, Dexin and Zhang, Xiangzheng and others},\n  booktitle={Proceedings of the 31st ACM International Conference on Multimedia},\n  pages={4219--4227},\n  year={2023}\n}\u003c/pre\u003e\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fyuxie11%2FR2D2","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fyuxie11%2FR2D2","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fyuxie11%2FR2D2/lists"}