{"id":19932120,"url":"https://github.com/amazon-science/prompt-pretraining","last_synced_at":"2025-07-19T10:38:15.854Z","repository":{"id":180419483,"uuid":"617276726","full_name":"amazon-science/prompt-pretraining","owner":"amazon-science","description":"Official implementation for the paper \"Prompt Pre-Training with Over Twenty-Thousand Classes for Open-Vocabulary Visual Recognition\"","archived":false,"fork":false,"pushed_at":"2024-05-03T20:49:08.000Z","size":16330,"stargazers_count":258,"open_issues_count":7,"forks_count":9,"subscribers_count":5,"default_branch":"main","last_synced_at":"2025-06-30T19:50:37.970Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/amazon-science.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":"CODE_OF_CONDUCT.md","threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2023-03-22T03:31:46.000Z","updated_at":"2025-06-03T07:47:44.000Z","dependencies_parsed_at":null,"dependency_job_id":"75379ed6-c60e-4947-a485-db31fda4662a","html_url":"https://github.com/amazon-science/prompt-pretraining","commit_stats":null,"previous_names":["amazon-science/prompt-pretraining"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/amazon-science/prompt-pretraining","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/amazon-science%2Fprompt-pretraining","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/amazon-science%2Fprompt-pretraining/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/amazon-science%2Fprompt-pretraining/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/amazon-science%2Fprompt-pretraining/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/amazon-science","download_url":"https://codeload.github.com/amazon-science/prompt-pretraining/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/amazon-science%2Fprompt-pretraining/sbom","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":265919078,"owners_count":23849296,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-11-12T23:09:10.149Z","updated_at":"2025-07-19T10:38:15.833Z","avatar_url":"https://github.com/amazon-science.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# Prompt Pre-Training with Over Twenty-Thousand Classes for Open-Vocabulary Visual Recognition\n\n\u003ch5 align=\"center\"\u003e\u003ci\u003e\"Scaling up prompt learning on ImageNet-21K achieves SOTA on 21 downstream datasets.\"\u003c/i\u003e\u003c/h5\u003e\n\n\u003e [**Prompt Pre-Training with Over Twenty-Thousand Classes for Open-Vocabulary Visual Recognition**](https://arxiv.org/abs/2304.04704)\u003cbr\u003e\n\u003e [Shuhuai Ren](https://renshuhuai-andy.github.io/), [Aston Zhang](https://www.astonzhang.com/), [Yi Zhu](https://bryanyzhu.github.io/), [Shuai Zhang](https://shuaizhang.tech/), [Shuai Zheng](https://szhengac.github.io/), [Mu Li](http://www.cs.cmu.edu/~muli/), [Alex Smola](https://alex.smola.org/), [Xu Sun](https://xusun.org/index.htm)\n\n\n[![paper](https://img.shields.io/badge/arXiv-Paper-\u003cCOLOR\u003e.svg)](https://arxiv.org/abs/2304.04704) \n[![Colab](https://colab.research.google.com/assets/colab-badge.svg)](\nhttps://colab.research.google.com/drive/1OEFw1GfKXogx8mdFS2pClLjPK3aPEZyY?usp=sharing)\n\n[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/prompt-pre-training-with-twenty-thousand/prompt-engineering-on-imagenet-21k)](https://paperswithcode.com/sota/prompt-engineering-on-imagenet-21k?p=prompt-pre-training-with-twenty-thousand) \n[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/prompt-pre-training-with-twenty-thousand/prompt-engineering-on-imagenet-a)](https://paperswithcode.com/sota/prompt-engineering-on-imagenet-a?p=prompt-pre-training-with-twenty-thousand)\n[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/prompt-pre-training-with-twenty-thousand/prompt-engineering-on-imagenet-r)](https://paperswithcode.com/sota/prompt-engineering-on-imagenet-r?p=prompt-pre-training-with-twenty-thousand)\n[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/prompt-pre-training-with-twenty-thousand/prompt-engineering-on-imagenet-s)](https://paperswithcode.com/sota/prompt-engineering-on-imagenet-s?p=prompt-pre-training-with-twenty-thousand)\n[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/prompt-pre-training-with-twenty-thousand/open-vocabulary-semantic-segmentation-on-coco)](https://paperswithcode.com/sota/open-vocabulary-semantic-segmentation-on-coco?p=prompt-pre-training-with-twenty-thousand)\n[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/prompt-pre-training-with-twenty-thousand/open-vocabulary-semantic-segmentation-on-5)](https://paperswithcode.com/sota/open-vocabulary-semantic-segmentation-on-5?p=prompt-pre-training-with-twenty-thousand) [![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/prompt-pre-training-with-twenty-thousand/open-vocabulary-object-detection-on-lvis-v1-0)](https://paperswithcode.com/sota/open-vocabulary-object-detection-on-lvis-v1-0?p=prompt-pre-training-with-twenty-thousand) \n\n# :rocket: News\n* **(Jul 11, 2023)** \n  * Inference demo for object detection in Jupyter. \n* **(May 31, 2023)** \n  * Inference demo for image classification in Google Colab. \n* **(Mar 22, 2023)** \n  * Codes for prompt pretraining (POMP) on ImageNet-21K, cross-dataset and cross-task evaluation.\n  * Checkpoints of pre-trained POMP prompts, segmentation backbones, and detection backbones.\n\u003chr /\u003e\n\n## Highlights\n\n![main figure](docs/main_figure.png)\n\n\n## Main Contributions\n\n1) We introduce a prompt pre-training method POMP, which fisrt enables prompt learning on large-scale datasets like ImageNet-21K with over twenty-thousand classes.\n2) POMP is memory and computation efficient. Compared with previous methods like CoOp, it achieves comparable accuracy on ImageNet-1K with only 19\\% GPU memory and 50\\% training time.\n3) POMP achieves new SOTAs on various open-vocabulary visual recognition datasets and tasks.\n\n## Installation \nFor installation and other package requirements, please follow the instructions detailed in [INSTALL.md](docs/INSTALL.md). \n\n## Data preparation\nPlease follow the instructions at [DATASETS.md](docs/DATASETS.md) to prepare all datasets.\n\n## Pre-trained Models\nPlease follow the instructions at [MODELS.md](docs/MODELS.md) to prepare all pre-trained models.\n\n## Training and Evaluation\nPlease refer to the [RUN.md](docs/RUN.md) for detailed instructions on training, evaluating and reproducing the results.\n\n\n\u003chr /\u003e\n\n## Contact\nIf you have any questions, please feel free to create an issue on this repository.\n\n## Citation\nIf you find this code useful for your research, please consider citing:\n```\n@article{ren2023pomp,\n  title={Prompt Pre-Training with Twenty-Thousand Classes for Open-Vocabulary Visual Recognition},\n  author={Ren, Shuhuai and Zhang, Aston and Zhu, Yi and Zhang, Shuai and Zheng, Shuai and Li, Mu and Smola, Alex and Sun, Xu},\n  journal={arXiv preprint arXiv:2304.04704},\n  year={2023}\n}\n```\n\n## Acknowledgements\n\nOur code is based on [CoOp](https://github.com/KaiyangZhou/CoOp), [MaPLe](https://github.com/muzairkhattak/multimodal-prompt-learning), [Dassl](https://github.com/KaiyangZhou/Dassl.pytorch), [Detic](https://github.com/facebookresearch/Detic) and [ZSSeg](https://github.com/MendelXu/zsseg.baseline) repositories. We thank the authors for releasing their code. \n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Famazon-science%2Fprompt-pretraining","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Famazon-science%2Fprompt-pretraining","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Famazon-science%2Fprompt-pretraining/lists"}