{"id":13737890,"url":"https://github.com/ZhangYuanhan-AI/Bamboo","last_synced_at":"2025-05-08T15:31:58.274Z","repository":{"id":38620455,"uuid":"469003163","full_name":"ZhangYuanhan-AI/Bamboo","owner":"ZhangYuanhan-AI","description":"Bamboo: 4 times larger than ImageNet; 2 time larger than Object365; Built by active learning.","archived":false,"fork":false,"pushed_at":"2024-04-07T02:55:17.000Z","size":5670,"stargazers_count":175,"open_issues_count":4,"forks_count":7,"subscribers_count":5,"default_branch":"main","last_synced_at":"2025-03-31T23:43:23.416Z","etag":null,"topics":["active-learning","dataset-generation","pre-training"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/ZhangYuanhan-AI.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null}},"created_at":"2022-03-12T06:43:40.000Z","updated_at":"2025-03-13T08:52:43.000Z","dependencies_parsed_at":"2024-01-07T17:10:57.465Z","dependency_job_id":"fa1ee69c-b7a4-4bc5-b1d3-8b14f6846bde","html_url":"https://github.com/ZhangYuanhan-AI/Bamboo","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ZhangYuanhan-AI%2FBamboo","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ZhangYuanhan-AI%2FBamboo/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ZhangYuanhan-AI%2FBamboo/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ZhangYuanhan-AI%2FBamboo/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/ZhangYuanhan-AI","download_url":"https://codeload.github.com/ZhangYuanhan-AI/Bamboo/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":253096209,"owners_count":21853559,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["active-learning","dataset-generation","pre-training"],"created_at":"2024-08-03T03:02:04.735Z","updated_at":"2025-05-08T15:31:57.794Z","avatar_url":"https://github.com/ZhangYuanhan-AI.png","language":"Python","funding_links":[],"categories":["Python"],"sub_categories":[],"readme":"![fig1](Figures/teaser.png)\n\n\u003cdiv align=\"center\"\u003e\n\n\u003cdiv\u003e\n    \u003ca href='https://zhangyuanhan-ai.github.io/' target='_blank'\u003eYuanhan Zhang\u003c/a\u003e\u003csup\u003e1\u003c/sup\u003e\u0026emsp;\n    \u003ca href='https://github.com/Davidzhangyuanhan/Bamboo' target='_blank'\u003eQinghong Sun\u003c/a\u003e\u003csup\u003e2\u003c/sup\u003e\u0026emsp;\n    \u003ca href='https://github.com/Davidzhangyuanhan/Bamboo' target='_blank'\u003eYichun Zhou\u003c/a\u003e\u003csup\u003e3\u003c/sup\u003e\u0026emsp;\n    \u003ca href='https://github.com/Davidzhangyuanhan/Bamboo' target='_blank'\u003eZexin He\u003c/a\u003e\u003csup\u003e3\u003c/sup\u003e\u0026emsp;\n    \u003ca href='https://scholar.google.com.hk/citations?user=ngPR1dIAAAAJ\u0026hl=zh-CN' target='_blank'\u003eZhenfei Yin\u003c/a\u003e\u003csup\u003e4\u003c/sup\u003e\u0026emsp;\n    \u003ca href='https://github.com/Davidzhangyuanhan/Bamboo' target='_blank'\u003eKun Wang\u003c/a\u003e\u003csup\u003e4\u003c/sup\u003e\u0026emsp; \u003cbr\u003e\n    \u003ca href='https://lucassheng.github.io/' target='_blank'\u003e Lv Sheng\u003c/a\u003e\u003csup\u003e3\u003c/sup\u003e\u0026emsp;\n    \u003ca href='http://mmlab.siat.ac.cn/yuqiao' target='_blank'\u003eYu Qiao\u003c/a\u003e\u003csup\u003e5\u003c/sup\u003e\u0026emsp;\n    \u003ca href='https://amandajshao.github.io/' target='_blank'\u003eJing Shao\u003c/a\u003e\u003csup\u003e4\u003c/sup\u003e\u0026emsp;\n    \u003ca href='https://liuziwei7.github.io/' target='_blank'\u003eZiwei Liu\u003c/a\u003e\u003csup\u003e1\u003c/sup\u003e\n\u003c/div\u003e\n\u003cdiv\u003e\n    \u003csup\u003e1\u003c/sup\u003eS-Lab, Nanyang Technological University\u0026emsp;\n    \u003csup\u003e2\u003c/sup\u003eBeijing University of Posts and Telecommunication\u0026emsp; \u003cbr\u003e\n    \u003csup\u003e3\u003c/sup\u003eBeihang University\u0026emsp;\n    \u003csup\u003e4\u003c/sup\u003eSenseTime Research\u0026emsp;\n    \u003csup\u003e5\u003c/sup\u003eShanghai AI Laboratory\n\u003c/div\u003e\n\n\u003cbr\u003e\n\n\u003cimg src=\"Figures/teaser_annimation.gif\" alt=\"Pineapple\" style=\"width:360px;height:200px;float:right;margin-top:10px\"\u003e\n\n\u003ch3\u003eTL;DR\u003c/h3\u003e\n\n\nBamboo is a mega-scale and information-dense dataset for classification and detection pre-training. It is built upon integrating 24 public datasets (e.g. ImagenNet, Places365, Object365, OpenImages) and added new annotations through active learning. Bamboo has 69M image classification annotations (\u003cspan style=\"color:#AE2011\"\u003e**4 times larger than ImageNet**\u003c/span\u003e) and 32M object bounding boxes (\u003cspan style=\"color:#AE2011\"\u003e**2 times larger than Object365**\u003c/span\u003e).\n\n\n---\n\n\u003cdiv\u003e\n    \u003ca href='https://arxiv.org/abs/2203.07845' target='_blank'\u003e[Paper]\u003c/a\u003e\n\u003c/div\u003e\n\u003c/div\u003e\n\n## Leaderboard\n[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/bamboo-building-mega-scale-vision-dataset/image-classification-on-dtd)](https://paperswithcode.com/sota/image-classification-on-dtd?p=bamboo-building-mega-scale-vision-dataset) :partying_face:!\\\n[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/bamboo-building-mega-scale-vision-dataset/image-classification-on-food-101-1)](https://paperswithcode.com/sota/image-classification-on-food-101-1?p=bamboo-building-mega-scale-vision-dataset) :partying_face:!\\\n[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/bamboo-building-mega-scale-vision-dataset/fine-grained-image-classification-on-sun397)](https://paperswithcode.com/sota/fine-grained-image-classification-on-sun397?p=bamboo-building-mega-scale-vision-dataset)\\\n[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/bamboo-building-mega-scale-vision-dataset/image-classification-on-flowers-102)](https://paperswithcode.com/sota/image-classification-on-flowers-102?p=bamboo-building-mega-scale-vision-dataset)\\\n[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/bamboo-building-mega-scale-vision-dataset/fine-grained-image-classification-on-caltech)](https://paperswithcode.com/sota/fine-grained-image-classification-on-caltech?p=bamboo-building-mega-scale-vision-dataset)\\\n[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/bamboo-building-mega-scale-vision-dataset/fine-grained-image-classification-on-oxford-1)](https://paperswithcode.com/sota/fine-grained-image-classification-on-oxford-1?p=bamboo-building-mega-scale-vision-dataset) \\\n[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/bamboo-building-mega-scale-vision-dataset/image-classification-on-cifar-100)](https://paperswithcode.com/sota/image-classification-on-cifar-100?p=bamboo-building-mega-scale-vision-dataset)\\\n[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/bamboo-building-mega-scale-vision-dataset/fine-grained-image-classification-on-stanford)](https://paperswithcode.com/sota/fine-grained-image-classification-on-stanford?p=bamboo-building-mega-scale-vision-dataset)\\\n[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/bamboo-building-mega-scale-vision-dataset/image-classification-on-cifar-10)](https://paperswithcode.com/sota/image-classification-on-cifar-10?p=bamboo-building-mega-scale-vision-dataset)\n\n## Updates\n[11/2022] We release [Bamboo-Det](https://entuedu-my.sharepoint.com/:u:/g/personal/yuanhan002_e_ntu_edu_sg/EbkbdpG3MnZKl5oQtcSPI1kBJAqjBQVxdH7F_KpG6UV3Bg?e=5bsd4S). \\\n[10/2022] We won the **first place** in [Computer Vision in the Wild Challenge(ImageNet-1K in Pre-training track)](https://computer-vision-in-the-wild.github.io/eccv-2022/). :partying_face:! \\\n[06/2022] We split Bamboo-CLS into 30 datasets that represent different realms (e.g. car, mammals, food and etc.) in the natural worlds: [HERE](./superclass/README.md) \\\n[06/2022] We release Bamboo-CLS with FC layer, it can classify 115,217 categories. \\\n[06/2022] We release our label system with many useful attributes!. \\\n[03/2022] Bamboo-CLS ResNet-50 and Bamboo-CLS ViT B/16 have been **released**. \\\n[03/2022] [arXiv](https://arxiv.org/abs/2203.07845) paper has been **released**.\n\n## About Bamboo\n\n### Downloads\n- Send your request to yuanhan002@e.ntu.edu.sg. The request should include your name and orgnization as follows. We will notify you by email as soon as possible.\n    ```\n    NAME: XXX\n    ORGANIZATION: XXX (Bamboo is only for academic research and non-commercial use)\n    ```\n\n### Label sytem\nWe provide the hierarchy for our label system at [HERE](https://drive.google.com/file/d/1x53MYBQvRl9Ii3ahYT6chwAfJ48kMFuy/view?usp=sharing). This JSON file includes the following **attrubutes** of each concept. We hope this information will be beneficial for your research.\n\nWe take concept/class ``dog`` as an example.\n- Load JSON file\n    ```\n    #input\n    with open('PATH-TO-JSON-FILE.json') as f:\n    bamboo = json.load(f)\n    print(bamboo.keys())\n    ```\n    ```\n    #output\n    'father2child', 'child2father', 'id2name', 'id2desc', 'id2desc_zh', 'id2name_zh'\n    ```\n- Check the ``id (n02084071)`` of the ``dog`` on [HERE](https://opengvlab.shlab.org.cn/bamboo/search).\n- Get the **attrubutes** you need.\n    - Hypernyms ``bamboo['child2father']['n02084071']``: domestic_animals, canine.\n    - Hyponyms ``bamboo['father2child']['n02084071']``: husky, griffon, shiba inu and etc.\n    - Description ``bamboo['id2desc']['n02084071']``: a member of the genus Canis (probably descended from the common wolf) that has been domesticated by man since prehistoric times; occurs in many breeds.\n    - Included in which public dataset ``bamboo['id2state']['n02084071']['academic']``: openimage, iWildCam2020, STL10, cifar10, iNat2021, ImageNet21K, coco, OpenImage, object365.\n\n\u003cimg src=\"Figures/json_annimation.gif\" alt=\"Pineapple\"\u003e\n\n### Meta File\n- [Class Name](https://drive.google.com/file/d/1qROHNRf9tu8SDUmyL8cIq159ZmvTBY6J/view?usp=share_link)\n- [Class id -\u003e id](https://drive.google.com/file/d/1s8lytaOVU7GrvGTSs8028rAlILXENSs3/view?usp=share_link)\n- [id -\u003e Class id](https://drive.google.com/file/d/164DiLaMX2iN9PwfZqmXF9UitOwEegcGW/view?usp=share_link)\n\n### Special meta file\nDownloading the whole dataset might be unnecessary for most purposes. We provide meta files based on the following dimension.\n- [ ] Class-wise (e.g. dog, car, boat and etc.)\n- [x] Superclass-wise (e.g. animal, transportation, structure and etc.): [HERE](./superclass/README.md)\n\n\n### How to download files from Google drives in the terminal?\n- Install ``gdown`` \n    ```\n    pip install gdown\n    ```\n- get the ``id`` of the files \n    Link, e.g. https://drive.google.com/file/d/1WEKQ_68Y9i9FzakvPYU6Yj5SOvkZCIEm/view?usp=sharing \\\n    id: 1WEKQ_68Y9i9FzakvPYU6Yj5SOvkZCIEm\n- Download \n    ```\n    gdown https://drive.google.com/uc?id=1WEKQ_68Y9i9FzakvPYU6Yj5SOvkZCIEm\n    ```\n\n\n## Model Zoo\n\n### Bamboo-CLS\n| Model     | Link                                                                                         | Data       | cifar10 | cifar100 | food  | pet   | flower | sun   | stanfordcar | dtd   | caltech | fgvc-aircraft | AVG       |\n|-----------|----------------------------------------------------------------------------------------------|------------|---------|----------|-------|-------|--------|-------|-------------|-------|---------|---------------|-----------|\n| ResNet-50 | Official                                                                                     | CLIP       |    88.7 |     70.3 |  86.4 |  88.2 |   96.1 |  73.3 |        78.3 |  76.4 |    89.6 |          49.1 | 79.64     |\n| ViT B/16  | Official                                                                                     | CLIP       |    96.2 |     83.1 |  92.8 |  93.1 |   98.1 |  78.4 |        86.7 |  79.2 |    94.7 |          59.5 | 86.18     |\n| ResNet-50 | [link](https://drive.google.com/file/d/1DrNT5gTK5ouB9c4VMzYpPuAFt-GWA9z3/view?usp=sharing) | Bamboo-CLS | 93.6   | 81.7    | 85.6 | 93.0 | 99.4  | 71.6 | 92.3       | 78.2 | 93.6   | 84.4          | **87.33** |\n| ViT B/16  | [link](https://drive.google.com/file/d/1JNyx81QfB5Fkrho6tBCFqoUYI-VLEvX6/view?usp=sharing) [link_with-FC](https://drive.google.com/file/d/1JNyx81QfB5Fkrho6tBCFqoUYI-VLEvX6/view?usp=sharing) | Bamboo-CLS |   98.5 |    91.0 | 93.3 | 95.3 |  99.7 | 79.5 |       93.9 | 81.9 |   94.8 |          88.8 | **91.65** |\n\n\n### Bamboo-DET \n|Dataset| Model | Link| VOC (AP50) | CITY (MR) | COCO (mmAP) |\n|--------|--------|--------|--------|--------|--------|\n|OpenImages|ResNet-50 + FPN| Official| 82.4 | 16.8| 37.4 |\n|Object365|ResNet-50 + FPN| Official| 86.4 | 14.7| 39.3 |\n|Bamboo-DET([Detectron2](https://github.com/facebookresearch/detectron2))|ResNet-50 + FPN| [link](https://entuedu-my.sharepoint.com/:u:/g/personal/yuanhan002_e_ntu_edu_sg/EbkbdpG3MnZKl5oQtcSPI1kBJAqjBQVxdH7F_KpG6UV3Bg?e=5bsd4S)| 87.5 | 12.6| 43.9 |\n\n\n\n## Getting Started\n\n### Installation\n```\n# Create conda environment\nconda create -n bamboo python=3.7\nconda activate bamboo\n\n# Install Pytorch\nconda install pytorch==1.8.0 torchvision==0.9.0 cudatoolkit=10.2 -c pytorch\n\n# Clone and install\ngit clone https://github.com/Davidzhangyuanhan/Bamboo.git\n```\n### Linear Probe\n#### Step 1: \nDownloading and organizing each downstream dataset as follows\n\n```\ndata\n├── flowers\n│   ├── train/\n│   ├── test/\n│   ├── train_meta.list\n│   ├── test_meta.list\n```\n#### Step 2: \nChanging root and meta in ``Bamboo-Benchmark/configs/100p/config_\\*.yaml``\n\n#### Step 3:\nWriting the path of the downloaded/your model config in ``Bamboo-Benchmark/configs/models_cfg/\\*.yaml``\n\n#### Step 4:\nWriting the name of the downloaded/your model in ``Bamboo-Benchmark/multi_run_100p.sh``\n\n#### Step 5:\n``sh Bamboo-Benchmark/multi_run_100p.sh``\n\n## Citation\nIf you use this code in your research, please kindly cite the following papers.\n\n```\n@misc{zhang2022bamboo,\n      title={Bamboo: Building Mega-Scale Vision Dataset Continually with Human-Machine Synergy}, \n      author={Yuanhan Zhang and Qinghong Sun and Yichun Zhou and Zexin He and Zhenfei Yin and Kun Wang and Lu Sheng and Yu Qiao and Jing Shao and Ziwei Liu},\n      year={2022},\n      eprint={2203.07845},\n      archivePrefix={arXiv},\n      primaryClass={cs.CV}\n}\n```\n\n## Acknowledgement\n\nThanks to Siyu Chen (https://github.com/Siyu-C) for implementing the Bamboo-Benchmark.\n\n\n\u003cdiv align=\"center\"\u003e\n\n[![Hits](https://hits.seeyoufarm.com/api/count/incr/badge.svg?url=https%3A%2F%2Fgithub.com%2FZhangYuanhan-AI%2FBamboo\u0026count_bg=%2379C83D\u0026title_bg=%23555555\u0026icon=\u0026icon_color=%23E7E7E7\u0026title=visitors\u0026edge_flat=false)](https://hits.seeyoufarm.com)\n\n\u003c/div\u003e\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FZhangYuanhan-AI%2FBamboo","html_url":"https://awesome.ecosyste.ms/projects/github.com%2FZhangYuanhan-AI%2FBamboo","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FZhangYuanhan-AI%2FBamboo/lists"}