{"id":28103453,"url":"https://github.com/westlake-ai/semireward","last_synced_at":"2025-05-13T20:37:53.695Z","repository":{"id":198420263,"uuid":"682708830","full_name":"Westlake-AI/SemiReward","owner":"Westlake-AI","description":"[ICLR 2024] SemiReward: A General Reward Model for Semi-supervised Learning","archived":false,"fork":false,"pushed_at":"2024-06-10T22:01:56.000Z","size":1181,"stargazers_count":50,"open_issues_count":0,"forks_count":2,"subscribers_count":2,"default_branch":"main","last_synced_at":"2024-06-11T00:50:01.670Z","etag":null,"topics":["audio-classification","cifar-100","computer-vision","esc-50","label-noise","machine-learning","natural-language-processing","regression","reward-model","semi-supervised-learning","transformer","vision-transformer","weakly-supervised-learning","yahoo-answers"],"latest_commit_sha":null,"homepage":"https://arxiv.org/abs/2310.03013","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/Westlake-AI.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGE_LOG.md","contributing":null,"funding":null,"license":"LICENSE.txt","code_of_conduct":"CODE_OF_CONDUCT.md","threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":"SECURITY.md","support":"SUPPORT.md","governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2023-08-24T18:50:39.000Z","updated_at":"2024-06-10T22:02:00.000Z","dependencies_parsed_at":"2024-06-11T00:22:27.239Z","dependency_job_id":null,"html_url":"https://github.com/Westlake-AI/SemiReward","commit_stats":null,"previous_names":["westlake-ai/semireward"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Westlake-AI%2FSemiReward","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Westlake-AI%2FSemiReward/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Westlake-AI%2FSemiReward/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Westlake-AI%2FSemiReward/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/Westlake-AI","download_url":"https://codeload.github.com/Westlake-AI/SemiReward/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":254022060,"owners_count":22001050,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["audio-classification","cifar-100","computer-vision","esc-50","label-noise","machine-learning","natural-language-processing","regression","reward-model","semi-supervised-learning","transformer","vision-transformer","weakly-supervised-learning","yahoo-answers"],"created_at":"2025-05-13T20:37:52.844Z","updated_at":"2025-05-13T20:37:53.686Z","avatar_url":"https://github.com/Westlake-AI.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cdiv id=\"top\"\u003e\u003c/div\u003e\n\u003c!--\n*** Thanks for checking out the Best-README-Template. If you have a suggestion\n*** that would make this better, please fork the repo and create a pull request\n*** or simply open an issue with the tag \"enhancement\".\n*** Don't forget to give the project a star!\n*** Thanks again! Now go create something AMAZING! :D\n--\u003e\n\n\u003c!-- PROJECT SHIELDS --\u003e\n\n\u003c!--\n*** I'm using markdown \"reference style\" links for readability.\n*** Reference links are enclosed in brackets [ ] instead of parentheses ( ).\n*** See the bottom of this document for the declaration of the reference variables\n*** for contributors-url, forks-url, etc. This is an optional, concise syntax you may use.\n*** https://www.markdownguide.org/basic-syntax/#reference-style-links\n--\u003e\n\n\u003c!-- [![Contributors][contributors-shield]][contributors-url]\n[![Forks][forks-shield]][forks-url]\n[![Stargazers][stars-shield]][stars-url]\n[![Issues][issues-shield]][issues-url] --\u003e\n\u003c!-- \n***[![MIT License][license-shield]][license-url]\n--\u003e\n\n\u003c!-- PROJECT LOGO --\u003e\n\n\u003cdiv align=\"center\"\u003e\n\u003ch2\u003e\u003ca href=\"https://arxiv.org/abs/2310.03013\"\u003eSemiReward: A General Reward Model for Semi-supervised Learning (ICLR 2024)\u003c/a\u003e \u003c/h2\u003e\n\n[Siyuan Li](https://lupin1998.github.io/)\u003csup\u003e\\*,1,2\u003c/sup\u003e, [Weiyang Jin](https://scholar.google.co.id/citations?hl=zh-CN\u0026user=cazmdIMAAAAJ)\u003csup\u003e\\*,1\u003c/sup\u003e, [Zedong Wang](https://zedongwang.netlify.app/)\u003csup\u003e1,2\u003c/sup\u003e, [Fang Wu](https://smiles724.github.io/)\u003csup\u003e1,2\u003c/sup\u003e, [Zicheng Liu](https://pone7.github.io/)\u003csup\u003e1,2\u003c/sup\u003e, [Chen Tan](https://chengtan9907.github.io/)\u003csup\u003e1,2\u003c/sup\u003e, [Stan Z. Li](https://scholar.google.com/citations?user=Y-nyLGIAAAAJ\u0026hl=zh-CN)\u003csup\u003e†,1\u003c/sup\u003e\n\n\u003csup\u003e1\u003c/sup\u003e[Westlake University](https://westlake.edu.cn/), \u003csup\u003e2\u003c/sup\u003e[Zhejiang University](https://www.zju.edu.cn/english/)\n\u003c/div\u003e\n\n\u003cp align=\"center\"\u003e\n\u003ca href=\"https://arxiv.org/abs/2310.03013\" alt=\"arXiv\"\u003e\n    \u003cimg src=\"https://img.shields.io/badge/arXiv-2310.03013-b31b1b.svg?style=flat\" /\u003e\u003c/a\u003e\n\u003ca href=\"https://github.com/Westlake-AI/SemiReward/blob/main/LICENSE.txt\" alt=\"license\"\u003e\n    \u003cimg src=\"https://img.shields.io/badge/license-Apache--2.0-%23B7A800\" /\u003e\u003c/a\u003e\n\u003ca href=\"https://openreview.net/forum?id=dnqPvUjyRI\" alt=\"Colab\"\u003e\n    \u003cimg src=\"https://img.shields.io/badge/openreview-SemiReward-blue\" /\u003e\u003c/a\u003e\n\u003c/p\u003e\n\nSemi-supervised Reward framework (SemiReward) is designed to predict reward scores to evaluate and filter out high-quality pseudo labels, which is pluggable to mainstream Semi-Supervised Learning (SSL) methods in wide task types and scenarios. The results and details are reported in [our paper](https://arxiv.org/abs/2310.03013). The implementations and models of **SemiReward** are based on **USB** codebase.\n_**USB** is a Pytorch-based Python package for SSL. It is easy-to-use/extend, *affordable* to small groups, and comprehensive for developing and evaluating SSL algorithms. USB provides the implementation of 14 SSL algorithms based on Consistency Regularization, and 15 tasks for evaluation from CV, NLP, and Audio domain. More details can be seen in [Semi-supervised Learning](https://github.com/microsoft/Semi-supervised-learning)._\n\n\u003cp align=\"center\"\u003e\n\u003cimg src=\"https://github-production-user-asset-6210df.s3.amazonaws.com/44519745/276408256-f860de7b-bb3c-42c3-8ef9-2f1a91dac55b.png\" width=100% height=100% \nclass=\"center\"\u003e\n\u003c/p\u003e\n\n\u003c!-- TABLE OF CONTENTS --\u003e\n\n\u003cdetails\u003e\n  \u003csummary\u003eTable of Contents\u003c/summary\u003e\n  \u003col\u003e\n    \u003cli\u003e\u003ca href=\"#news-and-updates\"\u003eNews and Updates\u003c/a\u003e\u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#intro\"\u003eIntroduction\u003c/a\u003e\u003c/li\u003e\n    \u003cli\u003e\n      \u003ca href=\"#getting-started\"\u003eGetting Started\u003c/a\u003e\n      \u003cul\u003e\n        \u003cli\u003e\u003ca href=\"#prerequisites\"\u003ePrerequisites\u003c/a\u003e\u003c/li\u003e\n        \u003cli\u003e\u003ca href=\"#installation\"\u003eInstallation\u003c/a\u003e\u003c/li\u003e\n      \u003c/ul\u003e\n    \u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#usage\"\u003eUsage\u003c/a\u003e\u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#contributing\"\u003eCommunity\u003c/a\u003e\u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#license\"\u003eLicense\u003c/a\u003e\u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#acknowledgments\"\u003eAcknowledgments\u003c/a\u003e\u003c/li\u003e\n  \u003c/ol\u003e\n\u003c/details\u003e\n\n\u003c!-- Introduction --\u003e\n\n## Introduction\n\nSemi-supervised learning (SSL) has witnessed great progress with various improvements in the self-training framework with pseudo labeling. The main challenge is how to distinguish high-quality pseudo labels against the confirmation bias. However, existing pseudo-label selection strategies are limited to pre-defined schemes or complex hand-crafted policies specially designed for classification, failing to achieve high-quality labels, fast convergence, and task versatility simultaneously. To these ends, we propose a Semi-supervised Reward framework (SemiReward) that predicts reward scores to evaluate and filter out high-quality pseudo labels, which is pluggable to mainstream SSL methods in wide task types and scenarios. To mitigate confirmation bias, SemiReward is trained online in two stages with a generator model and subsampling strategy. With classification and regression tasks on 13 standard SSL benchmarks of three modalities, extensive experiments verify that SemiReward achieves significant performance gains and faster convergence speeds upon Pseudo Label, FlexMatch, and Free/SoftMatch.\n\n\u003cp align=\"center\"\u003e\n\u003cimg src=\"https://github-production-user-asset-6210df.s3.amazonaws.com/44519745/276409517-7a7907f5-01b9-4953-818d-767c0dfb9c6b.png\" width=90% \nclass=\"center\"\u003e\n\u003c/p\u003e\n\n\u003c!-- News and Updates --\u003e\n\n## News and Updates\n\n- [01/16/2024] SemiReward v0.2.0 has been updated and accepted by [ICLR'2024](https://openreview.net/forum?id=dnqPvUjyRI).\n- [10/18/2023] SemiReward v0.1.0 has been released.\n\n\u003cp align=\"right\"\u003e(\u003ca href=\"#top\"\u003eback to top\u003c/a\u003e)\u003c/p\u003e\n\n\n## Getting Started\n\nFirst, you need to set up USB locally.\nTo get a local copy up, running follow these simple example steps.\n\n### Prerequisites\n\nUSB is built on pytorch, with torchvision, torchaudio, and transformers.\n\nTo install the required packages, you can create a conda environment:\n\n```sh\nconda create --name semireward python=3.8\n```\n\nthen use pip to install required packages:\n\n```sh\npip install -r requirements.txt\n```\n\nFrom now on, you can start use USB by typing \n\n```sh\npython train.py --c config/usb_cv/fixmatch/fixmatch_cifar100_200_0.yaml\n```\n\n### Installation\n\nUSB provide a Python package *semilearn* of USB for users who want to start training/testing the supported SSL algorithms on their data quickly:\n\n```sh\npip install semilearn\n```\n\nYou can also develop your own SSL algorithm and evaluate it by cloning SemiReward (USB):\n```sh\ngit clone https://github.com/Westlake-AI/SemiReward.git\n```\n\n\u003cp align=\"right\"\u003e(\u003ca href=\"#top\"\u003eback to top\u003c/a\u003e)\u003c/p\u003e\n\n\n### Prepare Datasets\n\nThe detailed instructions for downloading and processing are shown in [Dataset Download](./preprocess/). Please follow it to download datasets before running or developing algorithms.\n\n\u003cp align=\"right\"\u003e(\u003ca href=\"#top\"\u003eback to top\u003c/a\u003e)\u003c/p\u003e\n\n\n## Usage\n\n### Start with Docker\nThe following steps to train your own SemiReward model just as same with USB.\n\n**Step1: Check your environment**\n\nYou need to properly install Docker and nvidia driver first. To use GPU in a docker container\nYou also need to install nvidia-docker2 ([Installation Guide](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/install-guide.html#docker)).\nThen, Please check your CUDA version via `nvidia-smi`\n\n**Step2: Clone the project**\n\n```shell\ngit clone https://github.com/microsoft/Semi-supervised-learning.git\n```\n\n**Step3: Build the Docker image**\n\nBefore building the image, you may modify the [Dockerfile](Dockerfile) according to your CUDA version.\nThe CUDA version we use is 11.6. You can change the base image tag according to [this site](https://hub.docker.com/r/nvidia/cuda/tags).\nYou also need to change the `--extra-index-url` according to your CUDA version in order to install the correct version of Pytorch.\nYou can check the url through [Pytorch website](https://pytorch.org).\n\nUse this command to build the image\n\n```shell\ncd Semi-supervised-learning \u0026\u0026 docker build -t semilearn .\n```\n\nJob done. You can use the image you just built for your own project. Don't forget to use the argument `--gpu` when you want\nto use GPU in a container.\n\n### Training\n\nHere is an example to train one of baselines FlexMatch on CIFAR-100 with 200 labels. Training other supported algorithms (on other datasets with different label settings) can be specified by a config file:\n\n```sh\npython train.py --c config/usb_cv/flexmatch/flexmatch_cifar100_200_0.yaml\n```\n\nHere is an example to train FlexMatch with SemiReward on CIFAR-100 with 200 labels. Training other baselines with SemiReward can be specified by a config file:\n\n```sh\npython train.py --c config/SemiReward/usb_cv/flexmatch/flexmatch_cifar100_200_0.yaml\n```\nYou can change hyperparameters for SemiReward by configurations (.yaml files) like other baselines. If you want to change loss or something is fixed in our method for SemiReward, it is recommanded to open flie from:\n\n```sh\nsemilearn/algorithms/srflexmatch/srflexmatch.py\n```\n\n**Tips:** Semireward uses **4GPUs** for training by default. Also, for users in some areas of China, huggingface region locking occurs, so local pre-training weights need to be used when using the Bert and huBert models. Take the Bert model as an example, you need to focus on `./semilearn/datasets/collactors/nlp_collactor.py`, find line 102 to change it's address into your local folder for Bert. Also, in file `./semilearn/nets/bert/bert.py` line 13, it need to the same way to adjust.\n\n### Evaluation\n\nAfter training, you can check the evaluation performance on training logs, or running evaluation script:\n\n```\npython eval.py --dataset cifar100 --num_classes 100 --load_path /PATH/TO/CHECKPOINT\n```\n\n\u003cp align=\"center\"\u003e\n\u003cimg src=\"https://github.com/Westlake-AI/openmixup/assets/44519745/266f5667-9e5f-44c9-ba63-f3e8b733d5a9\" width=95% \nclass=\"center\"\u003e\n\u003c/p\u003e\n\n### Develop\n\nCheck the developing documentation for creating your own SSL algorithm!\n\n_For more examples, please refer to the [Documentation](https://example.com)_\n\n\u003cp align=\"right\"\u003e(\u003ca href=\"#top\"\u003eback to top\u003c/a\u003e)\u003c/p\u003e\n\n\n## Contributing\n\nIf you have any ideas to improve SemiReward, we welcome your contributions! Feel free to fork the repository and submit a pull request. Alternatively, you can open an issue and label it as \"enhancement.\" Don't forget to show your support by giving the project a star! Thank you once more!\n\n1. Fork the project\n2. Create your branch (`git checkout -b your_name/your_branch`)\n3. Commit your changes (`git commit -m 'Add some features'`)\n4. Push to the branch (`git push origin your_name/your_branch`)\n5. Open a Pull Request\n\n\n## License\n\nDistributed under the MIT License. See `LICENSE.txt` for more information.\n\n\u003cp align=\"right\"\u003e(\u003ca href=\"#top\"\u003eback to top\u003c/a\u003e)\u003c/p\u003e\n\n\n## Citation\n\nPlease consider citing us if you find this project helpful for your project/paper:\n\n```\n@inproceedings{iclr2024semireward,\n  title={SemiReward: A General Reward Model for Semi-supervised Learning},\n  author={Siyuan Li and Weiyang Jin and Zedong Wang and Fang Wu and Zicheng Liu and Cheng Tan and Stan Z. Li},\n  booktitle={International Conference on Learning Representations},\n  year={2024}\n}\n```\n\n\n## Acknowledgments\n\nSemiReward's implementation is mainly based on the following codebases. We gratefully thank the authors for their wonderful works:\n\n- [USB](https://github.com/microsoft/Semi-supervised-learning)\n- [TorchSSL](https://github.com/TorchSSL/TorchSSL)\n- [FixMatch](https://github.com/google-research/fixmatch)\n- [CoMatch](https://github.com/salesforce/CoMatch)\n- [SimMatch](https://github.com/KyleZheng1997/simmatch)\n- [HuggingFace](https://huggingface.co/docs/transformers/index)\n- [Pytorch Lighting](https://github.com/Lightning-AI/lightning)\n- [README Template](https://github.com/othneildrew/Best-README-Template)\n\n## Contribution and Contact\n\nFor adding new features, looking for helps, or reporting bugs associated with `SemiReward`, please open a [GitHub issue](https://github.com/Westlake-AI/SemiReward/issues) and [pull request](https://github.com/Westlake-AI/SemiReward/pulls) with the tag \"new features\" or \"help wanted\". Feel free to contact us through email if you have any questions.\n\n- Siyuan Li (lisiyuan@westlake.edu.cn), Westlake University \u0026 Zhejiang University\n- Weiyang Jin (wayneyjin@gmail.com), Westlake University \u0026 Beijing Jiaotong University\n\n\u003cp align=\"right\"\u003e(\u003ca href=\"#top\"\u003eback to top\u003c/a\u003e)\u003c/p\u003e\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fwestlake-ai%2Fsemireward","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fwestlake-ai%2Fsemireward","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fwestlake-ai%2Fsemireward/lists"}