{"id":50890867,"url":"https://github.com/VIPL-SLP/VAC_CSLR","last_synced_at":"2026-07-22T02:01:00.520Z","repository":{"id":37698531,"uuid":"355095060","full_name":"VIPL-SLP/VAC_CSLR","owner":"VIPL-SLP","description":"Visual Alignment Constraint for Continuous Sign Language Recognition. ( ICCV 2021)","archived":false,"fork":false,"pushed_at":"2023-03-23T04:23:59.000Z","size":365,"stargazers_count":131,"open_issues_count":26,"forks_count":20,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-05-27T11:57:50.600Z","etag":null,"topics":["continuous-sign-language","sequence-learning","sign-language-recognition"],"latest_commit_sha":null,"homepage":"https://openaccess.thecvf.com/content/ICCV2021/html/Min_Visual_Alignment_Constraint_for_Continuous_Sign_Language_Recognition_ICCV_2021_paper.html","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/VIPL-SLP.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null}},"created_at":"2021-04-06T07:18:43.000Z","updated_at":"2025-05-25T20:17:22.000Z","dependencies_parsed_at":"2024-03-19T08:44:25.210Z","dependency_job_id":null,"html_url":"https://github.com/VIPL-SLP/VAC_CSLR","commit_stats":null,"previous_names":["vipl-slp/vac_cslr","ycmin95/vac_cslr"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/VIPL-SLP/VAC_CSLR","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/VIPL-SLP%2FVAC_CSLR","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/VIPL-SLP%2FVAC_CSLR/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/VIPL-SLP%2FVAC_CSLR/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/VIPL-SLP%2FVAC_CSLR/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/VIPL-SLP","download_url":"https://codeload.github.com/VIPL-SLP/VAC_CSLR/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/VIPL-SLP%2FVAC_CSLR/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":35743464,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-07-20T02:08:10.276Z","status":"online","status_checked_at":"2026-07-22T02:00:06.236Z","response_time":124,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["continuous-sign-language","sequence-learning","sign-language-recognition"],"created_at":"2026-06-15T21:00:23.019Z","updated_at":"2026-07-22T02:01:00.514Z","avatar_url":"https://github.com/VIPL-SLP.png","language":"Python","funding_links":[],"categories":["🔤 Sign-to-Text Projects"],"sub_categories":[],"readme":"# VAC_CSLR\n[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/visual-alignment-constraint-for-continuous/sign-language-recognition-on-rwth-phoenix)](https://paperswithcode.com/sota/sign-language-recognition-on-rwth-phoenix?p=visual-alignment-constraint-for-continuous)\n\nThis repo holds the code of the paper: Visual Alignment Constraint for Continuous Sign Language Recognition.(ICCV 2021) [[paper]](https://openaccess.thecvf.com/content/ICCV2021/html/Min_Visual_Alignment_Constraint_for_Continuous_Sign_Language_Recognition_ICCV_2021_paper.html)\n\n\u003cimg src=\".\\framework.png\" alt=\"framework\" style=\"zoom: 80%;\" /\u003e\n\n---\n### Update (2022.05.14)\n\nIn recent experiments, we found an implementation improvement about the proposed method. In our early experiments, we adopt `nn.DataParallel` to parallel the visual feature extractor on multiple GPUs. However, only statistic updated on device 0 is kept during training ([Dataparallel](https://pytorch.org/docs/stable/generated/torch.nn.DataParallel.html)), which leads to unstable training results (results may be different when adopting different numbers of GPUs and batch sizes). Therefore, we adopt [syncBN](https://github.com/vacancy/Synchronized-BatchNorm-PyTorch) in this update, the training schedule can be shorten to 40 epochs, and the relevant results are also provided. Experimental results on other datasets will be provided in our future journal version.\n\n```python\nfrom modules.sync_batchnorm import convert_model\n\ndef model_to_device(self, model):\n    model = model.to(self.device.output_device)\n    if len(self.device.gpu_list) \u003e 1:\n        model.conv2d = nn.DataParallel(\n            model.conv2d,\n            device_ids=self.device.gpu_list,\n            output_device=self.device.output_device)\n    model = convert_model(model)\n    model.cuda()\n    return model\n```\n\nWith the provided code, the updated results are expected as:\n\n| Backbone                | WER on Dev | WER on Test |                       Pretrained model                       |\n| :---------------------- | :--------: | :---------: | :----------------------------------------------------------: |\n| ResNet18 (baseline)     |    23.8    |    25.4     | [[Baidu]](https://pan.baidu.com/s/17ernd4x3YIAEKpVa1rJqWA?pwd=iccv) [[GoogleDrive]](https://drive.google.com/file/d/1_yPOrVyxO2AJiLC6xOAPiGuPu41ov5Yg/view?usp=sharing) |\n| ResNet18+VAC (CTC only) |    21.5    |    22.1     | [[Baidu]](https://pan.baidu.com/s/1vDQyNrKM9Ar2ppvnCcohBA?pwd=VAC0) [[GoogleDrive]](https://drive.google.com/file/d/1etgf94fGvvIvR6c0VCXc8j2aFy5BsrZp/view?usp=sharing) |\n| ResNet18+VAC+SMKD       |  **19.8**  |  **20.5**   | [[Baidu]](https://pan.baidu.com/s/1jWT6FhxpD36fQilXZgyW9A?pwd=SMKD) [[GoogleDrive]](https://drive.google.com/file/d/1ULbB4qNdPhDjdKUX3JlgSYkQI2W3Lwm9/view?usp=sharing) |\n\nThe VAC result is corresponding to the setting of`loss_weights: SeqCTC: 1.0, ConvCTC: 1.0`. In addition to that, the VAC+SMKD adopt the setting of `model_args: share_classifier: True, weight_norm: True`.\n\nIf you find this repo useful in your research works, please consider cite our papers [VAC](https://openaccess.thecvf.com/content/ICCV2021/html/Min_Visual_Alignment_Constraint_for_Continuous_Sign_Language_Recognition_ICCV_2021_paper.html) and [SMKD](https://openaccess.thecvf.com/content/ICCV2021/html/Hao_Self-Mutual_Distillation_Learning_for_Continuous_Sign_Language_Recognition_ICCV_2021_paper.html).\n\n---\n### Prerequisites\n\n- This project is implemented in Pytorch (\u003e1.8). Thus please install Pytorch first.\n\n- ctcdecode==0.4 [[parlance/ctcdecode]](https://github.com/parlance/ctcdecode)，for beam search decode.\n\n- [Optional] sclite [[kaldi-asr/kaldi]](https://github.com/kaldi-asr/kaldi), install kaldi tool to get sclite for evaluation. After installation, create a soft link toward the sclite:    \n  `ln -s PATH_TO_KALDI/tools/sctk-2.4.10/bin/sclite ./software/sclite`\n  We also provide a python version evaluation tool for convenience, but sclite can provide more detailed statistics.\n\n- [Optional] [SeanNaren/warp-ctc](https://github.com/SeanNaren/warp-ctc) At the beginning of this research, we adopt warp-ctc for supervision, and we recently find that pytorch version CTC can reach similar results.\n\n### Data Preparation\n\n1. Download the RWTH-PHOENIX-Weather 2014 Dataset [[download link]](https://www-i6.informatik.rwth-aachen.de/~koller/RWTH-PHOENIX/). Our experiments based on phoenix-2014.v3.tar.gz.\n\n2. After finishing dataset download, extract it to ./dataset/phoenix, it is suggested to make a soft link toward downloaded dataset.   \n   `ln -s PATH_TO_DATASET/phoenix2014-release ./dataset/phoenix2014`\n\n3. The original image sequence is 210x260, we resize it to 256x256 for augmentation. Run the following command to generate gloss dict and resize image sequence.     \n\n   ```bash\n   cd ./preprocess\n   python data_preprocess.py --process-image --multiprocessing\n   ```\n\n### Inference\n\n​\tWe provide the pretrained models for inference, you can download them from:\n\n| Backbone | WER on Dev | WER on Test | Pretrained model                                             |\n| -------- | ---------- | ----------- | ------------------------------------------------------------ |\n| ResNet18 | 21.2%      | 22.3%       | [[Baidu]](https://pan.baidu.com/s/12WSc2Xhy7LSkLojh1XqY6g) (passwd: qi83)\u003cbr /\u003e[[Dropbox]](https://www.dropbox.com/s/zbas78emfz5m4bp/resnet18_slr_pretrained_distill25.pt?dl=0)     |\n\n​\tTo evaluate the pretrained model, run the command below：   \n`python main.py --load-weights resnet18_slr_pretrained.pt --phase test`\n\n​\t(When evaluating the SMKD pretrained model,  please modify the weight_norm and share_classifier in config files as True).\n\n### Training\n\nThe priorities of configuration files are: command line \u003e config file \u003e default values of argparse. To train the SLR model on phoenix14, run the command below:\n\n`python main.py --work-dir PATH_TO_SAVE_RESULTS --config PATH_TO_CONFIG_FILE --device AVAILABLE_GPUS`\n\n### Feature Extraction\n\nWe also provide feature extraction function to extract frame-wise features for other research purpose, which can be achieved by:\n\n`python main.py --load-weights PATH_TO_PRETRAINED_MODEL --phase features ` \n\n### To Do List\n\n- [x] Pure python implemented evaluation tools.\n- [x] WAR and WER calculation scripts.\n\n### Citation\n\nIf you find this repo useful in your research works, please consider citing:\n\n```latex\n@InProceedings{Min_2021_ICCV,\n    author    = {Min, Yuecong and Hao, Aiming and Chai, Xiujuan and Chen, Xilin},\n    title     = {Visual Alignment Constraint for Continuous Sign Language Recognition},\n    booktitle = {Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},\n    month     = {October},\n    year      = {2021},\n    pages     = {11542-11551}\n}\n```\n\nSelf-Mutual Distillation Learning for Continuous Sign Language Recognition [[paper]](https://openaccess.thecvf.com/content/ICCV2021/html/Hao_Self-Mutual_Distillation_Learning_for_Continuous_Sign_Language_Recognition_ICCV_2021_paper.html)\n\n```latex\n@InProceedings{Hao_2021_ICCV,\n    author    = {Hao, Aiming and Min, Yuecong and Chen, Xilin},\n    title     = {Self-Mutual Distillation Learning for Continuous Sign Language Recognition},\n    booktitle = {Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},\n    month     = {October},\n    year      = {2021},\n    pages     = {11303-11312}\n}\n```\n\n### Acknowledge\n\nWe appreciate the help from Runpeng Cui, Hao Zhou@[Rhythmblue](https://github.com/Rhythmblue) and Xinzhe Han@[GeraldHan](https://github.com/GeraldHan) :)\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FVIPL-SLP%2FVAC_CSLR","html_url":"https://awesome.ecosyste.ms/projects/github.com%2FVIPL-SLP%2FVAC_CSLR","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FVIPL-SLP%2FVAC_CSLR/lists"}