{"id":21299885,"url":"https://github.com/cwx-worst-one/eat","last_synced_at":"2025-04-05T15:04:53.527Z","repository":{"id":212065608,"uuid":"730622088","full_name":"cwx-worst-one/EAT","owner":"cwx-worst-one","description":" [IJCAI 2024]  EAT: Self-Supervised Pre-Training with Efficient Audio Transformer","archived":false,"fork":false,"pushed_at":"2024-12-23T06:12:15.000Z","size":5370,"stargazers_count":137,"open_issues_count":2,"forks_count":8,"subscribers_count":4,"default_branch":"main","last_synced_at":"2025-03-29T14:06:42.089Z","etag":null,"topics":["audio","audio-classification","deep-learning","eat","fairseq","pytorch","representation-learning","self-supervised-learning"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/cwx-worst-one.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":"CODE_OF_CONDUCT.md","threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2023-12-12T10:19:55.000Z","updated_at":"2025-03-26T06:16:51.000Z","dependencies_parsed_at":"2024-04-17T12:50:17.343Z","dependency_job_id":"22543ffb-3519-44ea-9c20-71d643e5739d","html_url":"https://github.com/cwx-worst-one/EAT","commit_stats":null,"previous_names":["cwx-worst-one/eat"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/cwx-worst-one%2FEAT","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/cwx-worst-one%2FEAT/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/cwx-worst-one%2FEAT/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/cwx-worst-one%2FEAT/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/cwx-worst-one","download_url":"https://codeload.github.com/cwx-worst-one/EAT/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":247353729,"owners_count":20925329,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["audio","audio-classification","deep-learning","eat","fairseq","pytorch","representation-learning","self-supervised-learning"],"created_at":"2024-11-21T15:06:31.241Z","updated_at":"2025-04-05T15:04:53.499Z","avatar_url":"https://github.com/cwx-worst-one.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003c!-- omit in toc --\u003e\n# EAT: Self-Supervised Pre-Training with Efficient Audio Transformer\n[![Platform](https://img.shields.io/badge/Platform-linux-lightgrey?logo=linux)](https://www.linux.org/)\n[![Python](https://img.shields.io/badge/Python-3.8%2B-orange?logo=python)](https://www.python.org/)\n[![Pytorch](https://img.shields.io/badge/PyTorch-1.13%2B-brightgree?logo=PyTorch)](https://pytorch.org/)\n[![arXiv](https://img.shields.io/badge/Arxiv-2401.03497-blueviolet?logo=arxiv)](https://arxiv.org/abs/2401.03497)\n[![fairseq](https://img.shields.io/badge/Fairseq-0.12.2-blue)](https://github.com/facebookresearch/fairseq)\n[![License](https://img.shields.io/badge/License-MIT-red.svg)](https://github.com/cwx-worst-one/EAT)\n\n**Guides**\n- [Requirements and Installation](#requirements-and-installation)\n- [Model Checkpoints](#model-checkpoints)\n- [Feature Extraction](#feature-extraction)\n- [Data Preparation](#data-preparation)\n- [Pre-Training](#pre-training)\n- [Fine-Tuning](#fine-tuning)\n- [Inference and Evaluation](#inference-and-evaluation)\n\n\n\u003c!-- omit in toc --\u003e\n## News :fire:\n- We release EAT-large (20 epochs) with SOTA performance on AS-2M, AS-20K, ESC-50 and SPC-2. \n- We have updated the checkpoints and code, and now EAT seamlessly supports variable-length audio throughout training, feature extraction, inference, and evaluation phases.\n\n\u003c!-- omit in toc --\u003e\n## Introduction \nEAT is an audio SSL model with high effectiveness and efficiency during self-supervised pre-training. You can find details in the paper [EAT: Self-Supervised Pre-Training with Efficient Audio Transformer](https://arxiv.org/abs/2401.03497). \n\n## Requirements and Installation\n\u003c!-- To run the EAT code, you have two options for setting up your environment: manual setup or using our Docker image. --\u003e\n\n\u003c!-- omit in toc --\u003e\n\u003c!-- #### Manual Environment Setup --\u003e\nThe minimum environment requirements are `Python \u003e= 3.8` and `PyTorch \u003e= 1.13`. You could find the versions of other dependencies we use in `requirements.txt`. \n```shell \ngit clone https://github.com/pytorch/fairseq\ncd fairseq\npip install --editable ./\ngit clone https://github.com/cwx-worst-one/EAT\n```\n\n\u003c!-- omit in toc --\u003e\n\u003c!-- #### Using Docker Image :whale:\nWe also provide a Docker image for an easier and more consistent setup. The Docker image will be released soon, containing all necessary dependencies pre-installed. --\u003e\n\n## Model Checkpoints\nYou could download the EAT-base (10 epochs) checkpoints by Google Drive. \n- AS-2M [Pre-trained](https://drive.google.com/file/d/10pklbY_fKraQUIBizSg1kv4lJXNWxpxl/view?usp=sharing)\n- AS-2M Pre-trained+[Fine-tuned](https://drive.google.com/file/d/1F07zN8N54rXU-szvKUlYaCFMCepc4wHR/view?usp=sharing) (AS-2M)\n- AS-2M Pre-trained+[Fine-tuned](https://drive.google.com/file/d/1fRX_Mgj4sHxV2F6AVfoqXObfgzFMnHRA/view?usp=sharing) (AS-20K)\n\n:warning: Due to the limited amount of AudioSet data we possess compared to other models, we highly **recommend** [pre-training](#pre-training) the EAT model with your own data, which would probably perform better than the given one.\n\n**Update!!!!!** :new:  (**RECOMMEND**)  \nWe have introduced two new variants of the EAT pre-training model and their fine-tuned versions, each designed to enhance performance through either extended pre-training epochs or scaling up the model size.  \n\nLinks for model checkpoints:  \n- [EAT-base_epoch30](https://drive.google.com/file/d/19hfzLgHCkyqTOYmHt8dqVa9nm-weBq4f/view?usp=sharing) (pre-training) \n- [EAT-base_epoch30](https://drive.google.com/file/d/1aCYiQmoZv_Gh1FxnR-CCWpNAp6DIJzn6/view?usp=sharing) (fine-tuning on AS-2M) \n- [EAT-large_epoch20](https://drive.google.com/file/d/1PEgriRvHsqrtLzlA478VemX7Q0ZGl889/view?usp=sharing) (pre-training)\n- [EAT-large_epoch20](https://drive.google.com/file/d/1b_f_nQAdjM1B6u72OFUtFiUu-4yM2shd/view?usp=sharing) (fine-tuning on AS-2M)  \n\nPerformance metrics:  \n|Model|Backbone|Parameters|Pre-training \u003cbr\u003e Epoch|AS-20K \u003cbr\u003e mAP(%)|AS-2M \u003cbr\u003e mAP(%)|\n|:-:|:-:|:-:|:-:|:-:|:-:|\n|EAT-base|ViT-B|88M|10|40.3 | 48.6|\n|EAT-base|ViT-B|88M|30|41.3 | 48.9|\n|EAT-large|ViT-L|309M|20|**42.0** | **49.5**|\n\n\n## Feature Extraction\nWe provide the script for extracting audio features from the last layer of EAT encoder. The features are stored in `.npy` format and the sample rate of the extracted features is ~50Hz. EAT could provide frame-level features and utterance-level features (denoted by the CLS token).  \nTo extract latent representations from audio clips, you could use our pre-trained [checkpoint](https://drive.google.com/file/d/19hfzLgHCkyqTOYmHt8dqVa9nm-weBq4f/view?usp=sharing), fine-tuned [checkpoint](https://drive.google.com/file/d/1aCYiQmoZv_Gh1FxnR-CCWpNAp6DIJzn6/view?usp=sharing) or your owns, then please run the script `feature_extract.sh` by:\n```bash\nbash EAT/scripts/feature_extract.sh \n``` \n\n## Data Preparation\nThe main dataset in our experiment is [AudioSet](https://research.google.com/audioset/). Regrettably, we are unable to release the data due to copyright restrictions. Data manifest is available at [here](https://drive.google.com/file/d/1LH2C0q3d4zndoR3-oGkVdYYqDCIdxIsm/view?usp=drive_link). We follow the file format in [wav2vec](https://github.com/facebookresearch/fairseq/tree/main/examples/wav2vec) and [data2vec](https://github.com/facebookresearch/fairseq/tree/main/examples/data2vec), where `.tsv` format file is for index while `.lbl` and `.csv` format files are specific for classification task.  You could modify the files for your own database. \n\n## Pre-Training \nOur codes are adapted from [Audio-MAE](https://github.com/facebookresearch/AudioMAE) and [data2vec](https://github.com/facebookresearch/fairseq/tree/main/examples/data2vec). We employ `pretraining_AS2M.yaml` as our default pre-training config. To pre-train the EAT model on Audioset, you could run the script `pretraining_AS2M.sh` by:\n```bash\nbash EAT/scripts/pretraining_AS2M.sh \n``` \nIf you need to pre-train the EAT model on other datasets where audio lengths are not fixed at 10 seconds, you can refer to the instructions in\n`feature_extract/readme.md`\n\n## Fine-Tuning\nWe employ `finetuning.yaml` as our default fine-tuning config. To fine-tune the EAT model in different downstream tasks, you could run the script `finetuning_{task}.sh`, where `{task}` includes `AS20K`, `AS2M`, `ESC50` and `SPCv2`. For example, you can fine-tune EAT on `AS20K` by executing: \n```bash\nbash EAT/scripts/finetuning_AS20K.sh\n``` \n\n## Inference and Evaluation\nFor inference on single AudioSet audio clip with fine-tuned models, you could use our EAT checkpoints fine-tuning on [AS-2M](https://drive.google.com/file/d/1F07zN8N54rXU-szvKUlYaCFMCepc4wHR/view?usp=sharing) (recommended) or [AS-20K](https://drive.google.com/file/d/1fRX_Mgj4sHxV2F6AVfoqXObfgzFMnHRA/view?usp=sharing)\nand run the script `inference.sh` by: \n```bash\nbash EAT/scripts/inference.sh \n``` \nAn example output is as follows:\n```\n# top_k_prediction = 12\n************ Acoustic Event Inference ************\nLABEL                          PREDICTION\nPercussion                     0.523\nDrum kit                       0.437\nVibraphone                     0.420\nDrum                           0.316\nMusic                          0.303\nSnare drum                     0.277\nGlockenspiel                   0.225\nMarimba, xylophone             0.223\nCymbal                         0.213\nBass drum                      0.207\nHi-hat                         0.196\nMallet percussion              0.170\n**************************************************\n```\n  \nFor comprehensive evaluation on the entire AudioSet eval dataset with fine-tuned EAT models, you could run the evaluation script `eval.sh` by:\n```bash\nbash EAT/scripts/eval.sh \n```\nThis script will give you the evaluation value of mAP on AudioSet test dataset. \nPer-class AP can be found under the path `./EAT/ap_log.txt`. You could also refer to our results of finetuned EAT models on evaluation set of Audioset under the path `./EAT/results`.\n\n\n\u003c!-- omit in toc --\u003e\n## Performance\nPre-training on AS-2M, EAT gains state-of-the-art (SOTA) performance on several audio and speech classification datasets including AS-20K, AS-2M, ESC-50 and SPC-2.    \n![Alt text](/src/EAT_performance.png)\n\n\u003c!-- omit in toc --\u003e\n## Efficiency\nEAT achieves a total pre-training time reduction of ~15x compared to BEATs and ~10x relative to Audio-MAE. It costs only 10 epochs during EAT's pre-training on AS-2M.    \n![Alt text](/src/EAT_efficiency.png)\n\n\u003c!-- omit in toc --\u003e\n## Experiment Logs\nWe report the experiment logs using [wandb](https://wandb.ai). We have published a  short WandB report detailing the training process and performance metrics of the EAT model. You could visit it [here](https://api.wandb.ai/links/wxc12/obqrpq36).\n\n\n\u003c!-- omit in toc --\u003e\n## TODO \n- [x] release the final EAT large\n- [x] update codes and checkpoints for friendly usage\n- [ ] release the docker image\n\n## Acknowledgement\nOur codebase is based on the awesome [Audio-MAE](https://github.com/facebookresearch/AudioMAE) and [data2vec](https://github.com/facebookresearch/fairseq/tree/main/examples/data2vec) repo. \n\n\n## Institutional Contributors\n|  Institution | Contribution |\n|:------|:-----|\n| [Shanghai Jiao Tong University](https://www.seiee.sjtu.edu.cn/) | Researchers; Computing power |\n| [Peng Cheng Laboratory](https://data-starcloud.pcl.ac.cn/) | Researchers; Computing power |\n\n\u003c!-- omit in toc --\u003e\n## Citation\nIf you find our EAT codes and models useful, please cite the following paper:\n```\n@article{chen2024eat,\n  title={EAT: Self-Supervised Pre-Training with Efficient Audio Transformer},\n  author={Chen, Wenxi and Liang, Yuzhe and Ma, Ziyang and Zheng, Zhisheng and Chen, Xie},\n  journal={arXiv preprint arXiv:2401.03497},\n  year={2024}\n}\n```\n\n\u003c!-- omit in toc --\u003e\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fcwx-worst-one%2Feat","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fcwx-worst-one%2Feat","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fcwx-worst-one%2Feat/lists"}