{"id":13689892,"url":"https://github.com/hirofumi0810/neural_sp","last_synced_at":"2025-05-02T06:31:29.324Z","repository":{"id":47932684,"uuid":"103009093","full_name":"hirofumi0810/neural_sp","owner":"hirofumi0810","description":"End-to-end ASR/LM implementation with PyTorch","archived":false,"fork":false,"pushed_at":"2021-08-30T12:06:20.000Z","size":9084,"stargazers_count":595,"open_issues_count":45,"forks_count":141,"subscribers_count":32,"default_branch":"master","last_synced_at":"2024-11-12T15:43:13.020Z","etag":null,"topics":["asr","attention","attention-mechanism","automatic-speech-recognition","ctc","language-model","language-modeling","pytorch","rnn-transducer","seq2seq","sequence-to-sequence","speech","speech-recognition","streaming","transformer","transformer-xl"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/hirofumi0810.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null}},"created_at":"2017-09-10T06:33:28.000Z","updated_at":"2024-11-11T12:54:59.000Z","dependencies_parsed_at":"2022-08-31T05:03:39.700Z","dependency_job_id":null,"html_url":"https://github.com/hirofumi0810/neural_sp","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hirofumi0810%2Fneural_sp","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hirofumi0810%2Fneural_sp/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hirofumi0810%2Fneural_sp/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hirofumi0810%2Fneural_sp/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/hirofumi0810","download_url":"https://codeload.github.com/hirofumi0810/neural_sp/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":251998462,"owners_count":21677987,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["asr","attention","attention-mechanism","automatic-speech-recognition","ctc","language-model","language-modeling","pytorch","rnn-transducer","seq2seq","sequence-to-sequence","speech","speech-recognition","streaming","transformer","transformer-xl"],"created_at":"2024-08-02T16:00:32.257Z","updated_at":"2025-05-02T06:31:28.843Z","avatar_url":"https://github.com/hirofumi0810.png","language":"Python","funding_links":[],"categories":["\u003ca name=\"toolkits\"\u003e\u003c/a\u003e 2. Toolkits","语音识别"],"sub_categories":["\u003ca name=\"paperlist\"\u003e\u003c/a\u003e 1.2 All Paper List","网络服务_其他"],"readme":"[![Build Status](https://travis-ci.org/hirofumi0810/neural_sp.svg?branch=master)](https://travis-ci.org/hirofumi0810/neural_sp)\n[![codecov](https://codecov.io/gh/hirofumi0810/neural_sp/branch/master/graph/badge.svg?token=wy0VD7e3bH)](https://codecov.io/gh/hirofumi0810/neural_sp)\n\n# NeuralSP: Neural network based Speech Processing\n\n## How to install\n```\ncd tools\nmake KALDI=/path/to/kaldi TOOL=/path/to/save/tools\n```\n\n## Key features\n### Corpus\n  - ASR\n    - AISHELL-1\n    - AISHELL-2\n    - AMI\n    - CSJ\n    - LaboroTVSpeech\n    - Librispeech\n    - Switchboard (+Fisher)\n    - TEDLIUM2/TEDLIUM3\n    - TIMIT\n    - WSJ\n\n  - LM\n    - Penn Tree Bank\n    - WikiText2\n\n### Front-end\n  - Frame stacking\n  - Sequence summary network [[link](https://www.isca-speech.org/archive/Interspeech_2018/abstracts/1438.html)]\n  - SpecAugment [[link](https://arxiv.org/abs/1904.08779)]\n  - Adaptive SpecAugment [[link](https://arxiv.org/abs/1912.05533)]\n\n### Encoder\n  - RNN encoder\n    - (CNN-)BLSTM, (CNN-)LSTM, (CNN-)BLGRU, (CNN-)LGRU\n    - Latency-controlled BRNN [[link](https://arxiv.org/abs/1510.08983)]\n    - Random state passing (RSP) [[link](https://arxiv.org/abs/1910.11455)]\n  - Transformer encoder [[link](https://arxiv.org/abs/1706.03762)]\n    - Chunk hopping mechanism [[link](https://arxiv.org/abs/1902.06450)]\n    - Relative positional encoding [[link](https://arxiv.org/abs/1901.02860)]\n    - Causal mask\n  - Conformer encoder [[link](https://arxiv.org/abs/2005.08100)]\n  - Time-depth separable (TDS) convolution encoder [[link](https://arxiv.org/abs/1904.02619)] [[line](https://arxiv.org/abs/2001.09727)]\n  - Gated CNN encoder (GLU) [[link](https://openreview.net/forum?id=Hyig0zb0Z)]\n\n### Connectionist Temporal Classification (CTC) decoder\n  - Beam search\n  - Shallow fusion\n  - Forced alignment\n\n### RNN-Transducer (RNN-T) decoder [[link](https://arxiv.org/abs/1211.3711)]\n  - Beam search\n  - Shallow fusion\n\n### Attention-based decoder\n  - RNN decoder\n    - Shallow fusion\n    - Cold fusion [[link](https://arxiv.org/abs/1708.06426)]\n    - Deep fusion [[link](https://arxiv.org/abs/1503.03535)]\n    - Forward-backward attention decoding [[link](https://www.isca-speech.org/archive/Interspeech_2018/abstracts/1160.html)]\n    - Ensemble decoding\n    - internal LM estimation [[link](https://arxiv.org/abs/2011.01991)]\n  - Attention type\n    - location-based\n    - content-based\n    - dot-product\n    - GMM attention\n  - Streaming RNN decoder specific\n    - Hard monotonic attention [[link](https://arxiv.org/abs/1704.00784)]\n    - Monotonic chunkwise attention (MoChA) [[link](https://arxiv.org/abs/1712.05382)]\n    - Delay constrained training (DeCoT) [[link](https://arxiv.org/abs/2004.05009)]\n    - Minimum latency training (MinLT) [[link](https://arxiv.org/abs/2004.05009)]\n    - CTC-synchronous training (CTC-ST) [[link](https://arxiv.org/abs/2005.04712)]\n  - Transformer decoder [[link](https://arxiv.org/abs/1706.03762)]\n  - Streaming Transformer decoder specific\n    - Monotonic Multihead Attention [[link](https://arxiv.org/abs/1909.12406)] [[link](https://arxiv.org/abs/2005.09394)]\n\n### Language model (LM)\n  - RNNLM (recurrent neural network language model)\n  - Gated convolutional LM [[link](https://arxiv.org/abs/1612.08083)]\n  - Transformer LM\n  - Transformer-XL LM [[link](https://arxiv.org/abs/1901.02860)]\n  - Adaptive softmax [[link](https://arxiv.org/abs/1609.04309)]\n\n### Output units\n  - Phoneme\n  - Grapheme\n  - Wordpiece (BPE, sentencepiece)\n  - Word\n  - Word-char mix\n\n### Multi-task learning (MTL)\nMulti-task learning (MTL) with different units are supported to alleviate data sparseness.\n  - Hybrid CTC/attention [[link](https://www.merl.com/publications/docs/TR2017-190.pdf)]\n  - Hierarchical Attention (e.g., word attention + character attention) [[link](http://sap.ist.i.kyoto-u.ac.jp/lab/bib/intl/INA-SLT18.pdf)]\n  - Hierarchical CTC (e.g., word CTC + character CTC) [[link](https://arxiv.org/abs/1711.10136)]\n  - Hierarchical CTC+Attention (e.g., word attention + character CTC) [[link](http://www.sap.ist.i.kyoto-u.ac.jp/lab/bib/intl/UEN-ICASSP18.pdf)]\n  - Forward-backward attention [[link](https://www.isca-speech.org/archive/Interspeech_2018/abstracts/1160.html)]\n  - LM objective\n\n\n## ASR Performance\n### AISHELL-1 (CER)\n| Model         | dev | test |\n| -----------   | --- | ---- |\n| Conformer LAS | 4.1 | 4.5  |\n| Transformer   | 5.0 | 5.4  |\n| Streaming MMA | 5.5 | 6.1  |\n\n### AISHELL-2 (CER)\n| Model         | test_android | test_ios | test_mic |\n| -----------   | ------------ | -------- | -------- |\n| Conformer LAS | 6.1          | 5.5      | 5.9      |\n\n### CSJ (WER)\n| Model          | eval1 | eval2 | eval3 |\n| -------------- | ----- | ----- | ----- |\n| Conformer LAS  | 5.7   | 4.4   | 4.9   |\n| BLSTM LAS      | 6.5   | 5.1   | 5.6   |\n| LC-BLSTM MoChA | 7.4   | 5.6   | 6.4   |\n\n### Switchboard 300h (WER)\n| Model     | SWB  | CH   |\n| --------- | ---- | ---- |\n| BLSTM LAS | 9.1  | 18.8 |\n\n### Switchboard+Fisher 2000h (WER)\n| Model     | SWB  | CH   |\n| --------- | ---- | ---- |\n| BLSTM LAS | 7.8  | 13.8 |\n\n### LaboroTVSpeech (CER)\n| Model          | dev_4k | dev   | tedx-jp-10k |\n| -------------- | ----- | -----  | -----       |\n| Conformer LAS  | 7.8   | 10.1   | 12.4        |\n\n### Librispeech (WER)\n| Model          | dev-clean | dev-other | test-clean | test-other |\n| -------------- | --------- | --------- | ---------- | ---------- |\n| Conformer LAS  | 1.9       | 4.6       | 2.1        | 4.9        |\n| Transformer    | 2.1       | 5.3       | 2.4        | 5.7        |\n| BLSTM LAS      | 2.5       | 7.2       | 2.6        | 7.5        |\n| BLSTM RNN-T    | 2.9       | 8.5       | 3.2        | 9.0        |\n| UniLSTM RNN-T  | 3.7       | 11.7      | 4.0        | 11.6       |\n| UniLSTM MoChA  | 4.1       | 11.0      | 4.2        | 11.2       |\n| LC-BLSTM RNN-T | 3.3       | 9.8       | 3.5        | 10.2       |\n| LC-BLSTM MoChA | 3.3       | 8.8       | 3.5        | 9.1        |\n| Streaming MMA  | 2.5       | 6.9       | 2.7        | 7.1        |\n\n### TEDLIUM2 (WER)\n| Model          | dev   | test |\n| -------------- | ----  | ---- |\n| Conformer LAS  |  7.0  |  6.8 |\n| BLSTM LAS      |  8.1  |  7.5 |\n| LC-BLSTM RNN-T |  8.0  |  7.7 |\n| LC-BLSTM MoChA | 10.3  |  8.6 |\n| UniLSTM RNN-T  | 10.7  | 10.7 |\n| UniLSTM MoChA  | 13.5  | 11.6 |\n\n### WSJ (WER)\n| Model     | test_dev93 | test_eval92 |\n| --------- | ---------- | ----------- |\n| BLSTM LAS | 8.8        | 6.2         |\n\n## LM Performance\n### Penn Tree Bank (PPL)\n| Model       | valid | test  |\n| ----------- | ----- | ----- |\n| RNNLM       | 87.99 | 86.06 |\n| + cache=100 | 79.58 | 79.12 |\n| + cache=500 | 77.36 | 76.94 |\n\n### WikiText2 (PPL)\n| Model        | valid  | test  |\n| ------------ | ------ | ----- |\n| RNNLM        | 104.53 | 98.73 |\n| + cache=100  | 90.86  | 85.87 |\n| + cache=2000 | 76.10  | 72.77 |\n\n\n## Reference\n- https://github.com/kaldi-asr/kaldi\n- https://github.com/espnet/espnet\n- https://github.com/awni/speech\n- https://github.com/HawkAaron/E2E-ASR\n\n## Dependency\n- https://github.com/SeanNaren/warp-ctc\n- https://github.com/HawkAaron/warp-transducer\n- https://github.com/1ytic/warp-rnnt\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fhirofumi0810%2Fneural_sp","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fhirofumi0810%2Fneural_sp","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fhirofumi0810%2Fneural_sp/lists"}