{"id":13596480,"url":"https://github.com/coqui-ai/open-speech-corpora","last_synced_at":"2026-01-27T03:36:44.571Z","repository":{"id":41245724,"uuid":"168543012","full_name":"coqui-ai/open-speech-corpora","owner":"coqui-ai","description":"💎 A list of accessible speech corpora for ASR, TTS, and other Speech Technologies","archived":false,"fork":false,"pushed_at":"2024-06-06T11:33:44.000Z","size":142,"stargazers_count":1382,"open_issues_count":169,"forks_count":149,"subscribers_count":55,"default_branch":"master","last_synced_at":"2026-01-12T18:48:33.667Z","etag":null,"topics":["speech-emotion-recognition","speech-processing","speech-recognition","speech-separation","speech-synthesis","speech-to-text","stt","text-to-speech","tts","voice-activity-detection","voice-cloning","voice-recognition"],"latest_commit_sha":null,"homepage":"","language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/coqui-ai.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":"CODE_OF_CONDUCT.md","threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2019-01-31T14:57:39.000Z","updated_at":"2026-01-12T08:01:54.000Z","dependencies_parsed_at":"2024-08-01T16:41:22.443Z","dependency_job_id":null,"html_url":"https://github.com/coqui-ai/open-speech-corpora","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/coqui-ai/open-speech-corpora","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/coqui-ai%2Fopen-speech-corpora","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/coqui-ai%2Fopen-speech-corpora/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/coqui-ai%2Fopen-speech-corpora/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/coqui-ai%2Fopen-speech-corpora/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/coqui-ai","download_url":"https://codeload.github.com/coqui-ai/open-speech-corpora/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/coqui-ai%2Fopen-speech-corpora/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":28799833,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-01-27T01:07:07.743Z","status":"online","status_checked_at":"2026-01-27T02:00:07.755Z","response_time":168,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["speech-emotion-recognition","speech-processing","speech-recognition","speech-separation","speech-synthesis","speech-to-text","stt","text-to-speech","tts","voice-activity-detection","voice-cloning","voice-recognition"],"created_at":"2024-08-01T16:02:28.908Z","updated_at":"2026-01-27T03:36:44.558Z","avatar_url":"https://github.com/coqui-ai.png","language":null,"funding_links":[],"categories":["Speech","Others","miscellaneous"],"sub_categories":[],"readme":"# 💎 Open Speech Corpora\n\nA list of open speech corpora for Speech Technology research and development.\n\nThis list has a preference for free (i.e. no $ cost) and truly open corpora (e.g. released under a [Creative Commons license](https://en.wikipedia.org/wiki/Creative_Commons_license) or a [Community Data License Agreement](https://en.wikipedia.org/wiki/Linux_Foundation#Community_Data_License_Agreement_%28CDLA%29)). Not all these corpora may meet those criteria, but all the following corpora are accessible and usable for research and/or commercial use.\n\nFeel free to propse additions to the list!\n\n*There's a long backlog of corpora to be added in the [Issues](https://github.com/coqui-ai/open-speech-corpora/issues), and Pull Requests are very welcome :)*\n\n## 📜 [CC-0](https://creativecommons.org/share-your-work/public-domain/cc0/)\n\n| CORPUS | LANGUAGES | # HOURS | # SPEAKERS | DOWNLOAD | LICENSE |\n| --- | --- | --- | --- | --- | --- |\n| Common Voice | Multilingual | \u003e15,000 hours (validated); \u003e20,000 hours (total) | Multi-speaker | \u003chttps://voice.mozilla.org/en/datasets\u003e | [CC-0](https://creativecommons.org/share-your-work/public-domain/cc0/) |\n| Yesno | Hebrew | 6 mins | one male | \u003chttp://www.openslr.org/1/\u003e | [CC-0](https://creativecommons.org/share-your-work/public-domain/cc0/) |\n| LJ Speech Corpus | English | ~24 hours | [one female](https://librivox.org/reader/11049) | \u003chttps://data.keithito.com/data/speech/LJSpeech-1.1.tar.bz2\u003e | [CC-0](https://creativecommons.org/share-your-work/public-domain/cc0/) |\n| NST Danish ASR Database | Danish | 229,992 utterances | 616 speakers | original: \u003chttps://www.nb.no/sprakbanken/en/resource-catalogue/oai-nb-no-sbr-19/\u003e, reorganized: \u003chttps://www.nb.no/sprakbanken/en/resource-catalogue/oai-nb-no-sbr-55/\u003e | [CC-0](https://creativecommons.org/publicdomain/zero/1.0/) |\n| NST Danish Dictation | Danish | 34,955 utterances | 151 speakers | \u003chttps://www.nb.no/sprakbanken/en/resource-catalogue/oai-nb-no-sbr-20/\u003e | [CC-0](https://creativecommons.org/publicdomain/zero/1.0/) |\n| NST Danish Speech Synthesis | Danish | 4,108 utterances | 1 male speaker | \u003chttps://www.nb.no/sprakbanken/en/resource-catalogue/oai-nb-no-sbr-21/\u003e | [CC-0](https://creativecommons.org/publicdomain/zero/1.0/) |\n| NST Swedish ASR Database | Swedish | 366,000 utterances | 1,000 speakers | original: \u003chttps://www.nb.no/sprakbanken/en/resource-catalogue/oai-nb-no-sbr-16/\u003e, reorganized: \u003chttps://www.nb.no/sprakbanken/en/resource-catalogue/oai-nb-no-sbr-56/\u003e | [CC-0](https://creativecommons.org/publicdomain/zero/1.0/) |\n| NST Swedish Dictation | Swedish | 45,620 utterances | 195 speakers | \u003chttps://www.nb.no/sprakbanken/en/resource-catalogue/oai-nb-no-sbr-17/\u003e | [CC-0](https://creativecommons.org/publicdomain/zero/1.0/) |\n| NST Swedish Speech Synthesis | Swedish | 5,279 utterances | 1 male speaker | \u003chttps://www.nb.no/sprakbanken/en/resource-catalogue/oai-nb-no-sbr-18/\u003e | [CC-0](https://creativecommons.org/publicdomain/zero/1.0/) |\n| NST Norwegian ASR Database | Norwegian | 359,760 utterances | 980 speakers | original: \u003chttps://www.nb.no/sprakbanken/en/resource-catalogue/oai-nb-no-sbr-13/\u003e, reorganized: \u003chttps://www.nb.no/sprakbanken/en/resource-catalogue/oai-nb-no-sbr-54/\u003e | [CC-0](https://creativecommons.org/publicdomain/zero/1.0/) |\n| NST Norwegian Dictation | Norwegian | 33,360 utterances | 144 speakers | \u003chttps://www.nb.no/sprakbanken/en/resource-catalogue/oai-nb-no-sbr-14/\u003e | [CC-0](https://creativecommons.org/publicdomain/zero/1.0/) |\n| NST Norwegian Speech Synthesis | Norwegian | 5,363 utterances | 1 male speaker | \u003chttps://www.nb.no/sprakbanken/en/resource-catalogue/oai-nb-no-sbr-15/\u003e | [CC-0](https://creativecommons.org/publicdomain/zero/1.0/) |\n| NB Tale – Speech Database for Norwegian | Norwegian | 7,600 utterances + ~12 hours | 380 speakers | \u003chttps://www.nb.no/sprakbanken/en/resource-catalogue/oai-nb-no-sbr-31/\u003e | [CC-0](https://creativecommons.org/publicdomain/zero/1.0/) |\n| Norwegian Parliamentary Speech Corpus (v0.1) | Norwegian | ~59 hours | 203 speakers | \u003chttps://www.nb.no/sprakbanken/en/resource-catalogue/oai-nb-no-sbr-58/\u003e | [CC-0](https://creativecommons.org/publicdomain/zero/1.0/) |\n| Wikimedia Commons Odia | Odia | ~8 hours | ~20 speakers | \u003chttps://commons.wikimedia.org/wiki/Category:Odia_pronunciation\u003e | mostly(?) [CC-0](https://creativecommons.org/share-your-work/public-domain/cc0/) |\n| Thorsten-21.02-neutral | German | ~24 hours | 1 male speaker | \u003chttps://www.Thorsten-Voice.de\u003e | [CC-0](https://creativecommons.org/share-your-work/public-domain/cc0/) |\n| Thorsten-21.06-emotional | German | 2.400 utterances (8 emotions) | 1 male speaker | \u003chttps://www.Thorsten-Voice.de\u003e | [CC-0](https://creativecommons.org/share-your-work/public-domain/cc0/) |\n\n## 📜 [CC-BY](https://creativecommons.org/licenses/by/4.0/)\n\n| CORPUS | LANGUAGES | # HOURS | # SPEAKERS | DOWNLOAD | LICENSE |\n| --- | --- | --- | --- | --- | --- |\n| ARU Speech Corpus | English (UK) | 720 utterances / speaker | 12 (6 femals; 6 male) | \u003chttp://datacat.liverpool.ac.uk/681/1/ARU_Speech_Corpus_v1_0.zip\u003e | [CC-BY 3.0](https://creativecommons.org/licenses/by/3.0/) |\n| Althingi Parliamentary Speech Corpus  | Icelandic | 542 hours and 25 minutes | 196 speakers | \u003chttp://www.malfong.is/index.php?dlid=73\u0026lang=en\u003e | [CC-BY 4.0](https://creativecommons.org/licenses/by/4.0/) |\n| Alþingisumræður Parliamentary Speech Corpus | Icelandic | ~21 hours | | \u003chttp://www.malfong.is/index.php?dlid=8\u0026lang=en\u003e | [CC-BY 3.0](https://creativecommons.org/licenses/by/3.0/) |\n| Hjal Corpus | Icelandic | ~41,000 recordings | 883 speakers | \u003chttp://www.malfong.is/index.php?dlid=5\u0026lang=en\u003e | [CC-BY 3.0](https://creativecommons.org/licenses/by/3.0/) |\n| The Malromur Corpus | Icelandic | 152 hours | 563 speakers | \u003chttp://www.malfong.is/index.php?dlid=65\u0026lang=en\u003e | [CC-BY 4.0](https://creativecommons.org/licenses/by/4.0/) |\n| Telecooperation German Corpus for Kinect | German | ~35 hours | ~180 speakers | \u003chttp://www.repository.voxforge1.org/downloads/de/german-speechdata-TUDa-2015.tar.gz\u003e | [CC-BY 2.0](https://creativecommons.org/licenses/by/2.0/) |\n| African Speech Technology English-English Speech Corpus | English | ~21 hours | | \u003chttps://repo.sadilar.org/handle/20.500.12185/283\u003e | [CC-BY 2.5 South Africa](https://creativecommons.org/licenses/by/2.5/za/legalcode) |\n| African Speech Technology isiXhosa Speech Corpus | isiXhosa | ~26 hours | | \u003chttps://repo.sadilar.org/handle/20.500.12185/305\u003e | [CC-BY 2.5 South Africa](https://creativecommons.org/licenses/by/2.5/za/legalcode) |\n| NCHLT Afrikaans | Afrikaans | 56 hours | 210 speakers (98 female / 112 male) | \u003chttps://repo.sadilar.org/handle/20.500.12185/280\u003e | CC-BY 3.0 |\n| NCHLT English | English | 56 hours | 210 speakers (100 female / 110 male) | \u003chttps://repo.sadilar.org/handle/20.500.12185/274\u003e | CC-BY 3.0 |\n| NCHLT isiNdebele | isiNdebele | 56 hours | 148 speakers (78 female / 70 male) | \u003chttps://repo.sadilar.org/handle/20.500.12185/272\u003e | CC-BY 3.0 |\n| NCHLT isiXhosa | isiXhosa | 56 hours | 209 speakers (106 female / 103 male) | \u003chttps://repo.sadilar.org/handle/20.500.12185/279\u003e | CC-BY 3.0 |\n| NCHLT isiZulu | isiZulu | 56 hours | 210 speakers (98 female / 112 male) | \u003chttps://repo.sadilar.org/handle/20.500.12185/275\u003e | CC-BY 3.0 |\n| NCHLT Sepedi | Sepedi | 56 hours | 210 speakers (100 female / 110 male) | \u003chttps://repo.sadilar.org/handle/20.500.12185/270\u003e | CC-BY 3.0 |\n| NCHLT Sesotho | Sesotho | 56 hours | 210 speakers (113 female / 97 male) | \u003chttps://repo.sadilar.org/handle/20.500.12185/278\u003e | CC-BY 3.0 |\n| NCHLT Setswana | Setswana | 56 hours | 210 speakers (109 female / 101 male) | \u003chttps://repo.sadilar.org/handle/20.500.12185/281\u003e | CC-BY 3.0 |\n| NCHLT Siswati | Siswati | 56 hours | 197 speakers (96 female / 101 male) | \u003chttps://repo.sadilar.org/handle/20.500.12185/271\u003e | CC-BY 3.0 |\n| NCHLT Tshivenda | Tshivenda | 56 hours | 208 speakers (83 female / 125 male) | \u003chttps://repo.sadilar.org/handle/20.500.12185/276\u003e | CC-BY 3.0 |\n| NCHLT Xitsonga | Xitsonga | 56 hours | 198 speakers (95 female/103 male) | \u003chttps://repo.sadilar.org/handle/20.500.12185/277\u003e | CC-BY 3.0 |\n| Lwazi II Cross-lingual Proper Name Corpus | Afrikaans; English; isiZulu; Sesotho | 2 hours 5 mins| 20 speakers | \u003chttps://repo.sadilar.org/handle/20.500.12185/445\u003e | CC-BY 3.0 |\n| Lwazi II Proper Name Call Routing Telephone Corpus | English | 2 hours 7 mins | | \u003chttps://repo.sadilar.org/handle/20.500.12185/448\u003e | CC-BY 3.0 |\n| Lwazi II Afrikaans Trajectory Tracking Corpus | Afrikaans | 4 hours | one male | \u003chttps://repo.sadilar.org/handle/20.500.12185/442\u003e | CC-BY 3.0 |\n| LibriSpeech | English | ~1000 hours | 2484 speakers (1201 female / 1283 male) | \u003chttp://www.openslr.org/12/\u003e | CC-BY 4.0 |\n| Zeroth-Korean | Korean | 52.8 hours | 115 speakers | \u003chttp://www.openslr.org/40/\u003e | CC-BY 4.0 |\n| Speech Commands | English | 17.8 hours  | \u003e1,000 speakers | \u003chttps://ai.googleblog.com/2017/08/launching-speech-commands-dataset.html\u003e | CC-BY 4.0 |\n| ParlamentParla | Catalan | 320 hours  |  | \u003chttps://www.openslr.org/59/\u003e | CC-BY 4.0 |\n|  SIWIS | French | ~10 hours  | one female | \u003chttp://datashare.is.ed.ac.uk/download/DS_10283_2353.zip\u003e | [CC-BY 4.0](https://creativecommons.org/licenses/by/4.0/legalcode) |\n|  VCTK | English | 44 hours | 109 speakers  | \u003chttp://datashare.is.ed.ac.uk/download/DS_10283_3443.zip\u003e | [CC-BY 4.0](https://creativecommons.org/licenses/by/4.0/legalcode) |\n|  LibriTTS | English | 586 hours | 2,456 speakers (1,185 female / 1,271 male)  | \u003chttp://www.openslr.org/60/\u003e | [CC-BY 4.0](https://creativecommons.org/licenses/by/4.0/legalcode) |\n|  Augmented LibriSpeech | Audio (English); Text (English, French) | 236 hours | | \u003chttps://persyval-platform.univ-grenoble-alpes.fr/datasets/DS91\u003e | [CC-BY 4.0](https://creativecommons.org/licenses/by/4.0/legalcode) |\n|  Helsinki Prosody Corpus | English | 262.5 hours | 1,230 speakers | \u003chttps://github.com/Helsinki-NLP/prosody\u003e | [CC-BY 4.0](https://creativecommons.org/licenses/by/4.0/legalcode) |\n|Tuva Speech Database | Norwegian | 24 hours | 40 speakers | https://www.nb.no/sprakbanken/show?serial=oai:nb.no:sbr-44\u0026lang= |  [CC-BY 4.0](https://creativecommons.org/licenses/by/4.0/legalcode) |\n| COERLL Kʼicheʼ corpus | Kʼicheʼ | 34 minutes | ? speakers | https://cl.indiana.edu/~ftyers/resources/utexas-kiche-audio.tar.gz | [CC-BY 4.0](https://creativecommons.org/licenses/by/4.0/legalcode) |\n| Timers and Such v0.1 | English (synthetic: US, real: various nationalities) | synthetic: 172 hours, real: 0.29 hours | 21 synthetic, 11 real | https://zenodo.org/record/4110812#.X9j0RmBOkYM | [CC-BY 4.0](https://creativecommons.org/licenses/by/4.0/legalcode) |\n| Large Corpus of Czech Parliament Plenary Hearings | Czech | 444 hours | | \u003chttps://lindat.mff.cuni.cz/repository/xmlui/handle/11234/1-3126\u003e | [CC-BY 4.0](https://creativecommons.org/licenses/by/4.0/legalcode) |\n\n## 📜 [CC-BY-SA](https://creativecommons.org/licenses/by-sa/4.0/)\n\n| CORPUS | LANGUAGES | # HOURS | # SPEAKERS | DOWNLOAD | LICENSE |\n| --- | --- | --- | --- | --- | --- |\n| Iban | Iban | 8 hours |  | \u003chttp://www.openslr.org/24/\u003e \u003chttps://github.com/sarahjuan/iban\u003e | CC-BY-SA 2.0 |\n| Vystadial 2013 | English; Czech | 41 hours; 15 hours |  | \u003chttp://www.openslr.org/6/\u003e | CC-BY-SA 3.0 US |\n| Vystadial 2016 Czech | Czech | 77 hours; includes Vystadial 2013 Czech | | \u003chttps://lindat.cz/repository/xmlui/handle/11234/1-1740\u003e | CC-BY-SA 4.0 |\n| Free Spoken Digit Dataset | English | 2,000 isolated digits | 4 speakers | \u003chttps://github.com/Jakobovski/free-spoken-digit-dataset\u003e | CC-BY-SA 4.0 |\n| Google Javanese | Javanese | 296 hours| 1019 speakers| \u003chttp://www.openslr.org/35/\u003e | CC-BY-SA 4.0 |\n| Google Nepali | Nepali | 165 hours| 527 speakers| \u003chttp://www.openslr.org/54/\u003e | CC-BY-SA 4.0 |\n| Google Bengali | Bengali | 229 hours| 508 speakers| \u003chttp://www.openslr.org/53/\u003e | CC-BY-SA 4.0 |\n| Google Sinhala | Sinhala | 224 hours| 478 speakers| \u003chttp://www.openslr.org/52/\u003e | CC-BY-SA 4.0 |\n| Google Sundanese | Sundanese | 333 hours| 542 speakers| \u003chttp://www.openslr.org/36/\u003e | CC-BY-SA 4.0 |\n| Spoken Wikipedia Corpus (SWC-2017) | English; German; Dutch | 182 hours; 249 hours; 79 hours | 395 speakers; 339 speakers; 145 speakers | \u003chttps://nats.gitlab.io/swc/\u003e | CC-BY-SA 4.0 |\n| Chuvash TTS | Chuvash | 4 hours | 1 speaker | \u003chttps://github.com/ftyers/Turkic_TTS\u003e | CC-BY-SA 4.0  |\n| Forschergeist | German | 2 hours | 2 speakers (1 female; 1 male) | female speaker: \u003chttps://goofy.zamia.org/zamia-speech/corpora/forschergeist/annettevogt-20180320-rec.tgz\u003e; male speaker: \u003chttps://goofy.zamia.org/zamia-speech/corpora/forschergeist/timpritlove-20180320-rec.tgz\u003e | CC-BY-SA 4.0  |\n| Malayalam Speech Corpus by [SMC](https://blog.smc.org.in/malayalam-speech-corpus/) | Malayalam | 1:36 hours | 75 speakers (3 female, 12 male, 60 unidentified) | https://releases.smc.org.in/msc-reviewed-speech/ | CC-BY-SA 4.0  |\n| Google Malayalam | Malayalam | 3.02 hours| 24 speakers| \u003chttp://www.openslr.org/63/\u003e | CC-BY-SA 4.0 |\n\n## 📜 [CC-BY-ND](https://creativecommons.org/licenses/by-nd/4.0/)\n\n| CORPUS | LANGUAGES | # HOURS | # SPEAKERS | DOWNLOAD | LICENSE |\n| --- | --- | --- | --- | --- | --- |\n| IBM Recorded Debates v1 | English | 5 hours | 10 speakers | \u003chttps://www.research.ibm.com/haifa/dept/vst/debating_data.shtml#Debate%20Speech%20Analysis\u003e | CC-BY-ND |\n| IBM Recorded Debates v2 | English | ~14 hours  | 14 speakers | \u003chttps://www.research.ibm.com/haifa/dept/vst/debating_data.shtml#Debate%20Speech%20Analysis\u003e | CC-BY-ND |\n\n## 📜 [CC-BY-NC](https://creativecommons.org/licenses/by-nc/4.0/)\n\n| CORPUS | LANGUAGES | # HOURS | # SPEAKERS | DOWNLOAD | LICENSE |\n| --- | --- | --- | --- | --- | --- |\n| TV3Parla | Catalan | 240 hours  |  | \u003chttp://laklak.eu/share/tv3_0.3.tar.gz\u003e | [CC-BY-NC 4.0](https://creativecommons.org/licenses/by-nc/4.0/) |\n| Russian Open STT Corpus | Russian | ~10,000 hours public, ~10,000 more upon request  |  | \u003chttps://github.com/snakers4/open_stt/#links\u003e | [CC-BY-NC 4.0](https://creativecommons.org/licenses/by-nc/4.0/) with some [exceptions](https://github.com/snakers4/open_stt/blob/master/LICENSE)|\n| Russian Open TTS Corpus | Russian | 145 hours  | 3 males | \u003chttps://github.com/snakers4/open_tts/#links\u003e | [CC-BY-NC 4.0](https://creativecommons.org/licenses/by-nc/4.0/) with some [expections](https://github.com/snakers4/open_tts/blob/master/LICENSE)|\n| OVM – Otázky Václava Moravce | Czech | 35 hours  |  | \u003chttps://lindat.mff.cuni.cz/repository/xmlui/handle/11858/00-097C-0000-000D-EC98-3\u003e | [CC-BY-NC 3.0](https://creativecommons.org/licenses/by-nc/3.0/) |\n\n## 📜 [CC-BY-NC-SA](https://creativecommons.org/licenses/by-nc-sa/4.0/)\n\n| CORPUS | LANGUAGES | # HOURS | # SPEAKERS | DOWNLOAD | LICENSE |\n| --- | --- | --- | --- | --- | --- |\n| CHiME-Home | English | 6.8 hours |  | \u003chttps://archive.org/details/chime-home\u003e | [CC-BY-NC-SA 3.0](https://creativecommons.org/licenses/by-nc-sa/3.0/) |\n| Cameroon Pidgin English Corpus | Cameroon Pidgin English | ~17 hours |  | \u003chttp://ota.ox.ac.uk/text/2563.zip\u003e | [CC-BY-NC-SA 3.0](https://creativecommons.org/licenses/by-nc-sa/3.0/) |\n\n## 📜 [CC-BY-NC-ND](https://creativecommons.org/licenses/by-nc-nd/4.0/)\n\n| CORPUS | LANGUAGES | # HOURS | # SPEAKERS | DOWNLOAD | LICENSE |\n| --- | --- | --- | --- | --- | --- |\n| Tatoeba-Eng | English | ~250 hours (rough estimate) | 6 speakers | \u003chttps://voice.mozilla.org/en/datasets\u003e | [CC-BY-NC 4.0](https://creativecommons.org/licenses/by-nc/4.0/) (some audio) / [CC-BY-NC-ND 3.0](https://creativecommons.org/licenses/by-nc-nd/3.0/) (most audio) / [CC-BY 2.0](https://creativecommons.org/licenses/by/2.0/) (all text) |\n| TED-LIUM | English | 118 hours | 685 speakers (36h female / 81h male) | \u003chttp://www.openslr.org/7/\u003e | [CC-BY-NC-ND 3.0](https://creativecommons.org/licenses/by-nc-nd/3.0/) |\n| TED-LIUM-2 | English | 207 hours | 1242 speakers (66h female / 141h male) | \u003chttp://www.openslr.org/19/\u003e | [CC-BY-NC-ND 3.0](https://creativecommons.org/licenses/by-nc-nd/3.0/) |\n| TED-LIUM-3 | English | 452 hours | 2028 speakers (134h female / 316h male) | \u003chttp://www.openslr.org/51/\u003e | [CC-BY-NC-ND 3.0](https://creativecommons.org/licenses/by-nc-nd/3.0/) |\n| Pansori TEDxKR | Korean | 3 hours | 41 speakers | \u003chttp://www.openslr.org/58/\u003e | [CC-BY-NC-ND 4.0](https://creativecommons.org/licenses/by-nc-nd/4.0/) |\n| Primewords Mandarin | Mandarin | 100 hours | 296 speakers | \u003chttp://www.openslr.org/47/\u003e | [CC-BY-NC-ND 4.0](https://creativecommons.org/licenses/by-nc-nd/4.0/)|\n| MuST-C v1.0 | Audio (English); Text (Dutch, French, German, Italian, Portuguese, Romanian, Russian, Spanish) | 408, 504, 492, 465, 442, 385, 432, 489 hours per language pair | | \u003chttps://ict.fbk.eu/must-c-release-v1-0/\u003e | [CC-BY-NC-ND 4.0](https://creativecommons.org/licenses/by-nc-nd/4.0/) |\n| Czech Parliament Meetings | Czech | 88 hours | | \u003chttps://lindat.mff.cuni.cz/repository/xmlui/handle/11858/00-097C-0000-0005-CF9C-4\u003e | [CC-BY-NC-ND 3.0](https://creativecommons.org/licenses/by-nc-nd/3.0/) |\n| BembaSpeech | Bemba | 24 hours | 17 speakers (9 male / 8 female) | \u003chttps://github.com/csikasote/BembaSpeech\u003e | [CC-BY-NC-ND 4.0](https://creativecommons.org/licenses/by-nc-nd/4.0/) |\n\n## 📜 [CDLA-Permissive](https://cdla.io/permissive-1-0/)\n\n| CORPUS | LANGUAGES | # HOURS | # SPEAKERS | DOWNLOAD | LICENSE |\n| --- | --- | --- | --- | --- | --- |\n| DiPCo | English | ~5 hours | 32 speakers (13 female; 19 male) | \u003chttps://s3.amazonaws.com/dipco/DiPCo.tgz\u003e | [CDLA-Permissive-1.0](https://cdla.io/permissive-1-0/) |\n\n## 📜 [GNU General Public License](https://www.gnu.org/licenses/gpl.html)\n\n| CORPUS | LANGUAGES | # HOURS | # SPEAKERS | DOWNLOAD | LICENSE |\n| --- | --- | --- | --- | --- | --- |\n| VoxForge | English | ~120 hours | ~2966 speakers | \u003chttp://www.repository.voxforge1.org/downloads/en/Trunk/Audio/Main/16kHz_16bit/\u003e \u003chttps://voice.mozilla.org/en/datasets\u003e | GNU-GPL 3.0 |\n| VoxForge | Russian |  | | \u003chttp://www.repository.voxforge1.org/downloads/ru/Trunk/Audio/Main/16kHz_16bit/\u003e \u003chttp://www.repository.voxforge1.org/downloads/Russian/Trunk/Audio/Main/16kHz_16bit/\u003e| GNU-GPL 3.0 |\n| VoxForge | German |  | | \u003chttp://www.repository.voxforge1.org/downloads/de/Trunk/Audio/Main/16kHz_16bit/\u003e | GNU-GPL 3.0 |\n\n\n## 📜 [Apache License](https://www.apache.org/licenses/LICENSE-2.0)\n\n| CORPUS | LANGUAGES | # HOURS | # SPEAKERS | DOWNLOAD | LICENSE |\n| --- | --- | --- | --- | --- | --- |\n| AISHELL-1 | Mandarin | 170 hours | 400 speakers | \u003chttp://www.openslr.org/33/\u003e | Apache 2.0 |\n| Tunisian_MSA | Modern Standard Arabic (Tunisia) | 11.2 hours | 118 speakers | \u003chttp://www.openslr.org/46/\u003e | Apache 2.0 |\n| African Accented French | French | 22 hours | 232 speakers | \u003chttp://www.openslr.org/57/\u003e | Apache 2.0 |\n| THCHS-30 | Mandarin Chinese | 33.57 hours (13,389 utterances) | 40 speakers (31 female; 9 male) | \u003chttp://www.openslr.org/18/\u003e | Apache 2.0 |\n| Living Audio Dataset - Dutch | Dutch | 57:49 min | 1 speaker | \u003chttps://github.com/Idlak/Living-Audio-Dataset\u003e | Apache 2.0 |\n| Living Audio Dataset - English | English | 50:50 min | 1 speaker | \u003chttps://github.com/Idlak/Living-Audio-Dataset\u003e | Apache 2.0 |\n| Living Audio Dataset - Irish | Irish | 61:56 min | 1 speaker | \u003chttps://github.com/Idlak/Living-Audio-Dataset\u003e | Apache 2.0 |\n| Living Audio Dataset - Russian | Russian | 34:58 min | 1 speaker | \u003chttps://github.com/Idlak/Living-Audio-Dataset\u003e | Apache 2.0 |\n\n\n\n## 📜 [MIT License](https://opensource.org/licenses/MIT)\n\n| CORPUS | LANGUAGES | # HOURS | # SPEAKERS | DOWNLOAD | LICENSE |\n| --- | --- | --- | --- | --- | --- |\n| ALFFA | Amharic;Hausa (paid); Swahili; Wolof |  |  | \u003chttp://www.openslr.org/25/\u003e \u003chttps://github.com/besacier/ALFFA_PUBLIC\u003e | MIT |\n\n\n## 📜 [BSD 3-Clause License](https://opensource.org/licenses/BSD-3-Clause)\n\n\n| CORPUS | LANGUAGES | # HOURS | # SPEAKERS | DOWNLOAD | LICENSE |\n| --- | --- | --- | --- | --- | --- |\n| M-AILABS German Corpus | German | 237 hours and 22 minutes |  | \u003chttp://www.caito.de/data/Training/stt_tts/de_DE.tgz\u003e | [M-AILABS LICENSE](https://www.caito.de/2019/01/the-m-ailabs-speech-dataset/) (a data-specific [BSD 3-Clause License](https://opensource.org/licenses/BSD-3-Clause))|\n| M-AILABS Queen's English Corpus | Queen's English | 45 hours and 35 minutes |  | \u003chttp://www.caito.de/data/Training/stt_tts/en_UK.tgz\u003e | [M-AILABS LICENSE](https://www.caito.de/2019/01/the-m-ailabs-speech-dataset/) (a data-specific [BSD 3-Clause License](https://opensource.org/licenses/BSD-3-Clause))|\n| M-AILABS US English Corpus | American English | 102 hours and 7 minutes |  | \u003chttp://www.caito.de/data/Training/stt_tts/en_US.tgz\u003e | [M-AILABS LICENSE](https://www.caito.de/2019/01/the-m-ailabs-speech-dataset/) (a data-specific [BSD 3-Clause License](https://opensource.org/licenses/BSD-3-Clause))|\n| M-AILABS Spanish Corpus | Spanish Spanish | 108 hours and 34 minutes |  | \u003chttp://www.caito.de/data/Training/stt_tts/es_ES.tgz\u003e | [M-AILABS LICENSE](https://www.caito.de/2019/01/the-m-ailabs-speech-dataset/) (a data-specific [BSD 3-Clause License](https://opensource.org/licenses/BSD-3-Clause))|\n| M-AILABS Italian Corpus | Italian | 127 hours and 40 minutes |  | \u003chttp://www.caito.de/data/Training/stt_tts/it_IT.tgz\u003e | [M-AILABS LICENSE](https://www.caito.de/2019/01/the-m-ailabs-speech-dataset/) (a data-specific [BSD 3-Clause License](https://opensource.org/licenses/BSD-3-Clause))|\n| M-AILABS Ukrainian Corpus | Ukrainian | 87 hours and 8 minutes |  | \u003chttp://www.caito.de/data/Training/stt_tts/uk_UK.tgz\u003e | [M-AILABS LICENSE](https://www.caito.de/2019/01/the-m-ailabs-speech-dataset/) (a data-specific [BSD 3-Clause License](https://opensource.org/licenses/BSD-3-Clause))|\n| M-AILABS Russian Corpus | Russian | 46 hours and 47 minutes |  | \u003chttp://www.caito.de/data/Training/stt_tts/ru_RU.tgz\u003e | [M-AILABS LICENSE](https://www.caito.de/2019/01/the-m-ailabs-speech-dataset/) (a data-specific [BSD 3-Clause License](https://opensource.org/licenses/BSD-3-Clause))|\n| M-AILABS French-v0.9 Corpus | French | 190 hours and 30 minutes |  | \u003chttp://www.caito.de/data/Training/stt_tts/fr_FR.tgz\u003e | [M-AILABS LICENSE](https://www.caito.de/2019/01/the-m-ailabs-speech-dataset/) (a data-specific [BSD 3-Clause License](https://opensource.org/licenses/BSD-3-Clause))|\n| M-AILABS Polish Corpus | Polish | 53 hours and 50 minutes |  | \u003chttp://www.caito.de/data/Training/stt_tts/pl_PL.tgz\u003e | [M-AILABS LICENSE](https://www.caito.de/2019/01/the-m-ailabs-speech-dataset/) (a data-specific [BSD 3-Clause License](https://opensource.org/licenses/BSD-3-Clause))|\n\n## 📜 [Custom License](https://en.wikipedia.org/wiki/Copyright)\n\n| CORPUS | LANGUAGES | # HOURS | # SPEAKERS | DOWNLOAD | LICENSE |\n| --- | --- | --- | --- | --- | --- |\n| Fluent Speech Commands Corpus | English | 19 hours (30,043 utterances) | 97 speakers | \u003chttp://fluent.ai:2052/jf8398hf30f0381738rucj3828chfdnchs.tar.gz\u003e | [Fluent Speech Commands Public License](https://groups.google.com/a/fluent.ai/forum/#!msg/fluent-speech-commands/MXh_7Y-3QC8/9i2pHPW9AwAJ) |\n| CMU Wilderness | 700 Langs | Alignments distributed without audio or text total:~14,000 hours; per lang: ~20 hours |  | \u003chttps://github.com/festvox/datasets-CMU_Wilderness\u003e | \u003chttps://live.bible.is/terms\u003e |\n| CHiME-5 | English | 50 hours | 48 speakers | \u003chttp://spandh.dcs.shef.ac.uk/chime_challenge/data.html\u003e | [CHiME-5 License](http://spandh.dcs.shef.ac.uk/chime_challenge/download.html) |\n| Fearless Steps Corpus | English | 19,000 hours (20 hours transcribed) | ~450 speakers | \u003chttps://fearless-steps.github.io/ChallengePhase3/#19k_Corpus_Access\u003e | [NASA Media Usage Guidelines](https://www.nasa.gov/multimedia/guidelines/index.html) |\n| Microsoft Speech Corpus (Indian languages) | Telugu; Tamil; Gujarati | | | \u003chttps://msropendata.com/datasets/7230b4b1-912d-400e-be58-f84e0512985e\u003e | [Microsoft Speech Corpus (Indian Languages) License](https://msropendata.com/datasets/7230b4b1-912d-400e-be58-f84e0512985e) |\n| Microsoft Speech Language Translation Corpus | English; Chinese; Japanese| | | \u003chttps://msropendata.com/datasets/54813518-4ea6-4c39-9bb2-b0d1e5f0c187\u003e | [Microsoft Research Data License Agreement](https://msrodr-api.azurewebsites.net//licenses/2f933be3-284d-500b-7ea3-2aa2fd0f1bb2/file) |\n| Hey Snips Corpus | English | 11K positive \"Hey Snips\" (~4.4 hours) and 87K negative (~89 hours) utterances | 2215 speakers (positive \u0026 negative) and 4028 speakers (negative only) | \u003chttps://research.snips.ai/datasets/keyword-spotting\u003e | [Snips Data License](https://github.com/snipsco/keyword-spotting-research-datasets/blob/master/LICENSE) |\n| Snips SLU Corpus | English; French | 1660 \"Smart Lights EN\" (~1.3 hours), 1286 \"Smart Speaker EN\" (~55 minutes), 1138 \"Smart Speaker FR\" (~50 minutes) utterances | English: 69 speakers; French: 30 speakers | \u003chttps://research.snips.ai/datasets/spoken-language-understanding\u003e | [Snips Data License](https://github.com/snipsco/keyword-spotting-research-datasets/blob/master/LICENSE) |\n| CMU Sphinx Group - AN4 | English | \"an4_clstk\"(~50 minutes) \"an4test_clstk\" (~6 minutes) | \"an4_clstk\": 21 female, 53 male \"an4test_clstk\": 3 female, 7 male | http://www.speech.cs.cmu.edu/databases/an4/an4_raw.bigendian.tar.gz | [AN4](http://www.speech.cs.cmu.edu/databases/an4/LICENSE.html) |\n| FT Speech | Danish | ~1,857 hours (1,017,244 utterances) | 434 speakers (176 female, 258 male) | \u003chttps://ftspeech.dk\u003e | [FT Speech License](https://ftspeech.dk/LICENSE.html) |\n| FalaBrasil-LAPS-Constituicao | Brazilian-Portuguese | 9 hours | 1 speaker | \u003chttps://drive.google.com/uc?export=download\u0026confirm=SrvW\u0026id=1Nf849u-27CYRzJqedLaI-FaZfMRO7FT\u003e | [\"Bases de áudio transcrito e bases de texto normalizadas (sem pontuação, com números escritos por extenso, etc.) disponibilizadas de forma gratuita* pelo Grupo FalaBrasil. [disponibilizadas de forma gratuita*] / Portanto, apenas as bases livres estão sendo disponibilizadas.\"](http://labvis.ufpa.br/falabrasil/downloads/) |\n| FalaBrasil-LaPSMail | Brazilian-Portuguese | 1 hour | 25 speakers | \u003chttps://drive.google.com/uc?export=download\u0026confirm=PecV\u0026id=1B_Vq8MDSE4fBQefVxqCGSl-EcKAcjJLb\u003e | [\"Bases de áudio transcrito e bases de texto normalizadas (sem pontuação, com números escritos por extenso, etc.) disponibilizadas de forma gratuita* pelo Grupo FalaBrasil. [disponibilizadas de forma gratuita*] / Portanto, apenas as bases livres estão sendo disponibilizadas.\"](http://labvis.ufpa.br/falabrasil/downloads/) |\n| FalaBrasil-LaPS Benchmark | Brazilian-Portuguese | 1 hour | 1 speaker | \u003chttps://drive.google.com/uc?export=download\u0026confirm=XFfF\u0026id=1nZ8L9nJTt4blFC0RGT9Y7XRu02aAvDIo\u003e | [\"Bases de áudio transcrito e bases de texto normalizadas (sem pontuação, com números escritos por extenso, etc.) disponibilizadas de forma gratuita* pelo Grupo FalaBrasil. [disponibilizadas de forma gratuita*] / Portanto, apenas as bases livres estão sendo disponibilizadas.\"](http://labvis.ufpa.br/falabrasil/downloads/) |\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fcoqui-ai%2Fopen-speech-corpora","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fcoqui-ai%2Fopen-speech-corpora","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fcoqui-ai%2Fopen-speech-corpora/lists"}