{"id":13700977,"url":"https://github.com/SuperKogito/SER-datasets","last_synced_at":"2025-05-04T20:31:38.558Z","repository":{"id":37827967,"uuid":"222908388","full_name":"SuperKogito/SER-datasets","owner":"SuperKogito","description":"A collection of datasets for the purpose of emotion recognition/detection in speech.","archived":false,"fork":false,"pushed_at":"2024-09-30T20:56:41.000Z","size":3903,"stargazers_count":321,"open_issues_count":18,"forks_count":44,"subscribers_count":15,"default_branch":"master","last_synced_at":"2025-04-04T11:05:35.075Z","etag":null,"topics":["audio","audio-datasets","datasets","emotions","emotions-recognition","multimodal-emotion-recognition","speech","speech-emotion-recognition"],"latest_commit_sha":null,"homepage":"https://superkogito.github.io/SER-datasets","language":"HTML","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/SuperKogito.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2019-11-20T10:10:59.000Z","updated_at":"2025-04-04T08:17:09.000Z","dependencies_parsed_at":"2024-01-14T20:49:27.229Z","dependency_job_id":"925abc50-be9f-4850-8083-540c45606c8a","html_url":"https://github.com/SuperKogito/SER-datasets","commit_stats":null,"previous_names":[],"tags_count":10,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/SuperKogito%2FSER-datasets","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/SuperKogito%2FSER-datasets/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/SuperKogito%2FSER-datasets/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/SuperKogito%2FSER-datasets/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/SuperKogito","download_url":"https://codeload.github.com/SuperKogito/SER-datasets/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":252395356,"owners_count":21741026,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["audio","audio-datasets","datasets","emotions","emotions-recognition","multimodal-emotion-recognition","speech","speech-emotion-recognition"],"created_at":"2024-08-02T20:01:12.714Z","updated_at":"2025-05-04T20:31:38.053Z","avatar_url":"https://github.com/SuperKogito.png","language":"HTML","funding_links":[],"categories":["HTML","Speech"],"sub_categories":[],"readme":"***Speech Emotion Recognition (SER) Datasets:*** *A collection of datasets (count=77) for the purpose of emotion recognition/detection in speech.\nThe table is chronologically ordered and includes a description of the content of each dataset along with the emotions included.\nThe table can be browsed, sorted and searched under https://superkogito.github.io/SER-datasets/*\n| Dataset                                                                                                                                           | Year            | Content                                                                                                                                                                                                                                                                          | Emotions                                                                                                                                                                                                                                                                     | Format                        | Size                 | Language                                                          | Paper                                                                                                                                                                                                                                                                                                                                                     | Access                    | License                                                                                                                                        |\n|:--------------------------------------------------------------------------------------------------------------------------------------------------|:----------------|:---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:------------------------------|:---------------------|:------------------------------------------------------------------|:----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|:--------------------------|:-----------------------------------------------------------------------------------------------------------------------------------------------|\n| \u003csub\u003e[nEmo](https://github.com/amu-cai/nEMO)\u003c/sub\u003e                                                                                                | \u003csub\u003e2024\u003c/sub\u003e | \u003csub\u003e3 hours of samples recorded with the participation of nine actors.\u003c/sub\u003e                                                                                                                                                                                                    | \u003csub\u003e6 emotions: anger, fear, happiness, sadness, surprised, and neutral.\u003c/sub\u003e                                                                                                                                                                                              | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e0.434 GB\u003c/sub\u003e  | \u003csub\u003ePolish\u003c/sub\u003e                                                 | \u003csub\u003e[nEMO: Dataset of Emotional Speech in Polish](https://arxiv.org/abs/2404.06292)\u003c/sub\u003e                                                                                                                                                                                                                                                                | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY 4.0](https://creativecommons.org/licenses/by/4.0/)\u003c/sub\u003e                                                                           |\n| \u003csub\u003e[MDER](https://ieee-dataport.org/documents/moroccan-dialect-emotion-recognition-dataset#files)\u003c/sub\u003e                                         | \u003csub\u003e2024\u003c/sub\u003e | \u003csub\u003e2000 voice records of people speaking Moroccan dialect.\u003c/sub\u003e                                                                                                                                                                                                               | \u003csub\u003e5 emotions: Neutral, Happy, Sad, Angry and Fearful.\u003c/sub\u003e                                                                                                                                                                                                               | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e0.187 GB\u003c/sub\u003e  | \u003csub\u003eArabic Moroccan\u003c/sub\u003e                                        | \u003csub\u003e--\u003c/sub\u003e                                                                                                                                                                                                                                                                                                                                             | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY 4.0](https://creativecommons.org/licenses/by/4.0/)\u003c/sub\u003e                                                                           |\n| \u003csub\u003e[EMOVOME](https://zenodo.org/records/10694370)\u003c/sub\u003e                                                                                         | \u003csub\u003e2024\u003c/sub\u003e | \u003csub\u003e999 spontaneous voice messages from 100 Spanish speakers, collected from real conversations on a messaging app.\u003c/sub\u003e                                                                                                                                                       | \u003csub\u003eValence \u0026 arrousal dimensions and 7 emotions: happiness, disgust, anger, surprise, fear, sadness, and neutral.\u003c/sub\u003e                                                                                                                                                    | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e--\u003c/sub\u003e        | \u003csub\u003eSpanish\u003c/sub\u003e                                                | \u003csub\u003e[EMOVOME Database: Advancing Emotion Recognition in Speech Beyond Staged Scenarios](https://arxiv.org/abs/2403.02167)\u003c/sub\u003e                                                                                                                                                                                                                          | \u003csub\u003ePartially open\u003c/sub\u003e | \u003csub\u003e[CC BY 4.0](https://creativecommons.org/licenses/by/4.0/)\u003c/sub\u003e                                                                           |\n| \u003csub\u003e[EMNS](http://www.openslr.org/136/)\u003c/sub\u003e                                                                                                    | \u003csub\u003e2023\u003c/sub\u003e | \u003csub\u003e1206 high quality labeled utterances by one female speaker (2-3 hours).\u003c/sub\u003e                                                                                                                                                                                               | \u003csub\u003eAnger, excitement, disgust, happiness, surprise, sadness, and neutral (plus sarcasm)\u003c/sub\u003e                                                                                                                                                                              | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e0.042 GB\u003c/sub\u003e  | \u003csub\u003eEnglish (British)\u003c/sub\u003e                                      | \u003csub\u003e[EMNS /Imz/ Corpus: An emotive single-speaker dataset for narrative storytelling in games, television and graphic novels](https://arxiv.org/abs/2305.13137)\u003c/sub\u003e                                                                                                                                                                                    | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[Apache 2.0](https://apache.org/licenses/LICENSE-2.0)\u003c/sub\u003e                                                                               |\n| \u003csub\u003e[CAVES](https://rds.westernsydney.edu.au/Institutes/MARCS/2024/Christopher_Davis/)\u003c/sub\u003e                                                     | \u003csub\u003e2023\u003c/sub\u003e | \u003csub\u003eFull hd visual recordings of 10 native cantonese speakers uttering 50 sentences.\u003c/sub\u003e                                                                                                                                                                                      | \u003csub\u003eAnger, happiness, sadness, surprise, fear, disgust and neutral\u003c/sub\u003e                                                                                                                                                                                                    | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e47 GB\u003c/sub\u003e     | \u003csub\u003eChinese (cantonese)\u003c/sub\u003e                                    | \u003csub\u003e[A Cantonese Audio-Visual Emotional Speech (CAVES) dataset](https://link.springer.com/article/10.3758/s13428-023-02270-7)\u003c/sub\u003e                                                                                                                                                                                                                      | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003eAvailable for research purposes only\u003c/sub\u003e                                                                                                |\n| \u003csub\u003e[BANSpEmo](https://data.mendeley.com/datasets/rdwn4bs5ky/2)\u003c/sub\u003e                                                                            | \u003csub\u003e2023\u003c/sub\u003e | \u003csub\u003e792 utterance recordings from 22 unprofessional speakers (11 males and 11 females) of six basic emotional reactions of two sets of sentences.\u003c/sub\u003e                                                                                                                         | \u003csub\u003eangry, disgusted, happy, surprised, sad, fear\u003c/sub\u003e                                                                                                                                                                                                                     | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e0.555 GB\u003c/sub\u003e  | \u003csub\u003eBangla\u003c/sub\u003e                                                 | \u003csub\u003e[BANSpEmo: A Bangla Emotional Speech Recognition Dataset](https://arxiv.org/abs/2312.14020)\u003c/sub\u003e                                                                                                                                                                                                                                                    | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY 4.0](https://creativecommons.org/licenses/by/4.0/)\u003c/sub\u003e                                                                           |\n| \u003csub\u003e[KBES](https://data.mendeley.com/datasets/vsn37ps3rx/4)\u003c/sub\u003e                                                                                | \u003csub\u003e2023\u003c/sub\u003e | \u003csub\u003e900 audio signals from 35 actors (20 females and 15 males). Each emotion is represented with two intensity levels (low \u0026 high)\u003c/sub\u003e                                                                                                                                        | \u003csub\u003eangry, disgusted, happy, neutral, sad\u003c/sub\u003e                                                                                                                                                                                                                             | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e0.337 GB\u003c/sub\u003e  | \u003csub\u003eBangla\u003c/sub\u003e                                                 | \u003csub\u003e[KBES: A dataset for realistic Bangla speech emotion recognition with intensity level](https://www.sciencedirect.com/science/article/pii/S2352340923008107)\u003c/sub\u003e                                                                                                                                                                                    | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY 4.0](https://creativecommons.org/licenses/by/4.0/)\u003c/sub\u003e                                                                           |\n| \u003csub\u003e[RESD](https://huggingface.co/datasets/Aniemore/resd_annotated)\u003c/sub\u003e                                                                        | \u003csub\u003e2022\u003c/sub\u003e | \u003csub\u003eRussian emotional speech dialogue dataset ~3.5 hours of actor-voiced dialogues, each ~3 minutes long, with speech files (16000 or 44100Hz), with speech-to-text transcripts\u003c/sub\u003e                                                                                           | \u003csub\u003eanger, disgust, fear, enthusiasm, happiness, neutral, sadness\u003c/sub\u003e                                                                                                                                                                                                     | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e0.48 GB\u003c/sub\u003e   | \u003csub\u003eRussian\u003c/sub\u003e                                                | \u003csub\u003e[EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark](https://arxiv.org/abs/2406.07162)\u003c/sub\u003e                                                                                                                                                                                                                         | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[MIT](https://choosealicense.com/licenses/mit/)\u003c/sub\u003e                                                                                     |\n| \u003csub\u003e[Hi, KIA](https://zenodo.org/records/7091465)\u003c/sub\u003e                                                                                          | \u003csub\u003e2022\u003c/sub\u003e | \u003csub\u003eA shared short Wakeup Word database focusing on perceived emotion in speech The dataset contains 488 Wakeup Word speech\u003c/sub\u003e                                                                                                                                               | \u003csub\u003eangry, happy, sad, neutral\u003c/sub\u003e                                                                                                                                                                                                                                        | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e0.75 GB\u003c/sub\u003e   | \u003csub\u003eKorean\u003c/sub\u003e                                                 | \u003csub\u003e[Hi, KIA: A Speech Emotion Recognition Dataset for Wake-Up Words](https://arxiv.org/abs/2211.03371)\u003c/sub\u003e                                                                                                                                                                                                                                            | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0/)\u003c/sub\u003e                                                                     |\n| \u003csub\u003e[Emozionalmente](https://zenodo.org/records/6569824)\u003c/sub\u003e                                                                                   | \u003csub\u003e2022\u003c/sub\u003e | \u003csub\u003e6902 labeled samples acted out by 431 amateur actors while verbalizing 18 different sentences\u003c/sub\u003e                                                                                                                                                                         | \u003csub\u003eanger, disgust, fear, joy, sadness, surprise, neutral\u003c/sub\u003e                                                                                                                                                                                                             | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e0.581 GB\u003c/sub\u003e  | \u003csub\u003eItalian\u003c/sub\u003e                                                | \u003csub\u003e--\u003c/sub\u003e                                                                                                                                                                                                                                                                                                                                             | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY 4.0](https://creativecommons.org/licenses/by/4.0/)\u003c/sub\u003e                                                                           |\n| \u003csub\u003e[BanglaSER](https://data.mendeley.com/datasets/t9h6p943xy/5)\u003c/sub\u003e                                                                           | \u003csub\u003e2022\u003c/sub\u003e | \u003csub\u003e1467 Bangla speech-audio recordings by 34 non-professional participating actors (17 male and 17 female) from diverse age groups between 19 and 47 years.\u003c/sub\u003e                                                                                                              | \u003csub\u003eangry, happy, neutral, sad, surprise\u003c/sub\u003e                                                                                                                                                                                                                              | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e0.425 GB\u003c/sub\u003e  | \u003csub\u003eBangla\u003c/sub\u003e                                                 | \u003csub\u003e[BanglaSER: A speech emotion recognition dataset for the Bangla language](https://www.sciencedirect.com/science/article/pii/S235234092200302X)\u003c/sub\u003e                                                                                                                                                                                                 | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY 4.0](https://creativecommons.org/licenses/by/4.0/)\u003c/sub\u003e                                                                           |\n| \u003csub\u003e[B-SER](https://data.mendeley.com/datasets/t9h6p943xy/3)\u003c/sub\u003e                                                                               | \u003csub\u003e2022\u003c/sub\u003e | \u003csub\u003e1224 speech-audio recordings by 34 non-professional participating actors (17 male and 17 female) from diverse age groups between 19 and 47 years.\u003c/sub\u003e                                                                                                                     | \u003csub\u003eangry, happy, sad and surprise\u003c/sub\u003e                                                                                                                                                                                                                                    | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e0.363 GB\u003c/sub\u003e  | \u003csub\u003eBangla\u003c/sub\u003e                                                 | \u003csub\u003e--\u003c/sub\u003e                                                                                                                                                                                                                                                                                                                                             | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY 4.0](https://creativecommons.org/licenses/by/4.0/)\u003c/sub\u003e                                                                           |\n| \u003csub\u003e[Kannada](https://zenodo.org/records/6345107)\u003c/sub\u003e                                                                                          | \u003csub\u003e2022\u003c/sub\u003e | \u003csub\u003e468 audio samples, six different sentences, pronounced by thirteen people (four male and nine female), in five basic emotions plus one neutral emotion\u003c/sub\u003e                                                                                                                | \u003csub\u003eAnger, Sadness, Surprise, Happiness, Fear, Neutral\u003c/sub\u003e                                                                                                                                                                                                                | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e0.1661 GB\u003c/sub\u003e | \u003csub\u003eKannada\u003c/sub\u003e                                                | \u003csub\u003e--\u003c/sub\u003e                                                                                                                                                                                                                                                                                                                                             | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY 4.0](https://creativecommons.org/licenses/by/4.0/)\u003c/sub\u003e                                                                           |\n| \u003csub\u003e[Quechua-SER](https://figshare.com/articles/media/Quechua_Collao_for_Speech_Emotion_Recognition/20292516)\u003c/sub\u003e                              | \u003csub\u003e2022\u003c/sub\u003e | \u003csub\u003e12420 audio recordings (~15 hours) and their transcriptions by 7 native speakers.\u003c/sub\u003e                                                                                                                                                                                     | \u003csub\u003eEmotional labels using dimensions: valence, arousal, and dominance.\u003c/sub\u003e                                                                                                                                                                                               | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e3.53 GB\u003c/sub\u003e   | \u003csub\u003eQuechua Collao\u003c/sub\u003e                                         | \u003csub\u003e[A speech corpus of Quechua Collao for automatic dimensional emotion recognition](https://www.nature.com/articles/s41597-022-01855-9)\u003c/sub\u003e                                                                                                                                                                                                          | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY 4.0](https://creativecommons.org/licenses/by/4.0/)\u003c/sub\u003e                                                                           |\n| \u003csub\u003e[MESD](https://data.mendeley.com/datasets/cy34mh68j9/5)\u003c/sub\u003e                                                                                | \u003csub\u003e2022\u003c/sub\u003e | \u003csub\u003e864 audio files of single-word emotional utterances with Mexican cultural shaping.\u003c/sub\u003e                                                                                                                                                                                    | \u003csub\u003e6 emotions provides single-word utterances for anger, disgust, fear, happiness, neutral, and sadness.\u003c/sub\u003e                                                                                                                                                             | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e0.097 GB\u003c/sub\u003e  | \u003csub\u003eSpanish (Mexican)\u003c/sub\u003e                                      | \u003csub\u003e[The Mexican Emotional Speech Database (MESD): elaboration and assessment based on machine learning](https://pubmed.ncbi.nlm.nih.gov/34891601/)\u003c/sub\u003e                                                                                                                                                                                                | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY 4.0](https://creativecommons.org/licenses/by/4.0/)\u003c/sub\u003e                                                                           |\n| \u003csub\u003e[SyntAct](https://zenodo.org/record/6573016#.ZAjy_9LMJpj)\u003c/sub\u003e                                                                              | \u003csub\u003e2022\u003c/sub\u003e | \u003csub\u003eSynthesized database with 997 utterances of three basic emotions and neutral expression based on rule-based manipulation for a diphone synthesizer which we release to the public \u003c/sub\u003e                                                                                    | \u003csub\u003e6 emotions: angry, bored, happy, neutral, sad and scared\u003c/sub\u003e                                                                                                                                                                                                          | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e0.941 GB\u003c/sub\u003e  | \u003csub\u003eGerman\u003c/sub\u003e                                                 | \u003csub\u003e[SyntAct: A Synthesized Database of Basic Emotions](http://felix.syntheticspeech.de/publications/synthetic_database.pdf)\u003c/sub\u003e                                                                                                                                                                                                                       | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY-SA 4.0](https://creativecommons.org/licenses/by/4.0)\u003c/sub\u003e                                                                         |\n| \u003csub\u003e[BEAT](https://drive.google.com/drive/folders/1EKuWH8q178QOtFUYaNohdkZbBHQYAmhL)\u003c/sub\u003e                                                       | \u003csub\u003e2022\u003c/sub\u003e | \u003csub\u003e76-Hour and 30-Speaker of 4 different languages: English (60h), Chinese (12h), Spanish (2h) and Japanese (2h).\u003c/sub\u003e                                                                                                                                                        | \u003csub\u003e8 emotions: happiness, anger, disgust, sadness, contempt, surprise, fear, and neutral\u003c/sub\u003e                                                                                                                                                                             | \u003csub\u003eAudio, Video\u003c/sub\u003e       | \u003csub\u003e42 GB\u003c/sub\u003e     | \u003csub\u003eEnglish, Chinese, Spanish, Japanese\u003c/sub\u003e                    | \u003csub\u003e[A Large-Scale Semantic and Emotional Multi-Modal Dataset for Conversational Gestures Synthesis](https://www.ecva.net/papers/eccv_2022/papers_ECCV/papers/136670605.pdf)\u003c/sub\u003e                                                                                                                                                                       | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003eNon-commercial license\u003c/sub\u003e                                                                                                              |\n| \u003csub\u003e[Dusha](https://github.com/salute-developers/golos/tree/master/dusha)\u003c/sub\u003e                                                                  | \u003csub\u003e2022\u003c/sub\u003e | \u003csub\u003e300000 audio recordings (~350 hours) of Russian speech, their transcripts and emotiomal labels. The dataset has two subsets: acted and real-life\u003c/sub\u003e                                                                                                                      | \u003csub\u003e4 emotions: angry, happy, sad and neutral. Arousal and valence metrics are also available.\u003c/sub\u003e                                                                                                                                                                        | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e58 GB\u003c/sub\u003e     | \u003csub\u003eRussian\u003c/sub\u003e                                                | \u003csub\u003e[Large Raw Emotional Dataset with Aggregation Mechanism](https://arxiv.org/abs/2212.12266)\u003c/sub\u003e                                                                                                                                                                                                                                                     | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[Public license with attribution and conditions reserved](https://github.com/salute-developers/golos/blob/master/license/en_us.pdf)\u003c/sub\u003e |\n| \u003csub\u003e[MAFW](https://mafw-database.github.io/MAFW/)\u003c/sub\u003e                                                                                          | \u003csub\u003e2022\u003c/sub\u003e | \u003csub\u003e10045 video-audio clips in the wild.\u003c/sub\u003e                                                                                                                                                                                                                                  | \u003csub\u003e11 single-label emotion categories (anger, disgust, fear, happiness, neutral, sadness, surprise, contempt, anxiety, helplessness, and disappointment) and 32 multi-label emotion categories.\u003c/sub\u003e                                                                      | \u003csub\u003eAudio, Video\u003c/sub\u003e       | \u003csub\u003e--\u003c/sub\u003e        | \u003csub\u003e--\u003c/sub\u003e                                                     | \u003csub\u003e[MAFW: A Large-scale, Multi-modal, Compound Affective Database for Dynamic Facial Expression Recognition in the Wild](https://arxiv.org/abs/2208.00847)\u003c/sub\u003e                                                                                                                                                                                        | \u003csub\u003eRestricted\u003c/sub\u003e     | \u003csub\u003eNon-commercial research purposes\u003c/sub\u003e                                                                                                    |\n| \u003csub\u003e[EMOVIE](https://viem-ccy.github.io/EMOVIE/dataset_release.html)\u003c/sub\u003e                                                                       | \u003csub\u003e2021\u003c/sub\u003e | \u003csub\u003e9724 samples with audio files and its emotion human-labeled annotation.\u003c/sub\u003e                                                                                                                                                                                               | \u003csub\u003ePolarity metrics (positive:+1, negative:-1)\u003c/sub\u003e                                                                                                                                                                                                                       | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e0.572 GB\u003c/sub\u003e  | \u003csub\u003eChinese (Mandarin)\u003c/sub\u003e                                     | \u003csub\u003e[EMOVIE: A Mandarin Emotion Speech Dataset with a Simple Emotional Text-to-Speech Model](https://arxiv.org/abs/2106.09317)\u003c/sub\u003e                                                                                                                                                                                                                     | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY-NC-SA 2.0](https://creativecommons.org/licenses/by-nc-sa/2.0/legalcode)\u003c/sub\u003e                                                      |\n| \u003csub\u003e[emoUERJ](https://zenodo.org/records/5427549)\u003c/sub\u003e                                                                                          | \u003csub\u003e2021\u003c/sub\u003e | \u003csub\u003eTen sentences from eight actors, equally divided between genders, and they were free to choose the phrases for record audios in four emotions (377 audios). \u003c/sub\u003e                                                                                                          | \u003csub\u003ehappiness, anger, sadness or neutral\u003c/sub\u003e                                                                                                                                                                                                                              | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e0.1051 GB\u003c/sub\u003e | \u003csub\u003ePortuguese (Brazilian)\u003c/sub\u003e                                 | \u003csub\u003e--\u003c/sub\u003e                                                                                                                                                                                                                                                                                                                                             | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY 4.0](https://creativecommons.org/licenses/by/4.0/)\u003c/sub\u003e                                                                           |\n| \u003csub\u003e[Thorsten-Voice Dataset 2021.06 emotional](https://zenodo.org/records/5525023)\u003c/sub\u003e                                                         | \u003csub\u003e2021\u003c/sub\u003e | \u003csub\u003e2400 normalized mono recordings by one person (Thorsten Müller) representing 300 sentences. \u003c/sub\u003e                                                                                                                                                                          | \u003csub\u003eAmusement, Disgust Anger, Suprise and Neutral (plus drunk, whispering and sleepy states)\u003c/sub\u003e                                                                                                                                                                          | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e0.399 GB\u003c/sub\u003e  | \u003csub\u003eGerman\u003c/sub\u003e                                                 | \u003csub\u003e--\u003c/sub\u003e                                                                                                                                                                                                                                                                                                                                             | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC0: Public Domain](https://creativecommons.org/publicdomain/zero/1.0/)\u003c/sub\u003e                                                            |\n| \u003csub\u003e[ASED](https://github.com/Ethio2021/ASED_V1)\u003c/sub\u003e                                                                                           | \u003csub\u003e2021\u003c/sub\u003e | \u003csub\u003e2474 recordings by 65 participants (25 females and 40 males)). Recordings were judged and rejected according to the opionion of eight judges.\u003c/sub\u003e                                                                                                                         | \u003csub\u003eFive emotions: anger, happiness, fear, sadness and neutral\u003c/sub\u003e                                                                                                                                                                                                        | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e0.135 GB\u003c/sub\u003e  | \u003csub\u003eAmharic\u003c/sub\u003e                                                | \u003csub\u003e[A New Amharic Speech Emotion Dataset and Classification Benchmark](https://dl.acm.org/doi/10.1145/3529759)\u003c/sub\u003e                                                                                                                                                                                                                                    | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e--\u003c/sub\u003e                                                                                                                                  |\n| \u003csub\u003e[ESCorpus-PE](https://zenodo.org/records/5793223)\u003c/sub\u003e                                                                                      | \u003csub\u003e2021\u003c/sub\u003e | \u003csub\u003eSpanish peruvian speech gathered from Spanish interviews, TV reports, political debate and testimonials. It contains 3749 utterances, 80 speakers (44 male and 36 female), created from Youtube audios\u003c/sub\u003e                                                                | \u003csub\u003eValence, Arousal and Dominance\u003c/sub\u003e                                                                                                                                                                                                                                    | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e1.9 GB\u003c/sub\u003e    | \u003csub\u003eSpanish (Peruvian)\u003c/sub\u003e                                     | \u003csub\u003e--\u003c/sub\u003e                                                                                                                                                                                                                                                                                                                                             | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0/)\u003c/sub\u003e                                                                     |\n| \u003csub\u003e[SUBSECO](https://zenodo.org/record/6339787)\u003c/sub\u003e                                                                                           | \u003csub\u003e2021\u003c/sub\u003e | \u003csub\u003e7000 sentence-level utterances of the Bangla language, 20 professional actors (10 males and 10 females), recordings, 10 sentences for 7 target emotions.\u003c/sub\u003e                                                                                                              | \u003csub\u003eAnger, Disgust, Fear, Happiness, Neutral, Sadness and Surprise\u003c/sub\u003e                                                                                                                                                                                                    | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e1.7 GB\u003c/sub\u003e    | \u003csub\u003eEnglish\u003c/sub\u003e                                                | \u003csub\u003e[SUST Bangla Emotional Speech Corpus (SUBESCO): An audio-only emotional speech corpus for Bangla](https://doi.org/10.1371/journal.pone.0250173)\u003c/sub\u003e                                                                                                                                                                                                | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY 4.0](https://creativecommons.org/licenses/by/4.0/)\u003c/sub\u003e                                                                           |\n| \u003csub\u003e[Audio-Speech-Sentiment](https://www.kaggle.com/imsparsh/audio-speech-sentiment-analysis)\u003c/sub\u003e                                              | \u003csub\u003e2021\u003c/sub\u003e | \u003csub\u003eAudio Speech Sentiment Dataset\u003c/sub\u003e                                                                                                                                                                                                                                        | \u003csub\u003e4 emotions provides audio recordings of spoken sentences for anger, happiness, sadness, and neutral emotions.\u003c/sub\u003e                                                                                                                                                     | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e1.1 GB\u003c/sub\u003e    | \u003csub\u003eEnglish\u003c/sub\u003e                                                | \u003csub\u003e--\u003c/sub\u003e                                                                                                                                                                                                                                                                                                                                             | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC0: Public Domain](https://creativecommons.org/publicdomain/zero/1.0/)\u003c/sub\u003e                                                            |\n| \u003csub\u003e[LSSED](https://github.com/tobefans/LSSED)\u003c/sub\u003e                                                                                             | \u003csub\u003e2021\u003c/sub\u003e | \u003csub\u003eLSSED: A Large-Scale Dataset and Benchmark for Speech Emotion Recognition\u003c/sub\u003e                                                                                                                                                                                             | \u003csub\u003eAnger, happiness, sadness, disappointment, boredom, disgust, excitement, fear, surprise, normal, and other.\u003c/sub\u003e                                                                                                                                                       | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e90 GB\u003c/sub\u003e     | \u003csub\u003eEnglish\u003c/sub\u003e                                                | \u003csub\u003e[LSSED: A Large-Scale Spanish Emotional Speech Database for Speech Processing and Machine Learning](https://arxiv.org/abs/2102.01754)\u003c/sub\u003e                                                                                                                                                                                                          | \u003csub\u003eRestricted\u003c/sub\u003e     | \u003csub\u003e[-](https://github.com/tobefans/LSSED/blob/main/EULA.pdf)\u003c/sub\u003e                                                                           |\n| \u003csub\u003e[MLEnd](https://www.kaggle.com/datasets/jesusrequena/mlend-spoken-numerals)\u003c/sub\u003e                                                            | \u003csub\u003e2021\u003c/sub\u003e | \u003csub\u003e~32700 audio recordings files produced by 154 speakers. Each audio recording corresponds to one English numeral (from \"zero\" to \"billion\")\u003c/sub\u003e                                                                                                                            | \u003csub\u003eIntonations: neutral, bored, excited and question\u003c/sub\u003e                                                                                                                                                                                                                 | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e2.27 GB\u003c/sub\u003e   | \u003csub\u003e--\u003c/sub\u003e                                                     | \u003csub\u003e--\u003c/sub\u003e                                                                                                                                                                                                                                                                                                                                             | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003eUnknown\u003c/sub\u003e                                                                                                                             |\n| \u003csub\u003e[ASVP-ESD](https://www.kaggle.com/datasets/dejolilandry/asvpesdspeech-nonspeech-emotional-utterances)\u003c/sub\u003e                                  | \u003csub\u003e2021\u003c/sub\u003e | \u003csub\u003e~13285 audio files collected from movies, tv shows and youtube containing speech and non-speech.\u003c/sub\u003e                                                                                                                                                                      | \u003csub\u003e12 different natural emotions (boredom, neutral, happiness, sadness, anger, fear, surprise, disgust, excitement, pleasure, pain, disappointment) with 2 levels of intensity.\u003c/sub\u003e                                                                                      | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e2 GB\u003c/sub\u003e      | \u003csub\u003eChinese, English, French, Russian and others\u003c/sub\u003e           | \u003csub\u003e--\u003c/sub\u003e                                                                                                                                                                                                                                                                                                                                             | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003eUnknown\u003c/sub\u003e                                                                                                                             |\n| \u003csub\u003e[ESD](https://hltsingapore.github.io/ESD/)\u003c/sub\u003e                                                                                             | \u003csub\u003e2021\u003c/sub\u003e | \u003csub\u003e29 hours, 3500 sentences, by 10 native English speakers and 10 native Chinese speakers.\u003c/sub\u003e                                                                                                                                                                               | \u003csub\u003e5 emotions: angry, happy, neutral, sad, and surprise.\u003c/sub\u003e                                                                                                                                                                                                             | \u003csub\u003eAudio,  Text\u003c/sub\u003e       | \u003csub\u003e2.4 GB\u003c/sub\u003e    | \u003csub\u003eChinese, English\u003c/sub\u003e                                       | \u003csub\u003e[Seen And Unseen Emotional Style Transfer For Voice Conversion With A New Emotional Speech Dataset](https://arxiv.org/pdf/2010.14794.pdf)\u003c/sub\u003e                                                                                                                                                                                                      | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003eAcademic License\u003c/sub\u003e                                                                                                                    |\n| \u003csub\u003e[MuSe-CAR](https://zenodo.org/record/4134758)\u003c/sub\u003e                                                                                          | \u003csub\u003e2021\u003c/sub\u003e | \u003csub\u003e40 hours, 6,000+ recordings of 25,000+ sentences by 70+ English speakers (see db link for details).\u003c/sub\u003e                                                                                                                                                                   | \u003csub\u003econtinuous emotion dimensions characterized using valence, arousal, and trustworthiness.\u003c/sub\u003e                                                                                                                                                                          | \u003csub\u003eAudio, Video, Text\u003c/sub\u003e | \u003csub\u003e15 GB\u003c/sub\u003e     | \u003csub\u003eEnglish\u003c/sub\u003e                                                | \u003csub\u003e[The Multimodal Sentiment Analysis in Car Reviews (MuSe-CaR) Dataset: Collection, Insights and Improvements](https://arxiv.org/pdf/2101.06053.pdf)\u003c/sub\u003e                                                                                                                                                                                             | \u003csub\u003eRestricted\u003c/sub\u003e     | \u003csub\u003eAcademic License \u0026 Commercial License\u003c/sub\u003e                                                                                               |\n| \u003csub\u003e[THAI SER](https://github.com/vistec-AI/dataset-releases/releases/tag/v1)\u003c/sub\u003e                                                              | \u003csub\u003e2021\u003c/sub\u003e | \u003csub\u003eThe recordings are 41 hours, 36 minutes long (27,854 utterances), and were performed by 200 professional actors (112 female, 88 male).\u003c/sub\u003e                                                                                                                                | \u003csub\u003e5 main emotions assigned to actors: Neutral, Anger, Happiness, Sadness, and Frustration.\u003c/sub\u003e                                                                                                                                                                          | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e12 GB\u003c/sub\u003e     | \u003csub\u003eThai\u003c/sub\u003e                                                   | \u003csub\u003e--\u003c/sub\u003e                                                                                                                                                                                                                                                                                                                                             | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0)\u003c/sub\u003e                                                                      |\n| \u003csub\u003e[French Emotional Speech Database - Oréau](https://zenodo.org/records/4405783#.Yqjq_9JBxph)\u003c/sub\u003e                                            | \u003csub\u003e2020\u003c/sub\u003e | \u003csub\u003e79 utterances with 10 to 13 utterances pro emotion by 32 non-professional speakers.\u003c/sub\u003e                                                                                                                                                                                   | \u003csub\u003e7 emotions: sadness, anger, disgust, fear, surprise, joy, neutral.\u003c/sub\u003e                                                                                                                                                                                                | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e0.264 GB\u003c/sub\u003e  | \u003csub\u003eFrench\u003c/sub\u003e                                                 | \u003csub\u003e--\u003c/sub\u003e                                                                                                                                                                                                                                                                                                                                             | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY 4.0](https://creativecommons.org/licenses/by/4.0/)\u003c/sub\u003e                                                                           |\n| \u003csub\u003e[Att-HACK ](http://www.openslr.org/88/)\u003c/sub\u003e                                                                                                | \u003csub\u003e2020\u003c/sub\u003e | \u003csub\u003e25 speakers interpreting 100 utterances in 4 social attitudes, with 3-5 repetitions each per attitude for a total of around 30 hours of speech.\u003c/sub\u003e                                                                                                                       | \u003csub\u003eexpressive speech in French, 100 phrases with multiple versions (3 to 5) in four social attitudes (friendly, distant, dominant and seductive).\u003c/sub\u003e                                                                                                                    | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e6.6 GB\u003c/sub\u003e    | \u003csub\u003eFrench\u003c/sub\u003e                                                 | \u003csub\u003e[Att-HACK: An Expressive Speech Database with Social Attitudes](https://arxiv.org/abs/2004.04410)\u003c/sub\u003e                                                                                                                                                                                                                                              | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY-NC-ND 4.0](https://creativecommons.org/licenses/by-nc-nd/4.0/)\u003c/sub\u003e                                                               |\n| \u003csub\u003e[MSP-Podcast corpus](https://ecs.utdallas.edu/research/researchlabs/msp-lab/MSP-Podcast.html)\u003c/sub\u003e                                          | \u003csub\u003e2020\u003c/sub\u003e | \u003csub\u003e100 hours by over 100 speakers (see db link for details).\u003c/sub\u003e                                                                                                                                                                                                             | \u003csub\u003eThis corpus is annotated with emotional labels using attribute-based descriptors (activation, dominance and valence) and categorical labels (anger, happiness, sadness, disgust, surprised, fear, contempt, neutral and other).\u003c/sub\u003e                                   | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e13.4 GB\u003c/sub\u003e   | \u003csub\u003eEnglish\u003c/sub\u003e                                                | \u003csub\u003e[The MSP-Conversation Corpus](http://www.interspeech2020.org/index.php?m=content\u0026c=index\u0026a=show\u0026catid=290\u0026id=684)\u003c/sub\u003e                                                                                                                                                                                                                              | \u003csub\u003eRestricted\u003c/sub\u003e     | \u003csub\u003eAcademic License \u0026 Commercial License\u003c/sub\u003e                                                                                               |\n| \u003csub\u003e[AISHELL-3](https://www.openslr.org/93/)\u003c/sub\u003e                                                                                               | \u003csub\u003e2020\u003c/sub\u003e | \u003csub\u003eRoughly 85 hours of emotion-neutral recordings spoken by 218 native Chinese mandarin speakers and total 88035 utterances.\u003c/sub\u003e                                                                                                                                             | \u003csub\u003eNeutral\u003c/sub\u003e                                                                                                                                                                                                                                                           | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e19 GB\u003c/sub\u003e     | \u003csub\u003eChinese (Mandarin)\u003c/sub\u003e                                     | \u003csub\u003e[AISHELL-3: A Multi-speaker Mandarin TTS Corpus and the Baselines](https://arxiv.org/abs/2010.11567)\u003c/sub\u003e                                                                                                                                                                                                                                           | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[Apache 2.0](https://apache.org/licenses/LICENSE-2.0)\u003c/sub\u003e                                                                               |\n| \u003csub\u003e[BEASC](https://doi.org/10.6084/m9.figshare.12498033)\u003c/sub\u003e                                                                                  | \u003csub\u003e2020\u003c/sub\u003e | \u003csub\u003eBangla Emotional Audio-Speech Corpus\u003c/sub\u003e                                                                                                                                                                                                                                  | \u003csub\u003e6 emotions provides Bangla spoken utterances for anger, happiness, sadness, fear, surprise, and neutral.\u003c/sub\u003e                                                                                                                                                          | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e9 GB\u003c/sub\u003e      | \u003csub\u003eBangla\u003c/sub\u003e                                                 | \u003csub\u003e[BEASC: Bangla Emotional Audio-Speech Corpus - A Speech Emotion Recognition Corpus for the Low-Resource Bangla Language](https://easy.dans.knaw.nl/ui/datasets/id/easy-dataset:236649)\u003c/sub\u003e                                                                                                                                                         | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY 4.0](https://creativecommons.org/licenses/by/4.0/)\u003c/sub\u003e                                                                           |\n| \u003csub\u003e[emotiontts open db](https://github.com/emotiontts/emotiontts_open_db)\u003c/sub\u003e                                                                 | \u003csub\u003e2020\u003c/sub\u003e | \u003csub\u003eRecordings and their associated transcriptions by a diverse group of speakers.\u003c/sub\u003e                                                                                                                                                                                        | \u003csub\u003e4 emotions: general, joy, anger, and sadness.\u003c/sub\u003e                                                                                                                                                                                                                     | \u003csub\u003eAudio, Text\u003c/sub\u003e        | \u003csub\u003e--\u003c/sub\u003e        | \u003csub\u003eKorean\u003c/sub\u003e                                                 | \u003csub\u003e--\u003c/sub\u003e                                                                                                                                                                                                                                                                                                                                             | \u003csub\u003ePartially open\u003c/sub\u003e | \u003csub\u003e[CC BY-NC-SA 4.0](https://creativecommons.org/licenses/by-nc-sa/4.0/)\u003c/sub\u003e                                                               |\n| \u003csub\u003e[URDU-Dataset](https://github.com/siddiquelatif/urdu-dataset)\u003c/sub\u003e                                                                          | \u003csub\u003e2020\u003c/sub\u003e | \u003csub\u003e400 utterances by 38 speakers (27 male and 11 female).\u003c/sub\u003e                                                                                                                                                                                                                | \u003csub\u003e4 emotions: angry, happy, neutral, and sad.\u003c/sub\u003e                                                                                                                                                                                                                       | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e0.072 GB\u003c/sub\u003e  | \u003csub\u003eUrdu\u003c/sub\u003e                                                   | \u003csub\u003e[Cross Lingual Speech Emotion Recognition: Urdu vs. Western Languages](https://arxiv.org/pdf/1812.10411.pdf)\u003c/sub\u003e                                                                                                                                                                                                                                   | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e--\u003c/sub\u003e                                                                                                                                  |\n| \u003csub\u003e[BAVED](https://www.kaggle.com/a13x10/basic-arabic-vocal-emotions-dataset)\u003c/sub\u003e                                                             | \u003csub\u003e2020\u003c/sub\u003e | \u003csub\u003e1935 recording by 61 speakers (45 male and 16 female).\u003c/sub\u003e                                                                                                                                                                                                                | \u003csub\u003e3 levels of emotion.\u003c/sub\u003e                                                                                                                                                                                                                                              | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e0.195 GB\u003c/sub\u003e  | \u003csub\u003eArabic\u003c/sub\u003e                                                 | \u003csub\u003e--\u003c/sub\u003e                                                                                                                                                                                                                                                                                                                                             | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e--\u003c/sub\u003e                                                                                                                                  |\n| \u003csub\u003e[VIVAE](https://zenodo.org/record/4066235)\u003c/sub\u003e                                                                                             | \u003csub\u003e2020\u003c/sub\u003e | \u003csub\u003enon-speech, 1085 audio file by 11 speakers.\u003c/sub\u003e                                                                                                                                                                                                                           | \u003csub\u003enon-speech 6 emotions: achievement, anger, fear, pain, pleasure, and surprise with 3 emotional intensities (low, moderate, strong, peak).\u003c/sub\u003e                                                                                                                         | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e0.0935 GB\u003c/sub\u003e | \u003csub\u003eNonverbal (English)\u003c/sub\u003e                                    | \u003csub\u003e[The Variably Intense Vocalizations of Affect and Emotion (VIVAE) corpus prompts new perspective on nonspeech perception](http://dx.doi.org/10.1037/emo0001048)\u003c/sub\u003e                                                                                                                                                                                | \u003csub\u003eRestricted\u003c/sub\u003e     | \u003csub\u003e[CC BY-NC-SA 4.0](https://creativecommons.org/licenses/by-nc-sa/4.0/)\u003c/sub\u003e                                                               |\n| \u003csub\u003e[VESUS](https://engineering.jhu.edu/nsa/vesus/)\u003c/sub\u003e                                                                                        | \u003csub\u003e2019\u003c/sub\u003e | \u003csub\u003e252 distinct phrases, each read by 10 actors totalling 6 hours of speech.\u003c/sub\u003e                                                                                                                                                                                             | \u003csub\u003e5 emotions: anger, happiness, sadness, fear and neutral.\u003c/sub\u003e                                                                                                                                                                                                          | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e--\u003c/sub\u003e        | \u003csub\u003eEnglish\u003c/sub\u003e                                                | \u003csub\u003e[VESUS: A Crowd-Annotated Database to Study Emotion Production and Perception in Spoken English](https://engineering.jhu.edu/nsa/wp-content/uploads/2019/10/IS191413.pdf)\u003c/sub\u003e                                                                                                                                                                      | \u003csub\u003eRestricted\u003c/sub\u003e     | \u003csub\u003eAcademic EULA\u003c/sub\u003e                                                                                                                       |\n| \u003csub\u003e[Morgan Emotional Speech Set](https://arxiv.org/abs/2403.02167)\u003c/sub\u003e                                                                        | \u003csub\u003e2019\u003c/sub\u003e | \u003csub\u003e999 spontaneous voice messages from 100 Spanish speakers, collected from real conversations on a messaging app.\u003c/sub\u003e                                                                                                                                                       | \u003csub\u003eValence \u0026 arrousal dimensions and 4 emotions: happiness, anger, sadness, and calmness.\u003c/sub\u003e                                                                                                                                                                            | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e0.192 GB\u003c/sub\u003e  | \u003csub\u003eEnglish\u003c/sub\u003e                                                | \u003csub\u003e[Categorical and Dimensional Ratings of Emotional Speech: Behavioral Findings From the Morgan Emotional Speech Set](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC7203525/)\u003c/sub\u003e                                                                                                                                                                     | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY 4.0](https://creativecommons.org/licenses/by/4.0/)\u003c/sub\u003e                                                                           |\n| \u003csub\u003e[PMEmo](https://github.com/HuiZhangDB/PMEmo)\u003c/sub\u003e                                                                                           | \u003csub\u003e2019\u003c/sub\u003e | \u003csub\u003eDataset containing emotion annotations of 794 songs as well as the simultaneous electrodermal activity (EDA) signals. A Music Emotion Experiment was well-designed for collecting the affective-annotated music corpus of high quality, which recruited 457 subjects.\u003c/sub\u003e | \u003csub\u003eValence, Arousal\u003c/sub\u003e                                                                                                                                                                                                                                                  | \u003csub\u003eAudio, EDA\u003c/sub\u003e         | \u003csub\u003e1.3 GB\u003c/sub\u003e    | \u003csub\u003eChinese, English\u003c/sub\u003e                                       | \u003csub\u003e[The PMEmo Dataset for Music Emotion Recognition](https://dl.acm.org/doi/10.1145/3206025.3206037)\u003c/sub\u003e                                                                                                                                                                                                                                              | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY-SA 4.0](https://huisblog.cn/PMEmo//)\u003c/sub\u003e                                                                                         |\n| \u003csub\u003e[SEWA](https://db.sewaproject.eu/)\u003c/sub\u003e                                                                                                     | \u003csub\u003e2019\u003c/sub\u003e | \u003csub\u003emore than 2000 minutes of audio-visual data of 398 people (201 male and 197 female) coming from 6 cultures.\u003c/sub\u003e                                                                                                                                                           | \u003csub\u003eemotions are characterized using valence and arousal.\u003c/sub\u003e                                                                                                                                                                                                             | \u003csub\u003eAudio, Video\u003c/sub\u003e       | \u003csub\u003e--\u003c/sub\u003e        | \u003csub\u003eChinese, English, German, Greek, Hungarian and Serbian\u003c/sub\u003e | \u003csub\u003e[SEWA DB: A Rich Database for Audio-Visual Emotion and Sentiment Research in the Wild](https://arxiv.org/pdf/1901.02839.pdf)\u003c/sub\u003e                                                                                                                                                                                                                   | \u003csub\u003eRestricted\u003c/sub\u003e     | \u003csub\u003e[SEWA EULA](https://db.sewaproject.eu/media/doc/eula.pdf)\u003c/sub\u003e                                                                           |\n| \u003csub\u003e[MELD](https://affective-meld.github.io/)\u003c/sub\u003e                                                                                              | \u003csub\u003e2019\u003c/sub\u003e | \u003csub\u003e1400 dialogues and 14000 utterances from Friends TV series  by multiple speakers.\u003c/sub\u003e                                                                                                                                                                                     | \u003csub\u003e7 emotions: Anger, disgust, sadness, joy, neutral, surprise and fear.  MELD also has sentiment (positive, negative and neutral) annotation  for each utterance.\u003c/sub\u003e                                                                                                   | \u003csub\u003eAudio, Video, Text\u003c/sub\u003e | \u003csub\u003e10.1 GB\u003c/sub\u003e   | \u003csub\u003eEnglish\u003c/sub\u003e                                                | \u003csub\u003e[MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations](https://arxiv.org/pdf/1810.02508.pdf)\u003c/sub\u003e                                                                                                                                                                                                                        | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[MELD: GPL-3.0 License](https://github.com/declare-lab/MELD/blob/master/LICENSE)\u003c/sub\u003e                                                    |\n| \u003csub\u003e[ShEMO](https://github.com/mansourehk/ShEMO)\u003c/sub\u003e                                                                                           | \u003csub\u003e2019\u003c/sub\u003e | \u003csub\u003e3000 semi-natural utterances, equivalent to 3 hours and 25 minutes of speech data from online radio plays by 87 native-Persian speakers.\u003c/sub\u003e                                                                                                                              | \u003csub\u003e6 emotions: anger, fear, happiness, sadness, neutral and surprise.\u003c/sub\u003e                                                                                                                                                                                                | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e0.101 GB\u003c/sub\u003e  | \u003csub\u003ePersian\u003c/sub\u003e                                                | \u003csub\u003e[ShEMO: a large-scale validated database for Persian speech emotion detection](https://link.springer.com/article/10.1007/s10579-018-9427-x)\u003c/sub\u003e                                                                                                                                                                                                    | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e--\u003c/sub\u003e                                                                                                                                  |\n| \u003csub\u003e[DEMoS](https://zenodo.org/record/2544829)\u003c/sub\u003e                                                                                             | \u003csub\u003e2019\u003c/sub\u003e | \u003csub\u003e9365 emotional and 332 neutral samples produced by 68 native speakers (23 females, 45 males).\u003c/sub\u003e                                                                                                                                                                         | \u003csub\u003e7/6 emotions: anger, sadness, happiness, fear, surprise, disgust, and the secondary emotion guilt.\u003c/sub\u003e                                                                                                                                                                | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e2.5 GB\u003c/sub\u003e    | \u003csub\u003eItalian\u003c/sub\u003e                                                | \u003csub\u003e[DEMoS: An Italian emotional speech corpus. Elicitation methods, machine learning, and perception](https://link.springer.com/epdf/10.1007/s10579-019-09450-y?author_access_token=5pf0w_D4k9z28TM6n4PbVPe4RwlQNchNByi7wbcMAY5hiA-aXzXNbZYfsMDDq2CdHD-w5ArAxIwlsk2nC_26pSyEAcu1xlKJ1c9m3JZj2ZlFmlVoCZUTcG3Hq2_2ozMLo3Hq3Y0CHzLdTxihQwch5Q%3D%3D)\u003c/sub\u003e | \u003csub\u003eRestricted\u003c/sub\u003e     | \u003csub\u003eEULA: End User License Agreement\u003c/sub\u003e                                                                                                    |\n| \u003csub\u003e[AESDD](http://m3c.web.auth.gr/research/aesdd-speech-emotion-recognition/)\u003c/sub\u003e                                                             | \u003csub\u003e2018\u003c/sub\u003e | \u003csub\u003earound 500 utterances by a diverse group of actors (over 5 actors) siumlating various emotions.\u003c/sub\u003e                                                                                                                                                                       | \u003csub\u003e5 emotions: anger, disgust, fear, happiness, and sadness.\u003c/sub\u003e                                                                                                                                                                                                         | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e0.392 GB\u003c/sub\u003e  | \u003csub\u003eGreek\u003c/sub\u003e                                                  | \u003csub\u003e[Speech Emotion Recognition for Performance Interaction](https://www.researchgate.net/publication/326005164_Speech_Emotion_Recognition_for_Performance_Interaction)\u003c/sub\u003e                                                                                                                                                                            | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e--\u003c/sub\u003e                                                                                                                                  |\n| \u003csub\u003e[Emov-DB](https://mega.nz/#F!KBp32apT!gLIgyWf9iQ-yqnWFUFuUHg!mYwUnI4K)\u003c/sub\u003e                                                                 | \u003csub\u003e2018\u003c/sub\u003e | \u003csub\u003eRecordings for 4 speakers- 2 males and 2 females.\u003c/sub\u003e                                                                                                                                                                                                                     | \u003csub\u003eThe emotional styles are neutral, sleepiness, anger, disgust and amused.\u003c/sub\u003e                                                                                                                                                                                          | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e5.88 GB\u003c/sub\u003e   | \u003csub\u003eEnglish\u003c/sub\u003e                                                | \u003csub\u003e[The emotional voices database: Towards controlling the emotion dimension in voice generation systems](https://arxiv.org/pdf/1806.09514.pdf)\u003c/sub\u003e                                                                                                                                                                                                   | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e--\u003c/sub\u003e                                                                                                                                  |\n| \u003csub\u003e[OMG Emotion](https://www2.informatik.uni-hamburg.de/wtm/OMG-EmotionChallenge/)\u003c/sub\u003e                                                        | \u003csub\u003e2018\u003c/sub\u003e | \u003csub\u003e420 relatively long emotion videos with an average length of 1 minute, collected from a variety of Youtube channels.\u003c/sub\u003e                                                                                                                                                  | \u003csub\u003e7 emotions:anger, disgust, fear, happy, sad, surprise and neutral. Plus valence, arousal.\u003c/sub\u003e                                                                                                                                                                         | \u003csub\u003eAudio, Video\u003c/sub\u003e       | \u003csub\u003e--\u003c/sub\u003e        | \u003csub\u003eEnglish\u003c/sub\u003e                                                | \u003csub\u003e[The OMG-Emotion Behavior Dataset](https://arxiv.org/abs/1803.05434)\u003c/sub\u003e                                                                                                                                                                                                                                                                           | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY-NC-SA 3.0](https://creativecommons.org/licenses/by-nc-sa/3.0/)\u003c/sub\u003e                                                               |\n| \u003csub\u003e[RAVDESS](https://zenodo.org/record/1188976#.XrC7a5NKjOR)\u003c/sub\u003e                                                                              | \u003csub\u003e2018\u003c/sub\u003e | \u003csub\u003e7356 recordings by 24 actors.\u003c/sub\u003e                                                                                                                                                                                                                                         | \u003csub\u003e7 emotions: calm, happy, sad, angry, fearful, surprise, and disgust\u003c/sub\u003e                                                                                                                                                                                               | \u003csub\u003eAudio, Video\u003c/sub\u003e       | \u003csub\u003e24.8 GB\u003c/sub\u003e   | \u003csub\u003eEnglish\u003c/sub\u003e                                                | \u003csub\u003e[The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS): A dynamic, multimodal set of facial and vocal expressions in North American English](https://journals.plos.org/plosone/article?id=10.1371/journal.pone.0196391)\u003c/sub\u003e                                                                                                     | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY-NC-SA 4.0](https://creativecommons.org/licenses/by-nc-sa/4.0/)\u003c/sub\u003e                                                               |\n| \u003csub\u003e[JL corpus](https://www.kaggle.com/tli725/jl-corpus)\u003c/sub\u003e                                                                                   | \u003csub\u003e2018\u003c/sub\u003e | \u003csub\u003e2400 recording of 240 sentences by 4 actors (2 males and 2 females).\u003c/sub\u003e                                                                                                                                                                                                  | \u003csub\u003e5 primary emotions: angry, sad, neutral, happy, excited. 5 secondary emotions: anxious, apologetic, pensive, worried, enthusiastic.\u003c/sub\u003e                                                                                                                               | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e1.9 GB\u003c/sub\u003e    | \u003csub\u003eEnglish\u003c/sub\u003e                                                | \u003csub\u003e[An Open Source Emotional Speech Corpus for Human Robot Interaction Applications](https://www.isca-speech.org/archive/Interspeech_2018/pdfs/1349.pdf)\u003c/sub\u003e                                                                                                                                                                                          | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC0 1.0](https://creativecommons.org/publicdomain/zero/1.0/)\u003c/sub\u003e                                                                       |\n| \u003csub\u003e[CaFE](https://zenodo.org/record/1478765)\u003c/sub\u003e                                                                                              | \u003csub\u003e2018\u003c/sub\u003e | \u003csub\u003e6 different sentences by 12 speakers (6 fmelaes + 6 males).\u003c/sub\u003e                                                                                                                                                                                                           | \u003csub\u003e7 emotions: happy, sad, angry, fearful, surprise, disgust and neutral. Each emotion is acted in 2 different intensities.\u003c/sub\u003e                                                                                                                                          | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e2 GB\u003c/sub\u003e      | \u003csub\u003eFrench (Canadian)\u003c/sub\u003e                                      | \u003csub\u003e--\u003c/sub\u003e                                                                                                                                                                                                                                                                                                                                             | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY-NC-SA 4.0](https://creativecommons.org/licenses/by-nc-sa/4.0/)\u003c/sub\u003e                                                               |\n| \u003csub\u003e[EmoFilm](https://zenodo.org/record/1326428)\u003c/sub\u003e                                                                                           | \u003csub\u003e2018\u003c/sub\u003e | \u003csub\u003e1115 audio instances sentences extracted from various films.\u003c/sub\u003e                                                                                                                                                                                                          | \u003csub\u003e5 emotions: anger, contempt, happiness, fear, and sadness.\u003c/sub\u003e                                                                                                                                                                                                        | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e0.277 GB\u003c/sub\u003e  | \u003csub\u003eEnglish, Italian, Spanish\u003c/sub\u003e                              | \u003csub\u003e[Categorical vs Dimensional Perception of Italian Emotional Speech](https://pdfs.semanticscholar.org/e70e/fcf7f5b4c366a7b7e2c16267d7f7691a5391.pdf)\u003c/sub\u003e                                                                                                                                                                                            | \u003csub\u003eRestricted\u003c/sub\u003e     | \u003csub\u003eEULA: End User License Agreement\u003c/sub\u003e                                                                                                    |\n| \u003csub\u003e[ANAD](https://www.kaggle.com/suso172/arabic-natural-audio-dataset)\u003c/sub\u003e                                                                    | \u003csub\u003e2018\u003c/sub\u003e | \u003csub\u003e1384 recording by multiple speakers.\u003c/sub\u003e                                                                                                                                                                                                                                  | \u003csub\u003e3 emotions: angry, happy, surprised.\u003c/sub\u003e                                                                                                                                                                                                                              | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e2 GB\u003c/sub\u003e      | \u003csub\u003eArabic\u003c/sub\u003e                                                 | \u003csub\u003e[Arabic Natural Audio Dataset](https://data.mendeley.com/datasets/xm232yxf7t/1)\u003c/sub\u003e                                                                                                                                                                                                                                                                | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY-NC-SA 4.0](https://creativecommons.org/licenses/by-nc-sa/4.0/)\u003c/sub\u003e                                                               |\n| \u003csub\u003e[EmoSynth](https://zenodo.org/record/3727593)\u003c/sub\u003e                                                                                          | \u003csub\u003e2018\u003c/sub\u003e | \u003csub\u003e144 audio file labelled by 40 listeners.\u003c/sub\u003e                                                                                                                                                                                                                              | \u003csub\u003eEmotion (no speech) defined in regard of valence and arousal.\u003c/sub\u003e                                                                                                                                                                                                     | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e0.1034 GB\u003c/sub\u003e | \u003csub\u003e--\u003c/sub\u003e                                                     | \u003csub\u003e[The Perceived Emotion of Isolated Synthetic Audio: The EmoSynth Dataset and Results](https://dl.acm.org/doi/10.1145/3243274.3243277)\u003c/sub\u003e                                                                                                                                                                                                          | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY 4.0](https://creativecommons.org/licenses/by/4.0/)\u003c/sub\u003e                                                                           |\n| \u003csub\u003e[CMU-MOSEI](https://www.amir-zadeh.com/datasets)\u003c/sub\u003e                                                                                       | \u003csub\u003e2018\u003c/sub\u003e | \u003csub\u003e65 hours of annotated video from more than 1000 speakers and 250 topics.\u003c/sub\u003e                                                                                                                                                                                              | \u003csub\u003e6 Emotion (happiness, sadness, anger,fear, disgust, surprise) + Likert scale.\u003c/sub\u003e                                                                                                                                                                                     | \u003csub\u003eAudio, Video\u003c/sub\u003e       | \u003csub\u003e190.1 GB\u003c/sub\u003e  | \u003csub\u003eEnglish\u003c/sub\u003e                                                | \u003csub\u003e[Multi-attention Recurrent Network for Human Communication Comprehension](https://arxiv.org/pdf/1802.00923.pdf)\u003c/sub\u003e                                                                                                                                                                                                                                | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CMU-MOSEI License](https://github.com/A2Zadeh/CMU-MultimodalSDK/blob/master/LICENSE.txt)\u003c/sub\u003e                                           |\n| \u003csub\u003e[VERBO](https://sites.google.com/view/verbodatabase/home)\u003c/sub\u003e                                                                              | \u003csub\u003e2018\u003c/sub\u003e | \u003csub\u003e14 different phrases by 12 speakers (6 female + 6 male) for a total of 1167 recordings.\u003c/sub\u003e                                                                                                                                                                               | \u003csub\u003e7 emotions: Happiness, Disgust, Fear, Neutral, Anger, Surprise, Sadness\u003c/sub\u003e                                                                                                                                                                                           | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e--\u003c/sub\u003e        | \u003csub\u003ePortuguese\u003c/sub\u003e                                             | \u003csub\u003e[VERBO: Voice Emotion Recognition dataBase in Portuguese Language](https://thescipub.com/pdf/jcssp.2018.1420.1430.pdf)\u003c/sub\u003e                                                                                                                                                                                                                         | \u003csub\u003eRestricted\u003c/sub\u003e     | \u003csub\u003eAvailable for research purposes only\u003c/sub\u003e                                                                                                |\n| \u003csub\u003e[CMU-MOSI](https://www.amir-zadeh.com/datasets)\u003c/sub\u003e                                                                                        | \u003csub\u003e2017\u003c/sub\u003e | \u003csub\u003e2199 opinion utterances with annotated sentiment.\u003c/sub\u003e                                                                                                                                                                                                                     | \u003csub\u003eSentiment annotated between very negative to very positive in seven Likert steps.\u003c/sub\u003e                                                                                                                                                                                 | \u003csub\u003eAudio, Video\u003c/sub\u003e       | \u003csub\u003e4.3 GB\u003c/sub\u003e    | \u003csub\u003eEnglish\u003c/sub\u003e                                                | \u003csub\u003e[Multi-attention Recurrent Network for Human Communication Comprehension](https://arxiv.org/pdf/1802.00923.pdf)\u003c/sub\u003e                                                                                                                                                                                                                                | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CMU-MOSI License](https://github.com/A2Zadeh/CMU-MultimodalSDK/blob/master/LICENSE.txt)\u003c/sub\u003e                                            |\n| \u003csub\u003e[MSP-IMPROV](https://ecs.utdallas.edu/research/researchlabs/msp-lab/MSP-Improv.html)\u003c/sub\u003e                                                   | \u003csub\u003e2017\u003c/sub\u003e | \u003csub\u003e20 sentences by 12 actors.\u003c/sub\u003e                                                                                                                                                                                                                                            | \u003csub\u003e4 emotions: angry, sad, happy, neutral, other, without agreement\u003c/sub\u003e                                                                                                                                                                                                  | \u003csub\u003eAudio, Video\u003c/sub\u003e       | \u003csub\u003e3.4 GB\u003c/sub\u003e    | \u003csub\u003eEnglish\u003c/sub\u003e                                                | \u003csub\u003e[MSP-IMPROV: An Acted Corpus of Dyadic Interactions to Study Emotion Perception](https://ecs.utdallas.edu/research/researchlabs/msp-lab/publications/Busso_2017.pdf)\u003c/sub\u003e                                                                                                                                                                           | \u003csub\u003eRestricted\u003c/sub\u003e     | \u003csub\u003eAcademic License \u0026 Commercial License\u003c/sub\u003e                                                                                               |\n| \u003csub\u003e[CREMA-D](https://github.com/CheyneyComputerScience/CREMA-D)\u003c/sub\u003e                                                                           | \u003csub\u003e2017\u003c/sub\u003e | \u003csub\u003e7442 clip of 12 sentences spoken by 91 actors (48 males and 43 females).\u003c/sub\u003e                                                                                                                                                                                              | \u003csub\u003e6 emotions: angry, disgusted, fearful, happy, neutral, and sad\u003c/sub\u003e                                                                                                                                                                                                    | \u003csub\u003eAudio, Video\u003c/sub\u003e       | \u003csub\u003e0.607 GB\u003c/sub\u003e  | \u003csub\u003eEnglish\u003c/sub\u003e                                                | \u003csub\u003e[CREMA-D: Crowd-sourced Emotional Multimodal Actors Dataset](https://www.ncbi.nlm.nih.gov/pmc/articles/PMC4313618/)\u003c/sub\u003e                                                                                                                                                                                                                            | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[Open Database License \u0026 Database Content License](https://github.com/CheyneyComputerScience/CREMA-D/blob/master/LICENSE.txt)\u003c/sub\u003e       |\n| \u003csub\u003e[Example emotion videos used in investigation of emotion perception in schizophrenia](https://espace.library.uq.edu.au/view/UQ:446541)\u003c/sub\u003e | \u003csub\u003e2017\u003c/sub\u003e | \u003csub\u003e6 videos:Two example videos from each emotion category (angry, happy and neutral) by one female speaker.\u003c/sub\u003e                                                                                                                                                              | \u003csub\u003e3 emotions: angry, happy and neutral.\u003c/sub\u003e                                                                                                                                                                                                                             | \u003csub\u003eAudio, Video\u003c/sub\u003e       | \u003csub\u003e0.063 GB\u003c/sub\u003e  | \u003csub\u003eEnglish\u003c/sub\u003e                                                | \u003csub\u003e--\u003c/sub\u003e                                                                                                                                                                                                                                                                                                                                             | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[Permitted Non-commercial Re-use with Acknowledgment](https://guides.library.uq.edu.au/deposit_your_data/terms_and_conditions)\u003c/sub\u003e      |\n| \u003csub\u003e[EMOVO](http://voice.fub.it/activities/corpora/emovo/index.html)\u003c/sub\u003e                                                                       | \u003csub\u003e2014\u003c/sub\u003e | \u003csub\u003e6 actors  who  played  14  sentences.\u003c/sub\u003e                                                                                                                                                                                                                                 | \u003csub\u003e6 emotions: disgust, fear, anger, joy, surprise, sadness.\u003c/sub\u003e                                                                                                                                                                                                         | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e0.355 GB\u003c/sub\u003e  | \u003csub\u003eItalian\u003c/sub\u003e                                                | \u003csub\u003e[EMOVO Corpus: an Italian Emotional Speech Database](https://core.ac.uk/download/pdf/53857389.pdf)\u003c/sub\u003e                                                                                                                                                                                                                                             | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e--\u003c/sub\u003e                                                                                                                                  |\n| \u003csub\u003e[RECOLA](https://diuf.unifr.ch/main/diva/recola/download.html)\u003c/sub\u003e                                                                         | \u003csub\u003e2013\u003c/sub\u003e | \u003csub\u003e3.8 hours of recordings by 46 participants.\u003c/sub\u003e                                                                                                                                                                                                                           | \u003csub\u003enegative and positive sentiment (valence and arousal).\u003c/sub\u003e                                                                                                                                                                                                            | \u003csub\u003eAudio, Video\u003c/sub\u003e       | \u003csub\u003e--\u003c/sub\u003e        | \u003csub\u003e--\u003c/sub\u003e                                                     | \u003csub\u003e[Introducing the RECOLA Multimodal Corpus of Remote Collaborative and Affective Interactions](https://drive.google.com/file/d/0B2V_I9XKBODhNENKUnZWNFdVXzQ/view)\u003c/sub\u003e                                                                                                                                                                               | \u003csub\u003eRestricted\u003c/sub\u003e     | \u003csub\u003eAcademic License \u0026 Commercial License\u003c/sub\u003e                                                                                               |\n| \u003csub\u003e[GEMEP corpus](https://www.unige.ch/cisa/gemep)\u003c/sub\u003e                                                                                        | \u003csub\u003e2012\u003c/sub\u003e | \u003csub\u003eVideos10 actors portraying 10 states.\u003c/sub\u003e                                                                                                                                                                                                                                 | \u003csub\u003e12 emotions: amusement, anxiety, cold anger (irritation), despair, hot anger (rage),  fear (panic), interest, joy (elation), pleasure(sensory), pride, relief, and sadness. Plus, 5 additional emotions: admiration, contempt, disgust, surprise, and tenderness.\u003c/sub\u003e | \u003csub\u003eAudio, Video\u003c/sub\u003e       | \u003csub\u003e--\u003c/sub\u003e        | \u003csub\u003eFrench\u003c/sub\u003e                                                 | \u003csub\u003e[Introducing the Geneva Multimodal Expression Corpus for Experimental Research on Emotion Perception](https://www.researchgate.net/publication/51796867_Introducing_the_Geneva_Multimodal_Expression_Corpus_for_Experimental_Research_on_Emotion_Perception)\u003c/sub\u003e                                                                                   | \u003csub\u003eRestricted\u003c/sub\u003e     | \u003csub\u003e--\u003c/sub\u003e                                                                                                                                  |\n| \u003csub\u003e[OGVC](https://sites.google.com/site/ogcorpus/home/en)\u003c/sub\u003e                                                                                 | \u003csub\u003e2012\u003c/sub\u003e | \u003csub\u003e9114 spontaneous utterances and 2656 acted utterances by 4 professional actors (two male and two female).\u003c/sub\u003e                                                                                                                                                             | \u003csub\u003e9 emotional states: fear, surprise, sadness, disgust, anger, anticipation, joy, acceptance and the neutral state.\u003c/sub\u003e                                                                                                                                                 | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e5.3 GB\u003c/sub\u003e    | \u003csub\u003eJapanese\u003c/sub\u003e                                               | \u003csub\u003e[Naturalistic emotional speech collectionparadigm with online game and its psychological and acoustical assessment](https://www.jstage.jst.go.jp/article/ast/33/6/33_E1175/_pdf)\u003c/sub\u003e                                                                                                                                                               | \u003csub\u003eRestricted\u003c/sub\u003e     | \u003csub\u003e--\u003c/sub\u003e                                                                                                                                  |\n| \u003csub\u003e[LEGO corpus](https://www.ultes.eu/ressources/lego-spoken-dialogue-corpus/)\u003c/sub\u003e                                                            | \u003csub\u003e2012\u003c/sub\u003e | \u003csub\u003e347 dialogs with 9,083 system-user exchanges.\u003c/sub\u003e                                                                                                                                                                                                                         | \u003csub\u003eEmotions classified as garbage, non-angry, slightly angry and very angry.\u003c/sub\u003e                                                                                                                                                                                         | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e1.1 GB\u003c/sub\u003e    | \u003csub\u003e--\u003c/sub\u003e                                                     | \u003csub\u003e[A Parameterized and Annotated Spoken Dialog Corpus of the CMU Let’s Go Bus Information System](http://www.lrec-conf.org/proceedings/lrec2012/pdf/333_Paper.pdf)\u003c/sub\u003e                                                                                                                                                                               | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003eLicense available with the data. Free of charges for research purposes only.\u003c/sub\u003e                                                        |\n| \u003csub\u003e[SEMAINE](https://semaine-db.eu/)\u003c/sub\u003e                                                                                                      | \u003csub\u003e2012\u003c/sub\u003e | \u003csub\u003e95 dyadic conversations from 21 subjects. Each subject converses with another playing one of four characters with emotions.\u003c/sub\u003e                                                                                                                                           | \u003csub\u003e5 FeelTrace annotations: activation, valence, dominance, power, intensity\u003c/sub\u003e                                                                                                                                                                                         | \u003csub\u003eAudio, Video, Text\u003c/sub\u003e | \u003csub\u003e104 GB\u003c/sub\u003e    | \u003csub\u003eEnglish\u003c/sub\u003e                                                | \u003csub\u003e[The SEMAINE Database: Annotated Multimodal Records of Emotionally Colored Conversations between a Person and a Limited Agent](https://ieeexplore.ieee.org/document/5959155)\u003c/sub\u003e                                                                                                                                                                   | \u003csub\u003eRestricted\u003c/sub\u003e     | \u003csub\u003eAcademic EULA\u003c/sub\u003e                                                                                                                       |\n| \u003csub\u003e[SAVEE](http://kahlan.eps.surrey.ac.uk/savee/Database.html)\u003c/sub\u003e                                                                            | \u003csub\u003e2011\u003c/sub\u003e | \u003csub\u003e480 British English utterances by 4 males actors.\u003c/sub\u003e                                                                                                                                                                                                                     | \u003csub\u003e7 emotions: anger, disgust, fear, happiness, sadness, surprise and neutral.\u003c/sub\u003e                                                                                                                                                                                       | \u003csub\u003eAudio, Video\u003c/sub\u003e       | \u003csub\u003e--\u003c/sub\u003e        | \u003csub\u003eEnglish (British)\u003c/sub\u003e                                      | \u003csub\u003e[Multimodal Emotion Recognition](http://personal.ee.surrey.ac.uk/Personal/P.Jackson/pub/ma10/HaqJackson_MachineAudition10_approved.pdf)\u003c/sub\u003e                                                                                                                                                                                                        | \u003csub\u003eRestricted\u003c/sub\u003e     | \u003csub\u003eFree of charges for research purposes only.\u003c/sub\u003e                                                                                         |\n| \u003csub\u003e[TESS](https://tspace.library.utoronto.ca/handle/1807/24487)\u003c/sub\u003e                                                                           | \u003csub\u003e2010\u003c/sub\u003e | \u003csub\u003e2800 recording by 2 actresses.\u003c/sub\u003e                                                                                                                                                                                                                                        | \u003csub\u003e7 emotions: anger, disgust, fear, happiness, pleasant surprise, sadness, and neutral.\u003c/sub\u003e                                                                                                                                                                             | \u003csub\u003eAudio\u003c/sub\u003e              | \u003csub\u003e--\u003c/sub\u003e        | \u003csub\u003eEnglish\u003c/sub\u003e                                                | \u003csub\u003e[BEHAVIOURAL FINDINGS FROM THE TORONTO EMOTIONAL SPEECH SET](https://www.semanticscholar.org/paper/BEHAVIOURAL-FINDINGS-FROM-THE-TORONTO-EMOTIONAL-SET-Dupuis-Pichora-Fuller/d7f746b3aee801a353b6929a65d9a34a68e71c6f/figure/2)\u003c/sub\u003e                                                                                                                | \u003csub\u003eOpen\u003c/sub\u003e           | \u003csub\u003e[CC BY-NC-ND 4.0](https://creativecommons.org/licenses/by-nc-nd/4.0/)\u003c/sub\u003e                                                               |\n| \u003csub\u003e[EEKK](https://metashare.ut.ee/repository/download/4d42d7a8463411e2a6e4005056b40024a19021a316b54b7fb707757d43d1a889/)\u003c/sub\u003e                  | \u003csub\u003e2007\u003c/sub\u003e | \u003csub\u003e26 text passage read by 10 speakers.\u003c/sub\u003e                                                                                                                                                                                                                                  | \u003csub\u003e4 main emotions: joy, sadness, anger and n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FSuperKogito%2FSER-datasets","html_url":"https://awesome.ecosyste.ms/projects/github.com%2FSuperKogito%2FSER-datasets","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FSuperKogito%2FSER-datasets/lists"}