{"id":18668346,"url":"https://github.com/andythefactory/romanian-nlp-datasets","last_synced_at":"2026-02-07T08:31:58.800Z","repository":{"id":168790820,"uuid":"644589555","full_name":"AndyTheFactory/romanian-nlp-datasets","owner":"AndyTheFactory","description":"A list of Romanian NLP Datasets","archived":false,"fork":false,"pushed_at":"2025-12-02T21:20:55.000Z","size":56,"stargazers_count":53,"open_issues_count":0,"forks_count":10,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-12-05T21:41:34.763Z","etag":null,"topics":["nlp","nlp-data","nlp-dataset","nlp-datasets","nlp-resources","romanian","romanian-language"],"latest_commit_sha":null,"homepage":"","language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"cc0-1.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/AndyTheFactory.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2023-05-23T20:54:23.000Z","updated_at":"2025-12-02T21:20:59.000Z","dependencies_parsed_at":"2024-11-07T08:42:40.486Z","dependency_job_id":"80d801a4-3ad4-4235-966a-fabab308f0c7","html_url":"https://github.com/AndyTheFactory/romanian-nlp-datasets","commit_stats":null,"previous_names":["andythefactory/romanian-nlp-datasets"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/AndyTheFactory/romanian-nlp-datasets","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/AndyTheFactory%2Fromanian-nlp-datasets","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/AndyTheFactory%2Fromanian-nlp-datasets/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/AndyTheFactory%2Fromanian-nlp-datasets/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/AndyTheFactory%2Fromanian-nlp-datasets/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/AndyTheFactory","download_url":"https://codeload.github.com/AndyTheFactory/romanian-nlp-datasets/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/AndyTheFactory%2Fromanian-nlp-datasets/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":29190190,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-02-07T07:37:03.739Z","status":"ssl_error","status_checked_at":"2026-02-07T07:37:03.029Z","response_time":63,"last_error":"SSL_connect returned=1 errno=0 peeraddr=140.82.121.6:443 state=error: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["nlp","nlp-data","nlp-dataset","nlp-datasets","nlp-resources","romanian","romanian-language"],"created_at":"2024-11-07T08:42:19.198Z","updated_at":"2026-02-07T08:31:58.794Z","avatar_url":"https://github.com/AndyTheFactory.png","language":null,"funding_links":[],"categories":[],"sub_categories":[],"readme":"[![Awesome](https://awesome.re/badge-flat2.svg)](https://awesome.re)\n\n# A list of Romanian NLP Datasets\nA curated list of open source and open access Romanian Language NLP Datasets.\nFor the moment we don't add parallel copora to the list.\n\nFor additions or any other changes please submit a pull request.\n\nTable of contents\n=================\n\n\u003c!--ts--\u003e\n   * [Unlabeled text Corpora](#unlabeled-text-corpora)\n   * [Semantic Textual Similarity / Paraphrasing](#semantic-textual-similarity--paraphrasing)\n   * [Natural Language Inference](#natural-language-inference)\n   * [Summarization](#summarization)\n   * [Dialect and regional speech identification](#dialect-and-regional-speech-identification)\n   * [Named Entity Recognition (NER)](#named-entity-recognition-ner)\n   * [Autorship Attribution](#autorship-attribution)\n   * [Sentiment Analysis](#sentiment-analysis)\n   * [Dependency Parsing](#dependency-parsing)\n   * [Diacritics Restoration / Grammar Correction](#diacritics-restoration--grammar-correction)\n   * [Fake News / Clickbait / Satirical News](#fake-news--clickbait--satirical-news)\n   * [Offensive Language](#offensive-language)\n   * [Questions and Answering](#questions-and-answers)\n   * [Spelling, Dictionaries and Gramatical Errors](#spelling-and-gramatical-errors)\n   * [Automatic Speech Recognition (ASR)](#automatic-speech-recognition)\n\n\u003c!--te--\u003e\n\n## Unlabeled text Corpora\n\n\n* [❄️FuLG dataset ❄️](https://huggingface.co/datasets/faur-ai/fulg)\n\u003e     The FuLG dataset is a comprehensive Romanian language corpus comprising\n\u003e     150 billion tokens, carefully extracted from Common Crawl. \n[![arXiv](https://img.shields.io/badge/arXiv-2004.06165-f9f107.svg)](https://arxiv.org/abs/2407.13657)\n\n* [🌐 Oscar Common Crawl dataset 🌐](https://huggingface.co/datasets/oscar-corpus/OSCAR-2201)\n\u003e     Part of a large multilanguage corpus originated from Common Crawl.\n\u003e     It's a raw, unannotated corpus. It has roughly 50 GB of Romanian text\n\u003e     in 4.5 million documnets. For details check its homepage \n\u003e     and the paper\n \n[![arXiv](https://img.shields.io/badge/arXiv-2004.06165-f9f107.svg)](https://arxiv.org/abs/2004.06165)\n[![Homepage](https://img.shields.io/badge/oscar%20homepage-6ca1f0)](https://oscar-project.org/)\n\n* [📚 CC-100 📚](https://data.statmt.org/cc-100/)\n\u003e      Similar to Oscar, part of a multilanguage corpus also based on Common Crawl\n\u003e      from 2018. Romanian text is 16GB large\n\n[![arXiv](https://img.shields.io/badge/arXiv-1911.00359-f9f107.svg)](https://arxiv.org/abs/1911.00359)\n[![Homepage](https://img.shields.io/badge/cc-100%20homepage-6ca1f0)](https://data.statmt.org/cc-100/)\n\n* [🌍 Wikipedia Corpus 🌍](https://dumps.wikimedia.org/rowiki/)\n\u003e       Romanian language wikipedia dump. \n  \n* [📰⚖️ RoTex Collection 📰⚖️](https://github.com/aleris/ReadME-RoTex-Corpus-Builder)\n\u003e       A collection of varoius unannotated corpora collected around 2018-2019.\n\u003e       Includes books, scraped newspapers and juridical documents  \n\u003e \n* [📖 Romanian Language Repository 📖](https://github.com/lmidriganciochina/romaniancorpus)\n\u003e       A collection of written and spoken text from various\n\u003e       sources: Articles, Fairy tales, Fiction, History, Theatre, News\n \n* [🏛️ MARCELL Legislative Corpus 🏛️](https://elrc-share.eu/repository/browse/marcell-romanian-legislative-subcorpus-v2/2da548428b9d11eb9c1a00155d026706ce94a6b59ffc4b0e9fb5cd9cebe6889e/)\n\u003e      Romanian national legilation from  1881 to 2021. The corpus\n\u003e      includes mainly: governmental decisions, ministerial orders,\n\u003e      decisions, decrees and laws.\n\u003e      Automatically annotated for Named Entities\n\n[![ACL](https://img.shields.io/badge/ACL%20Anthology-ed1c24.svg)](http://www.lrec-conf.org/proceedings/lrec2020/pdf/2020.lrec-1.464.pdf)\n[![Homepage](https://img.shields.io/badge/marcell%20homepage-6ca1f0)](https://marcell-project.eu/)\n\n* [🦠 COVID-19 Tweets 🐦](https://github.com/UBC-NLP/megacov)\n\u003e    Mega-COV is a billion-scale dataset from Twitter for studying COVID-19. It is\n\u003e    available in over 100+ languages, Romanian being one of them. Tweets need\n\u003e    to be rehydrated\n\n[![arXiv](https://img.shields.io/badge/arXiv-2005.06012-f9f107.svg)](https://arxiv.org/abs/2005.06012)\n[![Medium](https://img.shields.io/badge/Medium-12100E?style=for-the-badge\u0026logo=medium\u0026logoColor=white)](https://mumageed.medium.com/billion-scale-investigation-of-covid-19-impact-on-human-communication-in-104-languages-874b5a37beac)\n\n* [COVIDSentiRO](https://github.com/Alegzandra/KES-2023/tree/main/datasets/COVIDSentiRO)\n\u003e    A corpus of Romanian tweets related to COVID and vaccination against COVID, created and\n\u003e    collected between January 2021 and February 2022. It contains 19319 tweets.\n\n* [📜 Minutes of the Sittings of the Chamber of Deputies of Romania 📜](https://elrc-share.eu/repository/browse/monolingual-corpus-from-minutes-of-the-sittings-of-the-chamber-of-deputies-of-romania-2016-2018-processed/759806e22e1311e9a4d400155d02670657928f90efa64c9dab1b177c2186bf6c/)\n\u003e     Minutes of the Sittings of the Chamber of Deputies of Romania (2016-2018)\n\u003e     Unannotated corpus\n  \n* [🔊 Minutes of the Sittings of the Romanian Parliament 🔊](https://elrc-share.eu/repository/browse/romanian-parliament-transcripts-1996-2018-processed/779b85aee4de11e9913100155d026706ac0a5e38c6824010a537363b37b6bd0f/)\n\u003e     contains 500k+ instances of speech from the parliament podium from\n\u003e     1996 to 2018. Sentence splitting and deduplication onm sentence level\n\u003e     have been applied as processing steps\n\u003e     Unannotated corpus\n  \n* [🗣️ Romanian Presidential Discourses 🗣️](https://github.com/grrrrah/RomanianPresidentialDiscourses)\n\u003e     Romanian presidential discouses (1990-2020) split in 4 files\n\u003e     one for each president. Unannotated corpus\n  \n  \n* [🎭 Culture Domain Corpus 🎭](https://elrc-share.eu/repository/browse/monolingual-romanian-corpus-in-the-culture-domain-processed/a1d6c98e1d5911e9b7d400155d026706fd71a90bf1df4bacb3f174edccb6e9b9/)\n\u003e     Monolingual Romanian corpus, including content from public websites related to culture\n\n* [⚖️ Law Domain Corpus ⚖️](https://elrc-share.eu/repository/browse/monolingual-romanian-corpus-in-the-law-domain/ee9f6b0289f611e6bfe700155d0205029e7b188412ab4e56bf6fd1d7d9e8b033/)\n\u003e    Monolingual (ron) corpus, containing 38063991 tokens and 854096 lexical types in the law domain.\n\n* [🏢 Public Administration Domain Corpus 🏢](https://elrc-share.eu/repository/browse/monolingual-romanian-corpus-in-the-public-administration-domain-processed/32b0c234327311e8b7d400155d026706afa147904c554dd1bad4764bd4a7aaed/)\n\u003e    Monolingual Romanian corpus, containing 360833 sentences (9064764 words) in the public administration domain.\n  \n* [📋 New Civil Procedure Code 📋](https://elrc-share.eu/repository/browse/romanian-new-civil-procedure-code-processed/e4d8e13046ff11e8b7d400155d026706010657f248274e6286aeb6488d8a2ee6/)\n\u003e    The New Civil Procedure Code in Romanian (monolingual) comprising 297888 words.\n\n* [⚖️ New Criminal Code ⚖️](https://elrc-share.eu/repository/browse/noul-cod-penal/100b7e38d25111ea913100155d0267066e48d760d0d24b39a4900b2b09864a02/)\n\u003e     The Romanian updated criminal code: text with law content.\n\n* [📰 Romanian News Articles Dataset 📰](https://github.com/mhakan20/RomanianNewsArticlesDataset)\n\u003e   news articles dataset from romanian newssites\n\u003e   title, summary and article\n\n* [📰 Old Newspapers 📰](https://www.kaggle.com/datasets/alvations/old-newspapers)\n\u003e   multi-language corpus from online available news sources.\n\u003e   It contains also 43mil words in Romanian language from Twitter, Blogs and Newspapers\n\n[![Homepage](https://img.shields.io/badge/hc%20corpora%20homepage-6ca1f0)](http://corpora.epizy.com/corpora.html?i=1)\n\n* [📚 ELTeC-Rom 📚](https://github.com/COST-ELTeC/ELTeC-rom)\n\u003e   The Romanian novel collection for ELTeC, the European Literary Text Collection\n\u003e   Sources: Biblioteca Metropolitana din Bucuresti, Biblioteca Universitara \"Mihai Eminescu\" din Iasi,\n\u003e   Biblioteca Judeteana din Botosani, personal micro-collections uploaded on Zenodo\n\u003e   under the following labels: \"Hajduks Library\"; \"RomanianNovel Library\"; \"CityMysteries Library\"; \"BibliotecaDHL_Iasi\" \n  \n* [📧 RO Business Emails 📧](https://huggingface.co/datasets/readerbench/ro-business-emails)\n\u003e   Public dataset of 1447 manually annotated Romanian business-oriented emails.\n\u003e   The corpus is annotated with 5 token-related labels, as well as 5 sequence-related classes\n\n[![MDPI](https://img.shields.io/badge/MDPI-Information-00813E.svg)](https://www.mdpi.com/2078-2489/14/6/321)\n\n\n* [📖RO-Stories📖](https://huggingface.co/datasets/readerbench/ro-stories)\n\u003e  The corpus consists of texts written by Romanian authors between 19th century and present, representing stories, short-stories, fairy tales and sketches.\n\u003e  The current version contains 19 authors, 1263 full texts and 12516 paragraphs of around 200 words each, preserving paragraphs integrity.\n\n* [📕ROST📕](https://www.kaggle.com/datasets/sandamariaavram/rost-romanian-stories-and-other-texts)\n\u003e  A dataset containing 400 Romanian texts written by 10 authors\n\u003e  The dataset contains stories, short stories, fairy tales, novels, articles, and sketches written by Ion Creangă,\n\u003e  Barbu Ştefănescu Delavrancea, Mihai Eminescu, Nicolae Filimon, Emil Gârleanu, Petre Ispirescu, Mihai Oltean, Emilia Plugaru, Liviu Rebreanu, Ioan Slavici.\n\n[![MDPI](https://img.shields.io/badge/MDPI-Information-00813E.svg)](https://www.mdpi.com/2227-7390/10/23/4589)\n\n* [🍳Romanian Cooking Recipes🍳](https://huggingface.co/datasets/BlackKakapo/recipes-ro)\n\u003e  891 Cooking Recipes in Romanian Language\n\n## Semantic Textual Similarity / Paraphrasing\n\n* [🔗 RO-STS 🔗](https://huggingface.co/datasets/ro_sts)\n\u003e   Semantic Textual Similarity dataset for the Romanian language\n\u003e   RO-STS contains 8,628 sentence pairs with their similarity scores\n\n[![NeurIPS](https://img.shields.io/badge/NeurIPS-ed1c24.svg)](https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/5f93f983524def3dca464469d2cf9f3e-Abstract-round1.html)\n\n* [📖 Romanian Bible Paraphrase Corpus 📖](https://huggingface.co/datasets/andyP/ro-paraphrase-bible)\n\u003e   A paraphrase corpus created from 10 different Romanian language Bible versions.\n\u003e   The final dataset contains 904,815 similar records and 218,977 non matching records, totaling 1,123,927\n\n\n* [🔄 Romanian paraphrase dataset 🔄](https://huggingface.co/datasets/BlackKakapo/paraphrase-ro)\n\u003e  Around ~100k examples of paraphrases. No clear explanation on how the dataset was built\n\n* [🌐 TaPaCo 🌐](https://huggingface.co/datasets/tapaco/viewer/ro/train)\n\u003e    A multi-language paraphrase corpus for 73 languages extracted from the Tatoeba database.\n\u003e    It has ~ 2000 romanian phrases totaling 941 paraphrase groups. \n\n[![ACL](https://img.shields.io/badge/ACL%20Anthology-ed1c24.svg)](https://aclanthology.org/2020.lrec-1.848/)\n[![Homepage](https://img.shields.io/badge/cc-100%20homepage-6ca1f0)](https://zenodo.org/record/3707949)\n\n## Natural Language Inference\n\n* [🧠 RONLI 🧠](https://github.com/Eduard6421/RONLI)\n\u003e   We introduce the first Romanian NLI corpus (RoNLI) comprising 58K training sentence pairs, which are obtained via distant supervision, and 6K validation and test sentence pairs, which are manually annotated with the correct labels.\n[![ACL](https://img.shields.io/badge/ACL%20Anthology-ed1c24.svg)](https://aclanthology.org/2024.acl-long.15/)\n\n\n\n* [~RO-NLI~](https://github.com/dumitrescustefan/RO-NLI)\n\u003e  The repository seems to be just an attempt at starting to build the dataset\n\n\n## Summarization\n* [📝 RO Text Summarization 📝](https://huggingface.co/datasets/readerbench/ro-text-summarization)\n\u003e   Around ~72k Full texts and their summary. Source seems to be news websites.\n\u003e   No description or explanation available\n\n## Dialect and regional speech identification\n* [🗣️ RoDia 🗣️](https://github.com/codrut2/RoDia)\n\u003e varied compilation of speech samples from five distinct regions of Romania, covering both urban and rural environments.\n\u003e Around 2800 records labeled with age, gender and type of dialect\n\n[![arXiv](https://img.shields.io/badge/arXiv-2309.03378-f9f107.svg)](https://arxiv.org/abs/2309.03378)\n\n* [🌍 MOROCO 🌍](https://github.com/butnaruandrei/MOROCO)\n\u003e MOROCO: The Moldavian and Romanian Dialectal Corpus\n\u003e The MOROCO data set contains Moldavian and Romanian samples of text collected from the news domain.\n\u003e The samples belong to one of the following six topics: culture, finance, politics, science, sports, tech\n\u003e totaling over 32.000 labeled records\n\n[![arXiv](https://img.shields.io/badge/arXiv-1901.06543-f9f107.svg)](https://arxiv.org/abs/1901.06543)\n\n  \n## Named Entity Recognition (NER)\n\n* [⚖️ LegalNERo ⚖️](https://huggingface.co/datasets/joelito/legalnero)\n* [🏷️ RONEC 🏷️](https://huggingface.co/datasets/ronec)\n* [🌐 WikiAnn 🌐](https://huggingface.co/datasets/wikiann)\n* [🏛️ SiMoNERo 🏛️](https://github.com/UniversalDependencies/UD_Romanian-SiMoNERo)\n* [📜 HistNERo 📜]()\n\u003e The dataset contains 323k tokens of text, covering more than half of the 19th century\n\u003e (i.e., 1817) until the late part of the 20th century (i.e., 1990).\n\u003e The samples belong to one of the following four historical regions of Romania,\n\u003e namely Bessarabia, Moldavia, Transylvania, and Wallachia.\n\n[![arXiv](https://img.shields.io/badge/arXiv-2405.00155-f9f107.svg)](https://arxiv.org/abs/2405.00155v1)\n\n\n\n## Autorship Attribution\n\n* [📚 ROST 📚](https://www.kaggle.com/datasets/sandamariaavram/rost-romanian-stories-and-other-texts)\n\n\n## Sentiment Analysis\n\n* [😊 RO_Sent 😊](https://huggingface.co/datasets/ro_sent)\n* [📖 Senti_Lex 📖](https://huggingface.co/datasets/senti_lex)\n* [🎭 LaROSeDa 🎭](https://huggingface.co/datasets/laroseda)\n* [🎬 SART 🎬](https://github.com/Alegzandra/KES-2023/tree/main/datasets/SART)\n* [❤️ RED ❤️](https://github.com/Alegzandra/RED-Romanian-Emotions-Dataset)\n* [🌐 Romanian Categorized Web Dataset 🌐](https://github.com/bogsio/RomanianCategorizedWebDataset)\n* [🎬 Romanian Sentiment Movie Reviews 🎬](https://www.kaggle.com/datasets/gringoandy/romanian-sentiment-movie-reviews)\n\n\n## Dependency Parsing\n\n* [🌳 CoNLL 2017 \u0026 2018 🌳](https://www.conll.org/previous-tasks)\n* [🔗 Deep Universal Dependencies 🔗](https://lindat.mff.cuni.cz/repository/xmlui/handle/11234/1-3720)\n* [📚 Curlicat Romanian Corpus 📚](https://elrc-share.eu/repository/browse/curlicat-romanian-corpus/8b6c8dca58ea11ed9c1a00155d026706fb03ef8b4c1847cfbe9cea869a82731e/)\n* [🌲 HamleDT 🌲](https://lindat.mff.cuni.cz/repository/xmlui/handle/11234/1-1508)\n* [📖 RoWordNet 📖](https://github.com/dumitrescustefan/RoWordNet)\n* [🌳 RoRefTrees 🌳](https://github.com/UniversalDependencies/UD_Romanian-RRT)\n\n## Diacritics Restoration / Grammar Correction\n\n* [✍️ Corpus for training and evaluating diacritics restoration systems ✍️](https://lindat.mff.cuni.cz/repository/xmlui/handle/11234/1-2607)\n* [✏️ RONACC ✏️](https://nextcloud.readerbench.com/index.php/s/9pwymesT5sycxoM)\n\n## Fake News / Clickbait / Satirical News\n\n* [❌ Fakerom ❌](https://www.tagtog.com/fakerom/fakerom)\n* [🎣 Clickbait dataset on Romanian SciTech News 🎣](https://github.com/ralucaginga/ClickbaitSciTechRO)\n* [😂 SaRoCo 😂](https://github.com/MihaelaGaman/SaRoCo)\n\n## Offensive Language\n\n* [🚫 RO-Offense 🚫](https://huggingface.co/datasets/readerbench/ro-offense)\n  \n* [📰 News RO-Offense 📰](https://huggingface.co/datasets/readerbench/news-ro-offense)\n\u003e manually annotated 4,052 comments on a Romanian local news website\n\u003e into one of the following classes: non-offensive, targeted insults,\n\u003e racist, homophobic, and sexist.\n\n[![arXiv](https://img.shields.io/badge/UTCluj-RoCHI-ac2820.svg)](http://rochi.utcluj.ro/articole/10/RoCHI2022-Cojocaru-A.pdf)\n\n* [📱 FB RO-Offense 📱](https://huggingface.co/datasets/readerbench/ro-fb-offense)\n\u003e 4455 organic generated comments from Facebook live broadcasts\n\u003e annotated not binary offensive language detection tasks and for fine-grained offensive language detection\n\n[![IEEE](https://img.shields.io/badge/IEEE-xplore-14303e.svg)](https://ieeexplore.ieee.org/document/10130824)\n\n* [🔍 RO-Offense-Sequences 🔍](https://huggingface.co/datasets/readerbench/ro-offense-sequences)\n\u003e 4800 Romanian comments annotated with offensive text spans\n\u003e Offensive span detection\n\n[![MDPI](https://img.shields.io/badge/MDPI-Information-00813e.svg)](https://www.mdpi.com/2078-2489/15/1/8)\n\n* [💢 Hate Speech RO 💢](https://github.com/andra-pumnea/hate-speech-ro)\n\u003e 3860 labeled hate speech records\n  \n* [🐦 ROFF 🐦](https://github.com/guzimanis/ROFF)\n\u003e Dataset consists of 5000 tweets, from which 924 were labeled as offensive (18.48 %)\n\u003e and 4076 tweets as non-offensive.\n\n[![ACL](https://img.shields.io/badge/ACL%20Anthology-ed1c24.svg)](https://aclanthology.org/2021.ranlp-1.102.pdf)\n  \n* [⚠️ CoRoSeOf ⚠️](https://github.com/DianaHoefels/CoRoSeOf)\n\u003e The corpus contains 39 245 tweets, annotated by multiple annotators, following the sexist label set of a recent study.\n\n[![ACL](https://img.shields.io/badge/ACL%20Anthology-ed1c24.svg)](https://aclanthology.org/2022.lrec-1.243.pdf)\n\n## Questions and Answers\n\n* [🧮 GSM8K RO 🧮](https://huggingface.co/datasets/BlackKakapo/gsm8k-ro)\n\u003e   This dataset is just the translation of the [gsm8k](https://huggingface.co/datasets/gsm8k) dataset.\n\u003e   GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems.\n\u003e   There is no information on the quality of the translation\n\n* [💻 ROCODE 💻](https://huggingface.co/datasets/cosmadrian/rocode)\n\u003e   RoCode, a competitive programming dataset, consisting of 2,642 problems written in Romanian,\n\u003e  11k solutions in C, C++ and Python and comprehensive testing suites for each problem. The purpose of RoCode\n\u003e  is to provide a benchmark for evaluating the code intelligence of language models trained on\n\u003e  Romanian / multilingual text as well as a fine-tuning set for pretrained Romanian models.\n\n[![arXiv](https://img.shields.io/badge/arXiv-2005.06012-f9f107.svg)](https://arxiv.org/abs/2402.13222)\n\n* [💻 RoITD 💻](https://huggingface.co/datasets/dragosnicolae555/RoITD)\n\u003e Romanian IT Dataset (RoITD) resembling SQuAD 1.1.\n\u003e RoITD consists of 9575 Romanian QA pairs formulated by crowd workers. QA pairs are based on 5043 articles from Romanian Wikipedia articles describing IT and household products.\n\u003e Of the total number of questions, 5103 are possible (i.e. the correct answer can be found within the paragraph) and 4472 are not possible (i.e. the given answer is a \"plausible answer\" and not correct)\n\n* [🩺 RoMedQA 🩺](https://github.com/ana-rogoz/MedQARo)\n\u003e The dataset comprises 102,646 high-quality QA pairs from real-world clinical records of 1,011 oncology patients (796 patients with breast cancer and 215 patients with lung cancer).\n\u003e The QA pairs are the results of a manual annotation process carried out by physicians specialized in oncology and radiotherapy\n\u003e RoMedQA includes 76,416 QA pairs about breast cancer patients and 26,230 about lung cancer patients, with questions grounded in medical case summaries (epicrises).\n\n[![arXiv](https://img.shields.io/badge/arXiv-2508.16390-f9f107.svg)](https://arxiv.org/abs/2508.16390v1)\n\n* [⚖️ JurRO ⚖️](https://github.com/craciuncg/GRAF)\n\u003e Romanian legal MCQA dataset, comprising 10,836 questions from three examinations.\n\u003e  Each entry essentially consists of a body in which a theoretical question is posed regarding\n\u003e a legal aspect, along with three possible answer choices labeled A, B, and C,\n\u003e out of which at mosttwo answers are correct.\n\n[![ACL](https://img.shields.io/badge/ACL%20Anthology-ed1c24.svg)](https://aclanthology.org/2025.findings-acl.659/)\n\n\n## Spelling, Dictionaries and Gramatical Errors\n\n* [✏️ Grammar-RO ✏️](https://huggingface.co/datasets/BlackKakapo/grammar-ro)\n\u003e  Synthetic dataset with ~1.9M records. Altered and correct statement as columns\n\n* [📖 RoAcReL 📖](https://huggingface.co/datasets/fmi-unibuc/RoAcReL)\n\u003e  Romanian Archaisms Regionalisms Lexicon containing ~ 1940 Word definitions\n\n* [🗺️ RoRuDi 🗺️](https://huggingface.co/datasets/fmi-unibuc/RoRuDi)\n\u003e Romanian Rules for Dialects - 1940 regionalisms, meanings and the region of provenience\n\n* [📚 RoLEX 📚](https://github.com/adrianastan/rolex)\n\u003e The dataset was developed mainly for speech processing applications, yet its applicability\n\u003e extends beyond this domain. RoLEX includes over 330,000 curated entries with information\n\u003e regarding lemma, morphosyntactic description, syllabification, lexical stress and phonemic transcription.\n\n[![Cambridge](https://img.shields.io/badge/Cambridge-Core-ef4135.svg)](https://adrianastan.com/papers/2022_NLE.pdf)\n\n## Automatic Speech Recognition\n\n* [🗣️ USPDATRO 🗣️](https://zenodo.org/records/7898233)\n\u003e Underrepresented Speech Dataset from Open Data. Duration of the dataset is 4h 18m 55s.\n\u003e Distribution according to platforms: 83% of the content comes from YouTube, 12% from SoundCloud, and 5% from Vimeo\n\u003e Dataset covers primarily under-represented speech groups (outside the 19–29 male category)\n\n[![MDPI](https://img.shields.io/badge/MDPI-Information-00813E.svg)](https://www.mdpi.com/2076-3417/14/19/9043)\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fandythefactory%2Fromanian-nlp-datasets","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fandythefactory%2Fromanian-nlp-datasets","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fandythefactory%2Fromanian-nlp-datasets/lists"}