Projects in Awesome Lists tagged with corpus
A curated list of projects in awesome lists tagged with corpus .
https://github.com/brightmart/nlp_chinese_corpus
大规模中文自然语言处理语料 Large Scale Chinese Corpus for NLP
bert chinese chinese-corpus chinese-dataset chinese-nlp corpus dataset language-model news nlp pretrain question-answering text-classification wiki word2vec
Last synced: 22 Feb 2026
https://github.com/dariusk/corpora
A collection of small corpuses of interesting data for the creation of bots and similar stuff.
Last synced: 14 May 2025
https://github.com/cluebenchmark/cluedatasetsearch
搜索所有中文NLP数据集,附常用英文NLP数据集
chinese corpus datasets knowledge-graph machine-reading-comprehension machine-translation match ner nlp qa sentiment-analysis text-classification text-similarity text-summarization
Last synced: 14 May 2025
https://github.com/CLUEbenchmark/CLUEDatasetSearch
搜索所有中文NLP数据集,附常用英文NLP数据集
chinese corpus datasets knowledge-graph machine-reading-comprehension machine-translation match ner nlp qa sentiment-analysis text-classification text-similarity text-summarization
Last synced: 28 Mar 2025
https://github.com/cluebenchmark/clue
中文语言理解测评基准 Chinese Language Understanding Evaluation Benchmark: datasets, baselines, pre-trained models, corpus and leaderboard
albert benchmark bert chinese chineseglue corpus dataset glue language-model nlu pretrained-models pytorch roberta tensorflow transformers
Last synced: 14 May 2025
https://github.com/CLUEbenchmark/CLUE
中文语言理解测评基准 Chinese Language Understanding Evaluation Benchmark: datasets, baselines, pre-trained models, corpus and leaderboard
albert benchmark bert chinese chineseglue corpus dataset glue language-model nlu pretrained-models pytorch roberta tensorflow transformers
Last synced: 28 Mar 2025
https://github.com/adbar/trafilatura
Python & command-line tool to gather text on the Web: web crawling/scraping, extraction of text, metadata, comments
article-extractor corpus corpus-builder corpus-tools crawler html-to-markdown html2text news news-aggregator news-crawler nlp readability rss-feed scraping tei text-cleaning text-extraction text-mining text-preprocessing web-scraping
Last synced: 24 Dec 2025
https://github.com/gunthercox/chatterbot-corpus
A multilingual dialog corpus
chatterbot corpus dialog language yaml
Last synced: 14 Feb 2026
https://github.com/NiuTrans/Classical-Modern
非常全的文言文(古文)-现代文平行语料
corpus parallel-corpus traditional-and-simplified-chinese traditional-chinese
Last synced: 09 May 2025
https://github.com/niutrans/classical-modern
非常全的文言文(古文)-现代文平行语料
corpus parallel-corpus traditional-and-simplified-chinese traditional-chinese
Last synced: 08 Apr 2025
https://github.com/chatopera/insuranceqa-corpus-zh
:helicopter: 保险行业语料库,聊天机器人
chatbot corpus dataset insurance insuranceqa-corpus-zh machine-learning natural-language-processing natural-language-understanding qasystem question-answering
Last synced: 15 May 2025
https://github.com/cluebenchmark/cluecorpus2020
Large-scale Pre-training Corpus for Chinese 100G 中文预训练语料
albert bert chinese chinese-corpus corpus datasets nlp pretrain roberta
Last synced: 26 Jan 2026
https://github.com/CLUEbenchmark/CLUECorpus2020
Large-scale Pre-training Corpus for Chinese 100G 中文预训练语料
albert bert chinese chinese-corpus corpus datasets nlp pretrain roberta
Last synced: 09 May 2025
https://github.com/quanteda/quanteda
An R package for the Quantitative Analysis of Textual Data
corpus natural-language-processing quanteda r text-analytics
Last synced: 16 May 2025
https://github.com/tensorlayer/seq2seq-chatbot
Chatbot in 200 lines of code using TensorLayer
bot chat chatbot corpus lstm nlp python rnn tensorflow tensorlayer
Last synced: 12 Apr 2025
https://github.com/soskek/bookcorpus
Crawl BookCorpus
bookcorpus corpus crawler nlp scraper
Last synced: 07 Oct 2025
https://github.com/CLUEbenchmark/CLUEPretrainedModels
高质量中文预训练模型集合:最先进大模型、最快小模型、相似度专门模型
albert bert chinese corpus dataset distillation pretrained-models roberta semantic-similarity sentence-analysis sentence-classification sentence-pairs text-classification
Last synced: 13 Apr 2025
https://github.com/cluebenchmark/cluepretrainedmodels
高质量中文预训练模型集合:最先进大模型、最快小模型、相似度专门模型
albert bert chinese corpus dataset distillation pretrained-models roberta semantic-similarity sentence-analysis sentence-classification sentence-pairs text-classification
Last synced: 04 Apr 2025
https://github.com/CBLUEbenchmark/CBLUE
中文医疗信息处理基准CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark
acl2022 benchmark biomedical-tasks chinese chineseblue corpus dataset evaluation
Last synced: 01 Apr 2025
https://github.com/chatopera/efaqa-corpus-zh
❤️Emotional First Aid Dataset, 心理咨询问答、聊天机器人语料库
corpus natural-language-processing natural-language-understanding psychology
Last synced: 16 May 2025
https://github.com/nonamestreet/weixin_public_corpus
微信公众号语料库
chinese-nlp corpora corpus linguistics natural-language-processing nlp wei-xin weixin weixin-data yu-liao yu-liao-ku
Last synced: 20 Nov 2025
https://github.com/crownpku/Small-Chinese-Corpus
Some useful Chinese corpus datasets 中文语料小数据
Last synced: 16 Nov 2025
https://github.com/crownpku/small-chinese-corpus
Some useful Chinese corpus datasets 中文语料小数据
Last synced: 04 Feb 2026
https://github.com/louisowen6/NLP_bahasa_resources
A Curated List of Dataset and Usable Library Resources for NLP in Bahasa Indonesia
bahasa-indonesia corpus corpus-linguistics dataset indonesian indonesian-language library natural-language-processing nlp nlp-bahasa-resources packages sentiment-analysis sentiment-analysis-dataset
Last synced: 13 Apr 2025
https://github.com/gair-nlp/mathpile
[NeurlPS D&B 2024] Generative AI for Math: MathPile
corpus language-model large-language-models math pre-training
Last synced: 16 May 2025
https://github.com/GAIR-NLP/MathPile
[NeurlPS D&B 2024] Generative AI for Math: MathPile
corpus language-model large-language-models math pre-training
Last synced: 22 Jul 2025
https://github.com/flairnlp/fundus
A very simple news crawler with a funny name
cc-news commoncrawl corpus corpus-tools crawler datasets image-classification image-extraction news-crawler news-scraping nlp python rss scraper sitemap text-extraction web-corpus web-scraping
Last synced: 08 Jan 2026
https://github.com/mesolitica/malaysian-dataset
We gather Malaysian dataset! https://malaysian-dataset.readthedocs.io/
bahasa-melayu corpus malay-dataset malaysia manglish text-mining
Last synced: 17 Jan 2026
https://github.com/flairNLP/fundus
A very simple news crawler with a funny name
cc-news commoncrawl corpus crawler news-crawler news-scraping nlp python rss scraper sitemap text-extraction web-corpus web-scraping
Last synced: 04 Mar 2025
https://github.com/grammarly/ua-gec
UA-GEC: Grammatical Error Correction and Fluency Corpus for the Ukrainian Language
corpus corpus-data corpus-tools dataset gec grammatical-error-correction natural-language-processing nlp-datasets ukrainian-language
Last synced: 21 Feb 2026
https://github.com/lil-lab/nlvr
Cornell NLVR and NLVR2 are natural language grounding datasets. Each example shows a visual input and a sentence describing it, and is annotated with the truth-value of the sentence.
computer-vision corpus machine-learning natural-language-processing
Last synced: 02 May 2025
https://github.com/helsinki-nlp/prosody
Helsinki Prosody Corpus and A System for Predicting Prosodic Prominence from Text
bert corpus dataset machine-learning natural-language-processing prosody pytorch sequence-labeling speech-synthesis
Last synced: 17 Jan 2026
https://github.com/strongcourage/fuzzing-corpus
My fuzzing corpus
corpus file-format fuzzing testsuite vulnerability
Last synced: 05 Apr 2026
https://github.com/kirralabs/indonesian-NLP-resources
data resource untuk NLP bahasa indonesia
corpus corpus-linguistics crawler dataset dependency-parser indonesian indonesian-language named-entity-recognition nlp parallel-corpus pos-tagging sentiment-analysis
Last synced: 15 Apr 2025
https://github.com/EdinburghNLP/code-docstring-corpus
Preprocessed Python functions and docstrings for automated code documentation (code2doc) and automated code generation (doc2code) tasks.
code-generation corpus docstrings documentation-generator neural-machine-translation
Last synced: 23 Mar 2025
https://github.com/franck-dernoncourt/pubmed-rct
PubMed 200k RCT dataset: a large dataset for sequential sentence classification.
corpus machine-learning medical nlp randomized-controlled-trials sentence-classification
Last synced: 06 Jan 2026
https://github.com/m1-llie/TUMCC
[IP&M 2022] Telegram地下市场中文黑话识别语料集。Telegram Underground Market Chinese Corpus. Paper: Identification of Chinese Dark Jargons in Telegram Underground Markets Using Context-Oriented and Linguistic Features (IP&M, 2022).
chinese corpus dataset telegram
Last synced: 15 May 2025
https://github.com/yohasebe/wp2txt
A command-line toolkit to extract text content and category data from Wikipedia dump files
corpus machine-learning nlp ruby wikipedia wikipedia-dump
Last synced: 04 Apr 2025
https://github.com/christos-c/bible-corpus
A multilingual parallel corpus created from translations of the Bible.
bible bible-corpus corpus multilingual translation
Last synced: 22 Mar 2025
https://github.com/srvk/how2-dataset
This repository contains code and metadata of How2 dataset
corpus dataset how2-dataset language machine-translation multimodality speech-recognition video
Last synced: 27 Mar 2025
https://github.com/scriptin/kanji-frequency
Kanji usage frequency data collected from various sources
cjk cjk-characters corpus corpus-linguistics data data-visualization frequency-lists japanese japanese-language kanji kanji-frequency
Last synced: 18 Jan 2026
https://github.com/pythainlp/lexicon-thai
คลังศัพท์ภาษาไทย
corpus lexicon-thai thai-language
Last synced: 07 Apr 2025
https://github.com/cluebenchmark/pyclue
Python toolkit for Chinese Language Understanding(CLUE) Evaluation benchmark
albert bert chinese-language chineseglue corpus evaluation-benchmark language-model roberta-wwm-ext tiny xlnet
Last synced: 13 Jul 2025
https://github.com/proycon/colibri-core
Colibri core is an NLP tool as well as a C++ and Python library for working with basic linguistic constructions such as n-grams and skipgrams (i.e patterns with one or more gaps, either of fixed or dynamic size) in a quick and memory-efficient way. At the core is the tool ``colibri-patternmodeller`` whi ch allows you to build, view, manipulate and query pattern models.
c-plus-plus computational-linguistics corpus library linguistics ngram ngrams nlp pattern-recognition python skipgram text-processing
Last synced: 12 Apr 2025
https://github.com/ikegami-yukino/dataset-list
lists of text corpus and more (mainly Japanese)
Last synced: 27 Dec 2025
https://github.com/candlewill/speech-corpus-collection
A Collection of Speech Corpus for ASR and TTS
Last synced: 25 Dec 2025
https://github.com/hironsan/ja.text8
Japanese text8 corpus for word embedding.
corpus deep-learning machine-learning natural-language-processing word2vec
Last synced: 11 Aug 2025
https://github.com/amir-zeldes/gum
Repository for the Georgetown University Multilayer Corpus (GUM)
annis annotations coreference corpus pos-tagging rhetorical-structure-theory treebank universal-dependencies
Last synced: 24 Dec 2025
https://github.com/islamAndAi/QURAN-NLP
Quran, Hadith, Translations, Tafaseer, Corpus Linguistics. Everything for NLP
ai chatbot corpus corpus-linguistics hadees hadith islam islamandai nlp quran search-engine tafaseer tafsir translation transliteration
Last synced: 12 Feb 2026
https://github.com/canclid/awesome-cantonese-nlp
A curated list of resources dedicated to Natural Language Processing (NLP) of Cantonese | 粵語 NLP
Last synced: 26 Feb 2026
https://github.com/GlobalMaksimum/sadedegel
A General Purpose NLP library for Turkish
acikhack2 ai artificial-intelligence bert binder corpus data-science deep-learning embeddings heroku machine-learning natural-language-processing neural-network neural-networks news-summarizer nlp python
Last synced: 03 May 2025
https://github.com/manifoldfinance/mev-corpus
MEV Data Corpus
blockchain corpus data ethereum flashbots mev miner-extracted-value
Last synced: 21 Jan 2026
https://github.com/chakki-works/coarij
Corpus of Annual Reports in Japan
corpus dataset finance natural-language-processing
Last synced: 26 Apr 2025
https://github.com/maxoodf/russian_news_corpus
Russian mass media stemmed texts corpus / Корпус лемматизированных (морфологически нормализованных) текстов российских СМИ
articles corpus machine-learning ml nlp nlp-machine-learning russian text word2vec
Last synced: 03 Aug 2025
https://github.com/philipperemy/japanese-words-to-vectors
Word2vec (word to vectors) approach for Japanese language using Gensim and Mecab.
corpus gensim japanese japanese-language wikipedia word2vec word2vec-algorithm
Last synced: 30 Apr 2025
https://github.com/isaacus-dev/open-australian-legal-corpus-creator
The code used to create and update the Open Australian Legal Corpus, the first and only multijurisdictional open corpus of Australian legislative and judicial documents.
australia corpus dataset datasets isaacus law legal open-data scraping web-scraping
Last synced: 11 Jul 2025
https://github.com/open-discourse/open-discourse
Open Discourse is the first fully comprehensive corpus of the plenary proceedings of the federal German Parliament (Bundestag).
bundestag corpus data hacktoberfest
Last synced: 14 Mar 2025
https://github.com/niutrans/languagecodes
We present a list of languages with their codes, families, regions and etc. We also present a list of multi-lingual corpora (with urls).
corpus language-codes multi-lingual
Last synced: 03 Mar 2026
https://github.com/megagonlabs/jrte-corpus
Japanese Realistic Textual Entailment Corpus (NLP 2020, LREC 2020)
corpus japanese-language natural-language-processing sentiment-polarity textual-entailment
Last synced: 23 Apr 2025
https://github.com/tommasoc80/EventStoryLine
Event StoryLine Corpus - annotated data, baselines and evaluation scripts, evaluation data.
Last synced: 21 Nov 2025
https://github.com/cyrta/voxceleb
mirror of VoxCeleb dataset - a large-scale speaker identification dataset
corpus dataset speaker speaker-identification speaker-recognition speaker-verification speech
Last synced: 11 May 2025
https://github.com/kgjerde/corporaexplorer
An R package for dynamic exploration of text collections
corpora corpus r shiny text-analysis
Last synced: 22 Oct 2025
https://github.com/zjunlp/iepile
IEPile: A Large-Scale Information Extraction Corpus
bilingual chinese corpus dataset english event-extraction ie iepie information-extraction instructions knowledge-graph large-language-models named-entity-recognition natural-language-processing relation-extraction
Last synced: 13 Jun 2025
https://github.com/proycon/folia
FoLiA: Format for Linguistic Annotation - FoLiA is a rich XML-based annotation format for the representation of language resources (including corpora) with linguistic annotations. A wide variety of linguistic annotations are supported, making FoLiA a useful format for NLP tasks and data interchange. Note that the actual Python library for processing FoLiA is implemented as part of PyNLPl, this contains higher-level tools that use the library as well as the full documentation, validation schemas, and set definitions
computational-linguistics corpus file-format folia language library linguistic-annotation-framework linguistics nlp python xml
Last synced: 07 May 2025
https://github.com/dumitrescustefan/ronec
Romanian Named Entity Corpus (RONEC) version 2.0
corpus named-entity-recognition ner romanian ronec
Last synced: 16 Jan 2026
https://github.com/howl-anderson/mitie_chinese_wikipedia_corpus
Pre-trained Wikipedia corpus by MITIE
chinese-corpus corpus mitie nlp nlp-machine-learning wikipeida
Last synced: 29 Jan 2026
https://github.com/sparkfish/shabby-pages
ShabbyPages is a state-of-the-art corpus of born-digital document images with both ground truth and distorted versions appropriate for use in training models to reverse distortions and recover to original denoised documents.
binarization born-digital computer-vision corpus data-science dataset denoising layout-detection
Last synced: 17 Aug 2025
https://github.com/superdoc-dev/docx-corpus
The largest open corpus of .docx files for document processing research
bun common-crawl corpus dataset document-processing docx machine-learning nlp typescript word-documents
Last synced: 12 Mar 2026
https://github.com/hailiang-wang/egret-wenda-corpus
A Public Corpus for Machine Learning
Last synced: 27 May 2026
https://github.com/penguincabinet/mama-katu-dm-corpus
The corpus of Japanese spam messages of invitation Mama Katu.
corpus database japanese mama-katu
Last synced: 25 Oct 2025
https://github.com/GermanT5/wikipedia2corpus
Wikipedia text corpus for self-supervised NLP model training
corpus german-nlp machine-learning nlp somajo wikipedia wikipedia-corpus
Last synced: 28 Mar 2025
https://github.com/liulalemx/felig-toolkit
A toolset for Amharic Language pre-processing. Includes an Amharic Stemmer, Transliterator, Stopword remover , Lexical analyzer, Corpus indexer and Term weighter.
amharic amharic-corpus amharic-nlp amharic-stemmer corpus lexical-analyzer linguistics stopword-removal transliterator
Last synced: 04 Sep 2025
https://github.com/tanloong/neosca
L2SCA & LCA fork: cross-platform, GUI, without Java dependency
constituency-parsing corpus corpus-analysis l2sca linguistics neosca nlp python syntactic-complexity tregex
Last synced: 07 May 2025
https://github.com/proiel/proiel-treebank
Official releases of the PROIEL treebank of ancient Indo-European languages
ancient-greek ancient-languages armenian corpus gothic2 language latin linguistics new-testament old-church-slavonic treebank
Last synced: 10 Jan 2026
https://github.com/ropenscilabs/tif
Text Interchange Formats
corpus natural-language-processing r r-package rstats term-frequency text-processing tokenizer
Last synced: 28 Oct 2025
https://github.com/gaussic/chinese-lyric-corpus
A Chinese lyric corpus which contains nearly 50,000 lyrics from 500 artists
chinese corpus lyrics text-generation
Last synced: 27 Feb 2026
https://github.com/hantang/data-corpus
语料数据和词库收集:中文、英文停用词,情感分析,分类词典,敏感词库(违禁词,审查词)。stop words, sentiment analysis, thesaurus, censorship/sensitive word
corpus nlp stopwords thesaurus
Last synced: 13 Feb 2026
https://github.com/canclid/canto-filter
粵文語料篩選器 Cantonese text filter
cantonese cantonese-language corpus corpus-data data nlp
Last synced: 27 Oct 2025
https://github.com/statico/aspen
🔎 📖 ✨ Custom, private search engine for text documents built with NextJS/React/ES6/ES7
corpus docker elasticsearch es6 es7 javascript nextjs plaintext plaintext-documents search search-engine
Last synced: 23 Apr 2025
https://github.com/timarkh/tsakorpus
Yet another search platform for linguistic corpora.
corpus corpus-linguistics corpus-tools elasticsearch flask language-documentation linguistic-corpora linguistics media-aligned-corpora parallel-corpora
Last synced: 16 Jan 2026
https://github.com/webis-de/webis-tldr-17-corpus
Code for constructing TLDR corpus from Reddit dataset
Last synced: 11 Mar 2026
https://github.com/cynthia/kosentences
Large scale unannotated Korean corpus for unsupervised tasks. (e.g. Language modeling)
corpus datasets korean language-modeling nlp
Last synced: 26 Jul 2025
https://github.com/kathyreid/opensource-voice-tools
A repo listing known open source voice tools, ordered by where they sit in the voice stack
asr chatbot conversational-ui corpus speech speech-recognition stt tts voice
Last synced: 27 Mar 2026
https://github.com/megagonlabs/asdc
Accommodation Search Dialog Corpus (宿泊施設探索対話コーパス)
corpus dialog japanese-language
Last synced: 05 Mar 2026
https://github.com/instituutnederlandsetaal/OpenConvert
Text conversion tool (from e.g. Word, HTML, txt) to corpus formats TEI or FoLiA)
Last synced: 10 May 2025
https://github.com/dbklim/russian_subtitles_dataset
Preprocessing of the dataset of 347 subtitles for the TV series (thanks to Taiga Corpus) to build a word2vec model, JamSpell model, neural network training, chat bot training or in any other NLP task.
bot cnn corpus dataset lstm machine-learning ml natural-language-processing nlp nlu rnn russian subtitles text text-analysis text-processing word2vec
Last synced: 29 Apr 2025