An open API service indexing awesome lists of open source software.

Projects in Awesome Lists tagged with corpus

A curated list of projects in awesome lists tagged with corpus .

https://github.com/dariusk/corpora

A collection of small corpuses of interesting data for the creation of bots and similar stuff.

bots corpus language words

Last synced: 14 May 2025

https://github.com/wainshine/chinese-names-corpus

中文人名语料库。人名生成器。中文姓名,姓氏,名字,称呼,日本人名,翻译人名,英文人名。可用于中文分词、人名实体识别。

corpus dataset dict names ner

Last synced: 28 Jan 2026

https://github.com/cluebenchmark/clue

中文语言理解测评基准 Chinese Language Understanding Evaluation Benchmark: datasets, baselines, pre-trained models, corpus and leaderboard

albert benchmark bert chinese chineseglue corpus dataset glue language-model nlu pretrained-models pytorch roberta tensorflow transformers

Last synced: 14 May 2025

https://github.com/wainshine/Chinese-Names-Corpus

中文人名语料库。人名生成器。中文姓名,姓氏,名字,称呼,日本人名,翻译人名,英文人名。可用于中文分词、人名实体识别。

corpus dataset dict names ner

Last synced: 25 Mar 2025

https://github.com/CLUEbenchmark/CLUE

中文语言理解测评基准 Chinese Language Understanding Evaluation Benchmark: datasets, baselines, pre-trained models, corpus and leaderboard

albert benchmark bert chinese chineseglue corpus dataset glue language-model nlu pretrained-models pytorch roberta tensorflow transformers

Last synced: 28 Mar 2025

https://github.com/lucasjinreal/weibo_terminater

Final Weibo Crawler Scrap Anything From Weibo, comments, weibo contents, followers, anything. The Terminator

chatbot chinese corpus scraper sina weibo

Last synced: 15 May 2025

https://github.com/candlewill/dialog_corpus

用于训练中英文对话系统的语料库 Datasets for Training Chatbot System

chatbot corpus dataset dialog system

Last synced: 15 May 2025

https://github.com/gunthercox/chatterbot-corpus

A multilingual dialog corpus

chatterbot corpus dialog language yaml

Last synced: 14 Feb 2026

https://github.com/NiuTrans/Classical-Modern

非常全的文言文(古文)-现代文平行语料

corpus parallel-corpus traditional-and-simplified-chinese traditional-chinese

Last synced: 09 May 2025

https://github.com/niutrans/classical-modern

非常全的文言文(古文)-现代文平行语料

corpus parallel-corpus traditional-and-simplified-chinese traditional-chinese

Last synced: 08 Apr 2025

https://github.com/wainshine/company-names-corpus

公司名语料库。机构名语料库。公司简称,缩写,品牌词,企业名。可用于中文分词、机构名实体识别。

company corpus dataset dict ner

Last synced: 28 Jan 2026

https://github.com/wainshine/Company-Names-Corpus

公司名语料库。机构名语料库。公司简称,缩写,品牌词,企业名。可用于中文分词、机构名实体识别。

company corpus dataset dict ner

Last synced: 30 Mar 2025

https://github.com/cluebenchmark/cluecorpus2020

Large-scale Pre-training Corpus for Chinese 100G 中文预训练语料

albert bert chinese chinese-corpus corpus datasets nlp pretrain roberta

Last synced: 26 Jan 2026

https://github.com/CLUEbenchmark/CLUECorpus2020

Large-scale Pre-training Corpus for Chinese 100G 中文预训练语料

albert bert chinese chinese-corpus corpus datasets nlp pretrain roberta

Last synced: 09 May 2025

https://github.com/quanteda/quanteda

An R package for the Quantitative Analysis of Textual Data

corpus natural-language-processing quanteda r text-analytics

Last synced: 16 May 2025

https://github.com/tensorlayer/seq2seq-chatbot

Chatbot in 200 lines of code using TensorLayer

bot chat chatbot corpus lstm nlp python rnn tensorflow tensorlayer

Last synced: 12 Apr 2025

https://github.com/soskek/bookcorpus

Crawl BookCorpus

bookcorpus corpus crawler nlp scraper

Last synced: 07 Oct 2025

https://github.com/CBLUEbenchmark/CBLUE

中文医疗信息处理基准CBLUE: A Chinese Biomedical Language Understanding Evaluation Benchmark

acl2022 benchmark biomedical-tasks chinese chineseblue corpus dataset evaluation

Last synced: 01 Apr 2025

https://github.com/chatopera/efaqa-corpus-zh

❤️Emotional First Aid Dataset, 心理咨询问答、聊天机器人语料库

corpus natural-language-processing natural-language-understanding psychology

Last synced: 16 May 2025

https://github.com/crownpku/Small-Chinese-Corpus

Some useful Chinese corpus datasets 中文语料小数据

chinese-nlp corpus

Last synced: 16 Nov 2025

https://github.com/crownpku/small-chinese-corpus

Some useful Chinese corpus datasets 中文语料小数据

chinese-nlp corpus

Last synced: 04 Feb 2026

https://github.com/gair-nlp/mathpile

[NeurlPS D&B 2024] Generative AI for Math: MathPile

corpus language-model large-language-models math pre-training

Last synced: 16 May 2025

https://github.com/GAIR-NLP/MathPile

[NeurlPS D&B 2024] Generative AI for Math: MathPile

corpus language-model large-language-models math pre-training

Last synced: 22 Jul 2025

https://github.com/mesolitica/malaysian-dataset

We gather Malaysian dataset! https://malaysian-dataset.readthedocs.io/

bahasa-melayu corpus malay-dataset malaysia manglish text-mining

Last synced: 17 Jan 2026

https://github.com/grammarly/ua-gec

UA-GEC: Grammatical Error Correction and Fluency Corpus for the Ukrainian Language

corpus corpus-data corpus-tools dataset gec grammatical-error-correction natural-language-processing nlp-datasets ukrainian-language

Last synced: 21 Feb 2026

https://github.com/lil-lab/nlvr

Cornell NLVR and NLVR2 are natural language grounding datasets. Each example shows a visual input and a sentence describing it, and is annotated with the truth-value of the sentence.

computer-vision corpus machine-learning natural-language-processing

Last synced: 02 May 2025

https://github.com/helsinki-nlp/prosody

Helsinki Prosody Corpus and A System for Predicting Prosodic Prominence from Text

bert corpus dataset machine-learning natural-language-processing prosody pytorch sequence-labeling speech-synthesis

Last synced: 17 Jan 2026

https://github.com/EdinburghNLP/code-docstring-corpus

Preprocessed Python functions and docstrings for automated code documentation (code2doc) and automated code generation (doc2code) tasks.

code-generation corpus docstrings documentation-generator neural-machine-translation

Last synced: 23 Mar 2025

https://github.com/franck-dernoncourt/pubmed-rct

PubMed 200k RCT dataset: a large dataset for sequential sentence classification.

corpus machine-learning medical nlp randomized-controlled-trials sentence-classification

Last synced: 06 Jan 2026

https://github.com/m1-llie/TUMCC

[IP&M 2022] Telegram地下市场中文黑话识别语料集。Telegram Underground Market Chinese Corpus. Paper: Identification of Chinese Dark Jargons in Telegram Underground Markets Using Context-Oriented and Linguistic Features (IP&M, 2022).

chinese corpus dataset telegram

Last synced: 15 May 2025

https://github.com/yohasebe/wp2txt

A command-line toolkit to extract text content and category data from Wikipedia dump files

corpus machine-learning nlp ruby wikipedia wikipedia-dump

Last synced: 04 Apr 2025

https://github.com/christos-c/bible-corpus

A multilingual parallel corpus created from translations of the Bible.

bible bible-corpus corpus multilingual translation

Last synced: 22 Mar 2025

https://github.com/srvk/how2-dataset

This repository contains code and metadata of How2 dataset

corpus dataset how2-dataset language machine-translation multimodality speech-recognition video

Last synced: 27 Mar 2025

https://github.com/pythainlp/lexicon-thai

คลังศัพท์ภาษาไทย

corpus lexicon-thai thai-language

Last synced: 07 Apr 2025

https://github.com/writecrow/ocr2text

Convert a PDF via OCR to a TXT file in UTF-8 encoding

batch converter corpus ocr pdf tesseract

Last synced: 21 Aug 2025

https://github.com/yutkin/lenta.ru-news-dataset

Corpus of Russian news articles collected from Lenta.Ru

asynchronous asyncio corpus dataset lenta lenta-ru news nlp parser python russian

Last synced: 05 Apr 2025

https://github.com/cluebenchmark/pyclue

Python toolkit for Chinese Language Understanding(CLUE) Evaluation benchmark

albert bert chinese-language chineseglue corpus evaluation-benchmark language-model roberta-wwm-ext tiny xlnet

Last synced: 13 Jul 2025

https://github.com/proycon/colibri-core

Colibri core is an NLP tool as well as a C++ and Python library for working with basic linguistic constructions such as n-grams and skipgrams (i.e patterns with one or more gaps, either of fixed or dynamic size) in a quick and memory-efficient way. At the core is the tool ``colibri-patternmodeller`` whi ch allows you to build, view, manipulate and query pattern models.

c-plus-plus computational-linguistics corpus library linguistics ngram ngrams nlp pattern-recognition python skipgram text-processing

Last synced: 12 Apr 2025

https://github.com/ikegami-yukino/dataset-list

lists of text corpus and more (mainly Japanese)

corpus dataset wtfpl

Last synced: 27 Dec 2025

https://github.com/candlewill/speech-corpus-collection

A Collection of Speech Corpus for ASR and TTS

asr corpus dataset tts

Last synced: 25 Dec 2025

https://github.com/hironsan/ja.text8

Japanese text8 corpus for word embedding.

corpus deep-learning machine-learning natural-language-processing word2vec

Last synced: 11 Aug 2025

https://github.com/amir-zeldes/gum

Repository for the Georgetown University Multilayer Corpus (GUM)

annis annotations coreference corpus pos-tagging rhetorical-structure-theory treebank universal-dependencies

Last synced: 24 Dec 2025

https://github.com/islamAndAi/QURAN-NLP

Quran, Hadith, Translations, Tafaseer, Corpus Linguistics. Everything for NLP

ai chatbot corpus corpus-linguistics hadees hadith islam islamandai nlp quran search-engine tafaseer tafsir translation transliteration

Last synced: 12 Feb 2026

https://github.com/canclid/awesome-cantonese-nlp

A curated list of resources dedicated to Natural Language Processing (NLP) of Cantonese | 粵語 NLP

cantonese corpora corpus nlp

Last synced: 26 Feb 2026

https://github.com/chakki-works/coarij

Corpus of Annual Reports in Japan

corpus dataset finance natural-language-processing

Last synced: 26 Apr 2025

https://github.com/maxoodf/russian_news_corpus

Russian mass media stemmed texts corpus / Корпус лемматизированных (морфологически нормализованных) текстов российских СМИ

articles corpus machine-learning ml nlp nlp-machine-learning russian text word2vec

Last synced: 03 Aug 2025

https://github.com/philipperemy/japanese-words-to-vectors

Word2vec (word to vectors) approach for Japanese language using Gensim and Mecab.

corpus gensim japanese japanese-language wikipedia word2vec word2vec-algorithm

Last synced: 30 Apr 2025

https://github.com/isaacus-dev/open-australian-legal-corpus-creator

The code used to create and update the Open Australian Legal Corpus, the first and only multijurisdictional open corpus of Australian legislative and judicial documents.

australia corpus dataset datasets isaacus law legal open-data scraping web-scraping

Last synced: 11 Jul 2025

https://github.com/open-discourse/open-discourse

Open Discourse is the first fully comprehensive corpus of the plenary proceedings of the federal German Parliament (Bundestag).

bundestag corpus data hacktoberfest

Last synced: 14 Mar 2025

https://github.com/niutrans/languagecodes

We present a list of languages with their codes, families, regions and etc. We also present a list of multi-lingual corpora (with urls).

corpus language-codes multi-lingual

Last synced: 03 Mar 2026

https://github.com/megagonlabs/jrte-corpus

Japanese Realistic Textual Entailment Corpus (NLP 2020, LREC 2020)

corpus japanese-language natural-language-processing sentiment-polarity textual-entailment

Last synced: 23 Apr 2025

https://github.com/tommasoc80/EventStoryLine

Event StoryLine Corpus - annotated data, baselines and evaluation scripts, evaluation data.

corpus evaluation storylines

Last synced: 21 Nov 2025

https://github.com/wainshine/book-names-corpus

图书名语料库。含部分电影、游戏名称。

book corpus dataset dict isbn

Last synced: 16 Feb 2026

https://github.com/wainshine/medical-names-corpus

医疗语料库。医疗机构名语料库。药品本位码。

corpus dataset dict medical

Last synced: 02 Aug 2025

https://github.com/cyrta/voxceleb

mirror of VoxCeleb dataset - a large-scale speaker identification dataset

corpus dataset speaker speaker-identification speaker-recognition speaker-verification speech

Last synced: 11 May 2025

https://github.com/kgjerde/corporaexplorer

An R package for dynamic exploration of text collections

corpora corpus r shiny text-analysis

Last synced: 22 Oct 2025

https://github.com/proycon/folia

FoLiA: Format for Linguistic Annotation - FoLiA is a rich XML-based annotation format for the representation of language resources (including corpora) with linguistic annotations. A wide variety of linguistic annotations are supported, making FoLiA a useful format for NLP tasks and data interchange. Note that the actual Python library for processing FoLiA is implemented as part of PyNLPl, this contains higher-level tools that use the library as well as the full documentation, validation schemas, and set definitions

computational-linguistics corpus file-format folia language library linguistic-annotation-framework linguistics nlp python xml

Last synced: 07 May 2025

https://github.com/dumitrescustefan/ronec

Romanian Named Entity Corpus (RONEC) version 2.0

corpus named-entity-recognition ner romanian ronec

Last synced: 16 Jan 2026

https://github.com/quasilyte/gocorpus

The code used to serve gocorpus application

analysis corpus data go gogrep golang query search statistics syntax

Last synced: 21 Apr 2025

https://github.com/wainshine/species-names-corpus

物种名称语料库。植物名,动物名。

corpus dataset dict

Last synced: 08 Mar 2026

https://github.com/sparkfish/shabby-pages

ShabbyPages is a state-of-the-art corpus of born-digital document images with both ground truth and distorted versions appropriate for use in training models to reverse distortions and recover to original denoised documents.

binarization born-digital computer-vision corpus data-science dataset denoising layout-detection

Last synced: 17 Aug 2025

https://github.com/superdoc-dev/docx-corpus

The largest open corpus of .docx files for document processing research

bun common-crawl corpus dataset document-processing docx machine-learning nlp typescript word-documents

Last synced: 12 Mar 2026

https://github.com/hailiang-wang/egret-wenda-corpus

A Public Corpus for Machine Learning

corpus corpus-data qa

Last synced: 27 May 2026

https://github.com/penguincabinet/mama-katu-dm-corpus

The corpus of Japanese spam messages of invitation Mama Katu.

corpus database japanese mama-katu

Last synced: 25 Oct 2025

https://github.com/GermanT5/wikipedia2corpus

Wikipedia text corpus for self-supervised NLP model training

corpus german-nlp machine-learning nlp somajo wikipedia wikipedia-corpus

Last synced: 28 Mar 2025

https://github.com/liulalemx/felig-toolkit

A toolset for Amharic Language pre-processing. Includes an Amharic Stemmer, Transliterator, Stopword remover , Lexical analyzer, Corpus indexer and Term weighter.

amharic amharic-corpus amharic-nlp amharic-stemmer corpus lexical-analyzer linguistics stopword-removal transliterator

Last synced: 04 Sep 2025

https://github.com/tanloong/neosca

L2SCA & LCA fork: cross-platform, GUI, without Java dependency

constituency-parsing corpus corpus-analysis l2sca linguistics neosca nlp python syntactic-complexity tregex

Last synced: 07 May 2025

https://github.com/proiel/proiel-treebank

Official releases of the PROIEL treebank of ancient Indo-European languages

ancient-greek ancient-languages armenian corpus gothic2 language latin linguistics new-testament old-church-slavonic treebank

Last synced: 10 Jan 2026

https://github.com/gaussic/chinese-lyric-corpus

A Chinese lyric corpus which contains nearly 50,000 lyrics from 500 artists

chinese corpus lyrics text-generation

Last synced: 27 Feb 2026

https://github.com/hantang/data-corpus

语料数据和词库收集:中文、英文停用词,情感分析,分类词典,敏感词库(违禁词,审查词)。stop words, sentiment analysis, thesaurus, censorship/sensitive word

corpus nlp stopwords thesaurus

Last synced: 13 Feb 2026

https://github.com/canclid/canto-filter

粵文語料篩選器 Cantonese text filter

cantonese cantonese-language corpus corpus-data data nlp

Last synced: 27 Oct 2025

https://github.com/hugovk/everyfinnishword

Every Finnish word

corpus finnish language words

Last synced: 30 Jan 2026

https://github.com/statico/aspen

🔎 📖 ✨ Custom, private search engine for text documents built with NextJS/React/ES6/ES7

corpus docker elasticsearch es6 es7 javascript nextjs plaintext plaintext-documents search search-engine

Last synced: 23 Apr 2025

https://github.com/webis-de/webis-tldr-17-corpus

Code for constructing TLDR corpus from Reddit dataset

corpus summarization

Last synced: 11 Mar 2026

https://github.com/cynthia/kosentences

Large scale unannotated Korean corpus for unsupervised tasks. (e.g. Language modeling)

corpus datasets korean language-modeling nlp

Last synced: 26 Jul 2025

https://github.com/kathyreid/opensource-voice-tools

A repo listing known open source voice tools, ordered by where they sit in the voice stack

asr chatbot conversational-ui corpus speech speech-recognition stt tts voice

Last synced: 27 Mar 2026

https://github.com/megagonlabs/asdc

Accommodation Search Dialog Corpus (宿泊施設探索対話コーパス)

corpus dialog japanese-language

Last synced: 05 Mar 2026

https://github.com/instituutnederlandsetaal/OpenConvert

Text conversion tool (from e.g. Word, HTML, txt) to corpus formats TEI or FoLiA)

conversion corpus

Last synced: 10 May 2025

https://github.com/dbklim/russian_subtitles_dataset

Preprocessing of the dataset of 347 subtitles for the TV series (thanks to Taiga Corpus) to build a word2vec model, JamSpell model, neural network training, chat bot training or in any other NLP task.

bot cnn corpus dataset lstm machine-learning ml natural-language-processing nlp nlu rnn russian subtitles text text-analysis text-processing word2vec

Last synced: 29 Apr 2025