{"id":62648,"url":"https://github.com/JackHCC/NLP-Bubble","name":"NLP-Bubble","description":"🖨 Natural Language Processing Learning Blog，a Study Bubble to recording learning.","projects_count":57,"last_synced_at":"2026-07-14T10:00:22.137Z","repository":{"id":112238230,"uuid":"451423733","full_name":"JackHCC/NLP-Bubble","owner":"JackHCC","description":"🖨 Natural Language Processing Learning Blog，a Study Bubble to recording learning.","archived":false,"fork":false,"pushed_at":"2022-07-01T13:02:34.000Z","size":2633,"stargazers_count":6,"open_issues_count":0,"forks_count":0,"subscribers_count":1,"default_branch":"master","last_synced_at":"2026-06-09T04:03:28.030Z","etag":null,"topics":["awesome","machine-learning","nlp"],"latest_commit_sha":null,"homepage":"","language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/JackHCC.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2022-01-24T10:46:05.000Z","updated_at":"2024-12-20T15:42:53.000Z","dependencies_parsed_at":"2023-05-07T03:17:48.104Z","dependency_job_id":null,"html_url":"https://github.com/JackHCC/NLP-Bubble","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/JackHCC/NLP-Bubble","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/JackHCC%2FNLP-Bubble","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/JackHCC%2FNLP-Bubble/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/JackHCC%2FNLP-Bubble/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/JackHCC%2FNLP-Bubble/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/JackHCC","download_url":"https://codeload.github.com/JackHCC/NLP-Bubble/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/JackHCC%2FNLP-Bubble/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":34801014,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-06-26T02:00:06.560Z","response_time":106,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"created_at":"2024-07-02T00:00:23.456Z","updated_at":"2026-07-14T10:00:22.137Z","primary_language":null,"list_of_lists":false,"displayable":true,"categories":["NLP Task","Interview","DataSet","Lessons/Books","Papers","Resource"],"sub_categories":["Other","Classes"],"readme":"# NLP-Bubble\n🖨 Natural Language Processing Learning Blog，a Study Bubble to recording learning.\n\n![](image/logo/NLP-Bubble-banner.png)\n\n💡 NLP Learning Record 💡\n\n\n\n## Lessons/Books\n\n- [Statistical Learning Method v1](https://blog.creativecc.cn/posts/Lesson-Statistical-Learning-Method.html)\n\n- [CS224N Natural Language Processing 2022](https://github.com/JackHCC/Awesome-DL-Models/tree/master/Docx/CS224N)\n- [CS224W Machine Learning with Graphs 2021](https://blog.creativecc.cn/posts/Lesson-CS224W-Machine-Learning-with-Graphs.html)\n\n\n\n## Papers\n\n- [Arxiv NLP Reporter](https://github.com/JackHCC/Arxiv-NLP-Reporter)\n  - [Web Reader](https://blog.creativecc.cn/Arxiv-NLP-Reporter/)\n\n- Reading Web\n\n  - [ACL anthology](https://www.aclweb.org/anthology/)\n  - [NeurIPS](https://papers.nips.cc) , ICML, ICLR\n\n  - [online preprint servers](https://arxiv.org)\n\n\n\n## DataSet\n\n### Classes\n\n- Linguistic Data Consortium\n  - [Linguistic Data Consortium (upenn.edu)](https://catalog.ldc.upenn.edu/)\n  - [Linguistics (stanford.edu)](https://linguistics.stanford.edu/resources/resources-corpora)\n- Machine translation\n  - [Statistical Machine Translation (statmt.org)](https://statmt.org/)\n- Dependency parsing: Universal Dependencies\n  - [Universal Dependencies](https://universaldependencies.org/)\n\n### Other\n\n- Awesome\n  - [NLPDataSet](https://github.com/liucongg/NLPDataSet)\n  - [nlp-datasets](https://github.com/niderhoff/nlp-datasets)\n\n- Platform\n  - [千言：中文开源数据集合](https://www.luge.ai/)\n  - [Papers With Code](https://paperswithcode.com/datasets)\n  - Kaggle \n  - [GLUE](https://gluebenchmark.com/tasks)\n\n\n- Blogs\n  - [Datasets for Natural Language Processing](https://machinelearningmastery.com/datasets-natural-language-processing/)\n  - [Sentiment Analysis](https://nlp.stanford.edu/sentiment/)\n  - [The bAbI](https://research.facebook.com/downloads/babi/)\n\n\n\n## NLP Task\n\n思维导图：\n\n![](../../../Blog/JackCC.Blog/hexo_blog/source/images/lesson/NLP_Task.png)\n\n常见的32项NLP任务以及对应的评测数据、评测指标、目前的SOTA结果（2020.05）以及对应的Paper与Code.\n\n| 任务                                           | 描述                | corpus/dataset                       | 评价指标                                   | SOTA                                       | Papers                                                       | Code                                                         |\n| ---------------------------------------------- | ------------------- | ------------------------------------ | ------------------------------------------ | ------------------------------------------ | ------------------------------------------------------------ | ------------------------------------------------------------ |\n| Chunking                                       | 组块分析            | Penn Treebank                        | F1                                         | 95.77                                      | [A Joint Many-Task Model: Growing a Neural Network for Multiple NLP Tasks](https://arxiv.org/pdf/1611.01587v5.pdf) | [Link](https://github.com/hassyGo/charNgram2vec)             |\n| Common sense reasoning                         | 常识推理            | Event2Mind                           | cross-entropy                              | 4.22                                       | [Event2Mind: Commonsense Inference on Events, Intents, and Reactions](https://www.dialog-21.ru/media/5090/fenogenovaasplusetal-010.pdf) | [Link](https://github.com/Alenush/russian_event2mind)        |\n| Parsing                                        | 句法分析            | Penn Treebank                        | F1                                         | 95.13                                      | [Constituency Parsing with a Self-Attentive Encoder](https://arxiv.org/pdf/1805.01052v1.pdf) | [Link](https://github.com/nikitakit/self-attentive-parser)   |\n| Coreference resolution                         | 指代消解            | CoNLL 2012                           | average F1                                 | 73                                         | [Higher-order Coreference Resolution with Coarse-to-fine Inference](https://arxiv.org/pdf/1804.05392v1.pdf) | [Link](https://github.com/kentonl/e2e-coref)                 |\n| Dependency parsing                             | 依存句法分析        | Penn Treebank                        | POS\u003cbr/\u003eUAS\u003cbr/\u003eLAS                        | 97.3\u003cbr/\u003e95.44\u003cbr/\u003e93.76                   | [Deep Biaffine Attention for Neural Dependency Parsing](https://arxiv.org/pdf/1611.01734v3.pdf) | [Link](https://github.com/PaddlePaddle/PaddleNLP/tree/develop/examples/dependency_parsing/ddparser) |\n| Task-Oriented Dialogue/Intent Detection        | 任务型对话/意图识别 | ATIS/Snips                           | accuracy                                   | 94.1  97.0                                 | [Slot-Gated Modeling for Joint Slot Filling and Intent Prediction](https://aclanthology.org/N18-2118.pdf) | [Link](https://github.com/MiuLab/SlotGated-SLU)              |\n| Task-Oriented Dialogue/Slot Filling            | 任务型对话/槽填充   | ATIS/Snips                           | F1                                         | 95.2\u003cbr/\u003e88.8                              | [Slot-Gated Modeling for Joint Slot Filling and Intent Prediction](https://aclanthology.org/N18-2118.pdf) | [Link](https://github.com/MiuLab/SlotGated-SLU)              |\n| Task-Oriented Dialogue/Dialogue State Tracking | 任务型对话/状态追踪 | DSTC2                                | Area\u003cbr/\u003eFood\u003cbr/\u003ePrice\u003cbr/\u003eJoint          | 90\u003cbr/\u003e84\u003cbr/\u003e92\u003cbr/\u003e72                    | [Dialogue Learning with Human Teaching and Feedback in End-to-End Trainable Task-Oriented Dialogue Systems](https://arxiv.org/pdf/1804.06512v1.pdf) | [Link](https://github.com/google-research-datasets/simulated-dialogue) |\n| Domain adaptation                              | 领域适配            | Multi-Domain Sentiment Dataset       | average \u003cbr/\u003eaccuracy                      | 79.15                                      | [Strong Baselines for Neural Semi-supervised Learning under Domain Shift](https://arxiv.org/pdf/1804.09530v1.pdf) | [Link](https://github.com/bplank/semi-supervised-baselines)  |\n| Entity Linking                                 | 实体链接            | AIDA CoNLL-YAGO                      | Micro-F1-strong\u003cbr/\u003eMacro-F1-strong        | 86.6 \u003cbr/\u003e89.4                             | [End-to-End Neural Entity Linking](https://arxiv.org/pdf/1808.07699v2.pdf) | [Link](https://github.com/dalab/end2end_neural_el)           |\n| Information Extraction                         | 信息抽取            | ReVerb45K                            | Precision\u003cbr/\u003eRecall\u003cbr/\u003eF1                | 62.7\u003cbr/\u003e84.4\u003cbr/\u003e81.9                     | [CESI: Canonicalizing Open Knowledge Bases using Embeddings and Side Information](https://arxiv.org/pdf/1902.00172v1.pdf) | [Link](https://github.com/malllabiisc/cesi)                  |\n| Grammatical Error Correction                   | 语法错误纠正        | JFLEG                                | GLEU                                       | 61.5                                       | [Near Human-Level Performance in Grammatical Error Correction with Hybrid Machine Translation](https://arxiv.org/pdf/1804.05945v1.pdf) | Link                                                         |\n| Language modeling                              | 语言模型            | Penn Treebank                        | Validation perplexity\u003cbr/\u003e Test perplexity | 48.33\u003cbr/\u003e47.69                            | [Breaking the Softmax Bottleneck: A High-Rank RNN Language Model](https://arxiv.org/pdf/1711.03953v4.pdf) | [Link](https://github.com/zihangdai/mos)                     |\n| Lexical Normalization                          | 词汇规范化          | LexNorm2015                          | F1\u003cbr/\u003ePrecision\u003cbr/\u003eRecall                | 86.39 \u003cbr/\u003e93.53 \u003cbr/\u003e80.26                | [MoNoise: Modeling Noise Using a Modular Normalization System](https://arxiv.org/pdf/1710.03476v1.pdf) | [Link](https://bitbucket.org/robvanderg/monoise)             |\n| Machine translation                            | 机器翻译            | WMT 2014 EN-DE                       | BLEU                                       | 35.0                                       | [Understanding Back-Translation at Scale](https://arxiv.org/pdf/1808.09381v2.pdf) | [Link](https://github.com/pytorch/fairseq)                   |\n| Multimodal Emotion Recognition                 | 多模态情感识别      | IEMOCAP                              | Accuracy                                   | 76.5                                       | [Multimodal Sentiment Analysis using Hierarchical Fusion with Context Modeling](https://arxiv.org/pdf/1806.06228v1.pdf) | [Link](https://github.com/SenticNet/hfusion)                 |\n| Multimodal Metaphor Recognition                | 多模态隐喻识别      | verb-noun pairs adjective-noun pairs | F1                                         | 0.75\u003cbr/\u003e0.79                              | [Black Holes and White Rabbits: Metaphor Identification with Visual Features](https://aclanthology.org/N16-1020.pdf) | Link                                                         |\n| Multimodal Sentiment Analysis                  | 多模态情感分析      | MOSI                                 | Accuracy                                   | 80.3                                       | [Context-Dependent Sentiment Analysis in User-Generated Videos](https://aclanthology.org/P17-1081.pdf) | [Link](https://github.com/senticnet/sc-lstm)                 |\n| Named entity recognition                       | 命名实体识别        | CoNLL 2003                           | F1                                         | 93.09                                      | [Contextual String Embeddings for Sequence Labeling](https://aclanthology.org/C18-1139.pdf) | [Link](https://github.com/zalandoresearch/flair)             |\n| Natural language inference                     | 自然语言推理        | SciTail                              | Accuracy                                   | 88.3                                       | [Improving Language Understanding by Generative Pre-Training](https://s3-us-west-2.amazonaws.com/openai-assets/research-covers/language-unsupervised/language_understanding_paper.pdf) | [Link](https://github.com/huggingface/transformers)          |\n| Part-of-speech tagging                         | 词性标注            | Penn Treebank                        | Accuracy                                   | 97.96                                      | [Morphosyntactic Tagging with a Meta-BiLSTM Model over Context Sensitive Token Encodings](https://arxiv.org/pdf/1805.08237v1.pdf) | [Link](https://github.com/google/meta_tagger)                |\n| Question answering                             | 问答                | CliCR                                | F1                                         | 33.9                                       | [CliCR: A Dataset of Clinical Case Reports for Machine Reading Comprehension](https://arxiv.org/pdf/1803.09720v1.pdf) | [Link](https://github.com/clips/clicr)                       |\n| Word segmentation                              | 分词                | VLSP 2013                            | F1                                         | 97.90                                      | [A Fast and Accurate Vietnamese Word Segmenter](https://arxiv.org/pdf/1709.06307v2.pdf) | [Link](https://github.com/datquocnguyen/RDRsegmenter)        |\n| Word Sense Disambiguation                      | 词义消歧            | SemEval 2015                         | F1                                         | 67.1                                       | [Word Sense Disambiguation: A Unified Evaluation Framework and Empirical Comparison](https://aclanthology.org/E17-1010.pdf) | Link                                                         |\n| Text classification                            | 文本分类            | AG News                              | Error rate                                 | 5.01                                       | [Universal Language Model Fine-tuning for Text Classification](https://arxiv.org/pdf/1801.06146v5.pdf) | [Link](https://github.com/fastai/fastai)                     |\n| Summarization                                  | 摘要                | Gigaword                             | ROUGE-1\u003cbr/\u003eROUGE-2\u003cbr/\u003eROUGE-L            | 37.04\u003cbr/\u003e19.03\u003cbr/\u003e34.46                  | [Retrieve, Rerank and Rewrite: Soft Template Based Neural Summarization](https://aclanthology.org/P18-1015.pdf) | Link                                                         |\n| Sentiment analysis                             | 情感分析            | IMDb                                 | Accuracy                                   | 95.4                                       | [Universal Language Model Fine-tuning for Text Classification](https://arxiv.org/pdf/1801.06146v5.pdf) | [Link](https://github.com/fastai/fastai)                     |\n| Semantic role labeling                         | 语义角色标注        | OntoNotes                            | F1                                         | 85.5                                       | [Jointly Predicting Predicates and Arguments in Neural Semantic Role Labeling](https://arxiv.org/pdf/1805.04787v2.pdf) | [Link](https://github.com/luheng/lsgn)                       |\n| Semantic parsing                               | 语义解析            | LDC2014T12                           | F1 Newswire\u003cbr/\u003eF1 Full                    | 0.71\u003cbr/\u003e0.66                              | [AMR Parsing with an Incremental Joint Model](https://arxiv.org/pdf/1909.04303v2.pdf) | [Link](https://github.com/jcyk/AMR-parser)                   |\n| Semantic textual similarity                    | 语义文本相似度      | SentEval                             | MRPC\u003cbr/\u003eSICK-R\u003cbr/\u003eSICK-E\u003cbr/\u003eSTS         | 78.6/84.4\u003cbr/\u003e0.888\u003cbr/\u003e87.8\u003cbr/\u003e78.9/78.6 | [Learning General Purpose Distributed Sentence Representations via Large Scale Multi-task Learning](https://arxiv.org/pdf/1804.00079v1.pdf) | [Link](https://github.com/facebookresearch/SentEval)         |\n| Relationship Extraction                        | 关系抽取            | New York Times Corpus                | P@10%\u003cbr/\u003eP@30%                            | 73.6\u003cbr/\u003e59.5                              | [RESIDE: Improving Distantly-Supervised Neural Relation Extraction using Side Information](https://arxiv.org/pdf/1812.04361v2.pdf) | [Link](https://github.com/malllabiisc/RESIDE)                |\n| Relation Prediction                            | 关系预测            | WN18RR                               | H@10\u003cbr/\u003eH@1\u003cbr/\u003eMRR                       | 59.02\u003cbr/\u003e45.37\u003cbr/\u003e49.83                  | [Predicting Semantic Relations using Global Graph Properties](https://arxiv.org/pdf/1808.08644v1.pdf) | [Link](https://github.com/yuvalpinter/m3gm)                  |\n\n\n\n## Resource\n\n- [NLP-Interview-Notes](https://github.com/km1994/NLP-Interview-Notes)\n- [Recommendation-Advertisement-Search](https://github.com/km1994/recommendation_advertisement_search)\n- [NLPer-Arsenal](https://github.com/TingFree/NLPer-Arsenal)\n- [AI-Surveys](https://github.com/KaiyuanGao/AI-Surveys)\n\n\n\n## Interview\n\n- [Machine Learning](./docs/interview/machine-learning.md)\n- [Deep Learning](./docs/interview/deep-learning.md)\n- [Word Embedding](./docs/interview/word-embedding.md)\n- [Transformer](./docs/interview/transformer.md)\n- [Bert](./docs/interview/bert.md)\n- [Reverse](./docs/interview/reverse-interview.md)\n\n\n\n© [JackHCC](https://github.com/JackHCC)\n\n\n\n","projects_url":"https://awesome.ecosyste.ms/api/v1/lists/jackhcc%2Fnlp-bubble/projects"}