{"id":13401293,"url":"https://github.com/taylor-arnold/cleanNLP","last_synced_at":"2025-03-14T07:31:30.692Z","repository":{"id":37663444,"uuid":"70943553","full_name":"taylor-arnold/cleanNLP","owner":"taylor-arnold","description":"R package providing annotators and a normalized data model for natural language processing","archived":false,"fork":false,"pushed_at":"2024-05-20T19:07:15.000Z","size":11727,"stargazers_count":212,"open_issues_count":2,"forks_count":36,"subscribers_count":18,"default_branch":"master","last_synced_at":"2024-12-23T00:39:02.704Z","etag":null,"topics":["corenlp","natural-language-processing","r-package","spacy"],"latest_commit_sha":null,"homepage":"","language":"R","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"lgpl-2.1","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/taylor-arnold.png","metadata":{"files":{"readme":"README.md","changelog":"NEWS.md","contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2016-10-14T20:06:13.000Z","updated_at":"2024-11-08T21:03:34.000Z","dependencies_parsed_at":"2024-09-30T05:40:54.776Z","dependency_job_id":"aca3e4ac-77e3-4a30-93fd-4f17a244b066","html_url":"https://github.com/taylor-arnold/cleanNLP","commit_stats":{"total_commits":137,"total_committers":14,"mean_commits":9.785714285714286,"dds":0.3941605839416058,"last_synced_commit":"0e6bf7d8f618b62a875a305fe0a038c6564b4641"},"previous_names":["taylor-arnold/cleannlp","statsmaths/cleannlp"],"tags_count":2,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/taylor-arnold%2FcleanNLP","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/taylor-arnold%2FcleanNLP/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/taylor-arnold%2FcleanNLP/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/taylor-arnold%2FcleanNLP/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/taylor-arnold","download_url":"https://codeload.github.com/taylor-arnold/cleanNLP/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":243541863,"owners_count":20307772,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["corenlp","natural-language-processing","r-package","spacy"],"created_at":"2024-07-30T19:01:01.077Z","updated_at":"2025-03-14T07:31:29.362Z","avatar_url":"https://github.com/taylor-arnold.png","language":"R","funding_links":[],"categories":["R"],"sub_categories":[],"readme":"## cleanNLP: A Tidy Data Model for Natural Language Processing\n\n**Author:** Taylor B. Arnold\u003cbr/\u003e\n**License:** [LGPL-2](https://opensource.org/licenses/LGPL-2.1)\n\n[![CRAN Version](http://www.r-pkg.org/badges/version-ago/cleanNLP)](https://CRAN.R-project.org/package=cleanNLP) \n\n## Overview\n\nThe **cleanNLP** package is designed to make it as painless as possible\nto turn raw text into feature-rich data frames. A minimal working example\nof using **cleanNLP** consists of loading the package, setting up the NLP\nbackend, initializing the backend, and running the function `cnlp_annotate`.\nThe output is given as a list of data frame objects (classed as an\n\"cnlp_annotation\"). Here is an example using the udpipe backend:\n\n```{r}\nlibrary(cleanNLP)\ncnlp_init_udpipe()\n\nannotation \u003c- cnlp_annotate(input = c(\n        \"Here is the first text. It is short.\",\n        \"Here's the second. It is short too!\",\n        \"The third text is the shortest.\"\n))\nlapply(annotation, head)\n```\n```\n$token\n  doc_id sid tid token token_with_ws lemma  upos xpos\n1      1   1   1  Here         Here   here   ADV   RB\n2      1   1   2    is           is     be   AUX  VBZ\n3      1   1   3   the          the    the   DET   DT\n4      1   1   4 first        first  first   ADJ   JJ\n5      1   1   5  text          text  text  NOUN   NN\n6      1   1   6     .            .      . PUNCT    .\n                                                  feats tid_source relation\n1                                          PronType=Dem          0     root\n2 Mood=Ind|Number=Sing|Person=3|Tense=Pres|VerbForm=Fin          1      cop\n3                             Definite=Def|PronType=Art          5      det\n4                                Degree=Pos|NumType=Ord          5     amod\n5                                           Number=Sing          1    nsubj\n6                                                  \u003cNA\u003e          1    punct\n\n$document\n  doc_id\n1      1\n2      2\n3      3\n```\n\nThe `token` output table breaks the text into tokens, provides lemmatized\nforms of the words, part of speech tags, and dependency relationships. Two\nshort case-studies are linked to from the repository to show sample usage of\nthe library:\n\n- [State of the Union Addresses](https://statsmaths.github.io/cleanNLP/state-of-union.html)\n- [Exploring Wikipedia Data](https://statsmaths.github.io/cleanNLP/wikipedia.html)\n\nPlease see the notes below, and the official package documentation on\n[CRAN](https://cran.r-project.org/web/packages/cleanNLP/), for more options\nto control the way that text is parsed.\n\n## Installation\n\nYou can download the package from within R directly from CRAN:\n\n```{r}\ninstall.packages(\"cleanNLP\")\n```\n\nAfter installation, you should be able to use the udpipe backend (as used\nthe minimal example and case-studies above; model files will be installed\nautomatically) or the stringi backend without any additional setup. **For most\nusers, we find that these out-of-the-box solutions are a good starting point.**\nIn order to use the two Python backends, you must install the associated\n`cleannlp` python module. We recommend and support the Python 3.7 version of\n[Anaconda Python](https://www.anaconda.com/distribution/#download-section).\nAfter obtaining Python, install the module by running pip in a terminal:\n\n```{py}\npip install cleannlp\n```\n\nOnce installed, running the respective backend initialization functions will\nprovide further instructions for download the required models.\n\n## API Overview\n\n### V3\n\nThere have been numerous changes to the package in the newly released version 3.0.0.\nThese changes, while requiring some changes to existing code, have been carefully\ndesigned to make the package easier to both install and use. The three most important\nchanges include:\n\n- The object returned by `cnlp_annotate` is now a named list. Users can access its\nelements with the dollar sign operator. Functions such as `cnlp_get_token`\nand `cnlp_get_dependency` are no longer needed or included.\n- The dependencies are now attached to the tokens table to make them easier to use.\n\nIf you are running into any issues with the package, first make sure you are using\nupdated materials (mostly available from links within this repository).\n\n### Backends\n\nThe cleanNLP package is designed to allow users to make use of various NLP\nannotation algorithms without having to worry (too much) about the output\nformat, which is standardizes at best as possible. There are three backends\ncurrently available, each with their own pros and cons. They are:\n\n- **stringi**: a fast parser that only requires the stringi package,\nbut produces only tokenized text\n- **udpipe**: a parser with no external dependencies that produces\ntokens, lemmas, part of speech tags, and dependency relationships. The\nrecommended starting point given its balance between ease of use and\nfunctionality. It also supports the widest range of natural languages.\n- **spacy**: based on the Python library, a more feature complete parser\nthat included named entity recognition and word embeddings. It does require\na working Python installation and some other set-up. Recommended for users\nwho are familiar with Python or plan to make heavy use of the package.\n\nThe final backend (spacy) require some additional setup,\nnamely installing Python and the associated Python library, as documented above.\nTo select the desired backend, simply initialize the model prior to running the\nannotation.\n\n```{r}\ncnlp_init_stringi(locale=\"en_GB\")\ncnlp_init_udpipe(model_name=\"english\")\ncnlp_init_spacy(model_name=\"en\")\n```\n\nThe code above explicitly sets the default/English model. You can use a\ndifferent model/language when starting the model. For udpipe the models will\nbe downloaded automatically. For spacy the following helper\nfunctions are available:\n\n```{r}\ncnlp_download_spacy(model_name=\"en\")\n```\n\nSimply change the model name or language code to download alternative models.\n\n## Citation\n\nIf you make use of the toolkit in your work, please cite the following paper.\n\n```\n@article{,\n  title   = \"A Tidy Data Model for Natural Language Processing Using cleanNLP\",\n  author  = \"Arnold, Taylor B\",\n  journal = \"R Journal\",\n  volume  = \"9\",\n  number  = \"2\",\n  year    = \"2017\"\n}\n```\n\nPlease, however, note that the library has evolved since the paper was published.\nFor specific help with the package's API please check the updated documents\nlinked to from this site.\n\n## Note\n\nPlease note that this project is released with a\n[Contributor Code of Conduct](CONDUCT.md). By participating in this project\nyou agree to abide by its terms.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Ftaylor-arnold%2FcleanNLP","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Ftaylor-arnold%2FcleanNLP","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Ftaylor-arnold%2FcleanNLP/lists"}