{"id":20949020,"url":"https://github.com/intuitionengineeringteam/chars2vec","last_synced_at":"2025-04-05T17:07:17.868Z","repository":{"id":38712195,"uuid":"172519084","full_name":"IntuitionEngineeringTeam/chars2vec","owner":"IntuitionEngineeringTeam","description":"Character-based word embeddings model based on RNN for handling real world texts","archived":false,"fork":false,"pushed_at":"2023-10-09T10:33:00.000Z","size":7939,"stargazers_count":174,"open_issues_count":12,"forks_count":37,"subscribers_count":4,"default_branch":"master","last_synced_at":"2025-03-29T16:06:51.936Z","etag":null,"topics":["embeddings","language-model","natural-language-processing","natural-language-understanding"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/IntuitionEngineeringTeam.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2019-02-25T14:15:48.000Z","updated_at":"2025-03-26T01:54:51.000Z","dependencies_parsed_at":"2024-06-19T04:12:41.826Z","dependency_job_id":null,"html_url":"https://github.com/IntuitionEngineeringTeam/chars2vec","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/IntuitionEngineeringTeam%2Fchars2vec","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/IntuitionEngineeringTeam%2Fchars2vec/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/IntuitionEngineeringTeam%2Fchars2vec/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/IntuitionEngineeringTeam%2Fchars2vec/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/IntuitionEngineeringTeam","download_url":"https://codeload.github.com/IntuitionEngineeringTeam/chars2vec/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":247369952,"owners_count":20927928,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["embeddings","language-model","natural-language-processing","natural-language-understanding"],"created_at":"2024-11-19T00:29:15.909Z","updated_at":"2025-04-05T17:07:17.849Z","avatar_url":"https://github.com/IntuitionEngineeringTeam.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# chars2vec\n\n#### Character-based word embeddings model based on RNN\n\n\nChars2vec library could be very useful if you are dealing with the texts \ncontaining abbreviations, slang, typos, or some other specific textual dataset. \nChars2vec language model is based on the symbolic representation of words – \nthe model maps each word to a vector of a fixed length. \nThese vector representations are obtained with a custom neural network while \nthe latter is being trained on pairs of similar and non-similar words. \nThis custom neural net includes LSTM, reading sequences of characters in words, as its part. \nThe model maps similarly written words to proximal vectors. \nThis approach enables creation of an embedding in vector space for any sequence of characters. \nChars2vec models does not keep any dictionary of embeddings, \nbut generates embedding vectors inplace using pretrained model. \n\nThere are pretrained models of dimensions 50, 100, 150, 200 and 300 for the English language.\nThe library provides convenient user API to train a model for an arbitrary set of characters. \nRead more details about the architecture of [Chars2vec: \nCharacter-based language model for handling real world texts with spelling \nerrors and human slang](https://hackernoon.com/chars2vec-character-based-language-model-for-handling-real-world-texts-with-spelling-errors-and-a3e4053a147d) in Hacker Noon.\n\n#### Model available for Python 2.7 and 3.0+.\n\n### Installation\n\n\u003ch5\u003e 1. Build and install from source \u003c/h5\u003e\nDownload project source and run in your command line\n\n~~~shell\n\u003e\u003e python setup.py install\n~~~\n\n\u003ch5\u003e 2. Via pip \u003c/h5\u003e\nRun in your command line\n\n~~~shell\n\u003e\u003e pip install chars2vec\n~~~\n\n### Usage\n\nFunction `chars2vec.load_model(str path)` initializes the model from directory \nand returns `chars2vec.Chars2Vec` object.\nThere are 5 pretrained English model with dimensions: 50, 100, 150, 200 and 300.\nTo load this pretrained models:\n\n~~~python\nimport chars2vec\n\n# Load Inutition Engineering pretrained model\n# Models names: 'eng_50', 'eng_100', 'eng_150', 'eng_200', 'eng_300'\nc2v_model = chars2vec.load_model('eng_50')\n~~~ \nMethod `chars2vec.Chars2Vec.vectorize_words(words)` returns `numpy.ndarray` of shape `(n_words, dim)` with word embeddings.\n\n~~~python\nwords = ['list', 'of', 'words']\n\n# Create word embeddings\nword_embeddings = c2v_model.vectorize_words(words)\n~~~\n\n### Training\n\nFunction `chars2vec.train_model(int emb_dim, X_train, y_train, model_chars)` \ncreates and trains new chars2vec model and returns `chars2vec.Chars2Vec` object.\n\nParameter `emb_dim` is a dimension of the model. \n\nParameter `X_train` is a list or numpy.ndarray of word pairs.\nParameter `y_train` is a list or numpy.ndarray of target values that describe the proximity of words.\n\nTraining set (`X_train`, `y_train`) consists of pairs of \"similar\" and \"not similar\" words; \na pair of \"similar\" words is labeled with 0 target value, and a pair of \"not similar\" with 1. \n\nParameter `model_chars` is a list of chars for the model.\nCharacters which are not in the `model_chars`\nlist will be ignored by the model. \n\nRead more about chars2vec training and generation of training dataset in \n[article about chars2vec](https://hackernoon.com/chars2vec-character-based-language-model-for-handling-real-world-texts-with-spelling-errors-and-a3e4053a147d).\n\nFunction `chars2vec.save_model(c2v_model, str path_to_model)` saves the trained model \nto the directory.\n\n\n~~~python\nimport chars2vec\n\ndim = 50\npath_to_model = 'path/to/model/directory'\n\nX_train = [('mecbanizing', 'mechanizing'), # similar words, target is equal 0\n           ('dicovery', 'dis7overy'), # similar words, target is equal 0\n           ('prot$oplasmatic', 'prtoplasmatic'), # similar words, target is equal 0\n           ('copulateng', 'lzateful'), # not similar words, target is equal 1\n           ('estry', 'evadin6'), # not similar words, target is equal 1\n           ('cirrfosis', 'afear') # not similar words, target is equal 1\n          ]\n\ny_train = [0, 0, 0, 1, 1, 1]\n\nmodel_chars = ['!', '\"', '#', '$', '%', '\u0026', \"'\", '(', ')', '*', '+', ',', '-', '.',\n               '/', '0', '1', '2', '3', '4', '5', '6', '7', '8', '9', ':', ';', '\u003c',\n               '=', '\u003e', '?', '@', '_', 'a', 'b', 'c', 'd', 'e', 'f', 'g', 'h', 'i',\n               'j', 'k', 'l', 'm', 'n', 'o', 'p', 'q', 'r', 's', 't', 'u', 'v', 'w',\n               'x', 'y', 'z']\n\n# Create and train chars2vec model using given training data\nmy_c2v_model = chars2vec.train_model(dim, X_train, y_train, model_chars)\n\n# Save your pretrained model\nchars2vec.save_model(my_c2v_model, path_to_model)\n\n# Load your pretrained model \nc2v_model = chars2vec.load_model(path_to_model)\n~~~\n\nFull code examples for usage and training models see in\n`example_usage.py` and `example_training.py` files.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fintuitionengineeringteam%2Fchars2vec","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fintuitionengineeringteam%2Fchars2vec","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fintuitionengineeringteam%2Fchars2vec/lists"}