{"id":13800952,"url":"https://github.com/IBM/MAX-Word-Embedding-Generator","last_synced_at":"2025-05-13T10:30:40.400Z","repository":{"id":66065079,"uuid":"140601608","full_name":"IBM/MAX-Word-Embedding-Generator","owner":"IBM","description":"Generate embedding vectors from text files","archived":true,"fork":false,"pushed_at":"2020-04-07T14:17:15.000Z","size":41,"stargazers_count":8,"open_issues_count":2,"forks_count":17,"subscribers_count":23,"default_branch":"master","last_synced_at":"2024-08-04T00:05:39.935Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/IBM.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null}},"created_at":"2018-07-11T16:25:34.000Z","updated_at":"2024-07-22T19:15:45.000Z","dependencies_parsed_at":"2023-02-20T19:00:57.960Z","dependency_job_id":null,"html_url":"https://github.com/IBM/MAX-Word-Embedding-Generator","commit_stats":null,"previous_names":[],"tags_count":1,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/IBM%2FMAX-Word-Embedding-Generator","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/IBM%2FMAX-Word-Embedding-Generator/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/IBM%2FMAX-Word-Embedding-Generator/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/IBM%2FMAX-Word-Embedding-Generator/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/IBM","download_url":"https://codeload.github.com/IBM/MAX-Word-Embedding-Generator/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":225198945,"owners_count":17437003,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-08-04T00:01:17.948Z","updated_at":"2024-11-18T15:31:23.476Z","avatar_url":"https://github.com/IBM.png","language":"Python","funding_links":[],"categories":["Data \u0026 AI"],"sub_categories":[],"readme":"# IBM Code Model Asset Exchange: Word Embedding Generator\n\nThis repository contains code to generate word embeddings using the Swivel algorithm on [IBM Watson Machine Learning](https://www.ibm.com/cloud/machine-learning). This model is part of the [IBM Code Model Asset Exchange](https://developer.ibm.com/code/exchanges/models/).\n\nMachine learning algorithms usually expect numeric inputs. When a data scientist wants to use text to create a machine learning model, they must first find a way to represent their text as a vector of numbers. These vectors are called word embeddings. The Swivel algorithm is a frequency-based word embedding that uses a co-occurence matrix. The idea here is that words that have similar meanings tend to occur together in a text corpus. As a result, words that have similar meanings will have vector representations that are closer than those of unrelated words.\n\nThis demo contains scripts to run the Swivel algorithm on a preprocessed Wikipedia text corpus.\nFor instructions on generating word embeddings on your own text corpus see the instructions in the\n[original repository here](https://github.com/tensorflow/models/tree/master/research/swivel).\n\n## Model Metadata\n| Domain | Application | Industry  | Framework | Training Data | Input Data Format |\n| ------------- | --------  | -------- | --------- | --------- | -------------- |\n| Text/NLP | Natural Language | General | TensorFlow | [Any Text Corpus (e.g. Wiki Dump)](https://dumps.wikimedia.org/backup-index.html) | Text |\n\n# References #\n[1]\u003ca name=\"ref1\"\u003e\u003c/a\u003e N. Shazeer, R. Doherty, C. Evans, C. Waterson., [\"Swivel: Improving Embeddings\nby Noticing What's Missing\"](https://arxiv.org/pdf/1602.02215.pdf) arXiv preprint arXiv:1602.02215 (2016)\n\n## Licenses\n\n| Component | License | Link  |\n| ------------- | --------  | -------- |\n| This repository | [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0) | [LICENSE](LICENSE) |\n| Model Code (3rd party) | [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0) | [TensorFlow Models](https://github.com/tensorflow/models/blob/master/LICENSE)|\n|Data|[CC BY-SA 3.0](https://en.wikipedia.org/wiki/Wikipedia:Copyrights)|[Wikipedia Text Dump](https://dumps.wikimedia.org/backup-index.html)|\n\n# Quickstart\n\n## Prerequisites\n\n* This experiment requires a provisioned instance of IBM Watson Machine Learning service.\n\n### Setup an IBM Cloud Object Storage (COS) account\n- Create an IBM Cloud Object Storage account if you don't have one (https://www.ibm.com/cloud/storage)\n- Create credentials for either reading and writing or just reading\n\t- From the bluemix console page (https://console.bluemix.net/dashboard/apps/), choose `Cloud Object Storage`\n\t- On the left side, click the `service credentials`\n\t- Click on the `new credentials` button to create new credentials\n\t- In the `Add New Credentials` popup, use this parameter `{\"HMAC\":true}` in the `Add Inline Configuration...`\n\t- When you create the credentials, copy the `access_key_id` and `secret_access_key` values.\n\t- Make a note of the endpoint url\n\t\t- On the left side of the window, click on `Endpoint`\n\t\t- Copy the relevant public or private endpoint. [I choose the us-geo private endpoint].\n- In addition setup your [AWS S3 command line](https://aws.amazon.com/cli/) which can be used to create buckets and/or add files to COS.\n   - Export `AWS_ACCESS_KEY_ID` with your COS `access_key_id` and `AWS_SECRET_ACCESS_KEY` with your COS `secret_access_key`\n\n### Setup IBM CLI \u0026 ML CLI\n\n- Install [IBM Cloud CLI](https://console.bluemix.net/docs/cli/reference/ibmcloud/download_cli.html#install_use)\n  - Login using `bx login` or `bx login --sso` if within IBM\n- Install [ML CLI Plugin](https://dataplatform.ibm.com/docs/content/analyze-data/ml_dlaas_environment.html)\n  - After install, check if there is any plugins that need update\n    - `bx plugin update`\n  - Make sure to setup the various environment variables correctly:\n    - `ML_INSTANCE`, `ML_USERNAME`, `ML_PASSWORD`, `ML_ENV`\n\n## Training the model\n\nThe `train.sh` utility script will deploy the experiment to WML and start the training as a `training-run`\n\n```\ntrain.sh\n```\n\nAfter the train is started, it should print the training-id that is going to be necessary for steps below\n\n```\nStarting to train ...\nOK\nModel-ID is 'training-GCtN_YRig'\n```\n\n### Monitor the  training run\n\n- To list the training runs - `bx ml list training-runs`\n- To monitor a specific training run - `bx ml show training-runs \u003ctraining-id\u003e`\n- To monitor the output (stdout) from the training run - `bx ml monitor training-runs \u003ctraining-id\u003e`\n\t- This will print the first couple of lines, and may time out.\n\n## Exploring the embeddings\nThe `demo.sh` utility script will download the results from the bucket, convert the embeddings into binary vector format, and run a python application\nto explore the embeddings:\n```\ndemo.sh\n```\n\nWhen querying a single word, the results will list words that are similar in meaning.\n```\nquery\u003e dog\ndog\ndogs\ncat\n```\n\nIt is also possible to query to complete an analogy. (e.g. A _man_ is to a _woman_ as a _king_ is to... )\n```\nquery\u003e man woman king\nking\nqueen\nprincess\n```\n\n## Resources and Contributions\n   \nIf you are interested in contributing to the Model Asset Exchange project or have any queries, please follow the instructions [here](https://github.com/CODAIT/max-central-repo).","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FIBM%2FMAX-Word-Embedding-Generator","html_url":"https://awesome.ecosyste.ms/projects/github.com%2FIBM%2FMAX-Word-Embedding-Generator","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FIBM%2FMAX-Word-Embedding-Generator/lists"}