{"id":16271307,"url":"https://github.com/ikergarcia1996/metavec","last_synced_at":"2025-09-17T20:31:39.990Z","repository":{"id":53892414,"uuid":"402019059","full_name":"ikergarcia1996/MetaVec","owner":"ikergarcia1996","description":"A monolingual and cross-lingual meta-embedding generation and evaluation framework","archived":false,"fork":false,"pushed_at":"2022-04-29T15:37:23.000Z","size":71,"stargazers_count":80,"open_issues_count":1,"forks_count":5,"subscribers_count":3,"default_branch":"main","last_synced_at":"2025-04-01T21:48:09.882Z","etag":null,"topics":["embedding","embedding-evaluation","embedding-models","embedding-vectors","embeddings","emnlp2021","fasttext","fasttext-embeddings","meta-embedding","meta-embeddings","word2vec"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"gpl-3.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/ikergarcia1996.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null}},"created_at":"2021-09-01T10:23:18.000Z","updated_at":"2024-10-09T01:34:41.000Z","dependencies_parsed_at":"2022-08-13T03:30:44.145Z","dependency_job_id":null,"html_url":"https://github.com/ikergarcia1996/MetaVec","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/ikergarcia1996/MetaVec","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ikergarcia1996%2FMetaVec","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ikergarcia1996%2FMetaVec/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ikergarcia1996%2FMetaVec/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ikergarcia1996%2FMetaVec/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/ikergarcia1996","download_url":"https://codeload.github.com/ikergarcia1996/MetaVec/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ikergarcia1996%2FMetaVec/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":275658692,"owners_count":25504776,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-09-17T02:00:09.119Z","response_time":84,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["embedding","embedding-evaluation","embedding-models","embedding-vectors","embeddings","emnlp2021","fasttext","fasttext-embeddings","meta-embedding","meta-embeddings","word2vec"],"created_at":"2024-10-10T18:13:15.109Z","updated_at":"2025-09-17T20:31:39.689Z","avatar_url":"https://github.com/ikergarcia1996.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# MetaVec: The Best Word Embeddings To Date\nMetaVec is a monolingual and cross-lingual meta-embedding generation framework.  \nMetaVec outperforms every previously proposed meta-embedding generation method. Our best-meta embedding achieves the best-published results in a wide range of intrinsic evaluation tasks. \n\n![Intrinsic Evaluation Average Results](Results.png \"MetaVec Results\")\n____\n\n# Download MetaVec\n\nYou can download our pre-computed meta-embedding. \nThis meta embeddings combines [FastText](https://fasttext.cc/docs/en/english-vectors.html), [Numberbatch](https://github.com/commonsense/conceptnet-numberbatch), [JOINTChyb](http://ixa2.si.ehu.es/ukb/bilingual_embeddings.html) and [Paragram](https://github.com/jwieting/paragram-word).\n\n| Click to Download                                                   | Words     | Dimensions | Size   | Link                                                     |\n|---------------------------------------------------------------------|-----------|------------|--------|----------------------------------------------------------|\n| [MetaVec](https://adimen.si.ehu.es/~igarcia/embeddings/MetaVec.zip) | 4,573,185 | 300        | 11.8GB | https://adimen.si.ehu.es/~igarcia/embeddings/MetaVec.zip |\n\n### Reduced Vocabulary versions\nMetaVec is a very large embedding. We provide reduced vocabulary versions of the Meta-Embedding. The vector representations are the same,\nbut we include only a subset of the words in the vocabulary.\n\n| Click to Download                                                               | Words     | Dimensions | Size   |                                                                                                                                                   |\n|---------------------------------------------------------------------------------|-----------|------------|--------|---------------------------------------------------------------------------------------------------------------------------------------------------|\n| [MetaVec 2M]( https://adimen.si.ehu.es/~igarcia/embeddings/MetaVec_2M.zip)      | 1,999,995 | 300        | 4.8GB  | Only words in [crawl-300d-2M.vec](https://fasttext.cc/docs/en/english-vectors.html) (FastText Common Crawl) vocabulary                            |\n| [MetaVec 1M](https://adimen.si.ehu.es/~igarcia/embeddings/MetaVec_1M.zip)       | 830,063   | 300        | 2GB    | Only words in  [wiki-news-300d-1M.vec](https://fasttext.cc/docs/en/english-vectors.html)  (FastText Wikipedia) vocabulary                         |\n| [MetaVec 0.2M](https://adimen.si.ehu.es/~igarcia/embeddings/MetaVec_200000.zip) | 186,647   | 300        | 0.45GB | Only words in the 200,000 Most Common English Words List according to the  [Google's Trillion Words Corpus](https://books.google.com/ngrams/info) |\n\n# Citation\n```\n@inproceedings{garcia-ferrero-etal-2021-benchmarking-meta,\n    title = \"Benchmarking Meta-embeddings: What Works and What Does Not\",\n    author = \"Garc{\\'\\i}a-Ferrero, Iker  and\n      Agerri, Rodrigo  and\n      Rigau, German\",\n    booktitle = \"Findings of the Association for Computational Linguistics: EMNLP 2021\",\n    month = nov,\n    year = \"2021\",\n    address = \"Punta Cana, Dominican Republic\",\n    publisher = \"Association for Computational Linguistics\",\n    url = \"https://aclanthology.org/2021.findings-emnlp.333\",\n    pages = \"3957--3972\",\n```\n\n____\n# Usage\n## Installation \n\n### Generate meta-embeddings\nIf you only want to generate meta-embeddings you need:\n- Python 3 (3.7 or greater)\n- Numpy\n- tqdm\n- Cupy, if you have an Nvidia GPU you can install cupy to use the GPU instead of the CPU for faster matrix operations: https://docs.cupy.dev/en/stable/install.html\n\n### Evaluation framework\nIf you want to run the evaluation framework we have prepared a conda environment file that will install all the required dependencies. \nThis file will create a conda environment called \"metavec\" with the required dependencies to generate meta-embeddings, run the intrinsic and run the extrinsic evaluation framework. \n```\nconda env create -f environment.yml\nconda activate metavec\n```\n\nIf you don't want to use conda you can inspect the \"environment.yml\" file to see the required dependencies and manually install them.\n\n### Get third party requeriments\nIf you just want to generate meta-embeddings you only need to clone the [VecMap repository](https://github.com/artetxem/vecmap) inside the MetaVec directory. The following command will do that for you:\n```\nsh get_third_party.sh\n```\n\nIf you want to generate meta-embedding and evaluate them you need to clone the [Vecmap](https://github.com/artetxem/vecmap), [word embedding benchmarks](https://github.com/kudkudak/word-embeddings-benchmarks) and [Jiant](https://github.com/nyu-mll/jiant-v1-legacy) repositories. For the last two ones, we will download a modified version to run the same configuration used in our paper. The script will also install and download the nltk and spacy required packages for evaluation. To run this command you first need to install the [evaluation framework dependencies](#evaluation-framework).\n```\nsh get_third_party.sh all\n```\n\n## Generate a Meta-Embedding\n\nHere is an example command to generate a meta-embedding using FastText, JointcHYB, paragram and numberbatch as source embeddings.   \n- embeddings: Path of the source embeddings you want to combine to generate a meta-embedding\n- rotate_to: Path to the embedding to which all source embedding will be aligned using VecMap, it doesn't need to be one of the source embeddings. \n- output_path: Path where the generated meta-embedding will be saved\n```\npython run_metavec.py \\\n--embeddings embeddings/crawl-300d-2M.vec embeddings/JOINTC-HYB-ENES.emb embeddings/paragram_ws353.vec embeddings/numberbatch-en.txt \\\n--rotate_to embeddings/crawl-300d-2M.vec \\\n--output_path embeddings/MetaVec.vec\n```\n\nSee [embeddings/README.md](embeddings/README.md) for instructions on how to download the source embeddings that we test in our paper.\n\n## Run the Unified Evaluation Framework\n\n### Intrinsic Evaluation\nThe intrinsic evaluation uses the Word Embedding Benchmark toolkit: https://github.com/kudkudak/word-embeddings-benchmarks\n\n\nEvaluate a Word Embedding\n```\npython instrinsic_evaluation.py -i embeddings/MetaVec.vec\n```\n\nEvaluate all the Word Embeddings in a directory\n\n```\npython instrinsic_evaluation.py -d embeddings/\n```\n\nIf you want to set a custom output directory for the evaluation results use the \"--output_dir path\" argument.\n\n## Extrinsic Evaluation (GLUE)\nThe extrinsic evaluation uses the Jiant-V1 toolkit: https://github.com/nyu-mll/jiant-v1-legacy\n\nEvaluate a Word Embedding\n```\npython extrinsic_evaluation.py -i embeddings/MetaVec.vec\n```\n\nEvaluate all the Word Embeddings in a directory\n\n```\npython extrinsic_evaluation.py -d embeddings/\n```\n\nIf you want to set a custom output directory for the evaluation results use the \"--output_dir path\" argument.\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fikergarcia1996%2Fmetavec","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fikergarcia1996%2Fmetavec","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fikergarcia1996%2Fmetavec/lists"}