{"id":20456546,"url":"https://github.com/krassowski/easy-entrez","last_synced_at":"2025-09-21T20:31:39.629Z","repository":{"id":44706625,"uuid":"272182307","full_name":"krassowski/easy-entrez","owner":"krassowski","description":"Retrieve PubMed articles, text-mining annotations, or molecular data from \u003e35 Entrez databases via easy to use Python package - built on top of Entrez E-utilities API.","archived":false,"fork":false,"pushed_at":"2023-11-02T21:59:16.000Z","size":123,"stargazers_count":72,"open_issues_count":7,"forks_count":6,"subscribers_count":4,"default_branch":"main","last_synced_at":"2025-01-03T09:04:54.864Z","etag":null,"topics":["entrez","entrez-eutilities","eutilities","gene-annotations","literature-mining","literature-search","meta-analysis","pubmed","pubmed-central"],"latest_commit_sha":null,"homepage":"https://easy-entrez.readthedocs.io/en/latest/","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"lgpl-3.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/krassowski.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null}},"created_at":"2020-06-14T10:48:26.000Z","updated_at":"2024-12-30T23:05:36.000Z","dependencies_parsed_at":"2023-02-12T10:55:13.278Z","dependency_job_id":"7d0b2d8b-d1fa-4165-a722-e6352b9d1ed1","html_url":"https://github.com/krassowski/easy-entrez","commit_stats":{"total_commits":83,"total_committers":2,"mean_commits":41.5,"dds":"0.37349397590361444","last_synced_commit":"b7d1f12532350a1608dff6ac66d7c7ce4675b6e3"},"previous_names":[],"tags_count":6,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/krassowski%2Feasy-entrez","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/krassowski%2Feasy-entrez/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/krassowski%2Feasy-entrez/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/krassowski%2Feasy-entrez/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/krassowski","download_url":"https://codeload.github.com/krassowski/easy-entrez/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":233791071,"owners_count":18730768,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["entrez","entrez-eutilities","eutilities","gene-annotations","literature-mining","literature-search","meta-analysis","pubmed","pubmed-central"],"created_at":"2024-11-15T11:23:02.253Z","updated_at":"2025-09-21T20:31:34.349Z","avatar_url":"https://github.com/krassowski.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# easy-entrez\n\n![Tests](https://github.com/krassowski/easy-entrez/workflows/tests/badge.svg)\n![CodeQL](https://github.com/krassowski/easy-entrez/workflows/CodeQL/badge.svg)\n[![Documentation Status](https://readthedocs.org/projects/easy-entrez/badge/?version=latest)](https://easy-entrez.readthedocs.io/en/latest/?badge=latest)\n[![DOI](https://zenodo.org/badge/272182307.svg)](https://zenodo.org/badge/latestdoi/272182307)\n![Python](https://img.shields.io/badge/python-3.7%20%7C%203.8%20%7C%203.9%20%7C%203.10%20%7C%203.11%20%7C%203.12-blue)\n\nPython REST API for Entrez E-Utilities, aiming to  be easy to use and reliable.\n\nEasy-entrez:\n\n- makes common tasks easy thanks to simple Pythonic API,\n- is typed and integrates well with mypy,\n- is tested on Windows, Mac and Linux across Python 3.7 to 3.12,\n- is limited in scope, allowing to focus on the reliability of the core code,\n- does not use the stateful API as it is [error-prone](https://gitlab.com/ncbipy/entrezpy/-/issues/7) as seen on example of the alternative *entrezpy*.\n\n### Examples\n\n```python\nfrom easy_entrez import EntrezAPI\n\nentrez_api = EntrezAPI(\n    'your-tool-name',\n    'e@mail.com',\n    # optional\n    return_type='json'\n)\n\n# find up to 10 000 results for cancer in human\nresult = entrez_api.search('cancer AND human[organism]', max_results=10_000)\n\n# data will be populated with JSON or XML (depending on the `return_type` value)\nresult.data\n```\n\nSee more in the [Demo notebook](./Demo.ipynb) and [documentation](https://easy-entrez.readthedocs.io/en/latest).\n\nFor a real-world example (i.e. used for [this publication](https://www.frontiersin.org/articles/10.3389/fgene.2020.610798/full)) see notebooks in [multi-omics-state-of-the-field](https://github.com/krassowski/multi-omics-state-of-the-field) repository.\n\n#### Fetching genes for a variant from dbSNP\n\nFetch the SNP record for `rs6311`:\n\n```python\nrs6311 = entrez_api.fetch(['rs6311'], max_results=1, database='snp').data[0]\nrs6311\n```\n\nDisplay the result:\n\n```python\nfrom easy_entrez.parsing import xml_to_string\n\nprint(xml_to_string(rs6311))\n```\n\nFind the gene names for `rs6311`:\n\n```python\nnamespaces = {'ns0': 'https://www.ncbi.nlm.nih.gov/SNP/docsum'}\ngenes = [\n    name.text\n    for name in rs6311.findall('.//ns0:GENE_E/ns0:NAME', namespaces)\n]\nprint(genes)\n```\n\n\u003e `['HTR2A']`\n\nFetch data for multiple variants at once:\n\n```python\nresult = entrez_api.fetch(['rs6311', 'rs662138'], max_results=10, database='snp')\ngene_names = {\n    'rs' + document_summary.get('uid'): [\n        element.text\n        for element in document_summary.findall('.//ns0:GENE_E/ns0:NAME', namespaces)\n    ]\n    for document_summary in result.data\n}\nprint(gene_names)\n```\n\n\u003e `{'rs6311': ['HTR2A'], 'rs662138': ['SLC22A1']}`\n\n#### Obtaining the chromosomal position from SNP rsID number\n\n```python\nfrom pandas import DataFrame\n\nresult = entrez_api.fetch(['rs6311', 'rs662138'], max_results=10, database='snp')\n\nvariant_positions = DataFrame([\n    {\n        'id': 'rs' + document_summary.get('uid'),\n        'chromosome': chromosome,\n        'position': position\n    }\n    for document_summary in result.data\n    for chrom_and_position in document_summary.findall('.//ns0:CHRPOS', namespaces)\n    for chromosome, position in [chrom_and_position.text.split(':')]\n])\n\nvariant_positions\n```\n\n\u003e |    | id       |   chromosome |   position |\n\u003e |---:|:---------|-------------:|-----------:|\n\u003e |  0 | rs6311   |           13 |   46897343 |\n\u003e |  1 | rs662138 |            6 |  160143444 |\n\n\n#### Converting full variation/mutation data to tabular format\n\nParsing utilities can quickly extract the data to a `VariantSet` object\nholding pandas `DataFrame`s with coordinates and alternative alleles frequencies:\n\n```python\nfrom easy_entrez.parsing import parse_dbsnp_variants\n\nvariants = parse_dbsnp_variants(result)\nvariants\n```\n\n\u003e `\u003cVariantSet with 2 variants\u003e`\n\nTo get the coordinates:\n\n```python\nvariants.coordinates\n```\n\n\u003e | rs_id    | ref   | alts   |   chrom |       pos |   chrom_prev |   pos_prev | consequence                                                                  |\n\u003e |:---------|:------|:-------|--------:|----------:|-------------:|-----------:|:-----------------------------------------------------------------------------|\n\u003e| rs6311   | C     | A,T    |      13 |  46897343 |           13 |   47471478 | upstream_transcript_variant,intron_variant,genic_upstream_transcript_variant |\n\u003e| rs662138 | C     | G      |       6 | 160143444 |            6 |  160564476 | intron_variant                                                               |\n\nFor frequencies:\n\n```python\nvariants.alt_frequencies.head(5)  # using head to only display first 5 for brevity\n```\n\n\u003e |    | rs_id   | allele   |   source_frequency |   total_count | study       |     count |\n\u003e |---:|:--------|:---------|-------------------:|--------------:|:------------|----------:|\n\u003e |  0 | rs6311  | T        |           0.44349  |          2221 | 1000Genomes |   984.991 |\n\u003e |  1 | rs6311  | T        |           0.411261 |          1585 | ALSPAC      |   651.849 |\n\u003e |  2 | rs6311  | T        |           0.331696 |          1486 | Estonian    |   492.9   |\n\u003e |  3 | rs6311  | T        |           0.35     |            14 | GENOME_DK   |     4.9   |\n\u003e |  4 | rs6311  | T        |           0.402529 |         56309 | GnomAD      | 22666     |\n\n\n#### Obtaining the SNP rs ID number from chromosomal position\n\nYou can use the query string directly:\n\n```python\nresults = entrez_api.search(\n    '13[CHROMOSOME] AND human[ORGANISM] AND 31873085[POSITION]',\n    database='snp',\n    max_results=10\n)\nprint(results.data['esearchresult']['idlist'])\n```\n\n\u003e `['59296319', '17076752', '7336701', '4']`\n\nOr pass a dictionary (no validation of arguments is performed, `AND` conjunction is used):\n\n```python\nresults = entrez_api.search(\n    dict(chromosome=13, organism='human', position=31873085),\n    database='snp',\n    max_results=10\n)\nprint(results.data['esearchresult']['idlist'])\n```\n\n\u003e `['59296319', '17076752', '7336701', '4']`\n\nThe base position should use the latest genome assembly (GRCh38 at the time of writing);\nyou can use the position in previous assembly coordinates by replacing `POSITION` with `POSITION_GRCH37`.\nFor more information of the arguments accepted by the SNP database see the [entrez help page](https://www.ncbi.nlm.nih.gov/snp/docs/entrez_help/) on NCBI website.\n\n#### Obtaining amino acids change information for variants in given range\n\nFirst we search for dbSNP rs identifiers for variants in given region:\n\n```python\ndbsnp_ids = (\n    entrez_api\n    .search(\n        '12[CHROMOSOME] AND human[ORGANISM] AND 21178600:21178720[POSITION]',\n        database='snp',\n        max_results=100\n    )\n    .data\n    ['esearchresult']\n    ['idlist']\n)\n```\n\nThen fetch the variant data for identifiers:\n\n```python\nvariant_data = entrez_api.fetch(\n    ['rs' + rs_id for rs_id in dbsnp_ids],\n    max_results=10,\n    database='snp'\n)\n```\n\nAnd parse the data, extracting the HGVS out of summary:\n\n```python\nfrom easy_entrez.parsing import parse_dbsnp_variants\nfrom pandas import Series\n\n\ndef select_protein_hgvs(items):\n    return [\n        [sequence, hgvs]\n        for entry in items\n        for sequence, hgvs in [entry.split(':')]\n        if hgvs.startswith('p.')\n    ]\n\n\nprotein_hgvs = (\n    parse_dbsnp_variants(variant_data)\n    .summary\n    .HGVS\n    .apply(select_protein_hgvs)\n    .explode()\n    .dropna()\n    .apply(Series)\n    .rename(columns={0: 'sequence', 1: 'hgvs'})\n)\nprotein_hgvs.head()\n```\n\n\u003e | rs_id        | sequence    | hgvs        |\n\u003e |:-------------|:------------|:------------|\n\u003e | rs1940853486 | NP_006437.3 | p.Gly203Ter |\n\u003e | rs1940853414 | NP_006437.3 | p.Glu202Gly |\n\u003e | rs1940853378 | NP_006437.3 | p.Glu202Lys |\n\u003e | rs1940853299 | NP_006437.3 | p.Lys201Thr |\n\u003e | rs1940852987 | NP_006437.3 | p.Asp198Glu |\n\n#### Fetching more than 10 000 entries\n\nUse `in_batches_of` method to fetch more than 10k entries (e.g. `variant_ids`):\n\n```python\nsnps_result = (\n    entrez.api\n    .in_batches_of(1_000)\n    .fetch(variant_ids, max_results=5_000, database='snp')\n)\n```\n\nThe result is a dictionary with keys being identifiers used in each batch (because the Entrez API does not always return the indentifiers back) and values representing the result. You can use `parse_dbsnp_variants` directly on this dictionary.\n\n#### Find PubMed ID from DOI\n\nWhen searching GWAS catalog PMID is needed over DOI. You can covert one to the other using:\n\n```python\ndef doi_term(doi: str) -\u003e str:\n    \"\"\"Clean a DOI string by removing URL prefix.\"\"\"\n    doi = (\n        doi\n        .replace('http://', 'https://')\n        .replace('https://doi.org/', '')\n    )\n    return f'\"{doi}\"[Publisher ID]'\n\n\nresult = entrez_api.search(\n    doi_term('https://doi.org/10.3389/fcell.2021.626821'),\n    database='pubmed',\n    max_results=1\n)\nprint(result.data['esearchresult']['idlist'])\n```\n\n\u003e `['33834021']`\n\n### Installation\n\nRequires Python 3.6+ (though only 3.7+ is tested). Install with:\n\n```bash\npip install easy-entrez\n```\n\nIf you wish to enable (optional, `tqdm`-based) progress bars use:\n\n```bash\npip install easy-entrez[with_progress_bars]\n```\n\nIf you wish to enable (optional, `pandas`-based) parsing utilities use:\n\n```bash\npip install easy-entrez[with_parsing_utils]\n```\n\n### Contributing\n\nTo build the documentation locally:\n\n```bash\npip install -e .[docs]\nsphinx-build docs docs/_build\nopen docs/_build/index.html\n```\n\n### Alternatives\n\nYou might want to try:\n\n- [biopython.Entrez](https://biopython.org/docs/1.74/api/Bio.Entrez.html) - biopython is a heavy dependency, but probably good choice if you already use it\n- [pubmedpy](https://github.com/dhimmel/pubmedpy) - provides interesting utilities for parsing the responses\n- [entrez](https://github.com/jordibc/entrez) - appears to have a comparable scope but quite different API\n- [entrezpy](https://gitlab.com/ncbipy/entrezpy) - this one did not work well for me (hence this package), but may have improved since\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fkrassowski%2Feasy-entrez","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fkrassowski%2Feasy-entrez","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fkrassowski%2Feasy-entrez/lists"}