{"id":43171974,"url":"https://github.com/opencitations/oc_graphenricher","last_synced_at":"2026-02-01T02:35:42.029Z","repository":{"id":52287757,"uuid":"357128193","full_name":"opencitations/oc_graphenricher","owner":"opencitations","description":"A tool to enrich any OCDM compliant Knowledge Graph, finding new identifiers and deduplicating entities. ","archived":false,"fork":false,"pushed_at":"2024-03-06T11:39:33.000Z","size":274,"stargazers_count":10,"open_issues_count":0,"forks_count":3,"subscribers_count":4,"default_branch":"main","last_synced_at":"2025-04-24T17:49:11.406Z","etag":null,"topics":["deduplication","enrichment","instance-matching","knowledge-graph","opencitations","semantic-web"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"isc","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/opencitations.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null}},"created_at":"2021-04-12T09:10:19.000Z","updated_at":"2023-12-22T12:26:38.000Z","dependencies_parsed_at":"2022-09-07T04:42:01.831Z","dependency_job_id":"88b06093-4188-4541-99f7-91701e7e2ee0","html_url":"https://github.com/opencitations/oc_graphenricher","commit_stats":{"total_commits":30,"total_committers":6,"mean_commits":5.0,"dds":"0.33333333333333337","last_synced_commit":"ab7ffba5381b15b62a6ce3415694f16a3a57ad1a"},"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/opencitations/oc_graphenricher","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/opencitations%2Foc_graphenricher","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/opencitations%2Foc_graphenricher/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/opencitations%2Foc_graphenricher/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/opencitations%2Foc_graphenricher/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/opencitations","download_url":"https://codeload.github.com/opencitations/oc_graphenricher/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/opencitations%2Foc_graphenricher/sbom","scorecard":{"id":708995,"data":{"date":"2025-08-11","repo":{"name":"github.com/opencitations/oc_graphenricher","commit":"d944aeded806db4648ae00cdffa81134e5f4342a"},"scorecard":{"version":"v5.2.1-40-gf6ed084d","commit":"f6ed084d17c9236477efd66e5b258b9d4cc7b389"},"score":1.9,"checks":[{"name":"Maintained","score":0,"reason":"0 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 0","details":null,"documentation":{"short":"Determines if the project is \"actively maintained\".","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#maintained"}},{"name":"Token-Permissions","score":-1,"reason":"No tokens found","details":null,"documentation":{"short":"Determines if the project's workflows follow the principle of least privilege.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#token-permissions"}},{"name":"Packaging","score":-1,"reason":"packaging workflow not detected","details":["Warn: no GitHub/GitLab publishing workflow detected."],"documentation":{"short":"Determines if the project is published as a package that others can easily download, install, easily update, and uninstall.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#packaging"}},{"name":"Code-Review","score":2,"reason":"Found 4/20 approved changesets -- score normalized to 2","details":null,"documentation":{"short":"Determines if the project requires human code review before pull requests (aka merge requests) are merged.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#code-review"}},{"name":"Dangerous-Workflow","score":-1,"reason":"no workflows found","details":null,"documentation":{"short":"Determines if the project's GitHub Action workflows avoid dangerous patterns.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#dangerous-workflow"}},{"name":"Pinned-Dependencies","score":-1,"reason":"no dependencies found","details":null,"documentation":{"short":"Determines if the project has declared and pinned the dependencies of its build process.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#pinned-dependencies"}},{"name":"Binary-Artifacts","score":9,"reason":"binaries present in source code","details":["Warn: binary detected: dist/oc_graphenricher-0.2.5-py3-none-any.whl:1"],"documentation":{"short":"Determines if the project has generated executable (binary) artifacts in the source repository.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#binary-artifacts"}},{"name":"CII-Best-Practices","score":0,"reason":"no effort to earn an OpenSSF best practices badge detected","details":null,"documentation":{"short":"Determines if the project has an OpenSSF (formerly CII) Best Practices Badge.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#cii-best-practices"}},{"name":"Security-Policy","score":0,"reason":"security policy file not detected","details":["Warn: no security policy file detected","Warn: no security file to analyze","Warn: no security file to analyze","Warn: no security file to analyze"],"documentation":{"short":"Determines if the project has published a security policy.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#security-policy"}},{"name":"Fuzzing","score":0,"reason":"project is not fuzzed","details":["Warn: no fuzzer integrations found"],"documentation":{"short":"Determines if the project uses fuzzing.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#fuzzing"}},{"name":"License","score":10,"reason":"license file detected","details":["Info: project has a license file: LICENSE:0","Info: FSF or OSI recognized license: ISC License: LICENSE:0"],"documentation":{"short":"Determines if the project has defined a license.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#license"}},{"name":"Signed-Releases","score":-1,"reason":"no releases found","details":null,"documentation":{"short":"Determines if the project cryptographically signs release artifacts.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#signed-releases"}},{"name":"Branch-Protection","score":0,"reason":"branch protection not enabled on development/release branches","details":["Warn: branch protection not enabled for branch 'main'"],"documentation":{"short":"Determines if the default and release branches are protected with GitHub's branch protection settings.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#branch-protection"}},{"name":"Vulnerabilities","score":0,"reason":"13 existing vulnerabilities detected","details":["Warn: Project is vulnerable to: PYSEC-2024-230 / GHSA-248v-346w-9cwc","Warn: Project is vulnerable to: PYSEC-2024-60 / GHSA-jjg7-2v4v-x38h","Warn: Project is vulnerable to: GHSA-9hjg-9r4m-mvj7","Warn: Project is vulnerable to: GHSA-9wx4-h78v-vm56","Warn: Project is vulnerable to: PYSEC-2023-74 / GHSA-j8r2-6x86-q33q","Warn: Project is vulnerable to: PYSEC-2025-49 / GHSA-5rjg-fvgr-3xxf","Warn: Project is vulnerable to: GHSA-cx63-2mw6-8hw5","Warn: Project is vulnerable to: GHSA-g7vv-2v7x-gj9p","Warn: Project is vulnerable to: GHSA-34jh-p97f-mpxf","Warn: Project is vulnerable to: PYSEC-2023-212 / GHSA-g4mx-q9vg-27p4","Warn: Project is vulnerable to: GHSA-pq67-6m6q-mj2v","Warn: Project is vulnerable to: PYSEC-2021-108 / GHSA-q2q7-5pp4-w6pg","Warn: Project is vulnerable to: PYSEC-2023-192 / GHSA-v845-jxx5-vc9f"],"documentation":{"short":"Determines if the project has open, known unfixed vulnerabilities.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#vulnerabilities"}},{"name":"SAST","score":0,"reason":"SAST tool is not run on all commits -- score normalized to 0","details":["Warn: 0 commits out of 14 are checked with a SAST tool"],"documentation":{"short":"Determines if the project uses static code analysis.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#sast"}}]},"last_synced_at":"2025-08-22T07:35:49.193Z","repository_id":52287757,"created_at":"2025-08-22T07:35:49.193Z","updated_at":"2025-08-22T07:35:49.193Z"},"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":28965430,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-02-01T02:14:24.993Z","status":"ssl_error","status_checked_at":"2026-02-01T02:13:55.706Z","response_time":56,"last_error":"SSL_read: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["deduplication","enrichment","instance-matching","knowledge-graph","opencitations","semantic-web"],"created_at":"2026-02-01T02:35:41.410Z","updated_at":"2026-02-01T02:35:42.022Z","avatar_url":"https://github.com/opencitations.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cp align=\"center\"\u003e\n\n  \u003ch2 align=\"center\"\u003eGraphEnricher\u003c/h3\u003e\n  \u003cp align=\"center\"\u003e\n    A tool to enrich any \u003ca href=\"http://opencitations.net/model\"\u003eOCDM\u003c/a\u003e compliant Knowledge Graph, finding new identifiers\nand deduplicating entities.\n\u003c/p\u003e\n\n\u003c!-- TABLE OF CONTENTS --\u003e\n  \u003csummary\u003e\u003ch2 style=\"display: inline-block\"\u003eTable of Contents\u003c/h2\u003e\u003c/summary\u003e\n  \u003col\u003e\n    \u003cli\u003e\n      \u003ca href=\"#about-the-project\"\u003eAbout The Project\u003c/a\u003e\n    \u003c/li\u003e\n    \u003cli\u003e\n      \u003ca href=\"#getting-started\"\u003eGetting Started\u003c/a\u003e\n    \u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#usage\"\u003eUsage\u003c/a\u003e\u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#license\"\u003eLicense\u003c/a\u003e\u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#contact\"\u003eContact\u003c/a\u003e\u003c/li\u003e\n    \u003cli\u003e\u003ca href=\"#acknowledgements\"\u003eAcknowledgements\u003c/a\u003e\u003c/li\u003e\n  \u003c/ol\u003e\n\n\n\n\u003c!-- ABOUT THE PROJECT --\u003e\n## About The Project\n\nThis tool is divided in two part: an Enricher component responsible for finding new identifiers and adding them to the \ngraph set, and an InstanceMatching component responsible for deduplicating any entity that share the same identifier.\n### Enricher\nThe enricher iterates each Bibliographic Resources (BRs) contained in the graph set.\nFor each Bibliographic Resources (BRs) (avoiding issues and journals), get the list of the identifiers already\ncontained in the graph set and check if it already has a DOI, an ISSN, a Wikidata ID and an OpenAlex ID:\n- If an ISSN is specified, query Crossref to extract other ISSNs\n- If there's no DOI, query Crossref to get one by means of all the other data extracted\n- If there's no Wikidata ID, query Wikidata to get one by means of all the other identifiers\n- If there's no OpenAlex ID, query OpenAlex to get one by means of all the other identifiers\n\nAny new identifier found will be added to the Bibliographic Resource (BR).\n  \nThen, for each Agent Role (AR) related to the Bibliographic Resource (BR), get the list of all the identifier already contained in its linked Responsible Agent (RA) and:\n- If it doesn't have an ORCID, query ORCID to get it\n- If it doesn't have a VIAF, query VIAF to get it\n- If it doesn't have a Wikidata ID, query Wikidata by means of all the other identifier to get one\n- If the Responsible Agent (RA) is related to a publisher, query Crossref to get its ID by means of its DOI\n\nAny new identifier found will be added to the RA related to the AR.\n\nIn the end it will store a new graph set and its provenance.\n\nNB: even if it's not possible to have an identifier duplicated for the same entity, it's possible that in\nthe whole graph set you could find different identifiers that share the same schema and literal. For this\npurpose, you should use the **instancematching** module after you've enriched the graph set.\n\n### APIs and identifiers\nCurrently, there are 5 external API involved:\n- Crossref \n- ORCID\n- VIAF\n- WikiData\n- OpenAlex\n\nand we can discover the following identifiers:\n- DOI\n- ISSN\n- Crossref's publisher ID\n- ORCID\n- VIAF\n- Wikidata ID (by means of any other identifier, e.g.: PMID, VIAF, DOI, ...)\n- OpenAlex Work ID and Source ID (only for bibliographic resources and by means of other identifiers, e.g.: PMID, DOI, ...)\n\nIt's possible, anyway, to extend the class QueryInterface to add any other useful API.\n\n### Instance Matching\nThe instance matching process is articulated in three sequential step:\n- match the Responsible Agents (RAs)\n- match the Bibliographic Resources (BRs) \n- match the IDs\n\n#### Matching the Responsible Agents (RAs) \nDiscover all the Responsible Agents (RAs)  that share the same identifier's literal, creating a graph of\nthem. Then merge each connected component (cluster of Responsible Agents (RAs)  linked by the same identifier)\ninto one.\nFor each couple of Responsible Agent (RA) that are going to be merged, substitute the references of the\nResponsible Agent (RA) that will no longer exist, by removing the Responsible Agent (RA)\nfrom each of its referred Agent Role (AR) and add, instead, the merged one)\n\nIf the Responsible Agent (RA) linked by the Agent Role (AR) that will no longer exist is not linked by any\nother Agent Role (AR), then it will be marked as to be deleted, otherwise not.\n\nIn the end, generate the provenance and commit pending changes in the graph set\n\n#### Matching the Bibliographic Resources (BRs) \n\nDiscover all the Bibliographic Resources (BRs)  that share the same identifier's literal, creating a graph of them.\nThen merge each connected component (cluster of Bibliographic Resources (BR) linked by the same identifier) into one.\nFor each couple of Bibliographic Resources (BRs) that are going to be merged, merge also:\n - their containers by matching the proper type (issue of BR1 -\u003e issue of BR2)\n - their publisher\n\nIn the end, generate the provenance and commit pending changes in the graph set\n\n#### Matching the IDs\nDiscover all the IDs that share the same schema and literal, then merge all into one\nand substitute all the reference with the merged one.\n\nIn the end, generate the provenance and commit pending changes in the graph set\n\n\u003c!-- GETTING STARTED --\u003e\n## Getting Started\n\nTo get a local copy up and running follow these simple steps:\n1. install python \u003e= 3.8:\n\n```sudo apt install python3```\n\n2. Install oc_graphenricher via pip:\n```\npip install oc-graphenricher\n```\n\n### Installing from the sources\n1. Having already installed python, you can also install GraphEnricher via cloning this repository: \n```\ngit clone https://github.com/opencitations/oc_graphenricher`\ncd ./oc_graphenricher\n```\n2. install poetry:\n\n```pip install poetry```\n\n3. install all the dependencies:\n\n``` poetry install```\n\n4. build the package:\n\n```poetry build```\n\n5. install the package:\n\n```    pip install ./dist/oc_graphenricher-\u003cVERSION\u003e.tar.gz```\n\n6. run the tests (from the root of the project):\n\n```\npoetry run test\n```\n\n\u003c!-- USAGE EXAMPLES --\u003e\n## Usage\n\nIt's supposed to accept only graph set objects. To create one:\n\n```\ng = Graph()\ng = g.parse('../data/test_dump.ttl', format='nt11')\n\nreader = Reader()\ng_set = GraphSet(base_iri='https://w3id.org/oc/meta/')\nentities = reader.import_entities_from_graph(g_set, g, enable_validation=False, resp_agent='https://w3id.org/oc/meta/prov/pa/2')\n```\n\n### Enrichment\nAt this point, to run the enrichment phase:\n```\nfrom oc_graphenricher.enricher import Enricher\n\nenricher = GraphEnricher(g_set)\nenricher.enrich()\n```\nYou'll see the progress bar with an estimate of the time needed and the average time spent\nfor each Bibliographic Resource (BR) enriched. \n\n### Deduplication \nThen, having serialized the enriched graph set, and having read it again as the\n`g_set` object, to run the deduplication step do:\n\n```\nfrom oc_graphenricher.instancematching import InstanceMatching\n\nmatcher = InstanceMatching(g_set)\nmatcher.match()\n```\n\nThe match method will run sequentially:\n- deduplication of Responsible Agents (RAs)\n- deduplication of Bibliographic Resources (BRs)\n- deduplication of Identifiers (IDs)\n- save to file\n\nIf you need to, you can also deduplicate one of those independently of each other.\n\nTo deduplicate Responsible Agents (RAs):\n```\nfrom oc_graphenricher.instancematching import InstanceMatching\n\nmatcher = InstanceMatching(g_set)\nmatcher.instance_matching_ra()\nmatcher.save()\n```\n\nTo deduplicate Bibliographic Resources (BRs):\n```\nfrom oc_graphenricher.instancematching import InstanceMatching\n\nmatcher = InstanceMatching(g_set)\nmatcher.instance_matching_br()\nmatcher.save()\n```\nTo deduplicate Identifiers (IDs):\n```\nfrom oc_graphenricher.instancematching import InstanceMatching\n\nmatcher = InstanceMatching(g_set)\nmatcher.instance_matching_id()\nmatcher.save()\n```\n\n\n\n\n\n\u003c!-- LICENSE --\u003e\n## License\n\nDistributed under the ISC License. See `LICENSE` for more information.\n\n\n\n\u003c!-- CONTACT --\u003e\n## Contact\n\n#### Current maintainers of the repository\n\n- Arcangelo Massari - [@arcangelo7](https://github.com/arcangelo7) - arcangelo.massari@unibo.it\n- Arianna Moretti - [@ariannamorettj](https://github.com/ariannamorettj) - arianna.moretti4@unibo.it\n- Elia Rizzetto - [@eliarizzetto](https://github.com/eliarizzetto) - elia.rizzetto@studio.unibo.it\n- Silvio Peroni - [@essepuntato](https://github.com/essepuntato) - silvio.peroni@unibo.it\n\n\n#### Author of the repository and original code:\n- Gabriele Pisciotta - [@GaPisciotta](https://twitter.com/GaPisciotta) - ga.pisciotta@gmail.com\n\n\n\n\n\nProject Link: [https://github.com/opencitations/oc_graphenricher](https://github.com/opencitations/oc_graphenricher)\n\n\n\n\u003c!-- ACKNOWLEDGEMENTS --\u003e\n## Acknowledgements\nThis project has been developed as part of the \n[Wikipedia Citations in Wikidata](https://meta.wikimedia.org/wiki/Wikicite/grant/Wikipedia_Citations_in_Wikidata) \nresearch project, under the supervision of prof. Silvio Peroni.\n\n\n\n\n\u003c!-- MARKDOWN LINKS \u0026 IMAGES --\u003e\n\u003c!-- https://www.markdownguide.org/basic-syntax/#reference-style-links --\u003e\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fopencitations%2Foc_graphenricher","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fopencitations%2Foc_graphenricher","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fopencitations%2Foc_graphenricher/lists"}