{"id":13577830,"url":"https://github.com/plantnet/gbif-dl","last_synced_at":"2025-08-26T00:04:28.205Z","repository":{"id":47080927,"uuid":"320651964","full_name":"plantnet/gbif-dl","owner":"plantnet","description":"GBIF classification dataloaders","archived":false,"fork":false,"pushed_at":"2024-06-02T04:54:58.000Z","size":117,"stargazers_count":42,"open_issues_count":28,"forks_count":7,"subscribers_count":5,"default_branch":"master","last_synced_at":"2025-07-16T05:32:27.749Z","etag":null,"topics":["biodiversity","dataloader"],"latest_commit_sha":null,"homepage":"https://plantnet.github.io/gbif-dl/","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/plantnet.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2020-12-11T18:23:06.000Z","updated_at":"2025-05-08T08:03:03.000Z","dependencies_parsed_at":"2024-11-05T15:35:31.421Z","dependency_job_id":"69143097-3cf0-4cc5-9b44-10cfc989c1b4","html_url":"https://github.com/plantnet/gbif-dl","commit_stats":{"total_commits":63,"total_committers":4,"mean_commits":15.75,"dds":0.5238095238095238,"last_synced_commit":"5c00725e4e433dca28393b7f6f9b4058b8c6437e"},"previous_names":[],"tags_count":2,"template":false,"template_full_name":null,"purl":"pkg:github/plantnet/gbif-dl","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/plantnet%2Fgbif-dl","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/plantnet%2Fgbif-dl/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/plantnet%2Fgbif-dl/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/plantnet%2Fgbif-dl/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/plantnet","download_url":"https://codeload.github.com/plantnet/gbif-dl/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/plantnet%2Fgbif-dl/sbom","scorecard":{"id":737071,"data":{"date":"2025-08-11","repo":{"name":"github.com/plantnet/gbif-dl","commit":"5c00725e4e433dca28393b7f6f9b4058b8c6437e"},"scorecard":{"version":"v5.2.1-40-gf6ed084d","commit":"f6ed084d17c9236477efd66e5b258b9d4cc7b389"},"score":3.8,"checks":[{"name":"Code-Review","score":0,"reason":"Found 0/30 approved changesets -- score normalized to 0","details":null,"documentation":{"short":"Determines if the project requires human code review before pull requests (aka merge requests) are merged.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#code-review"}},{"name":"Packaging","score":-1,"reason":"packaging workflow not detected","details":["Warn: no GitHub/GitLab publishing workflow detected."],"documentation":{"short":"Determines if the project is published as a package that others can easily download, install, easily update, and uninstall.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#packaging"}},{"name":"Dangerous-Workflow","score":10,"reason":"no dangerous workflow patterns detected","details":null,"documentation":{"short":"Determines if the project's GitHub Action workflows avoid dangerous patterns.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#dangerous-workflow"}},{"name":"Maintained","score":0,"reason":"0 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 0","details":null,"documentation":{"short":"Determines if the project is \"actively maintained\".","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#maintained"}},{"name":"Binary-Artifacts","score":10,"reason":"no binaries found in the repo","details":null,"documentation":{"short":"Determines if the project has generated executable (binary) artifacts in the source repository.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#binary-artifacts"}},{"name":"Pinned-Dependencies","score":0,"reason":"dependency not pinned by hash detected -- score normalized to 0","details":["Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/docs.yml:12: update your workflow using https://app.stepsecurity.io/secureworkflow/plantnet/gbif-dl/docs.yml/master?enable=pin","Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/docs.yml:14: update your workflow using https://app.stepsecurity.io/secureworkflow/plantnet/gbif-dl/docs.yml/master?enable=pin","Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/test_black.yml:10: update your workflow using https://app.stepsecurity.io/secureworkflow/plantnet/gbif-dl/test_black.yml/master?enable=pin","Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/test_black.yml:12: update your workflow using https://app.stepsecurity.io/secureworkflow/plantnet/gbif-dl/test_black.yml/master?enable=pin","Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/test_unittests.yml:18: update your workflow using https://app.stepsecurity.io/secureworkflow/plantnet/gbif-dl/test_unittests.yml/master?enable=pin","Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/test_unittests.yml:20: update your workflow using https://app.stepsecurity.io/secureworkflow/plantnet/gbif-dl/test_unittests.yml/master?enable=pin","Warn: third-party GitHubAction not pinned by hash: .github/workflows/test_unittests.yml:45: update your workflow using https://app.stepsecurity.io/secureworkflow/plantnet/gbif-dl/test_unittests.yml/master?enable=pin","Warn: pipCommand not pinned by hash: .github/workflows/docs.yml:19","Warn: pipCommand not pinned by hash: .github/workflows/docs.yml:20","Warn: pipCommand not pinned by hash: .github/workflows/test_black.yml:18","Warn: pipCommand not pinned by hash: .github/workflows/test_unittests.yml:25","Warn: pipCommand not pinned by hash: .github/workflows/test_unittests.yml:26","Warn: pipCommand not pinned by hash: .github/workflows/test_unittests.yml:27","Info:   0 out of   6 GitHub-owned GitHubAction dependencies pinned","Info:   0 out of   1 third-party GitHubAction dependencies pinned","Info:   0 out of   6 pipCommand dependencies pinned"],"documentation":{"short":"Determines if the project has declared and pinned the dependencies of its build process.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#pinned-dependencies"}},{"name":"Token-Permissions","score":0,"reason":"detected GitHub workflow tokens with excessive permissions","details":["Warn: no topLevel permission defined: .github/workflows/docs.yml:1","Warn: no topLevel permission defined: .github/workflows/test_black.yml:1","Warn: no topLevel permission defined: .github/workflows/test_unittests.yml:1","Info: no jobLevel write permissions found"],"documentation":{"short":"Determines if the project's workflows follow the principle of least privilege.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#token-permissions"}},{"name":"CII-Best-Practices","score":0,"reason":"no effort to earn an OpenSSF best practices badge detected","details":null,"documentation":{"short":"Determines if the project has an OpenSSF (formerly CII) Best Practices Badge.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#cii-best-practices"}},{"name":"Security-Policy","score":0,"reason":"security policy file not detected","details":["Warn: no security policy file detected","Warn: no security file to analyze","Warn: no security file to analyze","Warn: no security file to analyze"],"documentation":{"short":"Determines if the project has published a security policy.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#security-policy"}},{"name":"Vulnerabilities","score":10,"reason":"0 existing vulnerabilities detected","details":null,"documentation":{"short":"Determines if the project has open, known unfixed vulnerabilities.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#vulnerabilities"}},{"name":"Fuzzing","score":0,"reason":"project is not fuzzed","details":["Warn: no fuzzer integrations found"],"documentation":{"short":"Determines if the project uses fuzzing.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#fuzzing"}},{"name":"License","score":10,"reason":"license file detected","details":["Info: project has a license file: LICENSE:0","Info: FSF or OSI recognized license: MIT License: LICENSE:0"],"documentation":{"short":"Determines if the project has defined a license.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#license"}},{"name":"Signed-Releases","score":-1,"reason":"no releases found","details":null,"documentation":{"short":"Determines if the project cryptographically signs release artifacts.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#signed-releases"}},{"name":"Branch-Protection","score":-1,"reason":"internal error: error during branchesHandler.setup: internal error: githubv4.Query: Resource not accessible by integration","details":null,"documentation":{"short":"Determines if the default and release branches are protected with GitHub's branch protection settings.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#branch-protection"}},{"name":"SAST","score":0,"reason":"SAST tool is not run on all commits -- score normalized to 0","details":["Warn: 0 commits out of 10 are checked with a SAST tool"],"documentation":{"short":"Determines if the project uses static code analysis.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#sast"}}]},"last_synced_at":"2025-08-22T16:10:55.271Z","repository_id":47080927,"created_at":"2025-08-22T16:10:55.271Z","updated_at":"2025-08-22T16:10:55.271Z"},"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":271740701,"owners_count":24812637,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-08-23T02:00:09.327Z","response_time":69,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["biodiversity","dataloader"],"created_at":"2024-08-01T15:01:24.700Z","updated_at":"2025-08-26T00:04:28.125Z","avatar_url":"https://github.com/plantnet.png","language":"Python","funding_links":[],"categories":["Python"],"sub_categories":[],"readme":"# gbif-dl 🌱 \u003e 💾\n\n[![Build Status](https://github.com/plantnet/gbif-dl/workflows/CI/badge.svg)](https://github.com/plantnet/gbif-dl/actions?query=workflow%3ACI+branch%3Amaster+event%3Apush)\n[![Supported Python versions](https://img.shields.io/pypi/pyversions/gbif-dl.svg)](https://pypi.python.org/pypi/gbif-dl)\n[![Documentation Status](https://img.shields.io/badge/docs-api-blue)](https://plantnet.github.io/gbif-dl/)\n\nthis package makes it simpler to obtain media data from the GBIF database to be used for training **machine learning classification** tasks. It wraps the [GBIF API](https://www.gbif.org/developer/summary) and supports directly querying the api to obtain and download a list of urls.\nExisting saved queries can also be obtained using the download api of GBIF simply by providing GBIF DOI key.\nThe package provides an efficient downloader that uses python asyncio modules to speed up downloading of many small files as typically occur in downloads.\n\n## Disclaimer\n\nUnlike GBIF occurrences that all have a creative common license (CC0, CC BY, or CC BY-NC), GBIF [does not give any official recommendation](https://data-blog.gbif.org/post/gbif-multimedia/) for licensing shared media files. The License fields are essentially free text filled in by the data provider. Data providers are strongly encouraged to set their licenses in a machine-readable format, but there is no guarantee. Thus, it is the responsibility of GBIF-DL users to set up the appropriate filters on the license field and to respect the conditions of use of these licenses.\n\nSince we are heading into relatively new territory regarding the use of GBIF media files for machine learning, it is currently unclear how publishers will feel about this when learning that their photos are used in this way. At Pl@ntNet, the choice to publish our images in GBIF has been carefully considered and we are fully aware that they would probably be used for this purpose. But we fully understand that some other data providers didn't think of this and we are interested to have open discussion around these aspects.\n\n## Installation\n\nInstallation can be done via pip.\n\n```\npip install gbif-dl\n```\n\n## Usage\n\nThe usage of `gbif-dl` helps users to create their own GBIF based media pipeline for training machine learning models. The package provides two core functionalities as followed:\n\n1. `gbif-dl.generators`: Generators provide image urls from the GBIF database given queries or a pre-defined URL.\n2. `gbif-dl.io`: Provides efficient media downloading to write the data to a storage device.\n\n### 1. Retrieve media urls from GBIF\n\n`gbif-dl` supports two ways to retrieve image urls. One is to use directly query the gbif api the `gbif_dl.api` module. This is suited for quickly retrieving smaller datasets that do not require extensive query parameters. Another way is to use already the [gbif download workflows](https://www.gbif.org/data-processing) which assemble a [Darwin Core Archives](https://github.com/gbif/ipt/wiki/DwCAHowToGuide) waiting on the gbif servers. These can be downloaded and parsed using the `gbif_dl.dwca` module as explained below.\n\n#### `gbif_dl.generators.api`: getting occurance media URLS by querying GBIF\n\nThe query supports all fields that are supported by the [GBIF occurance API](https://www.gbif.org/developer/occurrence#search). In the following example, we query three plants using the `speciesKey` of GBIF from the list of [top 1200 invasive plant species](https://www.cabi.org/ISC). Also, we are limiting the results by only retrieving results from [Plantnet](https://plantnet.org) _and_ [iNaturalist](https://www.inaturalist.org/). using the `datasetKey`.\n\nThe query is passed as a simple dictionary:\n\n```python\nqueries = {\n    \"speciesKey\": [\n        5352251, # \"Robinia pseudoacacia L\"\n        3190653, # \"Ailanthus altissima (Mill.) Swingle\"\n        3189866  # \"Acer negundo L\"\n    ],\n    \"datasetKey\": [\n        \"7a3679ef-5582-4aaa-81f0-8c2545cafc81\",  # plantnet\n        \"50c9509d-22c7-4a22-a47d-8c48425ef4a7\"  # inaturalist\n    ]\n}\n```\n\nGive this query, we can pass this to the `api.generate_urls` function which returns a python\ngenerator:\n\n```python\nimport gbif_dl\ndata_generator = gbif_dl.api.generate_urls(\n    queries=queries,\n    label=\"speciesKey\",\n)\n```\n\nAdditionally we have to specify the output `label` from the occurances which doesn't\nnecessarily have to be part of the query attributes. The `label` is later used to classify the results and store the data in hierachical structure: `label/image.jpg`.\n\nIterating over the generator now yields the media data returning a few thousand urls.\n\n```python\nfor i in data_generator:\n    print(i)\n```\n\neach return entry is a dictionary of media attributes, to be consumed by the downloader.\n\n```python\n{\n    'url': 'https://bs.plantnet.org/image/o/cfa25c7fb5cdf12719d1345769d3936d0ca73974',\n    'basename': 'fdcc3440ab0e3abf824a5c68c864b018cccfcd3b',\n    'label': '5352251'\n},\n{\n    'url': 'https://static.inaturalist.org/photos/58881180/original.jpeg?1577914533',\n    'basename': '7db818c0708ba859516353ff9b30ef942aca19de',\n    'label': '3189866'\n},\n{\n    'url': 'https://static.inaturalist.org/photos/58866788/original.jpeg?1577898729',\n    'basename': '58ae3ef46e59e9a06d67de09c8b7ef3b8db3c85a',\n    'label': '3189866'\n}\n```\n\n#### Balancing items\n\nVery often users won't be using all media downloads from a given query since this often results in datasets with heavily inbalanced number of samples per label. When generating urls from the API, users can specify certain additional attributes to influence the sampling process. For example, to balance the dataset by the dataset provider and by the species the following arguments can be used:\n\n- `split_streams_by`: splits the query into combination of several substreams where each stream represents the product of the query values. When combined with `nb_samples`, this produces a balanced dataset where each stream yields the same number of samples.\n- `nb_samples`: an integer that limits the total number of samples to be generated from the balanced streams. E.g, this can be used to just get `100` samples from the api. When set to `-1`, the minimum number of samples from all streams is used, hence this results in the **maximum number of balanced** sampled from all streams.\n\nIn the following example, we will receive a balanced dataset assembled from `3 species * 2 datasets = 6 streams` and only get minumum number of total samples from all 6 streams:\n\n```python\ndata_generator = gbif_dl.api.generate_urls(\n    queries=queries,\n    label=\"speciesKey\",\n    nb_samples=-1,\n    split_streams_by=[\"datasetKey\", \"speciesKey\"],\n)\n```\n\nFor other, more advanced, use-cases users can add more constraints:\n\n- `nb_samples_per_stream`: put a hard **limit** on the _maximum number of samples_ to be yielded by a stream.\n- `weighted_streams`: weights each stream by its original distribution. That way users can get a smaller subset of the data but keep the original **unbalanced** distribution of the data.\n\nThe following dataset consist of exactly 1000 samples for which the distribution of `speciesKey` is maintained from the full query of all samples. Furthermore, we only allow a maxmimum of 800 samples per species.\n\n```python\ndata_generator = gbifmediads.api.generate_urls(\n    queries=queries,\n    label=\"speciesKey\",\n    nb_samples=1000,\n    nb_samples_per_stream=800,\n    weighted_streams=True,\n    split_streams_by=[\"speciesKey\"],\n)\n```\n\n### Get URLS using Darwin Core Archives\n\nA url generator can also be created from a GBIF download link given a registered DOI or a GBIF download ID. In the following example we will be downloading and parse DWCA archive [that should yield the same results as in the query example above.](https://www.gbif.org/occurrence/download/0117522-200613084148143).\n\n- `dwca_root_path`: Set root path where to store the DWCA zip files. Defaults to None, which results in the creation of a temporary directory, If the path and DWCA archive already exist, it will not be downloaded again.\n\nThe following example creates a data_generator with the the same output class label as in the example above.\n\n```python\ndata_generator = gbif_dl.dwca.generate_urls(\n    \"10.15468/dl.vnm42s\", dwca_root_path=\"dwcas\", label=\"speciesKey\"\n)\n```\n\n### Downloading images to disk\n\nDownloading from a url generator can simply be done by running.\n\n```python\nstats = gbif_dl.io.download(data_generator, root=\"my_dataset\")\n```\n\nThe downloader provides very fast download speeds by using an async queue. Some fail-safe functionality can be provided by setting the number of `retries` to higher than 1.\n\n### Training Datasets\n\n#### PyTorch\n\n`gbif-dl` makes it simple to train a PyTorch image classification model by using e.g. `torchvision.ImageFolder`. Each item in the `data_generator` can be randomly assigned to a `train` or `test` subset using `random_subsets`. That way users can directly use the subsets.\n\n```python\nimport torchvision\ngbif_dl.io.download(data_generator, root=\"my_dataset\", random_subsets={'train': 0.9, 'test': 0.1})\ntrain_dataset = torchvision.datasets.ImageFolder(root='my_dataset/train', ...)\ntest_dataset = torchvision.datasets.ImageFolder(root='my_dataset/test', ...)\n```\n\n#### Tensorflow\n\nThe simpliest way to generate a `tf.data.Dataset` pipeline from a data generator is to use `tf.keras.preprocessing.image_dataset_from_directory`.\nSimilarily to the pytorch example, users just need to provide the root paths of the downloaded datasets.\n\n```python\nimport tensorflow as tf\ngbif_dl.io.download(data_generator, root=\"my_dataset\", random_subsets={'train': 0.9, 'test': 0.1})\ntrain_dataset = tf.keras.preprocessing.image_dataset_from_directory(root='my_dataset/train', label_mode=\"categorical\", labels=\"inferred\", *args, **kwargs)\ntest_dataset = tf.keras.preprocessing.image_dataset_from_directory(root='my_dataset/test', label_mode=\"categorical\", labels=\"inferred\", *args, **kwargs)\n```\n\n## FAQ\n\n#### Q: Downloading doesn't work from inside a jupyter notebook\n\nThis is a known issue of running asyncio code from within jupyter.\nPlease execute these lines before using gbif-dl\n\n```python\nimport nest_asyncio\nnest_asyncio.apply()\n```\n\n## License\n\nMIT\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fplantnet%2Fgbif-dl","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fplantnet%2Fgbif-dl","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fplantnet%2Fgbif-dl/lists"}