{"id":13678776,"url":"https://github.com/allenai/ir_datasets","last_synced_at":"2025-10-13T15:55:44.840Z","repository":{"id":38107399,"uuid":"310393570","full_name":"allenai/ir_datasets","owner":"allenai","description":"Provides a common interface to many IR ranking datasets.","archived":false,"fork":false,"pushed_at":"2025-09-03T10:48:22.000Z","size":2349,"stargazers_count":375,"open_issues_count":92,"forks_count":46,"subscribers_count":8,"default_branch":"master","last_synced_at":"2025-10-11T14:12:41.768Z","etag":null,"topics":["dataset","information-retrieval","ir"],"latest_commit_sha":null,"homepage":"https://ir-datasets.com/","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/allenai.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2020-11-05T19:09:09.000Z","updated_at":"2025-10-03T12:47:20.000Z","dependencies_parsed_at":"2025-09-03T12:19:05.109Z","dependency_job_id":"833df07a-2c2c-4446-9be0-4db0741f1c78","html_url":"https://github.com/allenai/ir_datasets","commit_stats":{"total_commits":372,"total_committers":14,"mean_commits":"26.571428571428573","dds":0.07795698924731187,"last_synced_commit":"77265cddf9be1f98147a5a7f037878c9b3b93d43"},"previous_names":[],"tags_count":23,"template":false,"template_full_name":null,"purl":"pkg:github/allenai/ir_datasets","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/allenai%2Fir_datasets","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/allenai%2Fir_datasets/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/allenai%2Fir_datasets/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/allenai%2Fir_datasets/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/allenai","download_url":"https://codeload.github.com/allenai/ir_datasets/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/allenai%2Fir_datasets/sbom","scorecard":{"id":185277,"data":{"date":"2025-08-11","repo":{"name":"github.com/allenai/ir_datasets","commit":"7d17396a3c8796fca4d62119ef9b3cfe00d6c67d"},"scorecard":{"version":"v5.2.1-40-gf6ed084d","commit":"f6ed084d17c9236477efd66e5b258b9d4cc7b389"},"score":2.8,"checks":[{"name":"Maintained","score":0,"reason":"1 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 0","details":null,"documentation":{"short":"Determines if the project is \"actively maintained\".","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#maintained"}},{"name":"Packaging","score":-1,"reason":"packaging workflow not detected","details":["Warn: no GitHub/GitLab publishing workflow detected."],"documentation":{"short":"Determines if the project is published as a package that others can easily download, install, easily update, and uninstall.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#packaging"}},{"name":"Code-Review","score":3,"reason":"Found 3/9 approved changesets -- score normalized to 3","details":null,"documentation":{"short":"Determines if the project requires human code review before pull requests (aka merge requests) are merged.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#code-review"}},{"name":"Dangerous-Workflow","score":10,"reason":"no dangerous workflow patterns detected","details":null,"documentation":{"short":"Determines if the project's GitHub Action workflows avoid dangerous patterns.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#dangerous-workflow"}},{"name":"Token-Permissions","score":0,"reason":"detected GitHub workflow tokens with excessive permissions","details":["Warn: no topLevel permission defined: .github/workflows/deploy.yml:1","Warn: no topLevel permission defined: .github/workflows/test.yml:1","Info: no jobLevel write permissions found"],"documentation":{"short":"Determines if the project's workflows follow the principle of least privilege.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#token-permissions"}},{"name":"CII-Best-Practices","score":0,"reason":"no effort to earn an OpenSSF best practices badge detected","details":null,"documentation":{"short":"Determines if the project has an OpenSSF (formerly CII) Best Practices Badge.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#cii-best-practices"}},{"name":"Binary-Artifacts","score":10,"reason":"no binaries found in the repo","details":null,"documentation":{"short":"Determines if the project has generated executable (binary) artifacts in the source repository.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#binary-artifacts"}},{"name":"Security-Policy","score":0,"reason":"security policy file not detected","details":["Warn: no security policy file detected","Warn: no security file to analyze","Warn: no security file to analyze","Warn: no security file to analyze"],"documentation":{"short":"Determines if the project has published a security policy.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#security-policy"}},{"name":"License","score":10,"reason":"license file detected","details":["Info: project has a license file: LICENSE:0","Info: FSF or OSI recognized license: Apache License 2.0: LICENSE:0"],"documentation":{"short":"Determines if the project has defined a license.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#license"}},{"name":"Branch-Protection","score":0,"reason":"branch protection not enabled on development/release branches","details":["Warn: branch protection not enabled for branch 'master'"],"documentation":{"short":"Determines if the default and release branches are protected with GitHub's branch protection settings.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#branch-protection"}},{"name":"Fuzzing","score":0,"reason":"project is not fuzzed","details":["Warn: no fuzzer integrations found"],"documentation":{"short":"Determines if the project uses fuzzing.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#fuzzing"}},{"name":"Signed-Releases","score":-1,"reason":"no releases found","details":null,"documentation":{"short":"Determines if the project cryptographically signs release artifacts.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#signed-releases"}},{"name":"Pinned-Dependencies","score":0,"reason":"dependency not pinned by hash detected -- score normalized to 0","details":["Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/deploy.yml:11: update your workflow using https://app.stepsecurity.io/secureworkflow/allenai/ir_datasets/deploy.yml/master?enable=pin","Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/deploy.yml:12: update your workflow using https://app.stepsecurity.io/secureworkflow/allenai/ir_datasets/deploy.yml/master?enable=pin","Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/test.yml:18: update your workflow using https://app.stepsecurity.io/secureworkflow/allenai/ir_datasets/test.yml/master?enable=pin","Warn: GitHub-owned GitHubAction not pinned by hash: .github/workflows/test.yml:21: update your workflow using https://app.stepsecurity.io/secureworkflow/allenai/ir_datasets/test.yml/master?enable=pin","Warn: pipCommand not pinned by hash: .github/workflows/deploy.yml:17","Warn: pipCommand not pinned by hash: .github/workflows/deploy.yml:18","Warn: pipCommand not pinned by hash: .github/workflows/test.yml:27","Warn: pipCommand not pinned by hash: .github/workflows/test.yml:28","Warn: pipCommand not pinned by hash: .github/workflows/test.yml:33","Info:   0 out of   4 GitHub-owned GitHubAction dependencies pinned","Info:   0 out of   5 pipCommand dependencies pinned"],"documentation":{"short":"Determines if the project has declared and pinned the dependencies of its build process.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#pinned-dependencies"}},{"name":"Vulnerabilities","score":0,"reason":"14 existing vulnerabilities detected","details":["Warn: Project is vulnerable to: GHSA-55x5-fj6c-h6m8","Warn: Project is vulnerable to: PYSEC-2021-19 / GHSA-jq4v-f5q6-mjqq","Warn: Project is vulnerable to: PYSEC-2020-62 / GHSA-pgww-xf46-h92r","Warn: Project is vulnerable to: PYSEC-2022-230 / GHSA-wrxv-2j5q-m38w","Warn: Project is vulnerable to: PYSEC-2021-856 / GHSA-5545-2q6w-2gh6","Warn: Project is vulnerable to: GHSA-6p56-wp2h-9hxr","Warn: Project is vulnerable to: PYSEC-2021-857 / GHSA-f7c7-j99h-c22f","Warn: Project is vulnerable to: GHSA-fpfv-jqm9-f5jm","Warn: Project is vulnerable to: PYSEC-2024-161","Warn: Project is vulnerable to: PYSEC-2021-142 / GHSA-8q59-q68h-6hv4","Warn: Project is vulnerable to: GHSA-9hjg-9r4m-mvj7","Warn: Project is vulnerable to: GHSA-9wx4-h78v-vm56","Warn: Project is vulnerable to: PYSEC-2023-74 / GHSA-j8r2-6x86-q33q","Warn: Project is vulnerable to: GHSA-g7vv-2v7x-gj9p"],"documentation":{"short":"Determines if the project has open, known unfixed vulnerabilities.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#vulnerabilities"}},{"name":"SAST","score":0,"reason":"SAST tool is not run on all commits -- score normalized to 0","details":["Warn: 0 commits out of 27 are checked with a SAST tool"],"documentation":{"short":"Determines if the project uses static code analysis.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#sast"}}]},"last_synced_at":"2025-08-16T19:42:04.403Z","repository_id":38107399,"created_at":"2025-08-16T19:42:04.403Z","updated_at":"2025-08-16T19:42:04.403Z"},"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":279014030,"owners_count":26085346,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-10-13T02:00:06.723Z","response_time":61,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["dataset","information-retrieval","ir"],"created_at":"2024-08-02T13:00:58.192Z","updated_at":"2025-10-13T15:55:44.816Z","avatar_url":"https://github.com/allenai.png","language":"Python","funding_links":[],"categories":["Python"],"sub_categories":[],"readme":"# ir_datasets\n\n`ir_datasets` is a python package that provides a common interface to many IR ad-hoc ranking\nbenchmarks, training datasets, etc.\n\nThe package takes care of downloading datasets (including documents, queries, relevance judgments,\netc.) when available from public sources. Instructions on how to obtain datasets are provided when\nthey are not publicly available.\n\n`ir_datasets` provides a common iterator format to allow them to be easily used in python. It\nattempts to provide the data in an unaltered form (i.e., keeping all fields and markup), while\nhandling differences in file formats, encoding, etc. Adapters provide extra functionality, e.g., to\nallow quick lookups of documents by ID.\n\nA command line interface is also available.\n\nYou can find a list of datasets and their features [here](https://ir-datasets.com/).\nWant a new dataset, added functionality, or a bug fixed? Feel free to post an issue or make a pull request! \n\n## Getting Started\n\nFor a quick start with the Python API, check out our Colab tutorials:\n[Python](https://colab.research.google.com/github/allenai/ir_datasets/blob/master/examples/ir_datasets.ipynb)\n[Command Line](https://colab.research.google.com/github/allenai/ir_datasets/blob/master/examples/ir_datasets_cli.ipynb)\n\nInstall via pip:\n\n```\npip install ir_datasets\n```\n\nIf you want the main branch, you install as such:\n\n```\npip install git+https://github.com/allenai/ir_datasets.git\n```\n\nIf you want to run an editable version locally:\n\n```\n$ git clone https://github.com/allenai/ir_datasets\n$ cd ir_datasets\n$ pip install -e .    \n```\n\nTested with python versions 3.7, 3.8, 3.9, and 3.10. (Mininum python version is 3.7.)\n\n## Features\n\n**Python and Command Line Interfaces**. Access datasts both through a simple Python API and\nvia the command line.\n\n```python\nimport ir_datasets\ndataset = ir_datasets.load('msmarco-passage/train')\n# Documents\nfor doc in dataset.docs_iter():\n    print(doc)\n# GenericDoc(doc_id='0', text='The presence of communication amid scientific minds was equa...\n# GenericDoc(doc_id='1', text='The Manhattan Project and its atomic bomb helped bring an en...\n# ...\n```\n\n```bash\nir_datasets export msmarco-passage/train docs | head -n2\n0 The presence of communication amid scientific minds was equally important to the success of the Manh...\n1 The Manhattan Project and its atomic bomb helped bring an end to World War II. Its legacy of peacefu...\n```\n\n**Automatically downloads source files** (when available). Will download and verify the source\nfiles for queries, documents, qrels, etc. when they are publicly available, as they are needed.\nA CI build checks weekly to ensure that all the downloadable content is available and correct:\n[![Downloadable Content](https://github.com/seanmacavaney/ir-datasets.com/actions/workflows/verify_downloads.yml/badge.svg)](https://github.com/seanmacavaney/ir-datasets.com/actions/workflows/verify_downloads.yml).\nWe mirror some troublesome files on [mirror.ir-datasets.com](https://mirror.ir-datasets.com/), and\nautomatically switch to the mirror when the original source is not available.\n\n```python\nimport ir_datasets\ndataset = ir_datasets.load('msmarco-passage/train')\nfor doc in dataset.docs_iter(): # Will download and extract MS-MARCO's collection.tar.gz the first time\n    ...\nfor query in dataset.queries_iter(): # Will download and extract MS-MARCO's queries.tar.gz the first time\n    ...\n```\n\n**Instructions for dataset access** (when not publicly available). Provides instructions on how\nto get a copy of the data when it is not publicly available online (e.g., when it requires a\ndata usage agreement).\n\n```python\nimport ir_datasets\ndataset = ir_datasets.load('trec-arabic')\nfor doc in dataset.docs_iter():\n    ...\n# Provides the following instructions:\n# The dataset is based on the Arabic Newswire corpus. It is available from the LDC via: \u003chttps://catalog.ldc.upenn.edu/LDC2001T55\u003e\n# To proceed, symlink the source file here: [gives path]\n```\n\n**Support for datasets big and small**. By using iterators, supports large datasets that may\nnot fit into system memory, such as ClueWeb.\n\n```python\nimport ir_datasets\ndataset = ir_datasets.load('clueweb09')\nfor doc in dataset.docs_iter():\n    ... # will iterate through all ~1B documents\n```\n\n**Fixes known dataset issues**. For instance, automatically corrects the document UTF-8 encoding\nproblem in the MS-MARCO passage collection.\n\n```python\nimport ir_datasets\ndataset = ir_datasets.load('msmarco-passage')\ndocstore = dataset.docs_store()\ndocstore.get('243').text\n# \"John Maynard Keynes, 1st Baron Keynes, CB, FBA (/ˈkeɪnz/ KAYNZ; 5 June 1883 – 21 April [SNIP]\"\n# Naïve UTF-8 decoding yields double-encoding artifacts like:\n# \"John Maynard Keynes, 1st Baron Keynes, CB, FBA (/Ë\\x88keÉªnz/ KAYNZ; 5 June 1883 â\\x80\\x93 21 April [SNIP]\"\n#                                                  ~~~~~~  ~~                       ~~~~~~~~~\n```\n\n**Fast Random Document Access.** Builds data structures that allow fast and efficient lookup of\ndocument content. For large datasets, such as ClueWeb, uses\n[checkpoint files](https://ir-datasets.com/clueweb_warc_checkpoints.md) to load documents from\nsource 40x faster than normal. Results are cached for even faster subsequent accesses.\n\n```python\nimport ir_datasets\ndataset = ir_datasets.load('clueweb12')\ndocstore = dataset.docs_store()\ndocstore.get_many(['clueweb12-0000tw-05-00014', 'clueweb12-0000tw-05-12119', 'clueweb12-0106wb-18-19516'])\n# {'clueweb12-0000tw-05-00014': ..., 'clueweb12-0000tw-05-12119': ..., 'clueweb12-0106wb-18-19516': ...}\n```\n\n**Fancy Iter Slicing.** Sometimes it's helpful to be able to select ranges of data (e.g., for processing\ndocument collections in parallel on multiple devices). Efficient implementations of slicing operations\nallow for much faster dataset partitioning than using `itertools.slice`.\n\n```python\nimport ir_datasets\ndataset = ir_datasets.load('clueweb12')\ndataset.docs_iter()[500:1000] # normal slicing behavior\n# WarcDoc(doc_id='clueweb12-0000tw-00-00502', ...), WarcDoc(doc_id='clueweb12-0000tw-00-00503', ...), ...\ndataset.docs_iter()[-10:-8] # includes negative indexing\n# WarcDoc(doc_id='clueweb12-1914wb-28-24245', ...), WarcDoc(doc_id='clueweb12-1914wb-28-24246', ...)\ndataset.docs_iter()[::100] # includes support for skip (only positive values)\n# WarcDoc(doc_id='clueweb12-0000tw-00-00000', ...), WarcDoc(doc_id='clueweb12-0000tw-00-00100', ...), ...\ndataset.docs_iter()[1/3:2/3] # supports proportional slicing (this takes the middle third of the collection)\n# WarcDoc(doc_id='clueweb12-0605wb-28-12714', ...), WarcDoc(doc_id='clueweb12-0605wb-28-12715', ...), ...\n```\n\n## Datasets\n\nAvailable datasets include:\n - [ANTIQUE](https://ir-datasets.com/antique.html)\n - [AQUAINT](https://ir-datasets.com/aquaint.html)\n - [BEIR (benchmark suite)](https://ir-datasets.com/beir.html)\n - [TREC CAR](https://ir-datasets.com/car.html)\n - [C4](https://ir-datasets.com/c4.html)\n - [ClueWeb09](https://ir-datasets.com/clueweb09.html)\n - [ClueWeb12](https://ir-datasets.com/clueweb12.html)\n - [CLIRMatrix](https://ir-datasets.com/clirmatrix.html)\n - [CodeSearchNet](https://ir-datasets.com/codesearchnet.html)\n - [CORD-19](https://ir-datasets.com/cord19.html)\n - [DPR Wiki100](https://ir-datasets.com/dpr-w100.html)\n - [GOV](https://ir-datasets.com/gov.html)\n - [GOV2](https://ir-datasets.com/gov2.html)\n - [HC4](https://ir-datasets.com/hc4.html)\n - [Highwire (TREC Genomics 2006-07)](https://ir-datasets.com/highwire.html)\n - [Medline](https://ir-datasets.com/medline.html)\n - [MSMARCO (document)](https://ir-datasets.com/msmarco-document.html)\n - [MSMARCO (passage)](https://ir-datasets.com/msmarco-passage.html)\n - [MSMARCO (QnA)](https://ir-datasets.com/msmarco-qna.html)\n - [Natural Questions](https://ir-datasets.com/natural-questions.html)\n - [NFCorpus (NutritionFacts)](https://ir-datasets.com/nfcorpus.html)\n - [NYT](https://ir-datasets.com/nyt.html)\n - [PubMed Central (TREC CDS)](https://ir-datasets.com/pmc.html)\n - [TREC Arabic](https://ir-datasets.com/trec-arabic.html)\n - [TREC Fair Ranking 2021](https://ir-datasets.com/trec-fair-2021.html)\n - [TREC Mandarin](https://ir-datasets.com/trec-mandarin.html)\n - [TREC Robust 2004](https://ir-datasets.com/trec-robust04.html)\n - [TREC Spanish](https://ir-datasets.com/trec-spanish.html)\n - [TripClick](https://ir-datasets.com/tripclick.html)\n - [Tweets 2013 (Internet Archive)](https://ir-datasets.com/tweets2013-ia.html)\n - [Vaswani](https://ir-datasets.com/vaswani.html)\n - [Washington Post](https://ir-datasets.com/wapo.html)\n - [WikIR](https://ir-datasets.com/wikir.html)\n\nThere are \"subsets\" under each dataset. For instance, `clueweb12/b13/trec-misinfo-2019` provides the\nqueries and judgments from the [2019 TREC misinformation track](https://trec.nist.gov/data/misinfo2019.html),\nand `msmarco-document/orcas` provides the [ORCAS dataset](https://microsoft.github.io/msmarco/ORCAS). They\ntend to be organized with the document collection at the top level.\n\nSee the ir_dataets docs ([ir_datasets.com](https://ir-datasets.com/)) for details about each\ndataset, its available subsets, and what data they provide.\n\n## Environment variables\n\n - `IR_DATASETS_HOME`: Home directory for ir_datasets data (default `~/.ir_datasets/`). Contains directories\n   for each top-level dataset.\n - `IR_DATASETS_TMP`: Temporary working directory (default `/tmp/ir_datasets/`).\n - `IR_DATASETS_DL_TIMEOUT`: Download stream read timeout, in seconds (default `15`). If no data is received\n   within this duration, the connection will be assumed to be dead, and another download may be attempted.\n - `IR_DATASETS_DL_TRIES`: Default number of download attempts before exception is thrown (default `3`).\n   When the server accepts Range requests, uses them. Otherwise, will download the entire file again\n - `IR_DATASETS_DL_DISABLE_PBAR`: Set to `true` to disable the progress bar for downloads. Useful in settings\n   where an interactive console is not available.\n - `IR_DATASETS_DL_SKIP_SSL`: Set to `true` to disable checking SSL certificates when downloading files.\n   Useful as a short-term solution when SSL certificates expire or are otherwise invalid. Note that this\n   does not disable hash verification of the downloaded content.\n - `IR_DATASETS_SKIP_DISK_FREE`: Set to `true` to disable checks for enough free space on disk before\n   downloading content or otherwise creating large files.\n - `IR_DATASETS_SMALL_FILE_SIZE`: The size of files that are considered \"small\", in bytes. Instructions for\n   linking small files rather then downloading them are not shown. Defaults to 5000000 (5MB).\n\n## Citing\n\nWhen using datasets provided by this package, be sure to properly cite them. Bibtex for each dataset\ncan be found on the [datasets documentation page](https://ir-datasets.com/).\n\nIf you use this tool, please cite [our SIGIR resource paper](https://arxiv.org/pdf/2103.02280.pdf):\n\n```\n@inproceedings{macavaney:sigir2021-irds,\n  author = {MacAvaney, Sean and Yates, Andrew and Feldman, Sergey and Downey, Doug and Cohan, Arman and Goharian, Nazli},\n  title = {Simplified Data Wrangling with ir_datasets},\n  year = {2021},\n  booktitle = {SIGIR}\n}\n```\n\n## Credits\n\nContributors to this repository:\n\n - Sean MacAvaney (University of Glasgow)\n - Shuo Sun (Johns Hopkins University)\n - Thomas Jänich (University of Glasgow)\n - Jan Heinrich Reimer (Martin Luther University Halle-Wittenberg)\n - Maik Fröbe (Martin Luther University Halle-Wittenberg)\n - Eugene Yang (Johns Hopkins University)\n - Augustin Godinot (NAVERLABS Europe, ENS Paris-Saclay)\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fallenai%2Fir_datasets","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fallenai%2Fir_datasets","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fallenai%2Fir_datasets/lists"}