{"id":51011992,"url":"https://github.com/daisybio/domainsplit","last_synced_at":"2026-06-21T04:01:35.427Z","repository":{"id":360973635,"uuid":"1239849713","full_name":"daisybio/domainsplit","owner":"daisybio","description":"A nextflow pipeline to split domain-domain interaction data into train, validation, and test while reducing data leakage and bias.","archived":false,"fork":false,"pushed_at":"2026-06-09T20:52:34.000Z","size":281,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-06-09T22:14:24.927Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":"Nextflow","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/daisybio.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":"docs/CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":"CITATIONS.md","codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2026-05-15T14:02:09.000Z","updated_at":"2026-05-28T16:13:03.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/daisybio/domainsplit","commit_stats":null,"previous_names":["daisybio/domainsplit"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/daisybio/domainsplit","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/daisybio%2Fdomainsplit","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/daisybio%2Fdomainsplit/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/daisybio%2Fdomainsplit/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/daisybio%2Fdomainsplit/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/daisybio","download_url":"https://codeload.github.com/daisybio/domainsplit/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/daisybio%2Fdomainsplit/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":34593129,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-06-21T02:00:05.568Z","response_time":54,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2026-06-21T04:01:31.559Z","updated_at":"2026-06-21T04:01:35.403Z","avatar_url":"https://github.com/daisybio.png","language":"Nextflow","funding_links":[],"categories":[],"sub_categories":[],"readme":"# daisybio/domainsplit\n\n[![Open in GitHub Codespaces](https://img.shields.io/badge/Open_In_GitHub_Codespaces-black?labelColor=grey\u0026logo=github)](https://github.com/codespaces/new/daisybio/domainsplit)\n[![GitHub Actions CI Status](https://github.com/daisybio/domainsplit/actions/workflows/nf-test.yml/badge.svg)](https://github.com/daisybio/domainsplit/actions/workflows/nf-test.yml)\n[![GitHub Actions Linting Status](https://github.com/daisybio/domainsplit/actions/workflows/linting.yml/badge.svg)](https://github.com/daisybio/domainsplit/actions/workflows/linting.yml)[![Cite with Zenodo](http://img.shields.io/badge/DOI-10.5281/zenodo.XXXXXXX-1073c8?labelColor=000000)](https://doi.org/10.5281/zenodo.XXXXXXX)\n[![nf-test](https://img.shields.io/badge/unit_tests-nf--test-337ab7.svg)](https://www.nf-test.com)\n\n[![Nextflow](https://img.shields.io/badge/version-%E2%89%A525.10.2-green?style=flat\u0026logo=nextflow\u0026logoColor=white\u0026color=%230DC09D\u0026link=https%3A%2F%2Fnextflow.io)](https://www.nextflow.io/)\n[![nf-core template version](https://img.shields.io/badge/nf--core_template-4.0.2-green?style=flat\u0026logo=nfcore\u0026logoColor=white\u0026color=%2324B064\u0026link=https%3A%2F%2Fnf-co.re)](https://github.com/nf-core/tools/releases/tag/4.0.2)\n[![run with conda](http://img.shields.io/badge/run%20with-conda-3EB049?labelColor=000000\u0026logo=anaconda)](https://docs.conda.io/en/latest/)\n[![run with docker](https://img.shields.io/badge/run%20with-docker-0db7ed?labelColor=000000\u0026logo=docker)](https://www.docker.com/)\n[![run with singularity](https://img.shields.io/badge/run%20with-singularity-1d355c.svg?labelColor=000000)](https://sylabs.io/docs/)\n[![Launch on Seqera Platform](https://img.shields.io/badge/Launch%20%F0%9F%9A%80-Seqera%20Platform-%234256e7)](https://cloud.seqera.io/launch?pipeline=https://github.com/daisybio/domainsplit)\n\n## Introduction\n\n**daisybio/domainsplit** is a bioinformatics pipeline that ...\nThis directory contains a Nextflow pipeline that downloads and processes all required public databases (3did, UniProt, Negatome, STRING, Pfam, pfam2go) to generate the `domainsplit.sqlite3` database of domain-domain interactions, then splits it into train/validation/test partitions using several leakage-reduction strategies.\n\n\u003c!-- TODO nf-core:\n   Complete this sentence with a 2-3 sentence summary of what types of data the pipeline ingests, a brief overview of the\n   major pipeline sections and the types of output it produces. You're giving an overview to someone new\n   to nf-core here, in 15-20 seconds. For an example, see https://github.com/nf-core/rnaseq/blob/master/README.md#introduction\n--\u003e\n\n\u003c!-- TODO nf-core: Include a figure that guides the user through the major workflow steps. Many nf-core\n     workflows use the \"tube map\" design for that. See https://nf-co.re/docs/community/brand/workflow-schematics#examples for examples.   --\u003e\n\u003c!-- TODO nf-core: Fill in short bullet-pointed list of the default steps in the pipeline --\u003e\n\n## Usage\n\n\u003e [!NOTE]\n\u003e If you are new to Nextflow and nf-core, please refer to [this page](https://nf-co.re/docs/get_started/environment_setup/overview) on how to set-up Nextflow. Make sure to [test your setup](https://nf-co.re/docs/get_started/run-your-first-pipeline) with `-profile test` before running the workflow on actual data.\n\nNow, you can run the pipeline using:\n\n\u003c!-- TODO nf-core: update the following command to include all required parameters for a minimal example --\u003e\n\n```bash\nnextflow run daisybio/domainsplit \\\n   -profile \u003cdocker/singularity/.../institute\u003e \\\n   --outdir \u003cOUTDIR\u003e\n```\n\n\u003e [!WARNING]\n\u003e Please provide pipeline parameters via the CLI or Nextflow `-params-file` option. Custom config files including those provided by the `-c` Nextflow option can be used to provide any configuration _**except for parameters**_; see [docs](https://nf-co.re/docs/running/run-pipelines#using-parameter-files).\n\n## Output\n\nThe main output is the SQLite database file `domainsplit.sqlite3` in the `results` directory, plus per-method, per-split databases under `results/split_databases/\u003cmethod\u003e/\u003csplit\u003e.sqlite3`.\n\n## Notes\n\n- Each process in `main.nf` is commented for clarity.\n- Some processes require external scripts (e.g., `mysql2sqlite` for 3did conversion).\n- You may need to adjust paths or URLs as needed.\n\n## Environment variables\n\nSome processes require secrets passed via environment variables. Copy `.env.example` to `.env` and fill in your values — the `.env` file is gitignored.\n\n| Variable   | Required             | Description                                                                                                     |\n| ---------- | -------------------- | --------------------------------------------------------------------------------------------------------------- |\n| `HF_TOKEN` | Yes (ESM embeddings) | HuggingFace access token for downloading ESM model weights. Obtain at \u003chttps://huggingface.co/settings/tokens\u003e. |\n\nOn a SLURM cluster, export the variable in your job script or pass it as a [Nextflow secret](https://www.nextflow.io/docs/latest/secrets.html):\n\n```bash\nnextflow secrets set HF_TOKEN \u003cyour_token\u003e\n```\n\n## Credits\n\ndaisybio/domainsplit was originally written by Konstantin Pelz, Christian Romberg, Chiara Thomas.\n\nWe thank the following people for their extensive assistance in the development of this pipeline:\n\n\u003c!-- TODO nf-core: If applicable, make list of people who have also contributed --\u003e\n\n## Contributions and Support\n\nIf you would like to contribute to this pipeline, please see the [contributing guidelines](docs/CONTRIBUTING.md).\n\n## Citations\n\n\u003c!-- TODO nf-core: Add citation for pipeline after first release. Uncomment lines below and update Zenodo doi and badge at the top of this file. --\u003e\n\u003c!-- If you use daisybio/domainsplit for your analysis, please cite it using the following doi: [10.5281/zenodo.XXXXXX](https://doi.org/10.5281/zenodo.XXXXXX) --\u003e\n\n\u003c!-- TODO nf-core: Add bibliography of tools and data used in your pipeline --\u003e\n\nAn extensive list of references for the tools used by the pipeline can be found in the [`CITATIONS.md`](CITATIONS.md) file.\n\nThis pipeline uses code and infrastructure developed and maintained by the [nf-core](https://nf-co.re) community, reused here under the [MIT license](https://github.com/nf-core/tools/blob/main/LICENSE).\n\n\u003e **The nf-core framework for community-curated bioinformatics pipelines.**\n\u003e\n\u003e Philip Ewels, Alexander Peltzer, Sven Fillinger, Harshil Patel, Johannes Alneberg, Andreas Wilm, Maxime Ulysse Garcia, Paolo Di Tommaso \u0026 Sven Nahnsen.\n\u003e\n\u003e _Nat Biotechnol._ 2020 Feb 13. doi: [10.1038/s41587-020-0439-x](https://dx.doi.org/10.1038/s41587-020-0439-x).\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fdaisybio%2Fdomainsplit","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fdaisybio%2Fdomainsplit","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fdaisybio%2Fdomainsplit/lists"}