{"id":47923667,"url":"https://github.com/ghga-de/nf-platypusindelcalling","last_synced_at":"2026-04-04T06:22:22.710Z","repository":{"id":68518664,"uuid":"492861735","full_name":"ghga-de/nf-platypusindelcalling","owner":"ghga-de","description":"This page is reserved for NextFlow based Indell Calling Workflow (with Platypus) from DKFZ","archived":false,"fork":false,"pushed_at":"2026-02-10T09:10:07.000Z","size":124741,"stargazers_count":2,"open_issues_count":6,"forks_count":0,"subscribers_count":6,"default_branch":"main","last_synced_at":"2026-02-10T14:52:21.418Z","etag":null,"topics":["annotation","annotations","dkfz","ghga","indel-calling","nextflow","nextflow-pipeline","odcf","pipeline","platypus","somatic-mutations","somatic-variants","variant-calling","workflow"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/ghga-de.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":"CODE_OF_CONDUCT.md","threat_model":null,"audit":null,"citation":"CITATIONS.md","codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2022-05-16T14:05:07.000Z","updated_at":"2025-10-31T16:04:30.000Z","dependencies_parsed_at":"2024-04-29T09:30:39.157Z","dependency_job_id":"9bd6538b-b7d8-4bcc-b7ec-338b26714640","html_url":"https://github.com/ghga-de/nf-platypusindelcalling","commit_stats":null,"previous_names":[],"tags_count":5,"template":false,"template_full_name":null,"purl":"pkg:github/ghga-de/nf-platypusindelcalling","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ghga-de%2Fnf-platypusindelcalling","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ghga-de%2Fnf-platypusindelcalling/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ghga-de%2Fnf-platypusindelcalling/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ghga-de%2Fnf-platypusindelcalling/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/ghga-de","download_url":"https://codeload.github.com/ghga-de/nf-platypusindelcalling/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/ghga-de%2Fnf-platypusindelcalling/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":31389822,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-04-04T04:26:24.776Z","status":"ssl_error","status_checked_at":"2026-04-04T04:23:34.147Z","response_time":60,"last_error":"SSL_read: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["annotation","annotations","dkfz","ghga","indel-calling","nextflow","nextflow-pipeline","odcf","pipeline","platypus","somatic-mutations","somatic-variants","variant-calling","workflow"],"created_at":"2026-04-04T06:22:22.073Z","updated_at":"2026-04-04T06:22:22.702Z","avatar_url":"https://github.com/ghga-de.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"[![Nextflow](https://img.shields.io/badge/nextflow%20DSL2-%E2%89%A521.10.3-23aa62.svg)](https://www.nextflow.io/)\n[![run with docker](https://img.shields.io/badge/run%20with-docker-0db7ed?logo=docker)](https://www.docker.com/)\n[![run with singularity](https://img.shields.io/badge/run%20with-singularity-1d355c.svg)](https://sylabs.io/docs/)\n\n\u003cp align=\"center\"\u003e\n    \u003cimg title=\"nf-platypusindelcalling workflow\" src=\"docs/images/nf-platypusindelcalling2.png\" width=70%\u003e\n\u003c/p\u003e\n\n## Introduction\n\n**nf-platypusindelcalling**:A Platypus-based insertion/deletion-detection workflow with extensive quality control additions. The workflow is based on DKFZ - ODCF OTP Indel Calling Pipeline.\n\nFor now, this workflow is only optimal to work in ODCF Cluster. The config file (conf/dkfz_cluster.config) can be used as an example. Running Annotation, DeepAnnotation, Filter and Tinda steps are optional and can be turned off using [runIndelAnnotation, runIndelDeepAnnotation, runIndelVCFFilter, runTinda] parameters sequential.\n\nThe pipeline is built using [Nextflow](https://www.nextflow.io), a workflow tool to run tasks across multiple compute infrastructures in a very portable manner. It uses Docker/Singularity containers making installation trivial and results highly reproducible. The [Nextflow DSL2](https://www.nextflow.io/docs/latest/dsl2.html) implementation of this pipeline uses one container per process which makes it much easier to maintain and update software dependencies.\n\nThis nextflow pipeline is the transition of [DKFZ-ODCF/IndelCallingWorkflow](https://github.com/DKFZ-ODCF/IndelCallingWorkflow).\n\n**Important Notice**: The whole workflow is only ready for DKFZ cluster users for now, It is strongly recommended to them to read whole documentation before usage. This workflow works better with nextflow/22.07.1-edge in the cluster, It is recommended to use \u003e22.07.1.\n\n## Pipeline summary\n\nThe pipeline has 6 main steps: Indel calling using platypus, basic annotations, deep annotations, filtering, sample swap check and multiqc report.\n\n1. Indel Calling:\n\n   Platypus ([`Platypus`](https://www.well.ox.ac.uk/research/research-groups/lunter-group/lunter-group/platypus-a-haplotype-based-variant-caller-for-next-generation-sequence-data))\n   : Platypus tool is used to call variants using local realignmnets and local assemblies. It can detect SNPs, MNPs, short indels, replacements, deletions up to several kb. It can be both used with WGS and WES. The tool has been thoroughly tested on data mapped with Stampy and BWA.\n\n2. Basic Annotations (--runIndelAnnotation True):\n\n   In-house scripts to annotate with several databases like gnomAD, dbSNP, and ExAC.\n\n   ANNOVAR ([`Annovar`](https://annovar.openbioinformatics.org/en/latest/))\n   : annotate_variation.pl is used to annotate variants. The tool makes classifications for intergenic, intogenic, nonsynoymous SNP, frameshift deletion or large-scale duplication regions.\n\n   ENSEMBL VEP(['ENSEBL VEP'](https://www.ensembl.org/info/docs/tools/vep/index.html)) :can also be used alternative to annovar. Gene annotations will be extracted.\n\n   Reliability and confidation annotations: It is an optional ste for mapability, hiseq, selfchain and repeat regions checks for reliability and confidence of those scores.\n\n3. Deep Annotation (--runIndelDeepAnnotation True):\n\n   If basic annotations are applied, an extra optional step for number of extra indel annotations like enhancer, cosmic, mirBASE, encode databases can be applied too.\n\n4. Filtering and Visualization (--runIndelVCFFilter True):\n\n   It is an optional step. Filtering is only required for the tumor samples with no-control and filtering can only be applied if basic annotation is performed.\n\n   Indel Extraction and Visualizations: INDELs can be extracted by certain minimum confidence level\n\n   Visualization and json reports: Extracted INDELs are visualized and analytics of INDEL categories are reported as JSON.\n\n5. Check Sample Swap (--runTinda True):\n\n   Canopy Based Clustering and Bias Filter, thi step can only be applied into the tumor samples with control.\n\n6. MultiQC (--skipmultiqc False):\n\n   Produces pipeline level analytics and reports.\n\n## Quick Start\n\n1. Install [`Nextflow`](https://www.nextflow.io/docs/latest/getstarted.html#installation) (`\u003e=21.10.3`)\n\n2. Install any of [`Docker`](https://docs.docker.com/engine/installation/) or [`Singularity`](https://www.sylabs.io/guides/3.0/user-guide/) (you can follow [this tutorial](https://singularity-tutorial.github.io/01-installation/))\n\n3. Download [Annovar](https://annovar.openbioinformatics.org/en/latest/user-guide/download/) and set-up suitable annotation table directory to perform annotation. Example:\n\n```console\nannotate_variation.pl -downdb wgEncodeGencodeBasicV19 humandb/ -build hg19\n```\n\nGene annotation is also possible with ENSEMBL VEP tool, for test purposes only, it can be used online. But for big analysis, it is recommended to either download cache file or use --download_cache flag in parameters.\n\nFollow the documentation [here](https://www.ensembl.org/info/docs/tools/vep/script/vep_cache.html#cache)\n\nExample:\n\nDownload [cache](https://ftp.ensembl.org/pub/release-110/variation/indexed_vep_cache/)\n\n```console\ncd $HOME/.vep\ncurl -O https://ftp.ensembl.org/pub/release-110/variation/indexed_vep_cache/homo_sapiens_vep_110_GRCh38.tar.gz\ntar xzf homo_sapiens_vep_110_GRCh38.tar.gz\n```\n\n4. Download the pipeline and test it on a minimal dataset with a single command:\n\n   ```console\n   git clone https://github.com/ghga-de/nf-platypusindelcalling.git\n   ```\n\nbefore run do this to bin directory, make it runnable!:\n\n```console\nchmod +x bin/*\n```\n\n```console\nnextflow run main.nf -profile test,YOURPROFILE --outdir \u003cOUTDIR\u003e --input \u003cSAMPLESHEET\u003e\n```\n\nNote that some form of configuration will be needed so that Nextflow knows how to fetch the required software. This is usually done in the form of a config profile (`YOURPROFILE` in the example command above). You can chain multiple config profiles in a comma-separated string.\n\n\u003e - The pipeline comes with config profiles called `docker` and `singularity` which instruct the pipeline to use the named tool for software management. For example, `-profile test,docker`.\n\u003e - Please check [nf-core/configs](https://github.com/nf-core/configs#documentation) to see if a custom config file to run nf-core pipelines already exists for your Institute. If so, you can simply use `-profile \u003cinstitute\u003e` in your command. This will enable either `docker` or `singularity` and set the appropriate execution settings for your local compute environment.\n\u003e - If you are using `singularity`, please use the [`nf-core download`](https://nf-co.re/tools/#downloading-pipelines-for-offline-use) command to download images first, before running the pipeline. Setting the [`NXF_SINGULARITY_CACHEDIR` or `singularity.cacheDir`](https://www.nextflow.io/docs/latest/singularity.html?#singularity-docker-hub) Nextflow options enables you to store and re-use the images from a central location for future pipeline runs.\n\n5. Simple test run\n\n   ```console\n   nextflow run main.nf --outdir results -profile singularity,dkfz_cluster_38\n   ```\n\n6. Start running your own analysis!\n\n   \u003c!-- TODO nf-core: Update the example \"typical command\" below used to run the pipeline --\u003e\n\n   ```console\n   nextflow run main.nf --input samplesheet.csv --outdir \u003cOUTDIR\u003e -profile \u003cdocker/singularity\u003e --config test/institute.config\n   ```\n\n## Samplesheet columns\n\n**sample**: The sample name will be tagged to the job\n\n**tumor**: The path to the tumor file\n\n**tumor_index**: The path to the tumor index file\n\n**control**: The path to the control file, if there is no control will be kept blank.\n\n**control_index**: The path to the control index file, if there is no control will be kept blank.\n\n## Data Requirements\n\nAnnotations are optional for the user.\nAll VCF and BED files need to be indexed with tabix and should be in the same folder!\n\nThe reference set bundle which is used in PCAWG study can be found and downloaded [here](https://dcc.icgc.org/api/v1/download?fn=/PCAWG/reference_data/pcawg-dkfz/dkfz-workflow-dependencies_150318_0951.tar.gz). (NOTE: only in hg19)\n\n**Basic Annotation Files**\n\n- dbSNP INDELs (vcf)\n- 1000K INDELs (vcf)\n- gnomAD Genome Sites for INDELs (vcf)\n- gnomAD Exome Sites for INDELs (vcf)\n- EVS variants (vcf)\n- ExAC variants (vcf)\n- Local Control files WGS (vcf)\n- Local Control files WES (vcf)\n\n**SNV Reliability Files**\n\n- UCSC Repeat Masker region (bed)\n- UCSC Mappability regions (bed)\n- UCSC Simple tandem repeat regions (bed)\n- UCSC DAC Black List regions (bed)\n- UCSC DUKE Excluded List regions (bed)\n- UCSC Hiseq Deep sequencing regions (bed)\n- UCSC Self Chain regions (bed)\n\n**Deep Annotation Files**\n\n- UCSC Enhangers (bed)\n- UCSC CpG islands (bed)\n- UCSC TFBS noncoding sites (bed)\n- UCSC Encode DNAse cluster (bed.gz)\n- snoRNAs miRBase (bed)\n- miRBase (bed)\n- Cosmic coding SNVs (bed)\n- miRNA target sites (bed)\n- Cgi Mountains (bed)\n- UCSC Phast Cons Elements (bed)\n- UCSC Encode TFBS (bed)\n\n## Reference Usage\n\nThis pipeline favors the use of [igenomes](https://support.illumina.com/sequencing/sequencing_software/igenome.html) and [refgenie](http://refgenie.databio.org/en/latest/overview/). Read the documentaton [here](https://nf-co.re/usage/reference_genomes) to learn more.\n\nFor igenomes usage: use genomes GRCh37 (--genome \"GRCh37\") or GRCh38 (--genome \"GRCh38\").\n\nFor refgenie usage: use genomes GRCh37 (--genome \"hg37\") or GRCh38 (--genome \"hg38\").\n\nIf not using igenomes or refgenie, --fasta, --fasta_fai, and --chr_prefix need to be spesifed! If --chr_sizes is not provided it will be automatically generated.\n\n## Annotation files\n\n## Documentation\n\nThe nf-platypusindelcalling pipeline comes with documentation about the pipeline [usage](https://github.com/ghga-de/nf-platypusindelcalling/blob/main/docs/usage.md) and [output](https://github.com/ghga-de/nf-platypusindelcalling/blob/main/docs/output.md).\n\nPlease read [usage](https://github.com/ghga-de/nf-platypusindelcalling/blob/main/docs/usage.md) document to learn how to perform sample analysis provided with this repository!\n\n## Credits\n\nnf-platypusindelcalling was originally translated from roddy-based pipeline by Kuebra Narci kuebra.narci@dkfz-heidelberg.de.\n\nThe pipeline is originally written in workflow management language Roddy. [Inspired github page](https://github.com/DKFZ-ODCF/IndelCallingWorkflow)\n\nThe Indel calling workflow was in the pan-cancer analysis of whole genomes (PCAWG) and can be cited in the following publication:\n\nPan-cancer analysis of whole genomes. The ICGC/TCGA Pan-Cancer Analysis of Whole Genomes Consortium. Nature volume 578, pages 82–93 (2020). DOI 10.1038/s41586-020-1969-6\n\nWe thank the following people for their extensive assistance in the development of this pipeline:\n\n- Nagarajan Paramasivam (@NagaComBio) n.paramasivam@dkfz.de\n\n**TODO**\n\n\u003c!-- TODO nf-core: If applicable, make list of people who have also contributed --\u003e\n\n## Contributions and Support\n\nIf you would like to contribute to this pipeline, please see the [contributing guidelines](.github/CONTRIBUTING.md).\n\n## Citations\n\n\u003c!-- If you use  nf-platypusindelcalling for your analysis, please cite it using the following doi: [10.5281/zenodo.XXXXXX](https://doi.org/10.5281/zenodo.XXXXXX) --\u003e\n\n\u003c!-- TODO nf-core: Add bibliography of tools and data used in your pipeline --\u003e\n\nAn extensive list of references for the tools used by the pipeline can be found in the [`CITATIONS.md`](CITATIONS.md) file.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fghga-de%2Fnf-platypusindelcalling","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fghga-de%2Fnf-platypusindelcalling","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fghga-de%2Fnf-platypusindelcalling/lists"}