{"id":16713133,"url":"https://github.com/adamtaranto/tirmite","last_synced_at":"2026-01-17T01:58:01.827Z","repository":{"id":47919179,"uuid":"100831217","full_name":"Adamtaranto/TIRmite","owner":"Adamtaranto","description":"Map TIR-pHMM models to genomic sequences for annotation of MITES and complete DNA-Transposons.","archived":false,"fork":false,"pushed_at":"2024-11-28T01:05:43.000Z","size":289,"stargazers_count":10,"open_issues_count":9,"forks_count":4,"subscribers_count":4,"default_branch":"main","last_synced_at":"2025-03-18T21:50:45.577Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"other","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/Adamtaranto.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE.txt","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2017-08-20T01:31:18.000Z","updated_at":"2025-01-22T16:08:55.000Z","dependencies_parsed_at":"2024-11-28T01:23:45.841Z","dependency_job_id":"90574348-d3d3-450e-a94f-d04a511eec8a","html_url":"https://github.com/Adamtaranto/TIRmite","commit_stats":{"total_commits":26,"total_committers":1,"mean_commits":26.0,"dds":0.0,"last_synced_commit":"e2e85d73fa1df3a1c5d9893f7b35bcb6f6a1558b"},"previous_names":[],"tags_count":6,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Adamtaranto%2FTIRmite","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Adamtaranto%2FTIRmite/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Adamtaranto%2FTIRmite/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Adamtaranto%2FTIRmite/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/Adamtaranto","download_url":"https://codeload.github.com/Adamtaranto/TIRmite/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":245116031,"owners_count":20563267,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-10-12T20:45:37.009Z","updated_at":"2026-01-17T01:58:01.802Z","avatar_url":"https://github.com/Adamtaranto.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"[![License: GPL v3](https://img.shields.io/badge/License-GPLv3-blue.svg)](https://www.gnu.org/licenses/gpl-3.0)\n[![PyPI version](https://badge.fury.io/py/TIRmite.svg)](https://badge.fury.io/py/TIRmite)\n[![codecov](https://codecov.io/gh/Adamtaranto/TIRmite/graph/badge.svg?token=DFEEPKDFZ0)](https://codecov.io/gh/Adamtaranto/TIRmite)\n[![install with bioconda](https://img.shields.io/badge/install%20with-bioconda-brightgreen.svg?style=flat)](http://bioconda.github.io/recipes/tirmite/README.html)\n[![Downloads](https://pepy.tech/badge/tirmite)](https://pepy.tech/project/tirmite)\n\n\u003cp align=\"center\"\u003e\n\u003cimg src=\"https://raw.githubusercontent.com/Adamtaranto/TIRmite/main/docs/tirmite_hexlogo.jpg\" width=\"256\" height=\"256\" title=\"tirmite_hex\" /\u003e\n\u003c/p\u003e\n\n# TIRmite\n\nAutonomous examples of transposons, belonging to many [distinct super-families](https://doi.org/10.1266/ggs.18-00024 \"Yes yes, except when they don't. Don't @ me, nerds.\"), share two common properties: A gene or genes encoding the mode of transposition; and terminal sequence features that are recognised by these gene products as the element boundaries.\n\nProper classification of transposons and grouping into families relies on both phylogeny of conserved sequences and conservation of transposition mechanism.\n\nHowever, not all TE instances are created equal — inhabiting the nulear soup of their host genome, where your brother's transposase is as good as your own, non-autonomous variants (lacking their own functional hardware) proliferate.\n\nMITEs are a classic example of this - derived from autonomous DNA elements with Terminal Inverted Repeats, they are Miniature Inverted-repeat Transposable Elements, sometimes little more than a pair of TIRs.\n\nWhen non-autonomous structural variants of a TE vastly outnumber their parent element, and include forms that capture novel genes (or other full transposons!), it becomes difficult to correctly cluster related elements based on the limited signal present in terminal sequences (TIRs, LTRs, etc).\n\n**TIRmite** employs profile Hidden Markov Models (HMMs) to model natural variation in transposon termini and recover divergent and degraded hits that are often missed by sequence-based aligners like BLAST.\n\nAn iterative pairing algorithm is then used to annotate cryptic transposon variants with variable internal sequence compositions.\n\nThe elements extracted by TIRmite generally represent structuaral variants derived from an autonomous ancestor and may be further clustered into families.\n\n# Table of contents\n\n* [About TIRmite](#about-tirmite)\n* [Options and usage](#options-and-usage)\n  * [Installing TIRmite](#installing-tirmite)\n  * [Example usage](#example-usage)\n  * [Standard options](#standard-options)\n* [Algorithm overview](#algorithm-overview)\n* [Contributing](#contributing)\n* [Issues](#issues)\n* [License](#license)\n\n## About TIRmite\n\nTIRmite will use profile-HMM models of Transposon Terminal Repeats for genome-wide annotation of transposon families. You can search for TE families with symmetrical termini (i.e. TIRs or LTRs) or asymmetrical elements with different conserved features at either end (i.e. Helitrons, Helentrons, and Starship elements).\n\nThree classes of output are produced:\n\n  1. All significant termini hit sequences are written to fasta (per query HMM).\n  2. Candidate elements comprised of paired termini are written to fasta (per query HMM).\n  3. Genomic annotations of candidate elements and, optionally, HMM hits\n  (paired and unpaired) are written as a single GFF3 file.\n\n## Options and usage\n\n### Installing TIRmite\n\nTIRmite requires Python \u003e= v3.8\n\nDependencies:\n\n* [HMMER3](http://hmmer.org)\n* [mafft](https://mafft.cbrc.jp/alignment/software/)\n* [BLAST+](ftp://ftp.ncbi.nlm.nih.gov/blast/executables/blast+/LATEST/) (Optional)\n\nYou can create a Conda environment with these dependencies using the `environment.yml` file in this repo.\n\n```bash\nconda env create -f environment.yml\n\nconda activate tirmite\n```\n\nInstallation options:\n\n1) `pip install` the latest development version directly from this repo.\n\n```bash\npip install git+https://github.com/Adamtaranto/TIRmite.git\n```\n\n2) Install latest release from PyPi.\n\n```bash\npip install tirmite\n```\n\n3) Install latest release (with dependencies) from Bioconda.\n\n```bash\nconda install -c bioconda tirmite\n```\n\nTest installation.\n\n```bash\n# Print version number and exit.\n% tirmite --version\ntirmite 1.3.0\n\n# Get usage information\n% tirmite --help\n```\n\n### Example usage\n\nFirst, you will need to build a pHMM of your element's terminal sequence/s.\n\nIf you have a draft TE model (i.e. from RepeatModeler or EDTA) and want to identify the TIR's or LTR's to use with TIRmite - I recommend using [*tSplit*](https://github.com/Adamtaranto/TE-splitter/) a tool for extraction of terminal repeats from complete transposons.\n\n\n1) Extract single TIR from sample element:\n\n```bash\n# Uses BLASTn to detect TIRs of min 40% identity and min 10 bp length\ntsplit TIR -i TIR_element.fa -d tsplit_results --minid 0.4 --method blastn --minterm 10 --splitmode external\n```\n\n2) Build a pHMM from the seed:\n\n```bash\nGENOME=\"genome.fa\" # Path to fasta containing one or more genomes to search for matches to seed sequence.\n\ntirmite seed --left-seed tsplit_results/TIR_element_tsplit_output.fasta --model-name MY_TIR --outdir MY_TIR_HMM --genome $GENOME --max-gap 10 --save-blast-hits --threads 8\n\n# Note: Setting `--flank-size 10` will output additional flanking bases outside the TIR, conservation in the flank accross many independent insertions may indicate your seed was truncated. Always check and adjust seed as required.\n```\n\n3) Use `nhmmer` to locate hits to the TIR-pHMM in a target genome.\n\n```bash\nHMMFILE=\"MY_TIR_HMM/MY_TIR.hmm\"\nNHMMERFILE=\"MY_TIR_nhmmer_hits.tab\"\nnhmmer --dna --cpu 8 --tblout $NHMMERFILE $HMMFILE $GENOME\n```\n\n**Custom DNA Matrices**\n\nNote: nhmmer can be supplied with custom DNA score matrices for assessing hmm match scores.\nStandard NCBI-BLAST matrices such as NUC.4.4 are compatible. (See: ftp://ftp.ncbi.nlm.nih.gov/blast/matrices/NUC.4.4)\n\n4) Use `tirmite pair` to identify valid TIR pairs. Outputs hits, elements, and annotations.\n\n```bash\ntirmite pair --genome $GENOME  --nhmmerFile $NHMMERFILE --hmmFile $HMMFILE --orientation F,R --mincov 0.4 --report all  --maxdist 20000 --stableReps 2 --outdir MY_TIR_PAIRING_OUTPUT --padlen 20 --maxeval 0.001 --gffOut --logfile\n```\n\n#### Legacy mode\n\nTIRmite `legacy` mode will take a TIR-pHMM and target genome fasta as input and run the full standard workflow, reporting all hits, valid pairings, and write GFF3 annotation file.\n\nNote: This usage will be phased out in a later release in favour of custom workflows.\n\n```bash\n# Use HMM search to pull more divergent TIR hits from your query genome.\n# TIR hits are paired in Fwd/Rev orientation\n# Fwd/Rev pairs must be within 20Kbp of each other\n# Hits must cover \u003e= 40% of the TIR-pHMM\ntirmite legacy --genome $GENOME --hmmFile $HMMFILE--orientation F,R \\\n--outdir results \\\n--stableReps 2 \\\n--report all \\\n--gffOut --maxdist 20000 --mincov 0.4\n```\n\nIf you don't have a HMM of your TIR, `tirmite legacy` can create one for you using an aligned sample of your TIR provided with `--alnFile`.\n\nTIRs should always be oriented 5\\`- 3\\` with the lefthand TIR.\n\nIn this example the two TIRs should be oriented to begin with \"GA\".\n\n5\\` **GA\\\u003e\\\u003e\\\u003e\\\u003e\\\u003e\\\u003e\\\u003e** ATGC \u003c\u003c\u003c\u003c\u003c\u003c\u003cTC 3\\`\n3\\` CT\u003e\u003e\u003e\u003e\u003e\u003e\u003e\u003e  TACG \u003c\u003c\u003c\u003c\u003c\u003c\u003cAG 5\\`\n\n### Standard options\n\nRun `tirmite --help` to view the program's most commonly used options:\n\n```code\ntirmite --help\nusage: tirmite [-h] [--version] COMMAND ...\n\nTIRmite: Transposon Terminal Repeat detection suite\n\npositional arguments:\n  COMMAND     Available subcommands\n    legacy    Original TIRmite workflow (HMM search + pairing)\n    seed      Build HMM models from seed sequences\n    pair      Pair precomputed nhmmer hits\n\noptions:\n  -h, --help  show this help message and exit\n  --version   show program's version number and exit\n\nAvailable subcommands:\n  legacy    Original TIRmite workflow (HMM search + pairing)\n  seed      Build HMM models from seed sequences\n  pair      Pair precomputed nhmmer hits\n\nExamples:\n  tirmite legacy --genome genome.fa --hmmFile model.hmm\n  tirmite seed --left-seed left.fa --model-name myTE --genome genome.fa\n  tirmite pair --genome genome.fa --nhmmerFile hits.out --hmmFile model.hmm\n```\n\n## Algorithm overview\n\n  1. Use nhmmer to query genome with termini HMMs.\n  2. Import all hits under *--maxeval* threshold.\n  3. For each significant terminus match, identify candidate partners, where:\n    - Hit is on the same sequence.\n    - Hit is in coreect relative orientation.\n    - Distance is \u003c= *--maxdist*.\n    - Hit length is \u003e= (model length * *--mincov* prop)\n  4. Rank candidate partners by distance downstream of positive-strand hits, and upstream of negative-strand hits.\n  5. Pair reciprocal top candidate hits.\n  6. For unpaired hits, find nearest unpaired candidate partner and check for reciprocity.\n  7. If the first unpaired candidate is non-reciprocal, check for 2nd-order reciprocity (is outbound top-candidate of current candidate reciprocal.)\n  8. Iterate steps 6-7 until all termini hits are paired OR number of iterations without new pairing exceeds *--stableReps*.\n\n## Contributing\n\nIf you would like to add a new feature or fix a bug, please see our [contribution guidelines](https://github.com/Adamtaranto/TIRmite?tab=contributing-ov-file#readme).\n\n- Open an issue\n- Fork the repo\n- Follow the dev env setup instructions below\n\n```bash\n# Clone this repo (or your own fork)\ngit clone https://github.com/Adamtaranto/TIRmite.git \u0026\u0026 cd TIRmite\n# Install custom conda env\nconda env create -f environment.yml\n# Activate conda env\nconda activate tirmite\n# Install an editable copy of the package\npip install -e '.[dev]'\n# Enable pre-commit checks\npre-commit install\n```\n\n## Issues\n\nSubmit feedback to the [Issue Tracker](https://github.com/Adamtaranto/TIRmite/issues)\n\n## License\n\nSoftware provided under GPL-3 license.\n\n## Star History\n\n[![Star History\nChart](https://api.star-history.com/svg?repos=adamtaranto/tirmite\u0026type=Date)](https://star-history.com/#adamtaranto/tirmite\u0026Date)\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fadamtaranto%2Ftirmite","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fadamtaranto%2Ftirmite","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fadamtaranto%2Ftirmite/lists"}