{"id":31322887,"url":"https://github.com/pinellolab/crispr-hawk","last_synced_at":"2026-06-08T06:33:03.436Z","repository":{"id":315064388,"uuid":"874151136","full_name":"pinellolab/CRISPR-HAWK","owner":"pinellolab","description":"Variant- and Haplotype-aware CRISPR guide design toolkit","archived":false,"fork":false,"pushed_at":"2026-05-12T17:12:50.000Z","size":75349,"stargazers_count":7,"open_issues_count":0,"forks_count":4,"subscribers_count":1,"default_branch":"main","last_synced_at":"2026-05-12T18:29:14.347Z","etag":null,"topics":["bioinformatics","crispr","crispr-design","genome-editing"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"other","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/pinellolab.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":"CODE_OF_CONDUCT.md","threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2024-10-17T10:41:20.000Z","updated_at":"2026-05-12T17:13:22.000Z","dependencies_parsed_at":"2025-10-20T14:39:29.498Z","dependency_job_id":null,"html_url":"https://github.com/pinellolab/CRISPR-HAWK","commit_stats":null,"previous_names":["manueltgn/crispr-hawk","pinellolab/crispr-hawk"],"tags_count":5,"template":false,"template_full_name":null,"purl":"pkg:github/pinellolab/CRISPR-HAWK","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/pinellolab%2FCRISPR-HAWK","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/pinellolab%2FCRISPR-HAWK/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/pinellolab%2FCRISPR-HAWK/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/pinellolab%2FCRISPR-HAWK/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/pinellolab","download_url":"https://codeload.github.com/pinellolab/CRISPR-HAWK/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/pinellolab%2FCRISPR-HAWK/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":34051768,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-06-08T02:00:07.615Z","response_time":111,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["bioinformatics","crispr","crispr-design","genome-editing"],"created_at":"2025-09-25T19:27:25.286Z","updated_at":"2026-06-08T06:33:03.430Z","avatar_url":"https://github.com/pinellolab.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"[![install with bioconda](https://img.shields.io/badge/install%20with-bioconda-brightgreen.svg?style=flat)](http://bioconda.github.io/recipes/crisprhawk/README.html)\n[![GitHub Release Version](https://img.shields.io/github/v/release/pinellolab/CRISPR-HAWK)](https://github.com/pinellolab/CRISPR-HAWK/releases)\n[![Build Status](https://github.com/pinellolab/CRISPR-HAWK/actions/workflows/python-package.yml/badge.svg)](https://github.com/pinellolab/CRISPR-HAWK/actions/workflows/python-package.yml)\n![license](https://img.shields.io/badge/license-AGPL--3.0-lightgrey)\n\n\n\u003cp align=\"center\"\u003e\n    \u003cimg src=\"assets/readme/logo.jpg\", alt=\"logo.jpg\"\u003e\n\u003c/p\u003e\n\nCRISPR-HAWK is a comprehensive and scalable tool for designing guide RNAs (gRNAs) and assessing genetic variants impact on on-target sites in CRISPR-Cas systems. Available as an offline tool with a user-friendly command-line interface, CRISPR-HAWK integrates large-scale human genetic variation datasets—such as the 1000 Genomes Project, the Human Genome Diversity Project (HGDP), and gnomAD—with orthogonal genomic annotations to systematically prioritize gRNAs targeting regions of interest. CRISPR-HAWK is Cas system-independent, supporting a wide range of nucleases including Cas9, SaCas9, Cpf1 (Cas12a), and others. It offers users full flexibility to define custom PAM sequences and guide lengths, enabling compatibility with emerging CRISPR technologies and tailored experimental requirements. The tool accounts for both single-nucleotide variants (SNVs) and small insertions and deletions (indels), and it is capable of handling individual- and population-specific haplotypes. This makes CRISPR-HAWK particularly suitable for both personalized and population-wide gRNA design. CRISPR-HAWK automates the entire workflow—from variant-aware preprocessing to gRNA discovery—delivering comprehensive outputs including ranked tables, annotated sequences, and high-quality figures. Its modular design ensures easy integration with existing pipelines and tools, such as [CRISPRme](https://github.com/pinellolab/CRISPRme) or [CRISPRitz](https://github.com/pinellolab/CRISPRitz), for subsequent off-target prediction and analysis of prioritized gRNAs.\n\n## Table of Contents\n\n0 [System Requirements](#0-system-requirements)\n\u003cbr\u003e1 [Installation](#1-installation)\n\u003cbr\u003e\u0026nbsp;\u0026nbsp;1.1 [Install CRISPR-HAWK from Mamba/Conda](#11-install-crispr-hawk-from-mambaconda)\n\u003cbr\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;1.1.1 [Install Conda or Mamba](#111-install-conda-or-mamba)\n\u003cbr\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;1.1.2 [Install CRISPR-HAWK](#112-install-crispr-hawk)\n\u003cbr\u003e\u0026nbsp;\u0026nbsp;1.2 [Install CRISPR-HAWK from Docker](#12-install-crispr-hawk-from-docker)\n\u003cbr\u003e\u0026nbsp;\u0026nbsp;1.3 [Install CRISPR-HAWK from PyPI](#13-install-crispr-hawk-from-pypi)\n\u003cbr\u003e\u0026nbsp;\u0026nbsp;1.4 [Install CRISPR-HAWK from Source Code](#14-install-crispr-hawk-from-source-code)\n\u003cbr\u003e\u0026nbsp;\u0026nbsp;1.5 [Install External Software Dependencies](#15-install-external-software-dependencies)\n\u003cbr\u003e\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;1.5.1 [Install CRISPRitz (for Off-target Estimation)](#151-install-crispritz-for-off-target-estimation)\n\u003cbr\u003e2 [Usage](#2-usage)\n\u003cbr\u003e\u0026nbsp;\u0026nbsp;2.1 [General Syntax](#21-general-syntax)\n\u003cbr\u003e\u0026nbsp;\u0026nbsp;2.2 [Search](#22-search)\n\u003cbr\u003e\u0026nbsp;\u0026nbsp;2.3 [Convert gnomAD VCF](#23-convert-gnomad-vcf)\n\u003cbr\u003e\u0026nbsp;\u0026nbsp;2.4 [Prepare Data for CRISPRme](#24-prepare-data-for-crisprme)\n\u003cbr\u003e3 [Test](#3-test)\n\u003cbr\u003e\u0026nbsp;\u0026nbsp;3.1 [Quick Test After Installation](#31-quick-test-after-installation)\n\u003cbr\u003e\u0026nbsp;\u0026nbsp;3.2 [Run Full Test Suite with PyTest](#32-run-full-test-suite-with-pytest)\n\u003cbr\u003e\u0026nbsp;\u0026nbsp; 3.3 [Troubleshooting](#33-troubleshooting)\n\u003cbr\u003e\u0026nbsp;\u0026nbsp; 3.4 [Reporting Issues](#34-reporting-issues)\n\u003cbr\u003e4 [Citation](#4-citation)\n\u003cbr\u003e5 [Contacts](#5-contacts)\n\u003cbr\u003e6 [License](#6-license)\n\n## 0 System Requirements\n\nTo ensure optimal performance, CRISPR-HAWK requires the following system specifications:\n\n- **Operating System**:\n\u003cbr\u003emacOS or any modern Linux distribution (e.g., Ubuntu, CentOS)\n\n- **Required Disk Space**:\n\u003cbr\u003e 3.5 GB\n\n- **Minimum RAM**:\n\u003cbr\u003e16 GB — sufficient for standard use cases and small to medium-sized datasets\n\n- **Recommended RAM for Large-Scale Analyses**:\n\u003cbr\u003e32 GB or more — recommended for memory-intensive tasks such as:\n\n    - Scanning regions larger than 1 Mb\n\n    - Processing large-scale variant datasets (e.g., gnomAD data)\n\n\u003e 📝 **Note**: For optimal performance and stability, especially when dealing with large-scale variant datasets, ensure that your system meets or exceeds the recommended specifications.\n\n## 1 Installation\n\nThis section provides step-by-step instructions to install CRISPR-HAWK and external dependencies. Choose the method that best suits your environment and preferences:\n\n- **[Install CRISPR-HAWK from Mamba/conda](#11-install-crispr-hawk-from-mambaconda)** (recommended)\n\u003cbr\u003eBest for users seeking an isolated and reproducible environment with minimal manual dependency handling.\n\n- **[Install CRISPR-HAWK from Docker](#12-install-crispr-hawk-from-docker)**\n\u003cbr\u003eIdeal for users who prefer containerized deployments or want to avoid configuring the environment manually.\n\n- **[Install CRISPR-HAWK from PyPI](#13-install-crispr-hawk-from-pypi)**\n\u003cbr\u003eQuick option for Python users already working within a virtual environment. May require manual handling of some dependencies.\n\n- **[Install CRISPR-HAWK from source code](#14-install-crispr-hawk-from-source-code)**\n\u003cbr\u003eSuitable for developers or contributors who want full control over the codebase or plan to customize CRISPR-HAWK.\n\n\u003e 📝 **Note:** We recommend using the Mamba/Conda or Docker installation methods for most users, as they ensure the highest compatibility and stability across systems.\n\n### 1.1 Install CRISPR-HAWK from Mamba/Conda\n\n#### 1.1.1 Install Conda or Mamba\n\nBefore installing CRISPR-HAWK, ensure that either **Conda** or **Mamba** is installed on your system. Based on recommendations from the Bioconda community and performance testing during CRISPR-HAWK development, we **recommend using [Mamba](https://mamba.readthedocs.io/en/latest/index.html)** over Conda. Mamba is a fast, efficient drop-in replacement for Conda, built with a high-performance dependency solver in C++.\n\n**Installation Steps**\n\n**1. Install Conda or Mamba**\n\n* To install **Conda**, follow the official instructions:\n\u003cbr\u003e[Conda Installation Guide](https://docs.conda.io/projects/conda/en/latest/user-guide/install/index.html)\n\n* To install **Mamba**, follow the official instructions:\n\u003cbr\u003e[Mamba Installation Guide](https://mamba.readthedocs.io/en/latest/installation/mamba-installation.html)\n\n**2. Configure Bioconda Channels**\nOnce Mamba (or Conda) is installed, configure your environment with the appropriate channels used by CRISPR-HAWK:\n```bash\nmamba config --add channels bioconda\nmamba config --add channels defaults\nmamba config --add channels conda-forge\nmamba config --set channel_priority strict\n```\n\n\u003e 💡 **Tip**: If you are using Conda instead of Mamba, simply replace `mamba` with `conda` in the commands above.\n\nBy completing these steps your system will be correctly configured to install CRISPR-HAWK and all required dependencies via Bioconda.\n\n**Apple Silicon (M1/M2/M3) Support**\n\nIf you're using a Mac with Apple Silicon, follow these additional steps to ensure compatibility with Bioconda packages (which are primarily built for Intel)\n\n\u003e 💡 **Tip**: Not sure if your Mac uses Apple Silicon (M1, M2, or M3)? You can check by visiting Apple’s official support page: [Identify your Mac model and chip](https://support.apple.com/en-us/116943)\n\n**System-wide (Recommended)**\n\nMake sure [Rosetta](https://support.apple.com/en-us/102527) is installed:\n```zsh\nsoftwareupdate --install-rosetta\n```\n\nConfigure Mamba (or Conda) to prefer Intel (x86_64) builds:\n```zsh\nmamba config --add subdirs osx-64\n```\n\nThis will allow Bioconda to fetch compatible packages globally across all environments.\n\n**Environment-specific (Alternative)**\n\nYou can also enable Intel compatibility in a specific environment only:\n```zsh\nCONDA_SUBDIR=osx-64 mamba create -n crisprhawk-env -c bioconda crisprhawk\n```\n\n\u003e ⚠️ If you use this method, remember to prepend CONDA_SUBDIR=osx-64 to every future conda install command within this environment — or set the variable globally in your shell profile.\n\n\n#### 1.1.2 Install CRISPR-HAWK\n\nOnce Conda or Mamba is installed and the Bioconda channels are properly configured (see [Section 1.1.1](#111-install-conda-or-mamba)), you can install CRISPR-HAWK in a dedicated environment.\n\nWe strongly recommend creating a separate environment to ensure reproducibility and avoid dependency conflicts.\n\n**1. Create a new environment (recommended)**\n\n```bash\nmamba create -n crisprhawk-env python=3.8 -y\nmamba activate crisprhawk-env\n```\n\n\u003e 💡 **Tip**: CRISPR-HAWK currently requires **Python 3.8** for full compatibility with all dependencies.\n\n**2. Install CRISPR-HAWK**\n\n```bash\nmamba install crisprhawk\n```\n\nThis command will automatically resolve and install all required dependencies from Bioconda and conda-forge.\n\n\u003e 💡 **Tip**: If you are using Conda instead of Mamba, simply replace `mamba` with `conda`.\n\nAfter installation, confirm that CRISPR-HAWK is correctly installed:\n\n```bash\ncrisprhawk --help\n```\n\nIf the command runs successfully and displays the help message, the installation is complete.\n\nTo update to the latest version available on Bioconda:\n\n```bash\nmamba update crisprhawk\n```\n\n### 1.2 Install CRISPR-HAWK from Docker\n\nCRISPR-HAWK is also available through Docker, providing an isolated environment with all required dependencies pre-installed, including a dedicated environment for CRISPRitz.\n\nYou can build the Docker image directly from the provided `Dockerfile`:\n\n```bash\ndocker build -t crisprhawk .\n```\n\nOnce the built, the CRISPR-HAWK Docker image will be ready for use. To confirm the image is successfully installed, you can list all available Docker images by typing:\n```bash\ndocker images\n```\n\nLook for an entry similar to the following:\n```text\nREPOSITORY          TAG       IMAGE ID       CREATED        SIZE\npinellolab/crisprhawk latest    \u003cimage_id\u003e     \u003ctimestamp\u003e    \u003csize\u003e\n```\n\nYou are now ready to run CRISPR-HAWK using Docker.\n\n### 1.3 Install CRISPR-HAWK from PyPI\n\nYou can install the latest stable release of CRISPR-HAWK directly from [PyPI](https://pypi.org/project/crisprhawk/) using pip:\n```bash\npip install crisprhawk\n```\n\n\u003e 💡 **Tip**: To ensure a clean environment and avoid dependency conflicts, we recommend installing CRISPR-HAWK within a virtual environment (or using `mamba`/`conda`).\n\n**Upgrading to the Latest Version**\n\nTo upgrade CRISPR-HAWK to the newest release on PyPI, run:\n```bash\npip install -U crisprhawk\n```\n\n\u003e 💡 **Tip**: If you encounter permission issues, use the `--user` flag or ensure you are operating inside a virtual environment.\n\n### 1.4 Install CRISPR-HAWK from Source Code\n\nInstalling CRISPR-HAWK from source is ideal for developers, contributors, or users who wish to inspect or customize the codebase.\n\nThis method assumes you already have **Python 3.8** installed and accessible from your system’s environment.\n\n**Prerequisites**\n\n- Python **3.8** (strictly required)\n\n- `git`\n\n- A virtual environment (optional but recommended)\n\n**Installation Steps**\n\n**1. Clone the Repository**\n```bash\ngit clone https://github.com/pinellolab/CRISPR-HAWK.git\ncd CRISPR-HAWK\n```\n\n**2. (Optional) Create and Activate a Virtual Environment**\n```bash\nmamba create -n crisprhawk-env python=3.8 -y\nmamba activate crisprhawk-env\n```\n\n**3. Install CRISPR-HAWK and Its Dependencies**\n```bash\npip install .  # regular installation\npip install -e .  # development-mode installation\n```\n\nThe `.` tells `pip` to install the current directory as a package, including all dependencies specified in `setup.py` or `pyproject.toml`.\n\n**Quick Test**\n\nOnce installation is complete, verify that the command-line interface is working:\n```bash\ncrisprhawk -h\n```\n\nIf the help message is displayed correctly, CRISPR-HAWK is successfully installed and callable from any directory in your system.\n\n### 1.5 Install External Software Dependencies\n\nCRISPR-HAWK relies on a few external tools for certain optional features, such as additional **on-target scoring** and genome-wide **off-target nomination**. These dependencies are **not bundled** with the core CRISPR-HAWK installation and must be installed separately if you wish to enable the corresponding features.\n\n\u003e 📝 **Note**: External tools are used only for optional modules. If they are not installed, CRISPR-HAWK will still run the core pipeline (variant-aware gRNA search, scoring, annotation), but the corresponding scoring or off-target functionalities will be skipped.\n\n\u003e 💡 **Tip**: Installing these tools requires Mamba or Conda to be available on your system. If you haven't installed Conda or Mamba yet, refer to the instructions in [Section 1.1.1](#111-install-conda-or-mamba).\n\n#### 1.5.1 Install CRISPRitz (for Off-target Estimation)\n\n[CRISPRitz](https://github.com/pinellolab/CRISPRitz) is an efficient tool for nominating CRISPR-Cas off-target sites across large genomes accounting for mismatches and DNA/RNA bulges. CRISPR-HAWK uses it to enable **fast, high-throughput off-target estimation** in the reference genome for the identified candidate gRNAs.\n\n\u003e 📝 **Note**: CRISPRitz is required only if you plan to use the `--estimate-offtargets` feature in the `crisprhawk search` command.\n\n\u003e 🐧 **Linux note**: Off-target estimation through CRISPRitz is currently supported only on Linux-based operating systems.\n\n**Installation Steps**:\n\n1. Create a dedicated environment for CRISPRitz:\n```bash\nmamba create -n crispritz-crisprhawk python=3.8 -y\n```\n\n2. Install CRISPRitz (latest version):\n```bash\nmamba install -n crispritz-crisprhawk crispritz=2.6.6 -y\n```\n\n\u003e 💬 **Why use a separate environment?**\n\u003cbr\u003eThis prevents potential conflicts between CRISPR-HAWK's and CRISPRitz’s dependencies, and ensures better reproducibility.\n\n\n**Test your installation**\n\nYou can confirm that CRISPRitz is correctly installed and functional by running:\n```bash\nmamba run -n crispritz-crisprhawk crispritz.py\n```\n\n## 2 Usage\n\nCRISPR-HAWK provides multiple functionalities designed to support variant- and haplotype-aware CRISPR guide design, gRNA efficiency evaluation, and integration with downstream analysis pipelines. Each command serves a distinct role in the workflow.\n\n### 2.1 General Syntax\n```bash\ncrisprhawk \u003ccommand\u003e [options]\n```\n\nTo view available commands:\n```bash\ncrisprhawk --help\n```\n\nTo check version:\n```bash\ncrisprhawk --version\n```\n\n### 2.2 Search\n\nThe `crisprhawk search` command is the core functionality of CRISPR-HAWK, designed to identify and annotate candidate gRNAs in both reference and variant genomes.\nIt integrates variant-aware search, functional annotation, and predictive scoring to help you prioritize the most robust and context-aware guides for CRISPR editing.\n\nThe search includes:\n\n* Support for any Cas system (Cas9, Cpf1, SaCas9, etc.)\n* Compatibility with custom PAM sequences and guide lengths\n* Variant-aware design from individual or population-level VCF files (SNVs and indels)\n* Scoring using **Azimuth**, **RS3**, **CFDon**, **Elevation-on**, **DeepCpf1**, **PLM-CRISPR**, and optionally **CRISPRon** and **sgDesigner** when their dedicated environments are available\n* Functional and gene annotation using user-specified BED files \n* Optional estimation and reporting of **off-targets**\n* Output in detailed and structured reports (TSV, haplotype tables, off-target tables)\n\nUsage:\n```bash\ncrisprhawk search -f \u003cfasta-dir\u003e -r \u003cbedfile\u003e -v \u003cvcf-dir\u003e -p \u003cpam\u003e -g \u003cguide-length\u003e -o \u003coutput-dir\u003e\n```\n\n\u003e 📝 **Note**: All FASTA files in `\u003cfasta-dir\u003e` must be one per chromosome (e.g., chr1.fa, chr2.fa, etc.).\n\n\u003e 💡 **Tip**: Optional scoring modules such as CRISPRon and sgDesigner are automatically detected by CRISPR-HAWK through their dedicated environments. If these environments are not present, the corresponding scores are skipped without interrupting the analysis.\n\n---\n\n#### Required Arguments\n\n| Option                       | Description                                                                                                                                    |\n| ---------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------- |\n| `-f`, `--fasta \u003cFASTA-DIR\u003e` | Directory containing chromosome-separated FASTA files for the reference genome. All files will be used as reference input.                                                      |\n| `-r`, `--regions \u003cBED-FILE\u003e` | BED file defining the target regions where candidate gRNAs will be searched (e.g., promoters, exons, enhancers).                      |\n| `-v`, `--vcf \u003cVCF-DIR\u003e`      | *(Optional but recommended)* Directory containing per-chromosome VCF files for variant-aware guide design. If omitted, the tool performs reference-only analysis. |\n| `-p`, `--pam \u003cPAM\u003e`          | PAM sequence used to define valid gRNA targets (e.g., `NGG` for SpCas9, `TTTV` for Cpf1).                                                      |\n| `-g`, `--guide-len \u003cLENGTH\u003e` | Length of the guide RNA (excluding PAM), e.g., 20 for SpCas9.                                                                         |\n\n#### Optional Arguments\n\n| Option                                         | Description                                                                                                                                                                       |\n| ---------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |\n| `-i`, `--fasta-idx \u003cFAI\u003e`                      | Optional FASTA index file (`.fai`). If not provided, it will be computed automatically.                                                                                           |\n| `--right`                                      | By default, guides are extracted **upstream** of the PAM. Use this flag to extract them **downstream** (right side), useful for Cpf1 or other reverse-PAM systems.                |\n| `--no-filter`                                  | By default, only VCF variants with `FILTER == PASS` are considered. Use this flag to include **all variants**, regardless of filter status.                                       |\n| `--annotation \u003cBED1 BED2 ...\u003e`                 | Provide one or more BED files with **custom genomic** features (e.g., enhancers, DHS, regulatory elements). Must include a 4th column with annotation name. |\n| `--annotation-colnames \u003cname1 name2 ...\u003e`      | Custom column names for the annotations from the `--annotation` files. Must match the number and order of BED files.                                                      |\n| `--gene-annotation \u003cGENE-BED1 GENE-BED2 ...\u003e`  | One or more **gene annotation BED files** (9-column format). Must follow GENCODE-style structure, with gene name in the 9th column and gene feature (e.g., exon, UTR) in the 7th. |\n| `--gene-annotation-colnames \u003cname1 name2 ...\u003e` | Custom column names for gene annotations, matching the files in `--gene-annotation`.                                                                                              |\n| `--haplotype-table`                            | When enabled, a TSV file reporting haplotype-aware variants and their associated guide matches will be produced.                                                                  |\n| `--compute-elevation-score`                    | Compute Elevation and Elevation-On scores to evaluate guide efficiency. Requires that the combined length of the guide and PAM is exactly 23 bp, and that the guide is upstream of the PAM. *(Default: disabled)* |\n| `--candidate-guides \u003cSTR (format CHR:POS:STRAND)\u003e`                         | One or more genomic coordinates of candidate guides to analyze in detail with {TOOLNAME}. Each guide must be provided in the chromosome:position:strand format (e.g., chr1:123456:+). For each candidate guide, a dedicated subreport will be generated showing the guide and its alternative gRNAs side-by-side. If graphical reports are enabled, additional plots will visualize the impact of genetic variants on on-target efficiency using CFD and Elevation scores where applicable *(Default: no candidate guides)* |\n| `--graphical-reports`                          | Generate graphical reports to summarize findings. Includes a pie chart showing the distribution of guide types (reference, spacer+PAM alternative, spacer alternative, PAM alternative) and delta plots illustrating the impact of genetic diversity on guide efficiency and on-target activity. *(Default: disabled)* |\n| `-t`, `--threads \u003cINT\u003e`                        | Number of threads to use for parallel processing. Use `-t 0` to utilize **all available cores**. *(Default: 1)*                                                                   |\n| `-o`, `--outdir \u003cDIR\u003e`       | Output directory to store results, reports, and intermediate files. If not specified, defaults to the current working directory.               |\n| `--verbosity \u003cLEVEL\u003e`                          | Controls the verbosity of logs. Options: `0` = Silent, `1` = Normal, `2` = Verbose, `3` = Debug. *(Default: 1)*                                                                   |\n| `--debug`                                      | Enables **debug mode**, showing full stack traces and internal logs for troubleshooting.                                                                                          |\n\n---\n\nExample:\n\n```bash\ncrisprhawk search \\\n  -f hg38.fa \\\n  -r targets.bed \\\n  -v 1000G/ \\\n  -p NGG \\\n  -g 20 \\\n  -o results/ \\\n  --annotation regulatory_regions.bed \\\n  --gene-annotation gencode.bed \\\n  --haplotype-table \\\n  -t 8\n```\n\nThis will:\n\n* Search 20 bp gRNAs with an NGG PAM in `targets.bed`\n* Consider population variants from VCFs in `1000G/`\n* Annotate guides using custom regulatory and gene annotations \n* Run in parallel using 8 threads\n\n#### Off-targets Estimation (Optional)\n\n\u003e 🐧 **Linux-note**: Off-target estimation is **only available on Linux-based operating systems** and assumes you have successfully followed the installation instructions in the [Install CRISPRitz](#151-install-crispritz-for-off-target-estimation) section.\n\nCRISPR-HAWK supports genome-wide off-target nomination through integration with [CRISPRitz](https://github.com/pinellolab/CRISPRitz), enabling the identification of potential unintended gRNA binding sites across the reference genome.\n\n\u003e 📝 **Note**: Off-target nomination in CRISPR-HAWK is limited to the reference genome only, for performance and scalability reasons. If you need to estimate off-targets while accounting for genetic variants (e.g., SNVs, indels, population haplotypes), we recommend using [CRISPRme](https://github.com/pinellolab/CRISPRme) — a specialized, variant-aware off-target analysis tool.\n\nWhen enabled, the off-target module allows:\n\n* Comprehensive search of potential off-targets in the reference genome\n\n* Support for mismatches and DNA/RNA bulges\n\n* Output of structured off-target reports for guide prioritization\n\n*Enabling Off-Target Search*\n\nTo run off-target estimation, you must enable `--estimate-offtargets` and provide a pre-built genome index compatible with CRISPRitz.\n\n| Option                      | Description                                                                                          |\n| --------------------------- | ---------------------------------------------------------------------------------------------------- |\n| `--estimate-offtargets`     | Activates the off-target search pipeline for all candidate guides using CRISPRitz.                   |\n| `--crispritz-index \u003cDIR\u003e`   | Directory containing the CRISPRitz genome index. Must match the FASTA files provided with `--fasta`. |\n| `--mm \u003cINT\u003e`                | Max number of mismatches allowed in off-target search. *(Default: 4)*                                  |\n| `--bdna \u003cINT\u003e`              | Max number of DNA bulges allowed. *(Default: 0)*                                                       |\n| `--brna \u003cINT\u003e`              | Max number of RNA bulges allowed. *(Default: 0)*                                                       |\n| `--offtargets-annotation \u003cBED1 BED2 ...\u003e`               | Provide one or more BED files with **custom genomic** features (e.g., enhancers, DHS, regulatory elements) to annotate offtargets. Must include a 4th column with annotation name.                                                     |\n| `--offtargets-annotation-colnames \u003cname1 name2 ...\u003e`    | Custom column names for the offtargets annotations from the `--offtargets-annotation` files. Must match the number and order of BED files.\n\n\nExample:\n```bash\ncrisprhawk search \\\n  -f genome_fasta/ \\\n  -r targets.bed \\\n  -v 1000G/ \\\n  -p NGG \\\n  -g 20 \\\n  -o results/ \\\n  --estimate-offtargets \\\n  --crispritz-index genome_library/NGG_2_hg38/ \\\n  --mm 6 \\\n  --bdna 1 \\\n  --brna 1 \\\n  -t 8\n```\n\nThis will:\n\n* Design guides from `targets.bed`, using 1000G variants\n\n* Estimate genome-wide off-target sites (up to 6 mismatches, 1 bulge each)\n\n* Save off-targets in a detailed TSV file per guide\n\n* Use 8 threads for faster processing\n\n\n**Build a CRISPRitz Genome Index**\n\nTo use CRISPRitz for off-target estimation, you must first generate an indexed version of your reference genome. The `index-genome` command in CRISPRitz precomputes a genome-wide searchable index for a specific PAM and guide configuration — similar to building a BWA index. This index is essential for fast and efficient off-target search, especially when allowing bulges (RNA or DNA) and scanning across large genomes or thousands of guide RNAs. \n\n\u003e 💡 **Tip**: CRISPRitz indexes can be reused multiple times and do not need to be generated before each run\n\n*Required Inputs*\n\n1. **Output name of the genome index**\n  \u003cbr\u003eThis is the name of the directory that will store the generated index (e.g., `hg38_index/`).\n\n2. **Directory containing genome FASTA files**\n  \u003cbr\u003eThe genome must be split by chromosome — i.e., one file per chromosome (e.g., `chr1.fa`, `chr2.fa`, ...).\n\n3. **PAM configuration file**\n  \u003cbr\u003eA text file containing:\n\n    * The full gRNA+PAM string with `N`s representing the guide length\n    \u003cbr\u003e(e.g., `NNNNNNNNNNNNNNNNNNNNGG`)\n\n    * A space-separated number indicating the PAM length\n    \u003cbr\u003e(e.g., `3` for NGG where the PAM is 3 bp)\n\nExample file content (`pamNGG.txt`):\n```\nNNNNNNNNNNNNNNNNNNNNGG 3\n```\n\n4. Maximum number of bulges to index (`-bMax`)\n  \u003cbr\u003eThis sets the maximum number of **DNA and RNA bulges** that can be used in future off-target searches using this index.\n\n5. Number of threads (`-th`, optional)\nNumber of threads for parallelization during indexing.\n\n*Output*\n\nA new directory (named after the first argument, e.g., `hg38_index/`) containing:\n\n* Indexed `.bin` files (one per chromosome)\n\n* These include all candidate target sites for the selected PAM, including padding required to enable bulge-aware search.\n\n*Example Usage*\n```bash\ncrispritz.py index-genome hg19_ref hg19_ref/ pam/pamNGG.txt -bMax 2 -th 4\n```\n\n| Parameter        | Description                                                                   |\n| ---------------- | ----------------------------------------------------------------------------- |\n| `hg19_ref`       | Name of the output directory that will store the indexed genome.              |\n| `hg19_ref/`      | Folder containing FASTA files (one per chromosome).                           |\n| `pam/pamNGG.txt` | Path to the PAM specification file (see format above).                        |\n| `-bMax 2`        | Index will support searches with up to **2 RNA bulges** and **2 DNA bulges**. |\n| `-th 4`          | Use 4 threads during index creation (optional, increases performance).        |\n\n\n\u003e💡 **Tip**: Need more help generating a CRISPRitz index?\n\u003cbr\u003eRefer to the [CRISPRitz documentation on indexing](https://github.com/pinellolab/CRISPRitz#crispritz-installation-and-usage) for complete details, tips, and supported genome formats.\n\n### 2.3 Convert gnomAD VCF\n\nThe `crisprhawk convert-gnomad-vcf` command is a utility designed to convert gnomAD VCF files (version ≥ 3.1) into a format compatible with CRISPR-HAWK. This step is essential when working with large-scale population datasets (e.g., gnomAD v3.1, v4.1), ensuring proper variant normalization, filtering, and indexing.\n\nThis module:\n\n* Supports both `.vcf.bgz` and `.vcf.gz` formats\n\n* Extracts and preserves **sample-level** genotypes for population-aware haplotype modeling\n\n* Can optionally handle **joint allele frequency files** from gnomAD v4.1 (joint genome/exome releases)\n\n* Supports parallel processing for large-scale datasets\n\nUsage:\n```bash\ncrisprhawk convert-gnomad-vcf -d \u003cvcf-dir\u003e -o \u003coutput-dir\u003e\n```\n\n\u003e 📝 **Note**: All `.vcf.bgz` (or `.vcf.gz`) files in `\u003cvcf-dir\u003e` will be processed. Ensure tabix indices (`.tbi`) are present in the same folder.\n\n#### Required Arguments\n| Option                  | Description                                                                                               |\n| ----------------------- | --------------------------------------------------------------------------------------------------------- |\n| `-d`, `--vcf-dir \u003cDIR\u003e` | Directory containing per-chromosome **gnomAD VCF files** (compressed as `.vcf.bgz` or `.vcf.gz`).         |\n| `-o`, `--outdir \u003cDIR\u003e`  | Output directory for the converted VCF files. If not provided, defaults to the current working directory. |\n\n\n#### Optional Arguments\n| Option                  | Description                                                                                                           |\n| ----------------------- | --------------------------------------------------------------------------------------------------------------------- |\n| `--joint`               | Enable this flag when using **gnomAD v4.1 joint** genome/exome files. Ensures correct parsing of allele frequencies.  |\n| `--keep`                | Include **all variants**, regardless of the VCF `FILTER` field. By default, only `FILTER=PASS` variants are retained. |\n| `--suffix \u003cSUFFIX\u003e`     | Optional suffix to append to the output VCF filenames (e.g., `_converted`). Useful for traceability.                  |\n| `-t`, `--threads \u003cINT\u003e` | Number of threads to use. Set `-t 0` to use all available CPU cores. *(Default: 1)*                                   |\n| `--verbosity \u003cLEVEL\u003e`   | Logging verbosity. Options: `0` = Silent, `1` = Normal, `2` = Verbose, `3` = Debug *(Default: 1)*                     |\n| `--debug`               | Enable debug mode and print full error tracebacks.                                                                    |\n\n\n**Example**\n```bash\ncrisprhawk convert-gnomad-vcf \\\n  -d gnomad_v4.1/ \\\n  -o converted_vcfs/ \\\n  --suffix _crisprhawk \\\n  -t 4\n```\n\nThis command will:\n\n* Convert all `.vcf.bgz` files in `gnomad_v4.1/`\n\n* Retain only variants with `FILTER=PASS`\n\n* Append `_crisprhawk` to each output filename\n\n* Run using 4 threads\n\n\u003e 📝 **Note**: Make sure your input files are correctly indexed (`.tbi`) and are chromosome-specific. CRISPR-HAWK expects one VCF per chromosome.\n\n### 2.4 Prepare Data for CRISPRme\n\nThe `crisprhawk prepare-data-crisprme` command transforms a CRISPR-HAWK report into guide files compatible with **[CRISPRme](https://github.com/pinellolab/CRISPRme)**, enabling downstream variant- and haplotype-aware off-target prediction using CRISPRme's framework.\n\nThis module:\n\n* Extracts gRNA sequences from a CRISPR-HAWK report\n\n* Creates per-guide FASTA files required by CRISPRme\n\n* Optionally generates a **PAM specification file** compatible with the selected CRISPR system\n\nUsage:\n```bash\ncrisprhawk prepare-data-crisprme --report \u003ccrisprhawk-report\u003e -o \u003coutput-dir\u003e\n```\n\n#### Required Arguments\n\n| Option                         | Description                                                                                      |\n| ------------------------------ | ------------------------------------------------------------------------------------------------ |\n| `--report \u003cCRISPRHAWK-REPORT\u003e` | Path to a CRISPR-HAWK report file containing guide sequences.        |\n| `-o`, `--outdir \u003cDIR\u003e`         | Output directory for the CRISPRme-compatible guide files. Defaults to current working directory. |\n\n#### Optional Arguments\n\n| Option              | Description                                                                            |\n| ------------------- | -------------------------------------------------------------------------------------- |\n| `--create-pam-file` | If set, also generates a **PAM file** in CRISPRme format in the same output directory. |\n| `--debug`           | Enable debug mode and print full error tracebacks.                                     |\n\n**Output Structure**\n\nThe command will generate:\n\n* One FASTA file per gRNA (named using the guide label or sequence)\n\n* Optionally, a `pam.txt` file for CRISPRme if `--create-pam-file` is specified\n\nThese files can be directly used as input to CRISPRme's `--guide` and `--pam` options.\n\n**Example**\n\n```bash\ncrisprhawk prepare-data-crisprme \\\n  --report results/crisprhawk_report.tsv \\\n  -o crisprme_inputs/ \\\n  --create-pam-file\n```\n\nThis command will:\n\n* Parse all guides listed in the `crisprhawk_report.tsv` report\n\n* Write per-guide FASTA files to the `crisprme_inputs/` folder\n\n* Generate a PAM file (`pam.txt`) compatible with CRISPRme\n\n\u003e 💡 **Tip**: This is especially useful when transitioning from **on-target selection** in CRISPR-HAWK to **off-target analysis** in CRISPRme.\n\n## 3 Test\n\nOnce installed CRISPR‑HAWK, you can verify that everything is working as expected. Below are instructions for running the full test suite with pytest, as well as a quick smoke test to ensure basic functionality.\n\n### 3.1 Quick Test After Installation\n\nIf you just want a quick check that CRISPR‑HAWK is installed correctly and its CLI is working, try the following:\n\n```bash\ncrisprhawk --help\n```\n\nYou should see the usage message with available commands. If this works, then:\n\n```bash\ncrisprhawk search -h\n```\n\nThis should display help text for the `search` command and show all its options.\n\nIf both commands complete without error, your installation is likely successful.\n\n### 3.2 Run Full Test Suite with PyTest\n\nMake sure you’ve installed the development dependencies first (pytest, etc.). Then from the root of the repository, run:\n\n```bash\npytest\n```\n\nThis will run all tests in the `tests/` directory. If everything passes, CRISPR‑HAWK is behaving correctly end‑to‑end.\n\nYou can also run more specific tests:\n\n```bash\npytest tests/test_utils.py\npytest tests/test_utils.py::test_some_specific_function\n```\n\n### 3.3 Troubleshooting\n\n* **Import errors or missing modules when running `pytest`**\n  \u003cbr\u003eEnsure the package is correctly installed. We recommend installing in editable mode during development:\n\n  ```bash\n  pip install -e .[dev]\n  ```\n\n  Also check that the `src/` directory is properly referenced in your project structure.\n\n* **Command not found (`crisprhawk`)**\n  \u003cbr\u003eMake sure your virtual environment is activated (if used), or that the `bin/` path where `crisprhawk` is installed is in your system's `$PATH`.\n\n* **Unexpected behavior or crashes**\n  \u003cbr\u003eRun the command with increased verbosity or debug mode:\n\n  ```bash\n  crisprhawk \u003ccommand\u003e \u003cargs\u003e --verbosity 3 --debug\n  ```\n\n  This will print internal logs and full stack traces to help identify the issue.\n\n\n### 3.4 Reporting Issues\n\nIf you encounter a problem that isn't resolved by the above suggestions:\n\n1. **Check existing issues**\n   \u003cbr\u003eVisit the [GitHub Issues page](https://github.com/pinellolab/CRISPR-HAWK/issues) to see if the problem has already been reported or discussed.\n\n2. **Open a new issue**\n   \u003cbr\u003eIf your problem is not listed, please open a new issue and include the following:\n\n   * **Description** of the problem or unexpected behavior\n   * **Steps to reproduce** the error (including input data or command-line arguments, if applicable)\n   * **Platform details**:\n\n     * OS (e.g. Ubuntu 22.04, macOS 13)\n     * Python version\n     * CRISPR-HAWK version (or commit hash if using a development version)\n   * **Full error message and traceback**, if available (please use a code block for readability)\n   * Optionally, include the output of:\n\n     ```bash\n     crisprhawk --version\n     pip list\n     ```\n\n3. **Contact**\n   \u003cbr\u003eIf the issue involves sensitive data or requires direct contact, please email the maintainer(s) listed in the repository.\n\n\u003e 💬 Got suggestions or issues? We’d love to hear from you — your input helps us build a better CRISPR‑HAWK. Thank you!\n\n## 4 Citation\n\nIf you use CRISPR-HAWK in your research, please cite:\n\n\u003e Kumbara A, Tognon M, Carone G, Fontanesi A, Bombieri N, Giugno R, Pinello L. CRISPR-HAWK: Haplotype- and Variant-aware guide design toolkit for CRISPR-Cas. bioRxiv [Preprint]. 2025 Dec 27:2025.12.27.696698. doi: 10.64898/2025.12.27.696698. PMID: 41497669; PMCID: PMC12767513.\n\n## 5 Contacts\n\n* Manuel Tognon\n  \u003cbr\u003emanuel.tognon@univr.it\n\n* Rosalba Giugno\n  \u003cbr\u003erosalba.giugno@univr.it\n\n* Luca Pinello\n  \u003cbr\u003elpinello@mgh.harvard.edu\n\n## 6 License\n\nCRISPR-HAWK is licensed under the AGPL-3.0 license, which permits its use for academic research purposes only.\n\nFor any commercial or for-profit use, please contact the authors.\n\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fpinellolab%2Fcrispr-hawk","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fpinellolab%2Fcrispr-hawk","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fpinellolab%2Fcrispr-hawk/lists"}