{"id":22582474,"url":"https://github.com/paccmann/paccmann_kinase_binding_residues","last_synced_at":"2025-04-10T19:12:04.229Z","repository":{"id":41048887,"uuid":"387501914","full_name":"PaccMann/paccmann_kinase_binding_residues","owner":"PaccMann","description":"Comparison of active site and full kinase sequences for drug-target affinity prediction and molecular generation. Full paper: https://pubs.acs.org/doi/10.1021/acs.jcim.1c00889","archived":false,"fork":false,"pushed_at":"2022-10-03T10:55:10.000Z","size":7868,"stargazers_count":36,"open_issues_count":0,"forks_count":6,"subscribers_count":4,"default_branch":"master","last_synced_at":"2025-03-24T16:53:24.756Z","etag":null,"topics":["active-sites","kinases","knn","machine-learning","protein-ligand-interactions","proteochemometrics"],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/PaccMann.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null}},"created_at":"2021-07-19T14:57:14.000Z","updated_at":"2024-12-17T10:30:53.000Z","dependencies_parsed_at":"2023-01-19T02:45:50.624Z","dependency_job_id":null,"html_url":"https://github.com/PaccMann/paccmann_kinase_binding_residues","commit_stats":null,"previous_names":[],"tags_count":1,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/PaccMann%2Fpaccmann_kinase_binding_residues","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/PaccMann%2Fpaccmann_kinase_binding_residues/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/PaccMann%2Fpaccmann_kinase_binding_residues/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/PaccMann%2Fpaccmann_kinase_binding_residues/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/PaccMann","download_url":"https://codeload.github.com/PaccMann/paccmann_kinase_binding_residues/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":248279803,"owners_count":21077408,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["active-sites","kinases","knn","machine-learning","protein-ligand-interactions","proteochemometrics"],"created_at":"2024-12-08T06:10:20.011Z","updated_at":"2025-04-10T19:12:04.209Z","avatar_url":"https://github.com/PaccMann.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# Active sites outperform full proteins for modeling kinases\n[![Python package](https://github.com/PaccMann/paccmann_kinase_binding_residues/actions/workflows/python-package.yml/badge.svg)](https://github.com/PaccMann/paccmann_kinase_binding_residues/actions/workflows/python-package.yml)\n[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)\n[![Code style: black](https://img.shields.io/badge/code%20style-black-000000.svg)](https://github.com/psf/black)\n[![DOI:10.1021/acs.jcim.1c00889](https://zenodo.org/badge/DOI/10.1021/acs.jcim.1c00889.svg)](https://doi.org/10.1021/acs.jcim.1c00889)\n[![DOI:10.1021/acs.jcim.2c00840](https://zenodo.org/badge/DOI/10.1021/acs.jcim.2c00840.svg)](https://doi.org/10.1021/acs.jcim.2c00840)\n\n\n## Summary\nThis repository contains data \u0026 code for the JCIM paper: [Active Site Sequence Representations of Human Kinases Outperform Full Sequence Representations for Affinity Prediction and Inhibitor Generation: 3D Effects in a 1D Model](https://pubs.acs.org/doi/10.1021/acs.jcim.1c00889). We study the impact of different protein sequence representations for modeling human kinases. We find that using **active site residues yields superior performance to using full protein sequences for predicting binding affinity**. We also study the difference of active site vs. full sequence on de-novo design tasks. We generate kinase inhibitors directly from protein sequences with our previously developed hybrid-VAE (PaccMann\u003csup\u003eRL\u003c/sup\u003e) but find no major differences between both kinase representations.\n\n\n \u003cp align=\"center\"\u003e\n\u003cimg width=\"75%\" height=\"75%\" src=\"https://github.com/PaccMann/paccmann_kinase_binding_residues/blob/master/assets/full_vs_active_site.png\"\u003e\n \u003c/p\u003e\n\n## News\n- October 2022: Our report on comparing different definitions of kinase active sites has been published as a *letter* in the ACS [**Journal of Chemical Information \u0026 Modeling**](https://pubs.acs.org/doi/full/10.1021/acs.jcim.2c00840). \nTherein, we also propose several **novel protein sequence augmentation strategies**.\n- March 2022: We added a report with new experiments to compare our original active site sequence representations ([29 residues, see *Sheridan et al. (2012)*](https://pubs.acs.org/doi/abs/10.1021/ci900176y) with the 16 residues identified by [*Martin et al. (2010)](https://pubs.acs.org/doi/10.1021/ci200314j). (The report is now longer available since it has been superseded by the [JCIM letter](https://pubs.acs.org/doi/full/10.1021/acs.jcim.2c00840).)\n- January 2022: We are proud to be featured on [**the JCIM cover**](https://pubs.acs.org/toc/jcisd8/62/2) with _an AI artwork_! 👉  \u003cimg align=\"right\" width=\"25%\" height=\"25%\" src=\"https://github.com/PaccMann/paccmann_kinase_binding_residues/blob/master/assets/cover.jpg\"\u003e\n- December 2021: Our work has been published in the ACS [**Journal of Chemical Information \u0026 Modeling**](https://pubs.acs.org/doi/10.1021/acs.jcim.1c00889). \n- December 2021: The part about binding affinity prediction was presented at the [NeurIPS 2021 workshop on *Machine Learning for Structural Biology*](https://www.mlsb.io) and the [ELLIS Machine Learning for Molecule Discovery workshop](https://moleculediscovery.github.io/workshop2021/) alongside NeurIPS 2021.\n- November 2021: Our work has **won** the [🥇 #IOPP best poster award🥇](https://ioppublishing.org/twitter-conference/) in the category *Biomedical engineering* \n- July 2021: A preliminary version of our work was presented at the [Twitter #IOPPposter conference](https://ioppublishing.org/twitter-conference/) (see GIF below ⬇️))\n\n![Summary GIF](https://github.com/PaccMann/paccmann_kinase_binding_residues/blob/master/assets/summary.gif \"Summary GIF\")\n\n\n## Description\nThis repository facilitates the reproduction of the experiments conducted in the JCIM paper [`Active Site Sequence Representations of Human Kinases Outperform Full Sequence Representations for Affinity Prediction and Inhibitor Generation: 3D Effects in a 1D Model`](https://pubs.acs.org/doi/10.1021/acs.jcim.1c00889). We provide scripts to:\n1. Train and evaluate the BimodalMCA for drug-protein affinity prediction\n2. Evaluate the bimodal KNN affinity predictor either in a CV setting or on a plain train/test script\n3. Optimize a SMILES- or SELFIES-based molecular generative model to produce molecules with high binding affinities for a protein of interest (affinity is predicted with the KNN model).\n\n### Data\nThe preprocessed BindingDB data (CV and test data for ligand split and kinase split, data used for pretraining and affinity optimization) can be accessed on this [Box link](https://ibm.biz/active_site_data). We also release the aligned active site sequences (29 residues) for all kinases. If you use the data, please [cite](#citation) our work.\n\n\n## Installation\n\nThe core functionality of this repo is provided in the `pkbr` package which can be installed (in editable mode) by `pip install -e .`.\nIn order to execute the example scripts, we recommend setting up a conda environment:\n\n\n```console\nconda env create -f conda.yml\nconda activate pkbr\n```\nAfterwards you can execute the scripts to run the KNN, the BiMCA or the affinity optimization.\n\n## Code examples\nAll examples assume that you downloaded the data from [this Box link](https://ibm.biz/active_site_data) and stored it in a folder `data` in the root of this repo.\n\n### Running the KNN affinity predictor in a cross validation\n```sh\npython3 scripts/knn_cv.py -d data/ligand_split -tr train.csv -te validation.csv \\\n-f 10 -lp data/ligands.smi -kp data/human_kinases_active_site.smi -r my_cv_results\n```\n\n### Running KNN in single train/test split\n```sh\npython3 scripts/knn_test.py -t data/ligand_split/fold_0/validation.csv -test data/ligand_split/fold_0/test.csv \\\n-tlp data/ligands.smi -telp data/ligands.smi -kp data/human_kinases_active_site.smi -r my_results.csv\n```\n\n### Training the BiMCA model\n```sh\npython3 scripts/bimca_train.py \\\n\tdata/ligand_split/fold_0/train.csv data/ligand_split/fold_0/validation.csv data/ligand_split/fold_0/test.csv \\\n\thuman-kinase-alignment data/human_kinases_active_site.smi data/ligands.smi data/smiles_vocab.json \\\n\tmodels config/active_site.json -n my_as_model\n```\n\n### Evaluating the BiMCA model\n```sh\npython3 scripts/bimca_test.py \\\ndata/ligand_split/fold_0/validation.csv human-kinase-alignment data/human_kinases_sequence.smi data/ligands.smi \\\npath_to_your_trained_model\n\n```\n\n### Affinity optimization with SMILES generator\nTo execute this part you need to utilize the pretrained SMILES/SELFIES VAE stored under `data/models`\n```sh\npython scripts/gp_generation_smiles_knn.py \\\n    data/models/smiles_vae data/affinity_optimization/bindingdb_all_kinase_active_site.csv \\\n    data/affinity_optimization/example_active_site.smi \\\n    smiles_active_site_generator/ -s 42 -t 0.85 -r 10 -n 50 -s 42 -i 40 -c 80\n```\n\n### Affinity optimization with SELFIES generator\n```sh\npython scripts/gp_generation_selfies_knn.py \\\n    data/models/selfies_vae data/affinity_optimization/bindingdb_all_kinase_sequence.csv \\\n    data/affinity_optimization/example_sequence.smi \\\n    selfies_sequence_generator/ -s 42 -t 0.85 -r 10 -n 50 -s 42 -i 40 -c 80\n```\n\n## Choosing active-site sequences\nSee our new letter in [JCIM](https://pubs.acs.org/doi/full/10.1021/acs.jcim.2c00840). \n\n#### Definitions\n\u003cimg align=\"right\" width=\"60%\" height=\"60%\" src=\"https://github.com/PaccMann/paccmann_kinase_binding_residues/blob/master/assets/definitions.jpg\"\u003e\n\nHow to exactly define an \"active site\" is a critical choice. While we originally relied on\nthe definition by [Sheridan et al. (2009)](https://pubs.acs.org/doi/10.1021/ci900176y) we have compared now to the active site\ndefinition by [Martin et al. (2012)](https://pubs.acs.org/doi/full/10.1021/ci200314j) and a *Combined* definition that uses a total of\n35 residues from either definitions. This improves performance significantly, especially\nfor allosteric binders.\n\n\u003cimg align=\"right\" width=\"60%\" height=\"60%\" src=\"https://github.com/PaccMann/paccmann_kinase_binding_residues/blob/master/assets/augmentations.png\"\u003e\n\n#### Augmentations\n\nWe also devised novel protein sequence augmentation schemes by flipping/swapping\ncontiguous active-site subsequences that lie discontiguously in the full protein sequence.\n\nTo train a model with the new active site definition and the new augmentation, follow the [installation](#installation)\nsetup and download the data from [Box](https://ibm.biz/active_site_data).\n\n\n\nAfterwards run:\n```sh\npython3 scripts/bimca_train.py \\\n\tdata/ligand_split/fold_0/train.csv data/ligand_split/fold_0/validation.csv data/ligand_split/fold_0/test.csv \\\n\thuman-kinase-alignment data/human_kinases_active_site_combined.smi data/ligands.smi data/smiles_vocab.json \\\n\tmodels config/active_site_augment.json -n combined_augment\n```\nYou can modify `config/active_site_augment.json` to play with the hyperparameters and\nthe probability of the augmentations.\n\n\n\n## Citation\nIf you use this repo, our data, the proposed active site definitions or the sequence augmentation strategies in your projects, please cite the following:\n\n```bib\n@article{born2022active,\n\tauthor = {Born, Jannis and Huynh, Tien and Stroobants, Astrid and Cornell, Wendy D. and Manica, Matteo},\n\ttitle = {Active Site Sequence Representations of Human Kinases Outperform Full Sequence Representations for Affinity Prediction and Inhibitor Generation: 3D Effects in a 1D Model},\n\tjournal = {Journal of Chemical Information and Modeling},\n\tvolume = {62},\n\tnumber = {2},\n\tpages = {240-257},\n\tyear = {2022},\n\tdoi = {10.1021/acs.jcim.1c00889},\n\tnote ={PMID: 34905358},\n\tURL = {https://doi.org/10.1021/acs.jcim.1c00889}\n}\n\n@article{born2022on,\n        author = {Born, Jannis and Shoshan, Yoel and Huynh, Tien and Cornell, Wendy D. and Martin, Eric J. and Manica, Matteo},\n        title = {On the Choice of Active Site Sequences for Kinase-Ligand Affinity Prediction},\n        journal = {Journal of Chemical Information and Modeling},\n        volume = {62},\n        number = {18},\n        pages = {4295-4299},\n        year = {2022},\n        doi = {10.1021/acs.jcim.2c00840},\n        URL = {https://doi.org/10.1021/acs.jcim.2c00840}\n}\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fpaccmann%2Fpaccmann_kinase_binding_residues","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fpaccmann%2Fpaccmann_kinase_binding_residues","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fpaccmann%2Fpaccmann_kinase_binding_residues/lists"}