{"id":15140782,"url":"https://github.com/lsj2408/transformer-m","last_synced_at":"2026-03-13T19:03:22.065Z","repository":{"id":60810365,"uuid":"545406093","full_name":"lsj2408/Transformer-M","owner":"lsj2408","description":"[ICLR 2023] One Transformer Can Understand Both 2D \u0026 3D Molecular Data (official implementation)","archived":false,"fork":false,"pushed_at":"2023-03-31T12:36:34.000Z","size":5710,"stargazers_count":209,"open_issues_count":9,"forks_count":26,"subscribers_count":5,"default_branch":"main","last_synced_at":"2025-05-12T00:12:04.938Z","etag":null,"topics":["general-purpose-molecular-model","graph-neural-network","graph-transformer","molecular-modeling","molecule","transformer"],"latest_commit_sha":null,"homepage":"https://arxiv.org/abs/2210.01765","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/lsj2408.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2022-10-04T10:05:33.000Z","updated_at":"2025-05-02T16:17:27.000Z","dependencies_parsed_at":"2024-12-26T19:06:28.213Z","dependency_job_id":"f029517e-b3da-4edb-a005-888a3a529d1b","html_url":"https://github.com/lsj2408/Transformer-M","commit_stats":{"total_commits":19,"total_committers":2,"mean_commits":9.5,"dds":"0.052631578947368474","last_synced_commit":"45647e143a5282e0e97117969396446084bcf1ab"},"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/lsj2408/Transformer-M","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/lsj2408%2FTransformer-M","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/lsj2408%2FTransformer-M/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/lsj2408%2FTransformer-M/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/lsj2408%2FTransformer-M/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/lsj2408","download_url":"https://codeload.github.com/lsj2408/Transformer-M/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/lsj2408%2FTransformer-M/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":279061413,"owners_count":26095389,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-10-15T02:00:07.814Z","response_time":56,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["general-purpose-molecular-model","graph-neural-network","graph-transformer","molecular-modeling","molecule","transformer"],"created_at":"2024-09-26T08:41:14.736Z","updated_at":"2025-10-15T08:29:27.662Z","avatar_url":"https://github.com/lsj2408.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# One Transformer Can Understand Both 2D \u0026 3D Molecular Data\n[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/one-transformer-can-understand-both-2d-3d/graph-regression-on-pcqm4mv2-lsc)](https://paperswithcode.com/sota/graph-regression-on-pcqm4mv2-lsc?p=one-transformer-can-understand-both-2d-3d)\n\nThis repository is the official implementation of “[One Transformer Can Understand Both 2D \u0026 3D Molecular Data](https://arxiv.org/abs/2210.01765)”, based on the official implementation of [Graphormer](https://github.com/microsoft/Graphormer) and [Fairseq](https://github.com/facebookresearch/fairseq) in [PyTorch](https://github.com/pytorch/pytorch).\n\n\u003e One Transformer Can Understand Both 2D \u0026 3D Molecular Data\n\u003e\n\u003e *Shengjie Luo, Tianlang Chen\\*, Yixian Xu\\*, Shuxin Zheng, Tie-Yan Liu, Liwei Wang, Di He*\n\n## 🔥 News\n- **2023.03.31**: The fine-tuning code of QM9 has been released.\n- **2022.11.22**: Congratulations! Transformer-M has been used by **all Top-3 winners** in [**PCQM4Mv2 Track, 2nd OGB Large-Scale Challenge, NeurIPS 2022**](https://ogb.stanford.edu/neurips2022/results/#winners_pcqm4mv2)!\n  - 1st Place winner,    Team WeLoveGraphs from GraphCore,     [code](https://github.com/graphcore/ogb-lsc-pcqm4mv2) \u0026 [report](https://ogb.stanford.edu/paper/neurips2022/pcqm4mv2_WeLoveGraphs.pdf).\n  - co-2nd Place winner, Team VisNet from Microsoft,           [code](https://github.com/microsoft/ViSNet/tree/OGB-LSC%40NIPS2022) \u0026 [report](https://github.com/microsoft/ViSNet/blob/OGB-LSC%40NIPS2022/ViSNet_Tech_Report_OGB_LSC_NIPS22.pdf).\n  - co-2nd Place winner, Team NVIDIA-PCQM4Mv2 from NVIDIA,  [code](https://github.com/jfpuget/NVIDIA-PCQM4Mv2) \u0026 [report](https://ogb.stanford.edu/paper/neurips2022/pcqm4mv2_NVIDIA-PCQM4Mv2.pdf).\n- **2022.10.05**: Codes and model checkpoints are released!\n\n## Overview\n\n![arch](docs/arch.jpg)\n\nTransformer-M is a versatile and effective molecular model that can take molecular data of 2D or 3D formats as input and generate meaningful semantic representations. Using the standard Transformer as the backbone architecture, Transformer-M develops two separated channels to encode 2D and 3D structural information and incorporate them with the atom features in the network modules. When the input data is in a particular format, the corresponding channel will be activated, and the other will be disabled. Empirical results show that our Transformer-M can achieve strong performance on 2D and 3D tasks simultaneously, which is the first step toward general-purpose molecular models in chemistry.\n\n## Results on PCQM4Mv2, OGB Large-Scale Challenge\n\n![](docs/Table1.png)\n🚀**Note:**  **PCQM4Mv2** is also the benchmark dataset of the graph-level track in the **2nd OGB-LSC** at [**NeurIPS 2022 competition track**](https://ogb.stanford.edu/neurips2022/). As non-participants, we open source all the codes and model weights, and sincerely welcome participants to use our model. Looking forward to your feedback!\n\n## Installation\n\n- Clone this repository\n\n```shell\ngit clone https://github.com/lsj2408/Transformer-M.git\n```\n\n- Install the dependencies (Using [Anaconda](https://www.anaconda.com/), tested with CUDA version 11.0)\n\n```shell\ncd ./Transformer-M\nconda env create -f requirement.yaml\nconda activate Transformer-M\npip install torch==1.7.1+cu110 torchvision==0.8.2+cu110 torchaudio==0.7.2 -f https://download.pytorch.org/whl/torch_stable.html\npip install torch_geometric==1.6.3\npip install torch_scatter==2.0.7\npip install torch_sparse==0.6.9\npip install azureml-defaults\npip install rdkit-pypi cython\npython setup.py build_ext --inplace\npython setup_cython.py build_ext --inplace\npip install -e .\npip install --upgrade protobuf==3.20.1\npip install --upgrade tensorboard==2.9.1\npip install --upgrade tensorboardX==2.5.1\n```\n\n## Checkpoints\n\n| Model | File Size | Update Date  | Valid MAE on PCQM4Mv2 | Download Link                                            |\n| ----- | --------- | ------------ | --------------------- | -------------------------------------------------------- |\n| L12   | 189MB     | Oct 04, 2022 | 0.0785                | https://1drv.ms/u/s!AgZyC7AzHtDBdWUZttg6N2TsOxw?e=sUOhox |\n| L18   | 270MB     | Oct 04, 2022 | 0.0772                | https://1drv.ms/u/s!AgZyC7AzHtDBdrY59-_mP38jsCg?e=URoyUK |\n| L12_old | 189MB   | Mar 31, 2023 | 0.0787                | https://1drv.ms/u/s!AgZyC7AzHtDBesDk9tZK1yvbtzE?e=5H91Zq |\n\n```shell\n# create paths to checkpoints for evaluation\n\n# download the above model weights (L12.pt, L18.pt) to ./\nmkdir -p logs/L12\nmkdir -p logs/L18\nmv L12.pt logs/L12/\nmv L18.pt logs/L18/\n```\n\n## Datasets\n\n- Preprocessed data: [download link](https://1drv.ms/u/s!AgZyC7AzHtDBeIDqE61u1ZEMv_8?e=3g428e)\n\n  ```shell\n  # create paths to datasets for evaluation/training\n  \n  # download the above compressed datasets (pcqm4mv2-pos.zip) to ./\n  unzip pcqm4mv2-pos.zip -d ./datasets\n  ```\n\n- You can also directly execute the evaluation/training code to process data from scratch.\n\n## Evaluation\n\n```shell\nexport data_path='./datasets/pcq-pos'                # path to data\nexport save_path='./logs/{folder_to_checkpoints}'    # path to checkpoints, e.g., ./logs/L12\n\nexport layers=12                                     # set layers=18 for 18-layer model\nexport hidden_size=768                               # dimension of hidden layers\nexport ffn_size=768                                  # dimension of feed-forward layers\nexport num_head=32                                   # number of attention heads\nexport num_3d_bias_kernel=128                        # number of Gaussian Basis kernels\nexport batch_size=256                                # batch size for a single gpu\nexport dataset_name=\"PCQM4M-LSC-V2-3D\"\t\t\t\t   \nexport add_3d=\"true\"\nbash evaluate.sh\n```\n\n## Training\n\n```shell\n# L12. Valid MAE: 0.0785\nexport data_path='./datasets/pcq-pos'               # path to data\nexport save_path='./logs/'                          # path to logs\n\nexport lr=2e-4                                      # peak learning rate\nexport warmup_steps=150000                          # warmup steps\nexport total_steps=1500000                          # total steps\nexport layers=12                                    # set layers=18 for 18-layer model\nexport hidden_size=768                              # dimension of hidden layers\nexport ffn_size=768                                 # dimension of feed-forward layers\nexport num_head=32                                  # number of attention heads\nexport batch_size=256                               # batch size for a single gpu\nexport dropout=0.0\nexport act_dropout=0.1\nexport attn_dropout=0.1\nexport weight_decay=0.0\nexport droppath_prob=0.1                            # probability of stochastic depth\nexport noise_scale=0.2                              # noise scale\nexport mode_prob=\"0.2,0.2,0.6\"                      # mode distribution for {2D+3D, 2D, 3D}\nexport dataset_name=\"PCQM4M-LSC-V2-3D\"\nexport add_3d=\"true\"\nexport num_3d_bias_kernel=128                       # number of Gaussian Basis kernels\nbash train.sh\n```\n\nOur model is trained on 4 NVIDIA Tesla A100 GPUs (40GB). The time cost for an epoch is around 10 minutes.\n\n## Downstream Task -- (QM9)\nDownload the checkpoint: L12-old.pt\n```shell\nexport ckpt_path='./L12-old.pt'                # path to checkpoints\nbash finetune_qm9.sh\n```\n\n## Citation\n\nIf you find this work useful, please kindly cite following papers:\n\n```latex\n@article{luo2022one,\n  title={One Transformer Can Understand Both 2D \\\u0026 3D Molecular Data},\n  author={Luo, Shengjie and Chen, Tianlang and Xu, Yixian and Zheng, Shuxin and Liu, Tie-Yan and Wang, Liwei and He, Di},\n  journal={arXiv preprint arXiv:2210.01765},\n  year={2022}\n}\n\n@inproceedings{\n  ying2021do,\n  title={Do Transformers Really Perform Badly for Graph Representation?},\n  author={Chengxuan Ying and Tianle Cai and Shengjie Luo and Shuxin Zheng and Guolin Ke and Di He and Yanming Shen and Tie-Yan Liu},\n  booktitle={Thirty-Fifth Conference on Neural Information Processing Systems},\n  year={2021},\n  url={https://openreview.net/forum?id=OeWooOxFwDa}\n}\n\n@article{shi2022benchmarking,\n  title={Benchmarking Graphormer on Large-Scale Molecular Modeling Datasets},\n  author={Yu Shi and Shuxin Zheng and Guolin Ke and Yifei Shen and Jiacheng You and Jiyan He and Shengjie Luo and Chang Liu and Di He and Tie-Yan Liu},\n  journal={arXiv preprint arXiv:2203.04810},\n  year={2022},\n  url={https://arxiv.org/abs/2203.04810}\n}\n```\n\n## Contact\n\nShengjie Luo (luosj@stu.pku.edu.cn)\n\nSincerely appreciate your suggestions on our work!\n\n## License\n\nThis project is licensed under the terms of the MIT license. See [LICENSE](https://github.com/lsj2408/Transformer-M/blob/main/LICENSE) for additional details.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Flsj2408%2Ftransformer-m","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Flsj2408%2Ftransformer-m","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Flsj2408%2Ftransformer-m/lists"}