{"id":20194411,"url":"https://github.com/Tencent/Tencent-Hunyuan-Large","last_synced_at":"2025-05-07T04:31:16.819Z","repository":{"id":261224981,"uuid":"877138933","full_name":"Tencent/Tencent-Hunyuan-Large","owner":"Tencent","description":null,"archived":false,"fork":false,"pushed_at":"2024-12-06T08:15:56.000Z","size":1637,"stargazers_count":1493,"open_issues_count":16,"forks_count":100,"subscribers_count":26,"default_branch":"main","last_synced_at":"2025-04-10T04:53:33.415Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"other","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/Tencent.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE.txt","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2024-10-23T06:43:41.000Z","updated_at":"2025-04-09T10:38:02.000Z","dependencies_parsed_at":"2024-11-20T08:18:42.004Z","dependency_job_id":"01c3bbd8-4ed4-41b1-a615-fa2cfc7aa2f1","html_url":"https://github.com/Tencent/Tencent-Hunyuan-Large","commit_stats":null,"previous_names":["tencent/tencent-hunyuan-large"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Tencent%2FTencent-Hunyuan-Large","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Tencent%2FTencent-Hunyuan-Large/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Tencent%2FTencent-Hunyuan-Large/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/Tencent%2FTencent-Hunyuan-Large/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/Tencent","download_url":"https://codeload.github.com/Tencent/Tencent-Hunyuan-Large/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":252813842,"owners_count":21808388,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-11-14T04:02:28.487Z","updated_at":"2025-05-07T04:31:16.801Z","avatar_url":"https://github.com/Tencent.png","language":"Python","funding_links":[],"categories":["A01_文本生成_文本对话","Python"],"sub_categories":["大语言对话模型及数据"],"readme":"\u003cp align=\"left\"\u003e\n    \u003ca href=\"README_CN.md\"\u003e中文\u003c/a\u003e\u0026nbsp ｜ English\u003c/a\u003e\n\u003c/p\u003e\n\u003cbr\u003e\u003cbr\u003e\n\n\u003cp align=\"center\"\u003e\n \u003cimg src=\"https://dscache.tencent-cloud.cn/upload/uploader/hunyuan-64b418fd052c033b228e04bc77bbc4b54fd7f5bc.png\" width=\"400\"/\u003e \u003cbr\u003e\n\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\n\n\u003cp align=\"center\"\u003e\n    🫣\u0026nbsp\u003ca href=\"https://huggingface.co/tencent/Tencent-Hunyuan-Large\"\u003e\u003cb\u003eHugging Face\u003c/b\u003e\u003c/a\u003e\u0026nbsp\u0026nbsp |  \u0026nbsp\u0026nbsp🖥️\u0026nbsp\u0026nbsp\u003ca href=\"https://llm.hunyuan.tencent.com/\" style=\"color: red;\"\u003e\u003cb\u003eofficial website\u003c/b\u003e\u003c/a\u003e\u0026nbsp\u0026nbsp｜\u0026nbsp\u0026nbsp🕖\u0026nbsp\u0026nbsp \u003ca href=\"https://cloud.tencent.com/product/hunyuan\" \u003e\u003cb\u003eHunyuanAPI\u003c/b\u003e\u003c/a\u003e\u0026nbsp\u0026nbsp｜\u0026nbsp\u0026nbsp🐳\u0026nbsp\u0026nbsp \u003ca href=\"https://gitee.com/Tencent/Tencent-Hunyuan-Large\" \u003e\u003cb\u003eGitee\u003c/b\u003e\u003c/a\u003e\n\u003c/p\u003e\u003cp align=\"center\"\u003e\n    \u003ca href=\"https://arxiv.org/abs/2411.02265\" style=\"color: red;\"\u003e\u003cb\u003eTechnical Report\u003c/b\u003e\u003c/a\u003e\u0026nbsp\u0026nbsp｜\u0026nbsp\u0026nbsp \u003ca href=\"https://huggingface.co/spaces/tencent/Hunyuan-Large\"\u003e\u003cb\u003eDemo\u003c/b\u003e\u003c/a\u003e\u0026nbsp\u0026nbsp\u0026nbsp｜\u0026nbsp\u0026nbsp \u003ca href=\"https://cloud.tencent.com/document/product/851/112032\" style=\"color: red;\"\u003e\u003cb\u003eTencent Cloud TI\u003c/b\u003e\u003c/a\u003e\u0026nbsp\u0026nbsp\u0026nbsp\u003c/p\u003e\n\u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n\u003cp\u003e\n    \u003ctable align=\"center\"\u003e\n        \u003ctbody\u003e\n            \u003ctr\u003e\n                \u003ctd align=\"center\" colspan=\"3\"\u003e\u003cstrong\u003eDownload Models\u003c/strong\u003e\u003c/td\u003e\n            \u003c/tr\u003e\n            \u003ctr\u003e\n                \u003ctd align=\"center\" style=\"width: 100px;\" \u003eModels\u003c/td\u003e\n                \u003ctd align=\"center\" style=\"width: 500px;\"\u003eHuggingface Download URL\u003c/td\u003e\n                \u003ctd align=\"center\" style=\"width: 500px;\"\u003eTencent Cloud Download URL\u003c/td\u003e\n            \u003c/tr\u003e\n            \u003ctr\u003e\n                \u003ctd style=\"width: 100px;\"\u003eHunyuan-A52B-Instruct-FP8\u003c/td\u003e\n                \u003ctd style=\"width: 500px;\"\u003e\u003ca href=\"https://huggingface.co/tencent/Tencent-Hunyuan-Large/tree/main/Hunyuan-A52B-Instruct-FP8\" style=\"color: red;\"\u003eHunyuan-A52B-Instruct-FP8\u003c/a\u003e\u003c/td\u003e\n                \u003ctd style=\"width: 500px;\"\u003e\u003ca href=\"https://cdn-large-model.hunyuan.tencent.com/Hunyuan-A52B-Instruct-128k-fp8-20241116.zip\" style=\"color: red;\"\u003eHunyuan-A52B-Instruct-FP8\u003c/a\u003e\u003c/td\u003e\n            \u003c/tr\u003e\n            \u003ctr\u003e\n                \u003ctd style=\"width: 100px;\"\u003eHunyuan-A52B-Instruct\u003c/td\u003e\n                \u003ctd style=\"width: 500px;\"\u003e\u003ca href=\"https://huggingface.co/tencent/Tencent-Hunyuan-Large/tree/main/Hunyuan-A52B-Instruct\" style=\"color: red;\"\u003eHunyuan-A52B-Instruct\u003c/a\u003e\u003c/td\u003e\n                \u003ctd style=\"width: 500px;\"\u003e\u003ca href=\"https://cdn-large-model.hunyuan.tencent.com/Hunyuan-A52B-Instruct-128k-20241116.zip\" style=\"color: red;\"\u003eHunyuan-A52B-Instruct\u003c/a\u003e\u003c/td\u003e\n            \u003c/tr\u003e\n            \u003ctr\u003e\n                \u003ctd style=\"width: 100px;\"\u003eHunyuan-A52B-Pretrain\u003c/td\u003e\n                \u003ctd style=\"width: 500px;\"\u003e\u003ca href=\"https://huggingface.co/tencent/Tencent-Hunyuan-Large/tree/main/Hunyuan-A52B-Pretrain\" style=\"color: red;\"\u003eHunyuan-A52B-Pretrain\u003c/a\u003e\u003c/td\u003e\n                \u003ctd style=\"width: 500px;\"\u003e\u003ca href=\"https://cdn-large-model.hunyuan.tencent.com/Hunyuan-A52B-Pretrain-256k.zip\" style=\"color: red;\"\u003eHunyuan-A52B-Pretrain\u003c/a\u003e\u003c/td\u003e\n            \u003c/tr\u003e\n        \u003c/tbody\u003e\n    \u003c/table\u003e\n\u003c/p\u003e\n\n\u003cp\u003e\u003c/p\u003e\n\n\n## Model Introduction\n\nWith the rapid development of artificial intelligence technology, large language models (LLMs) have made significant progress in fields such as natural language processing, computer vision, and scientific tasks. However, as the scale of these models increases, optimizing resource consumption while maintaining high performance has become a key challenge. To address this challenge, we have explored Mixture of Experts (MoE) models. The currently unveiled Hunyuan-Large (Hunyuan-MoE-A52B) model is the largest open-source Transformer-based MoE model in the industry, featuring a total of 389 billion parameters and 52 billion active parameters. This is currently the largest open-source Transformer-based MoE model in the industry, featuring a total of 389 billion parameters and 52 billion active parameters. \n\nBy open-sourcing the Hunyuan-Large model and revealing related technical details, we hope to inspire more researchers with innovative ideas and collectively advance the progress and application of AI technology. We welcome you to join our open-source community to explore and optimize future AI models together!\n \n### Introduction to Technical Advantages\n\n#### Model\n- **High-Quality Synthetic Data**: By enhancing training with synthetic data, Hunyuan-Large can learn richer representations, handle long-context inputs, and generalize better to unseen data.\n\n- **KV Cache Compression**: Utilizes Grouped Query Attention (GQA) and Cross-Layer Attention (CLA) strategies to significantly reduce memory usage and computational overhead of KV caches, improving inference throughput.\n\n- **Expert-Specific Learning Rate Scaling**: Sets different learning rates for different experts to ensure each sub-model effectively learns from the data and contributes to overall performance.\n\n- **Long-Context Processing Capability**: The pre-trained model supports text sequences up to 256K, and the Instruct model supports up to 128K, significantly enhancing the ability to handle long-context tasks.\n\n- **Extensive Benchmarking**: Conducts extensive experiments across various languages and tasks to validate the practical effectiveness and safety of Hunyuan-Large.\n\n#### Inference Framework\n- This open-source release offers two inference backend options tailored for the Hunyuan-Large model: the popular [vLLM-backend](https://github.com/quinnrong94/vllm/tree/dev_hunyuan) and the TensorRT-LLM Backend. Both solutions include optimizations for enhanced performance. For instance, the introduction of a new CLA structure significantly reduces GPU memory usage, achieving a 50% savings in the KV-Cache portion, which ensures efficient handling of long text scenarios. Additionally, by employing FP8 quantization, we achieve a 50% reduction in memory usage compared to traditional FP16/BF16 quantization, while maintaining precision and resulting in a 70% increase in throughput. Meanwhile, by leveraging the efficient operators at the core of TRT-LLM, the performance of the TRT-LLM solution surpasses that of vLLM by over 30%. The TRT-LLM solution is widely used in Tencent's Hunyuan project. In this release, we are initially open-sourcing the vLLM solution, with plans to release the TRT-LLM solution in the near future.\n\n#### Training Framework\n\n- The Hunyuan-Large open-source model is fully compatible with the Hugging Face format, enabling researchers and developers to perform model fine-tuning using the hf-deepspeed framework. Additionally, we support training acceleration through the use of flash attention. To further assist in the adoption process, we have made the corresponding training scripts and model implementations publicly available to the community through this release, facilitating subsequent model training and fine-tuning operations based on these resources.\n\n\u0026nbsp;\n\n## Related News\n* 2024.11.25 Our self-developed long-context benchmark, i.e., PenguinScrolls, has been officially released! You can explore the project on [GitHub](https://github.com/Penguin-Scrolls/PenguinScrolls) and access the dataset on [Hugging Face](https://huggingface.co/datasets/Penguin-Scrolls/PenguinScrolls).\n* 2024.11.18 **Hunyuan-A52B-Instruct** and **Hunyuan-A52B-Instruct-FP8** model update. \n* 2024.11.5 [TI Platform](https://cloud.tencent.com/product/ti) has integrated Hunyuan-Large model already, you can easily train and deploy it in just a few steps. Visit [Chat with Hunyuan-Large](https://console.cloud.tencent.com/tione/v2/aimarket/detail/hunyuan_series?PublicAlgoGroupId=hunyuan-large-chat\u0026detailTab=demo) to experience real-time conversations with the model, and explore [Hunyuan-Large Best Practice on TI](https://cloud.tencent.com/document/product/851/112032) to create your own customized Hunyuan-Large model. \n* 2024.11.5 We have open-sourced **Hunyuan-A52B-Pretrain**, **Hunyuan-A52B-Instruct**, and **Hunyuan-A52B-Instruct-FP8** on Hugging Face. We also released a technical report and a training and inference operations manual, providing detailed information on the model's capabilities and the procedures for training and inference.\n\n\n\n\n## Benchmark Evaluation\n**Hunyuan-Large pre-trained model** achieves the best overall performance compared to both Dense and MoE based \ncompetitors having similar activated parameter sizes.  For aggregated benchmarks such as MMLU, MMLU-Pro, and CMMLU, \nHunyuan-Large consistently achieves the best performance, confirming its comprehensive abilities on aggregated tasks.\nHunyuan-Large also shows superior performance in commonsense understanding and reasoning, and classical NLP tasks \nsuch as QA and reading comprehension tasks (e.g., CommonsenseQA, PIQA and TriviaQA).  \nFor the mathematics capability, Hunyuan-Large outperforms all baselines in math datasets of GSM8K and MATH, \nand also gains the best results on CMATH in Chinese.We also observe that Hunyuan-Large achieves the overall \nbest performance in all Chinese tasks (e.g., CMMLU, C-Eval).\n\n| Model            | LLama3.1-405B | LLama3.1-70B | Mixtral-8x22B | DeepSeek-V2 | Hunyuan-Large |\n|------------------|---------------|--------------|---------------|-------------|---------------|\n| MMLU             | 85.2          | 79.3         | 77.8          | 78.5        | **88.4**          |\n| MMLU-Pro         | **61.6**          | 53.8         | 49.5          | -           | 60.2          |\n| BBH              | 85.9          | 81.6         | 78.9          | 78.9        | **86.3**          |\n| HellaSwag        | -             | -            | **88.7**      | 87.8        | 86.8          |\n| CommonsenseQA    | 85.8          | 84.1         | 82.4          | -           | **92.9**          |\n| WinoGrande       | 86.7          | 85.3         | 85.0          | 84.9        | **88.7**          |\n| PIQA             | -             | -            | 83.6          | 83.7        | **88.3**          |\n| NaturalQuestions | -             | -            | 39.6          | 38.7        | **52.8**          |\n| DROP             | 84.8          | 79.6         | 80.4          | 80.1        | **88.9**          |\n| ARC-C            | **96.1**          | 92.9         | 91.2          | 92.4        | 95.0          |\n| TriviaQA         | -             | -            | 82.1          | 79.9        | **89.2**          |\n| CMMLU            | -             | -            | 60.0          | 84.0        | **90.2**          |\n| C-Eval           | -             | -            | 59.6          | 81.7        | **91.9**          |\n| C3               | -             | -            | 71.4          | 77.4        | **82.3**          |\n| GSM8K            | 89.0          | 83.7         | 83.7          | 79.2        | **92.8**          |\n| MATH             | 53.8          | 41.4         | 42.5          | 43.6        | **69.8**          |\n| CMATH            | -             | -            | 72.3          | 78.7        | **91.3**          |\n| HumanEval        | 61.0          | 58.5         | 53.1          | 48.8        | **71.4**          |\n| MBPP             | **73.4**          | 68.6         | 64.2          | 66.6        | 72.6          |\n\n**Hunyuan-Large-Instruct** achieves consistent improvements on most types of tasks compared to LLMs having similar \nactivated parameters, indicating the effectiveness of our post-training.    Delving into the model performance \nin different categories of benchmarks, we find that our instruct model achieves the best performance on MMLU and MATH dataset.  \nNotably, on the MMLU dataset, our model demonstrates a significant improvement, outperforming the LLama3.1-405B model by 2.6%.   \nThis enhancement is not just marginal but indicative of the Hunyuan-Large-Instruct’s superior understanding and reasoning \ncapabilities across a wide array of language understanding tasks. The model’s prowess is further underscored in its performance \non the MATH dataset, where it surpasses the LLama3.1-405B by a notable margin of 3.6%.  \nRemarkably, this leap in accuracy is achieved with only 52 billion activated parameters, underscoring the efficiency of our model.\n\n| Model                | LLama3.1 405B Inst. | LLama3.1 70B Inst. | Mixtral 8x22B Inst. | DeepSeekV2.5 Chat | Hunyuan-Large Inst. |\n|----------------------|---------------------|--------------------|---------------------|-------------------|---------------------|\n| MMLU                 | 87.3                | 83.6               | 77.8                | 80.4              | **89.9**            |\n| CMMLU                | -                   | -                  | 61.0                | -                 | **90.4**            |\n| C-Eval               | -                   | -                  | 60.0                | -                 | **88.6**            |\n| BBH                  | -                   | -                  | 78.4                | 84.3              | **89.5**            |\n| HellaSwag            | -                   | -                  | 86.0                | **90.3**          | 88.5                |\n| ARC-C                | **96.9**            | 94.8               | 90.0                | -                 | 94.6                |\n| GPQA_diamond         | **51.1**            | 46.7               | -                   | -                 | 42.4                |\n| MATH                 | 73.8                | 68.0               | 49.8                | 74.7              | **77.4**            |\n| HumanEval            | 89.0                | 80.5               | 75.0                | 89.0              | **90.0**            |\n| AlignBench           | 6.0                 | 5.9                | 6.2                 | 8.0               | **8.3**             |\n| MT-Bench             | 9.1                 | 8.8                | 8.1                 | 9.0               | **9.4**             |\n| IFEval strict-prompt | **86.0**            | 83.6               | 71.2                | -                 | 85.0                |\n| Arena-Hard |  69.3            | 55.7               |  -                | 76.2                 | **81.8**            |\n| AlpacaEval-2.0 | 39.3            | 34.3               | 30.9                | 50.5                 | **51.8**            |\n\n\n\n\n\n\u0026nbsp;\n\n## Data\n\nHunyuan-Large provides model training-related processes. In this section, you can process the training data format for use in model training.\n\n### Training Data Format and Processing\n\nTraining data is processed into a messages format as follows. The default system prompt for training and inference is \"You are a helpful assistant.\" Below are examples of single-turn and multi-turn data:\n\n\n```python\n# Single-Turn Data Format\n\nmessages = [\n    {\"role\": \"system\", \"content\": \"You are a helpful assistant.\"},\n    {\"role\": \"user\", \"content\": \"海水为什么是咸的\" },\n    {\"role\": \"assistant\", \"content\": \"海水是咸的主要是因为其中含有许多溶解在水中的盐类和矿物质。这些盐类和矿物质来自于地球表面的岩石和土壤中的化学物质，随着时间的推移，它们被带到了海洋中。当海水蒸发时，水分蒸发掉了，但盐类和矿物质仍然留在水中，导致海水变得更加咸味。因此，海水的咸度是由其中的盐类和矿物质的含量决定的。\"}\n]\n\n# Multi-Turn Data Format\nmessages = [\n    {\"role\": \"system\", \"content\": \"You are a helpful assistant.\"},\n    {\"role\": \"user\", \"content\": \"知道重庆森林这部电影吗？\" },\n    {\"role\": \"assistant\", \"content\": \"知道呀，是一部由王家卫导演的片子。\" },\n    {\"role\": \"user\", \"content\": \" 而主演里更是有王菲，一上映便受到追捧。\" },\n    {\"role\": \"assistant\", \"content\": \"所以此片获得了第14届香港电影金像奖最佳影片奖。\"}\n]\n\nfrom tokenizers import ByteLevelBPETokenizer\nfrom transformers import AutoTokenizer\n\ntokenizer = AutoTokenizer.from_pretrained(\"./models\", use_fast=False, trust_remote_code=True)\n\nids = tokenizer.apply_chat_template(messages)\n\n```\n\n\n\nFor more usage references, see the `./models/test.py` file.\n\n\u0026nbsp;\n\n## Quick Start\n\nYou can quickly get started by referring to the content in the \u003ca href=\"examples/README.md\"\u003eQuick Start Guide\u003c/a\u003e.\n\n## Model Training\n\nTo simplify the Training process, HunyuanLLM provides a pre-built Docker image:\n\n [hunyuaninfer/hunyuan-large](https://hub.docker.com/repository/docker/hunyuaninfer/hunyuan-large/general). \n\n### Hardware Requirements\n\nTested on H20, without enabling `make_moe_param_leaf_module` and using `zero3+offload`, with a `max_seq_length` of 2048, full fine-tuning requires at least 32 GPUs, and LoRA fine-tuning requires at least 8 GPUs.\n\n### Training Performance\n\nWith the minimum configuration (8 GPUs for LoRA fine-tuning), `per_device_train_batch_size` is set to 1, and `gradient_accumulation_steps` is set to 1, resulting in approximately 35 seconds per iteration.\n\n### Launch Method\n\nRefer to: [HuggingFace Transformers Trainer](https://huggingface.co/docs/transformers/v4.19.2/en/main_classes/trainer)\n\n#### Single-Machine Training\n\nIn the `train` directory, execute:\n\n```sh\npip install -r requirements.txt\nbash train.sh\n```\n\n#### Multi-Machine Training\n\nTo start training on multiple machines, follow the steps below and ensure that all machines are within the same cluster.\n\n##### Configure Passwordless SSH Login Between Machines\n\nThe following steps use two machines as an example, with their IPs represented as `${ip1}` and `${ip2}`. These operations are performed within a Docker container.\n\nFirst, configure passwordless SSH between containers on each machine.\n\n\n```sh\nssh-keygen\t\t\t# Generate id_rsa and id_rsa.pub for passwordless login\nssh-keygen -t rsa -A    # Generate /etc/ssh/ssh_host_rsa_key and ssh_host_ecdsa_key for starting 'SSH listen' later\n/usr/sbin/sshd -p 36005 -o ListenAddress=0.0.0.0        # Start SSH listen\necho \"Port 36005\" \u003e ~/.ssh/config   # Change SSH connection port to 36005\npasswd root    # Set root password to avoid alerts from monitoring platforms\n```\n\n\nNote: The `36005` here is an example. You can choose any port, but ensure that the port is **open** and **not occupied by other processes**.\n\nNext, within the container on each machine, execute:\n\n```sh\ncat ~/.ssh/id_rsa.pub\n```\n\n**Copy the output SSH public key and paste it into the `~/.ssh/authorized_keys` file, with one public key per line. This must be done on every machine.** Ultimately, the `~/.ssh/authorized_keys` file on each machine should be identical and contain the public keys of all machines.\n\nIt's important to note that during multi-node training, the code executed on each node must be consistent. It is recommended to mount a shared network drive. If mounting a shared drive is not possible, you need to manually copy the dataset, scripts, and code to the same directory on all machines.\n\n##### Start Multi-Machine Training\n\nOnce the preparation steps are completed and dependencies are confirmed to be installed (if not, execute `pip install -r requirements.txt` to install), you can add the following configuration at the beginning of `train.sh`:\n\n```shell\nexport HOST_GPU_NUM=8\n# Current machine IP\nexport LOCAL_IP=${ip1}\n# Multi-node machine IPs, separated by commas\nexport NODE_IP_LIST=\"${ip1}:8,${ip2}:8\"\n# Number of machine nodes\nexport NODES=2\nexport NODE_NUM=$((${NODES} * ${HOST_GPU_NUM}))\n```\n\nNote: Replace `${ip1}` and `${ip2}` with the actual IP addresses!\n\nThen, on the machine with `${ip1}`, execute `bash train.sh` in the `train/` directory. Note that on the first run, you might see the following output:\n\n```ssh\nThe authenticity of host '[ip]:36005 ([ip]:36005)' can't be established.\nECDSA key fingerprint is xxxxxx.\nECDSA key fingerprint is MD5:xxxxxx.\nAre you sure you want to continue connecting (yes/no)?\n```\n\nAt this point, type `yes` to continue.\n\n##### Key Parameters\n\nThe key parameters in the script are as follows:\n\n- `--deepspeed`: This parameter should point to a DeepSpeed configuration file. The `train` folder provides three default DeepSpeed configuration files: `ds_zero2_no_offload.json`, `ds_zero3_no_offload.json`, `ds_zero3_offload.json`. The required GPU memory decreases in this order.\n- `--model_name_or_path`: The path to the HF pre-trained model. Ensure this path contains the `modeling_hunyuan.py` and `configuration_hunyuan.py` files; otherwise, it cannot be loaded.\n- `--tokenizer_name_or_path`: The path to the tokenizer folder. Ensure this path contains the `tokenization_hy.py` file; otherwise, it cannot be loaded.\n- `--train_data_file`: The path to the training file, which should be a JSONL file.\n- `--output_dir`: The output directory where logs, tensorboard files, and model weights will be stored.\n- `--per_device_train_batch_size`: The batch size per GPU.\n- `--gradient_accumulation_steps`: The number of gradient accumulation steps. The global batch size is `per_device_train_batch_size * gradient_accumulation_steps * dp_size`.\n- `--max_steps`: The total number of training steps.\n- `--save_steps`: The number of steps between saving checkpoints.\n- `--use_lora`: Whether to use LoRA for training. This also accepts `--lora_rank`, `--lora_alpha`, and `--lora_dropout` parameters. LoRA is applied by default to the 'q_proj', 'k_proj', 'v_proj', 'o_proj' parameters. If you need to change this, modify it in the code. Note: **When using LoRA for training, only the LoRA weights are saved, not the base model weights**. If you need to merge LoRA weights, see the \"LoRA Weight Merging\" section below.\n- `--make_moe_param_leaf_module`: When using zero3 and MoE training, treat the MoE module as a leaf module, meaning its parameters are not split by zero3. This option is expected to significantly increase memory usage.\n- `--gradient_checkpointing`: Enable gradient checkpointing.\n- `--train_attention_params_only`: Whether to train only the attention parameters.\n- `--learning_rate`: The maximum learning rate during training.\n- `--min_lr`: The minimum learning rate during training.\n- `--use_flash_attn`: 开启 flash-attention 进行训练加速\n\n**Note:**\n\n- If you want to continue training from a previously saved checkpoint instead of loading pre-trained weights, specify `--resume_from_checkpoint` with the path to the checkpoint from the previous training. Do not specify `--model_name_or_path`, as this will only load the weights and not the training state.\n- When continuing training from a checkpoint, there might be slight deviations in loss due to randomness introduced by some non-deterministic algorithms, which is considered normal. Refer to: [HuggingFace Transformers Trainer Randomness](https://huggingface.co/docs/transformers/main/en/perf_train_gpu_one#randomness)\n- When `--model_name_or_path` is specified, all model-related parameters will be ignored.\n- Samples within a batch will be padded to align with the longest sample in the batch, with each sample having a maximum length of `max_seq_length`. Any excess will be truncated.\n- If you encounter warnings about bias weights not being loaded, you can ignore them, as biases are not used in Hunyuan-Large.\n\n\n#### What to Do If Out of Memory?\n\nRefer to: [DeepSpeed Configuration](https://www.deepspeed.ai/docs/config-json/)\n\nYou can try modifying the DeepSpeed configuration by removing the auto attribute from these parameters and reducing their values:\n\n- `stage3_param_persistence_threshold`\n- `stage3_prefetch_bucket_size`\n- `stage3_max_reuse_distance`\n- `stage3_max_reuse_distance`\n\n#### Merging LoRA Models\n\nThe saved LoRA weights cannot be merged into the zero3 model during training because, with zero3 enabled, model weights are split across different data parallel ranks. If you want to merge LoRA weights into the base model, you can do so offline to obtain the merged weight file. Execute `merge_lora_weight.sh` to merge the LoRA weights with the base model weights. The parameters include:\n\n- `--base_model_path`: Directory of the base model weights\n- `--adapter_model_path`: Directory of the LoRA weights\n- `--output_path`: Directory to save the merged weights\n- `--save_dtype`: Data format for storing the merged weights, available options include: fp16, bf16, fp32\n\n\u0026nbsp;\n\n## Inference and Deployment\n\nHunyuanLLM uses TRT-LLM and vLLM for deployment. We are open sourcing the [vLLM-backend](https://github.com/quinnrong94/vllm/tree/dev_hunyuan) deployment (see Reasoning with vLLM), and the TRT-LLM deployment (see Reasoning with TRT-LLM) will be available in the near future.\n\n## Using TRT-LLM for Inference\n\nTo be opened\n\n## Using vLLM for Inference\n\n### Docker:\n\nTo simplify the deployment process, HunyuanLLM provides a pre-built Docker image:\n\n [hunyuaninfer/hunyuan-large](https://hub.docker.com/repository/docker/hunyuaninfer/hunyuan-large/general). You only need to download the model files and start the Docker container using the code below to begin model inference.\n\n```shell\ndocker run --name hunyuanLLM_infer -itd --privileged --user root --net=host --ipc=host --gpus=8 hunyuaninfer/hunyuan-large:infer-open-source\n```\n\nNote: Docker container privilege management. The above code uses privileged mode (`--privileged`) to start the Docker container, which grants the container higher privileges, increasing the risk of data leakage and cluster security threats. It is recommended to avoid using privileged mode unless necessary to reduce security risks. For scenarios where privileged mode is required, conduct a thorough security assessment and implement appropriate security monitoring and hardening measures.\n\n### Configure Passwordless SSH Login Between Machines\n\nThe following steps use two machines as an example, with their IPs represented as `${ip1}` and `${ip2}`. These operations are performed within a Docker container.\n\nFirst, run `passwd` on both machines to set a password, for example: `Tmp123,./`\n\nCopy `inference/login_ssh.py` into the container and execute the following command, ensuring the IP and password are correctly entered.\n\n```shell\npython3 login_ssh.py --ips ${ip1},${ip2} --port 36000 --password=Tmp123,./\n```\n\n**Note 📢: Before starting, be sure to verify multi-machine communication using VLLM's debugging script: https://docs.vllm.ai/en/latest/getting_started/debugging.html**\n\n### BF16 Deployment\n\nBF16 requires 16 H20 GPUs for deployment. After verifying that multi-machine communication is correct, execute the following steps:\n\nBefore running the commands, set the following environment variables:\n\n```shell\n${LOCAL_IP}: The IP corresponding to bond1 on the current machine\n${MODEL_PATH}: Path to the Hunyuan LLM model\n```\n\n#### Step 1: Start Ray\n\nRay is an open-source library for parallel and distributed Python. In this section, we use Ray to achieve multi-machine communication.\n\nRay Component Configuration Hardening: The default configuration of Ray components does not enable authentication mechanisms for service ports (e.g., 6379, 8265), posing risks of unauthorized access and command execution. It is recommended to deploy Ray components only in trusted internal network environments or ensure strict access control list (ACL) policies are implemented for these ports to prevent unauthorized network access.\n\nFirst, start Ray on each node (either in the background or by keeping the terminal running):\n\nOn the head node:\n```shell\nexport VLLM_HOST_IP=${LOCAL_IP}\nexport NCCL_SOCKET_IFNAME=bond1\nexport GLOO_SOCKET_IFNAME=bond1\nray start --block --head --node-ip-address=${LOCAL_IP} --port=6379\n```\n\nOn all worker nodes:\n\nNote: Replace `{HEAD NODE $LOCAL_IP}` with the actual `${LOCAL_IP}` of the head node.\n```shell\nexport VLLM_HOST_IP=${LOCAL_IP}\nexport NCCL_SOCKET_IFNAME=bond1\nexport GLOO_SOCKET_IFNAME=bond1\nray start --block --address={HEAD NODE $LOCAL_IP}:6379 --node-ip-address=${LOCAL_IP}\n```\nIf Ray fails to start, execute `ray stop` and then run the above commands again.\n\n#### Step 2: Execute Inference\n\n#### Method 1: Command Line Inference\n\nBelow is a code snippet demonstrating how to quickly request the chat model using `vLLM`:\n\nNote: vLLM Component Remote Code Execution Protection. In the code below, if the `trust-remote-code` configuration option of the vLLM component is enabled, it will allow loading and executing code from remote model repositories, which may lead to the execution of malicious code. Unless explicitly required by business needs, it is recommended to keep this configuration option disabled to reduce potential security threats.\n\n```python\nimport os\nfrom vllm import LLM, SamplingParams\n\nmodel_path=os.environ.get('MODEL_PATH')\n\nllm = LLM(model=model_path,\n        tokenizer=model_path,\n        trust_remote_code=True,\n        max_model_len=10240,\n        dtype='bfloat16',\n        tensor_parallel_size=16,\n        pipeline_parallel_size=1,\n        disable_log_stats=False,\n        gpu_memory_utilization=0.98,\n        disable_custom_all_reduce=True,\n        #distributed_executor_backend='ray',\n        enforce_eager=True,\n        max_num_seqs=8,\n        use_v2_block_manager=True,\n        quantization=None)\n\nprompts = [\"海水为什么是咸的\"]\n\nsampling_params = SamplingParams(\n    temperature=0.7, top_p=0.6, max_tokens=200, top_k=20, repetition_penalty=1.05)\n\noutputs = llm.generate(prompts, sampling_params)\n\n# Print the outputs.\nfor output in outputs:\n    prompt = output.prompt\n    generated_text = output.outputs[0].text\n    print(f\"Prompt: {prompt!r}, Generated text: {generated_text!r}\")\n```\n\n#### Method 2: Service-Based Inference\n\nBelow we demonstrate how to deploy the model using `vLLM` in a service-based manner and make requests.\n\nRun the following on the head node:\n\n```shell\nexport VLLM_HOST_IP=${LOCAL_IP}\nexport NCCL_SOCKET_IFNAME=bond1\nexport GLOO_SOCKET_IFNAME=bond1\n```\n\nNext, start the service by running:\n\n```shell\ncd inference\nsh run_server.sh\n```\n\n*Tips*: Troubleshooting, if you encounter the following error:\n\n```python\nray.exceptions.RaySystemError: System error: No module named 'transformers_modules' traceback: Traceback (most recent call last):\nModuleNotFoundError: No module named 'transformers_modules'\n```\n\nCopy the `~/.cache/huggingface/modules/` directory from the head node to the corresponding path on all worker nodes.\n\nAfter successfully running `run_server.sh`, execute the request script:\n\n```shell\nsh openapi.sh\n```\n\nBe sure to modify `${LOCAL_IP}` and `${MODEL_PATH}` in `openapi.sh` to values match the corresponding service.\n\n\n### Quantized Model Deployment:\n\nThis section describes the process of deploying a quantized model using vLLM.\n\nImage: The deployment image is the same as for BF16.\n\n#### Int8 Quantized Model Deployment:\n\nTo deploy the Int8-weight-only version of the Hunyuan-L model, simply set the environment variables in `run_server_int8.sh`:\n\n```shell\n${MODEL_PATH}: Path to the BF16 model\n${LOCAL_IP}: The IP corresponding to bond1 on the current machine\n```\n\nThen, start the Int8 service by running:\n\n```shell\nsh run_server_int8.sh\n```\n\nAfter successfully running `run_server_int8.sh`, execute the request script:\n\n```shell\nsh openapi.sh\n```\n\n#### FP8 Quantized Model Deployment:\n\nTo deploy the W8A8C8 version of the Hunyuan-L model, simply set the environment variables in `run_server_fp8.sh`:\n\n```shell\n${MODEL_PATH}: Path to the FP8 model\n${LOCAL_IP}: The IP corresponding to bond1 on the current machine\n```\n\nThen, start the FP8 service by running:\n\n```shell\nsh run_server_fp8.sh\n```\n\nAfter successfully running `run_server_fp8.sh`, execute the request script:\n\n```shell\nsh openapi.sh\n```\n\n#### FP8 BENCHMARK\n\nThis part introduces the Benchmark of Hunyuan Large Instruct FP8 quantitative model.\n\n| Dataset | BF16 | W8A8C8-FP8 |\n|---------|------|------------|\n| ARC-C   | 94.6 | 94.2       |\n| C-Eval  | 88.6 | 89.2       |\n| CMMLU   | 90.4 | 89.8       |\n| MMLU    | 89.9 | 88.9       |\n\n### Inference Performance\n\nThis section presents the efficiency test results of deploying various models (original and quantized) using vLLM, including inference speed (tokens/s) under different batch sizes.\n\n| Inference Framework | Model                                                                                                  | Number of GPUs (H20) | input_length | batch=1 | batch=4 |\n| ------------------- | ------------------------------------------------------------------------------------------------------ | -------------------- | ------------ |---------|---------|\n| vLLM                | Hunyuan-Large                                                                                              | 16                   | 2048         | 20.2    | 75.5    |\n| vLLM                | Hunyuan-Large(int8 weight only)                                                                            | 8                    | 2048         | 19.3    | 73.6    |\n| vLLM                | Hunyuan-Large(W8A8C8-FP8)                                                                                  | 8                    | 2048         | 19.8    | 74.9    |\n\n## Tokenizer\n\nThe tokenizer used in the HunYuan-Large model balances compression rate and effectiveness, ensuring that embeddings are sufficiently trained. The vocabulary includes 100K tokens integrated from tiktoken. Additionally, we trained an extra 29K Chinese tokens using a large amount of high-quality Chinese training data to enhance the model's Chinese capabilities and the tokenizer's compression rate. Combined, our new tokenizer improves the compression rate compared to the LLaMA3 tokenizer, increasing from 2.78 characters/token to 3.13 characters/token.\n\n## Hunyuan API\n\nYou can experience our Hunyuan-Large model on Tencent Cloud. For details, please visit: https://cloud.tencent.com/document/product/1729/97730.\n\n## Interactive Demo Web\n\nThe Hunyuan-Large web demo is now open. Visit https://huggingface.co/spaces/tencent/Hunyuan-Large to easily experience our model.\n\n## Training/Inference on TI\nTencent Cloud's [TI Platform](https://cloud.tencent.com/product/ti) is a comprehensive machine learning platform tailored for AI engineers. With the Hunyuan-Large model already integrated, you can easily train and deploy it in just a few steps. Visit [Chat with Hunyuan-Large](https://console.cloud.tencent.com/tione/v2/aimarket/detail/hunyuan_series?PublicAlgoGroupId=hunyuan-large-chat\u0026detailTab=demo) to experience real-time conversations with the model, and explore [Hunyuan-Large Best Practice on TI](https://cloud.tencent.com/document/product/851/112032) to create your own customized Hunyuan-Large model. \n\n\n## Citation\nIf you find our work helpful, feel free to give us a cite.\n\n```\n@misc{sun2024hunyuanlargeopensourcemoemodel,\n      title={Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent}, \n      author={Xingwu Sun and Yanfeng Chen and Yiqing Huang and Ruobing Xie and Jiaqi Zhu and Kai Zhang and Shuaipeng Li and Zhen Yang and Jonny Han and Xiaobo Shu and Jiahao Bu and Zhongzhi Chen and Xuemeng Huang and Fengzong Lian and Saiyong Yang and Jianfeng Yan and Yuyuan Zeng and Xiaoqin Ren and Chao Yu and Lulu Wu and Yue Mao and Tao Yang and Suncong Zheng and Kan Wu and Dian Jiao and Jinbao Xue and Xipeng Zhang and Decheng Wu and Kai Liu and Dengpeng Wu and Guanghui Xu and Shaohua Chen and Shuang Chen and Xiao Feng and Yigeng Hong and Junqiang Zheng and Chengcheng Xu and Zongwei Li and Xiong Kuang and Jianglu Hu and Yiqi Chen and Yuchi Deng and Guiyang Li and Ao Liu and Chenchen Zhang and Shihui Hu and Zilong Zhao and Zifan Wu and Yao Ding and Weichao Wang and Han Liu and Roberts Wang and Hao Fei and Peijie She and Ze Zhao and Xun Cao and Hai Wang and Fusheng Xiang and Mengyuan Huang and Zhiyuan Xiong and Bin Hu and Xuebin Hou and Lei Jiang and Jiajia Wu and Yaping Deng and Yi Shen and Qian Wang and Weijie Liu and Jie Liu and Meng Chen and Liang Dong and Weiwen Jia and Hu Chen and Feifei Liu and Rui Yuan and Huilin Xu and Zhenxiang Yan and Tengfei Cao and Zhichao Hu and Xinhua Feng and Dong Du and Tinghao She and Yangyu Tao and Feng Zhang and Jianchen Zhu and Chengzhong Xu and Xirui Li and Chong Zha and Wen Ouyang and Yinben Xia and Xiang Li and Zekun He and Rongpeng Chen and Jiawei Song and Ruibin Chen and Fan Jiang and Chongqing Zhao and Bo Wang and Hao Gong and Rong Gan and Winston Hu and Zhanhui Kang and Yong Yang and Yuhong Liu and Di Wang and Jie Jiang},\n      year={2024},\n      eprint={2411.02265},\n      archivePrefix={arXiv},\n      primaryClass={cs.CL},\n      url={https://arxiv.org/abs/2411.02265}, \n}\n```\n\u003cbr\u003e\n\n## Contact Us\n\nIf you would like to leave a message for our R\u0026D and product teams, Welcome to contact our open-source team . You can also contact us via email (hunyuan_opensource@tencent.com).\n\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FTencent%2FTencent-Hunyuan-Large","html_url":"https://awesome.ecosyste.ms/projects/github.com%2FTencent%2FTencent-Hunyuan-Large","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FTencent%2FTencent-Hunyuan-Large/lists"}