{"id":13393543,"url":"https://github.com/lm-sys/FastChat","last_synced_at":"2025-03-13T19:31:43.233Z","repository":{"id":149282184,"uuid":"615882673","full_name":"lm-sys/FastChat","owner":"lm-sys","description":"An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.","archived":false,"fork":false,"pushed_at":"2025-03-01T06:43:01.000Z","size":35315,"stargazers_count":38069,"open_issues_count":921,"forks_count":4652,"subscribers_count":355,"default_branch":"main","last_synced_at":"2025-03-10T22:30:20.852Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/lm-sys.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2023-03-19T00:18:02.000Z","updated_at":"2025-03-10T20:23:14.000Z","dependencies_parsed_at":"2023-10-16T11:30:55.316Z","dependency_job_id":"861bf0fd-4ce6-485d-b389-8979f62d789f","html_url":"https://github.com/lm-sys/FastChat","commit_stats":{"total_commits":877,"total_committers":243,"mean_commits":"3.6090534979423867","dds":0.6921322690992018,"last_synced_commit":"8664268cf63237cda8954531bca2a3ba6d3c0e1d"},"previous_names":[],"tags_count":23,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/lm-sys%2FFastChat","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/lm-sys%2FFastChat/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/lm-sys%2FFastChat/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/lm-sys%2FFastChat/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/lm-sys","download_url":"https://codeload.github.com/lm-sys/FastChat/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":243469179,"owners_count":20295702,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-07-30T17:00:55.691Z","updated_at":"2025-03-13T19:31:43.220Z","avatar_url":"https://github.com/lm-sys.png","language":"Python","funding_links":[],"categories":["Python","Fine-tuning","Models and their datasets","3 Reasoning Tasks","精选开源项目合集","文本创作类","Fine Tuned Models","Open-source Models","2 Foundation Models","Table of Contents","LLM部署","Web apps","Platforms \u0026 Frameworks","Model","精选文章","LLM-List","Datasets","[tatsu-lab/stanford_alpaca](https://github.com/tatsu-lab/stanford_alpaca)","Large Language Model","Uncategorized","Large Language Models (LLMs)","Developer Libraries, SDKs, and APIs","A01_文本生成_文本对话","LLM Deployment","7. Training \u0026 Fine-tuning Ecosystem","Textual Large Language Model Backbone","HarmonyOS","ChatGPT Integrated Projects","others","Large Language Model Data","Repos","[TavernAI/TavernAI](https://github.com/TavernAI/TavernAI)","Learning","Industry Strength Natural Language Processing","Chat Models","LLM Inference","Tools for deploying LLM","💻 Software for Large Language Models","Models and Projects","🖥 Local Deployment Tools","Model Serving Frameworks","🚀 Model Serving \u0026 Deployment","General LLM Benchmarks","Training","Personal Assistants \u0026 Conversational Agents","App","📋 List of Open-Source Projects","multi lang models","LLMs Eval","\u003ca name=\"Python\"\u003e\u003c/a\u003ePython","8. Inference Engines","AI Coding Assistants","Models","📊 AI Evaluation \u0026 Benchmarks","Runtime","LLM Serving / Inference","UIs","Chatbots \u0026 Virtual Companions","Communities \u0026 Organizations","Machine Learning"],"sub_categories":["Vicuna","3.7 Multimodal Reasoning","GPT开源平替机器人","**开源项目与本地部署**","2.1 Language Foundation Models","Autonomous Systems","LLM 评估工具","Hosted and self-hosted","Large Language Model","大语言模型训练-评估平台","Open-LLM","Other LLaMA-derived projects:","Uncategorized","Contribute to our Repository","Python","大语言对话模型及数据","Windows Manager","Fine-tuning Data","GPT开源平替机器人🔥🔥🔥","数据","Repositories","Chinese Support","Other Cloud Provider Credits","Ray + LLM","Server Deployment \u0026 High-Performance Inference","LangManus","Chatbots","LLM Infra and Optimization","Server / Production","Large Language Models","Usage Tips","Chatbot","Command-line(shell) interface"],"readme":"# FastChat\n| [**Demo**](https://lmarena.ai/) | [**Discord**](https://discord.gg/6GXcFg3TH8) | [**X**](https://x.com/lmsysorg) |\n\nFastChat is an open platform for training, serving, and evaluating large language model based chatbots.\n- FastChat powers Chatbot Arena ([lmarena.ai](https://lmarena.ai)), serving over 10 million chat requests for 70+ LLMs.\n- Chatbot Arena has collected over 1.5M human votes from side-by-side LLM battles to compile an online [LLM Elo leaderboard](https://lmarena.ai/?leaderboard).\n\nFastChat's core features include:\n- The training and evaluation code for state-of-the-art models (e.g., Vicuna, MT-Bench).\n- A distributed multi-model serving system with web UI and OpenAI-compatible RESTful APIs.\n\n## News\n- [2024/03] 🔥 We released Chatbot Arena technical [report](https://arxiv.org/abs/2403.04132).\n- [2023/09] We released **LMSYS-Chat-1M**, a large-scale real-world LLM conversation dataset. Read the [report](https://arxiv.org/abs/2309.11998).\n- [2023/08] We released **Vicuna v1.5** based on Llama 2 with 4K and 16K context lengths. Download [weights](#vicuna-weights).\n- [2023/07] We released **Chatbot Arena Conversations**, a dataset containing 33k conversations with human preferences. Download it [here](https://huggingface.co/datasets/lmsys/chatbot_arena_conversations).\n\n\u003cdetails\u003e\n\u003csummary\u003eMore\u003c/summary\u003e\n\n- [2023/08] We released **LongChat v1.5** based on Llama 2 with 32K context lengths. Download [weights](#longchat).\n- [2023/06] We introduced **MT-bench**, a challenging multi-turn question set for evaluating chatbots. Check out the blog [post](https://lmsys.org/blog/2023-06-22-leaderboard/).\n- [2023/06] We introduced **LongChat**, our long-context chatbots and evaluation tools. Check out the blog [post](https://lmsys.org/blog/2023-06-29-longchat/).\n- [2023/05] We introduced **Chatbot Arena** for battles among LLMs. Check out the blog [post](https://lmsys.org/blog/2023-05-03-arena).\n- [2023/03] We released **Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90% ChatGPT Quality**. Check out the blog [post](https://vicuna.lmsys.org).\n\n\u003c/details\u003e\n\n\u003ca href=\"https://lmarena.ai\"\u003e\u003cimg src=\"assets/demo_narrow.gif\" width=\"70%\"\u003e\u003c/a\u003e\n\n## Contents\n- [Install](#install)\n- [Model Weights](#model-weights)\n- [Inference with Command Line Interface](#inference-with-command-line-interface)\n- [Serving with Web GUI](#serving-with-web-gui)\n- [API](#api)\n- [Evaluation](#evaluation)\n- [Fine-tuning](#fine-tuning)\n- [Citation](#citation)\n\n## Install\n\n### Method 1: With pip\n\n```bash\npip3 install \"fschat[model_worker,webui]\"\n```\n\n### Method 2: From source\n\n1. Clone this repository and navigate to the FastChat folder\n```bash\ngit clone https://github.com/lm-sys/FastChat.git\ncd FastChat\n```\n\nIf you are running on Mac:\n```bash\nbrew install rust cmake\n```\n\n2. Install Package\n```bash\npip3 install --upgrade pip  # enable PEP 660 support\npip3 install -e \".[model_worker,webui]\"\n```\n\n## Model Weights\n### Vicuna Weights\n[Vicuna](https://lmsys.org/blog/2023-03-30-vicuna/) is based on Llama 2 and should be used under Llama's [model license](https://github.com/facebookresearch/llama/blob/main/LICENSE).\n\nYou can use the commands below to start chatting. It will automatically download the weights from Hugging Face repos.\nDownloaded weights are stored in a `.cache` folder in the user's home folder (e.g., `~/.cache/huggingface/hub/\u003cmodel_name\u003e`).\n\nSee more command options and how to handle out-of-memory in the \"Inference with Command Line Interface\" section below.\n\n**NOTE: `transformers\u003e=4.31` is required for 16K versions.**\n\n| Size | Chat Command | Hugging Face Repo |\n| ---  | --- | --- |\n| 7B   | `python3 -m fastchat.serve.cli --model-path lmsys/vicuna-7b-v1.5`  | [lmsys/vicuna-7b-v1.5](https://huggingface.co/lmsys/vicuna-7b-v1.5)   |\n| 7B-16k   | `python3 -m fastchat.serve.cli --model-path lmsys/vicuna-7b-v1.5-16k`  | [lmsys/vicuna-7b-v1.5-16k](https://huggingface.co/lmsys/vicuna-7b-v1.5-16k)   |\n| 13B  | `python3 -m fastchat.serve.cli --model-path lmsys/vicuna-13b-v1.5` | [lmsys/vicuna-13b-v1.5](https://huggingface.co/lmsys/vicuna-13b-v1.5) |\n| 13B-16k  | `python3 -m fastchat.serve.cli --model-path lmsys/vicuna-13b-v1.5-16k` | [lmsys/vicuna-13b-v1.5-16k](https://huggingface.co/lmsys/vicuna-13b-v1.5-16k) |\n| 33B  | `python3 -m fastchat.serve.cli --model-path lmsys/vicuna-33b-v1.3` | [lmsys/vicuna-33b-v1.3](https://huggingface.co/lmsys/vicuna-33b-v1.3) |\n\n**Old weights**: see [docs/vicuna_weights_version.md](docs/vicuna_weights_version.md) for all versions of weights and their differences.\n\n### Other Models\nBesides Vicuna, we also released two additional models: [LongChat](https://lmsys.org/blog/2023-06-29-longchat/) and FastChat-T5.\nYou can use the commands below to chat with them. They will automatically download the weights from Hugging Face repos.\n\n| Model | Chat Command | Hugging Face Repo |\n| ---  | --- | --- |\n| LongChat-7B   | `python3 -m fastchat.serve.cli --model-path lmsys/longchat-7b-32k-v1.5`  | [lmsys/longchat-7b-32k](https://huggingface.co/lmsys/longchat-7b-32k-v1.5)   |\n| FastChat-T5-3B   | `python3 -m fastchat.serve.cli --model-path lmsys/fastchat-t5-3b-v1.0`  | [lmsys/fastchat-t5-3b-v1.0](https://huggingface.co/lmsys/fastchat-t5-3b-v1.0) |\n\n## Inference with Command Line Interface\n\n\u003ca href=\"https://lmarena.ai\"\u003e\u003cimg src=\"assets/screenshot_cli.png\" width=\"70%\"\u003e\u003c/a\u003e\n\n(Experimental Feature: You can specify `--style rich` to enable rich text output and better text streaming quality for some non-ASCII content. This may not work properly on certain terminals.)\n\n#### Supported Models\nFastChat supports a wide range of models, including\nLLama 2, Vicuna, Alpaca, Baize, ChatGLM, Dolly, Falcon, FastChat-T5, GPT4ALL, Guanaco, MTP, OpenAssistant, OpenChat, RedPajama, StableLM, WizardLM, xDAN-AI and more.\n\nSee a complete list of supported models and instructions to add a new model [here](docs/model_support.md).\n\n#### Single GPU\nThe command below requires around 14GB of GPU memory for Vicuna-7B and 28GB of GPU memory for Vicuna-13B.\nSee the [\"Not Enough Memory\" section](#not-enough-memory) below if you do not have enough memory.\n`--model-path` can be a local folder or a Hugging Face repo name.\n```\npython3 -m fastchat.serve.cli --model-path lmsys/vicuna-7b-v1.5\n```\n\n#### Multiple GPUs\nYou can use model parallelism to aggregate GPU memory from multiple GPUs on the same machine. \n```\npython3 -m fastchat.serve.cli --model-path lmsys/vicuna-7b-v1.5 --num-gpus 2\n```\n\nTips:\nSometimes the \"auto\" device mapping strategy in huggingface/transformers does not perfectly balance the memory allocation across multiple GPUs.\nYou can use `--max-gpu-memory` to specify the maximum memory per GPU for storing model weights.\nThis allows it to allocate more memory for activations, so you can use longer context lengths or larger batch sizes. For example,\n\n```\npython3 -m fastchat.serve.cli --model-path lmsys/vicuna-7b-v1.5 --num-gpus 2 --max-gpu-memory 8GiB\n```\n\n#### CPU Only\nThis runs on the CPU only and does not require GPU. It requires around 30GB of CPU memory for Vicuna-7B and around 60GB of CPU memory for Vicuna-13B.\n```\npython3 -m fastchat.serve.cli --model-path lmsys/vicuna-7b-v1.5 --device cpu\n```\n\nUse Intel AI Accelerator AVX512_BF16/AMX to accelerate CPU inference.\n```\nCPU_ISA=amx python3 -m fastchat.serve.cli --model-path lmsys/vicuna-7b-v1.5 --device cpu\n```\n\n#### Metal Backend (Mac Computers with Apple Silicon or AMD GPUs)\nUse `--device mps` to enable GPU acceleration on Mac computers (requires torch \u003e= 2.0).\nUse `--load-8bit` to turn on 8-bit compression.\n```\npython3 -m fastchat.serve.cli --model-path lmsys/vicuna-7b-v1.5 --device mps --load-8bit\n```\nVicuna-7B can run on a 32GB M1 Macbook with 1 - 2 words / second.\n\n#### Intel XPU (Intel Data Center and Arc A-Series GPUs)\nInstall the [Intel Extension for PyTorch](https://intel.github.io/intel-extension-for-pytorch/xpu/latest/tutorials/installation.html). Set the OneAPI environment variables:\n```\nsource /opt/intel/oneapi/setvars.sh\n```\n\nUse `--device xpu` to enable XPU/GPU acceleration.\n```\npython3 -m fastchat.serve.cli --model-path lmsys/vicuna-7b-v1.5 --device xpu\n```\nVicuna-7B can run on an Intel Arc A770 16GB.\n\n#### Ascend NPU\nInstall the [Ascend PyTorch Adapter](https://github.com/Ascend/pytorch). Set the CANN environment variables:\n```\nsource /usr/local/Ascend/ascend-toolkit/set_env.sh\n```\n\nUse `--device npu` to enable NPU acceleration.\n```\npython3 -m fastchat.serve.cli --model-path lmsys/vicuna-7b-v1.5 --device npu\n```\nVicuna-7B/13B can run on an Ascend NPU.\n\n#### Not Enough Memory\nIf you do not have enough memory, you can enable 8-bit compression by adding `--load-8bit` to commands above.\nThis can reduce memory usage by around half with slightly degraded model quality.\nIt is compatible with the CPU, GPU, and Metal backend.\n\nVicuna-13B with 8-bit compression can run on a single GPU with 16 GB of VRAM, like an Nvidia RTX 3090, RTX 4080, T4, V100 (16GB), or an AMD RX 6800 XT.\n\n```\npython3 -m fastchat.serve.cli --model-path lmsys/vicuna-7b-v1.5 --load-8bit\n```\n\nIn addition to that, you can add `--cpu-offloading` to commands above to offload weights that don't fit on your GPU onto the CPU memory.\nThis requires 8-bit compression to be enabled and the bitsandbytes package to be installed, which is only available on linux operating systems.\n\n#### More Platforms and Quantization\n- For AMD GPU users, please install ROCm and [the ROCm version of PyTorch](https://pytorch.org/get-started/locally/) before you install FastChat. See also this [post](https://github.com/lm-sys/FastChat/issues/104#issuecomment-1613791563).\n- FastChat supports ExLlama V2. See [docs/exllama_v2.md](/docs/exllama_v2.md).\n- FastChat supports GPTQ 4bit inference with [GPTQ-for-LLaMa](https://github.com/qwopqwop200/GPTQ-for-LLaMa). See [docs/gptq.md](/docs/gptq.md).\n- FastChat supports AWQ 4bit inference with [mit-han-lab/llm-awq](https://github.com/mit-han-lab/llm-awq). See [docs/awq.md](/docs/awq.md).\n- [MLC LLM](https://mlc.ai/mlc-llm/), backed by [TVM Unity](https://github.com/apache/tvm/tree/unity) compiler, deploys Vicuna natively on phones, consumer-class GPUs and web browsers via Vulkan, Metal, CUDA and WebGPU.\n\n#### Use models from modelscope\nFor Chinese users, you can use models from www.modelscope.cn via specify the following environment variables.\n```bash\nexport FASTCHAT_USE_MODELSCOPE=True\n```\n\n## Serving with Web GUI\n\n\u003ca href=\"https://lmarena.ai\"\u003e\u003cimg src=\"assets/screenshot_gui.png\" width=\"70%\"\u003e\u003c/a\u003e\n\nTo serve using the web UI, you need three main components: web servers that interface with users, model workers that host one or more models, and a controller to coordinate the webserver and model workers. You can learn more about the architecture [here](docs/server_arch.md).\n\nHere are the commands to follow in your terminal:\n\n#### Launch the controller\n```bash\npython3 -m fastchat.serve.controller\n```\n\nThis controller manages the distributed workers.\n\n#### Launch the model worker(s)\n```bash\npython3 -m fastchat.serve.model_worker --model-path lmsys/vicuna-7b-v1.5\n```\nWait until the process finishes loading the model and you see \"Uvicorn running on ...\". The model worker will register itself to the controller .\n\nTo ensure that your model worker is connected to your controller properly, send a test message using the following command:\n```bash\npython3 -m fastchat.serve.test_message --model-name vicuna-7b-v1.5\n```\nYou will see a short output.\n\n#### Launch the Gradio web server\n```bash\npython3 -m fastchat.serve.gradio_web_server\n```\n\nThis is the user interface that users will interact with.\n\nBy following these steps, you will be able to serve your models using the web UI. You can open your browser and chat with a model now.\nIf the models do not show up, try to reboot the gradio web server.\n\n## Launch Chatbot Arena (side-by-side battle UI)\n\nCurrently, Chatbot Arena is powered by FastChat. Here is how you can launch an instance of Chatbot Arena locally.\n\nFastChat supports popular API-based models such as OpenAI, Anthropic, Gemini, Mistral and more. To add a custom API, please refer to the model support [doc](./docs/model_support.md). Below we take OpenAI models as an example.\n\nCreate a JSON configuration file `api_endpoint.json` with the api endpoints of the models you want to serve, for example:\n```\n{\n    \"gpt-4o-2024-05-13\": {\n        \"model_name\": \"gpt-4o-2024-05-13\",\n        \"api_base\": \"https://api.openai.com/v1\",\n        \"api_type\": \"openai\",\n        \"api_key\": [Insert API Key],\n        \"anony_only\": false\n    }\n}\n```\nFor Anthropic models, specify `\"api_type\": \"anthropic_message\"` with your Anthropic key. Similarly, for gemini model, specify `\"api_type\": \"gemini\"`. More details can be found in [api_provider.py](https://github.com/lm-sys/FastChat/blob/main/fastchat/serve/api_provider.py).\n\nTo serve your own model using local gpus, follow the instructions in [Serving with Web GUI](#serving-with-web-gui).\n\nNow you're ready to launch the server:\n```\npython3 -m fastchat.serve.gradio_web_server_multi --register-api-endpoint-file api_endpoint.json\n```\n\n#### (Optional): Advanced Features, Scalability, Third Party UI\n- You can register multiple model workers to a single controller, which can be used for serving a single model with higher throughput or serving multiple models at the same time. When doing so, please allocate different GPUs and ports for different model workers.\n```\n# worker 0\nCUDA_VISIBLE_DEVICES=0 python3 -m fastchat.serve.model_worker --model-path lmsys/vicuna-7b-v1.5 --controller http://localhost:21001 --port 31000 --worker http://localhost:31000\n# worker 1\nCUDA_VISIBLE_DEVICES=1 python3 -m fastchat.serve.model_worker --model-path lmsys/fastchat-t5-3b-v1.0 --controller http://localhost:21001 --port 31001 --worker http://localhost:31001\n```\n- You can also launch a multi-tab gradio server, which includes the Chatbot Arena tabs.\n```bash\npython3 -m fastchat.serve.gradio_web_server_multi\n```\n- The default model worker based on huggingface/transformers has great compatibility but can be slow. If you want high-throughput batched serving, you can try [vLLM integration](docs/vllm_integration.md).\n- If you want to host it on your own UI or third party UI, see [Third Party UI](docs/third_party_ui.md).\n\n## API\n### OpenAI-Compatible RESTful APIs \u0026 SDK\nFastChat provides OpenAI-compatible APIs for its supported models, so you can use FastChat as a local drop-in replacement for OpenAI APIs.\nThe FastChat server is compatible with both [openai-python](https://github.com/openai/openai-python) library and cURL commands.\nThe REST API is capable of being executed from Google Colab free tier, as demonstrated in the [FastChat_API_GoogleColab.ipynb](https://github.com/lm-sys/FastChat/blob/main/playground/FastChat_API_GoogleColab.ipynb) notebook, available in our repository.\nSee [docs/openai_api.md](docs/openai_api.md).\n\n### Hugging Face Generation APIs\nSee [fastchat/serve/huggingface_api.py](fastchat/serve/huggingface_api.py).\n\n### LangChain Integration\nSee [docs/langchain_integration](docs/langchain_integration.md).\n\n## Evaluation\nWe use MT-bench, a set of challenging multi-turn open-ended questions to evaluate models. \nTo automate the evaluation process, we prompt strong LLMs like GPT-4 to act as judges and assess the quality of the models' responses.\nSee instructions for running MT-bench at [fastchat/llm_judge](fastchat/llm_judge).\n\nMT-bench is the new recommended way to benchmark your models. If you are still looking for the old 80 questions used in the vicuna blog post, please go to [vicuna-blog-eval](https://github.com/lm-sys/vicuna-blog-eval).\n\n## Fine-tuning\n### Data\n\nVicuna is created by fine-tuning a Llama base model using approximately 125K user-shared conversations gathered from ShareGPT.com with public APIs. To ensure data quality, we convert the HTML back to markdown and filter out some inappropriate or low-quality samples. Additionally, we divide lengthy conversations into smaller segments that fit the model's maximum context length. For detailed instructions to clean the ShareGPT data, check out [here](docs/commands/data_cleaning.md).\n\nWe will not release the ShareGPT dataset. If you would like to try the fine-tuning code, you can run it with some dummy conversations in [dummy_conversation.json](data/dummy_conversation.json). You can follow the same format and plug in your own data.\n\n### Code and Hyperparameters\nOur code is based on [Stanford Alpaca](https://github.com/tatsu-lab/stanford_alpaca) with additional support for multi-turn conversations.\nWe use similar hyperparameters as the Stanford Alpaca.\n\n| Hyperparameter | Global Batch Size | Learning rate | Epochs | Max length | Weight decay |\n| --- | ---: | ---: | ---: | ---: | ---: |\n| Vicuna-13B | 128 | 2e-5 | 3 | 2048 | 0 |\n\n### Fine-tuning Vicuna-7B with Local GPUs\n\n- Install dependency\n```bash\npip3 install -e \".[train]\"\n```\n\n- You can use the following command to train Vicuna-7B with 4 x A100 (40GB). Update `--model_name_or_path` with the actual path to Llama weights and `--data_path` with the actual path to data.\n```bash\ntorchrun --nproc_per_node=4 --master_port=20001 fastchat/train/train_mem.py \\\n    --model_name_or_path meta-llama/Llama-2-7b-hf \\\n    --data_path data/dummy_conversation.json \\\n    --bf16 True \\\n    --output_dir output_vicuna \\\n    --num_train_epochs 3 \\\n    --per_device_train_batch_size 2 \\\n    --per_device_eval_batch_size 2 \\\n    --gradient_accumulation_steps 16 \\\n    --evaluation_strategy \"no\" \\\n    --save_strategy \"steps\" \\\n    --save_steps 1200 \\\n    --save_total_limit 10 \\\n    --learning_rate 2e-5 \\\n    --weight_decay 0. \\\n    --warmup_ratio 0.03 \\\n    --lr_scheduler_type \"cosine\" \\\n    --logging_steps 1 \\\n    --fsdp \"full_shard auto_wrap\" \\\n    --fsdp_transformer_layer_cls_to_wrap 'LlamaDecoderLayer' \\\n    --tf32 True \\\n    --model_max_length 2048 \\\n    --gradient_checkpointing True \\\n    --lazy_preprocess True\n```\n\nTips:\n- If you are using V100 which is not supported by FlashAttention, you can use the [memory-efficient attention](https://arxiv.org/abs/2112.05682) implemented in [xFormers](https://github.com/facebookresearch/xformers). Install xformers and replace `fastchat/train/train_mem.py` above with [fastchat/train/train_xformers.py](fastchat/train/train_xformers.py).\n- If you meet out-of-memory due to \"FSDP Warning: When using FSDP, it is efficient and recommended... \", see solutions [here](https://github.com/huggingface/transformers/issues/24724#issuecomment-1645189539).\n- If you meet out-of-memory during model saving, see solutions [here](https://github.com/pytorch/pytorch/issues/98823).\n- To turn on logging to popular experiment tracking tools such as Tensorboard, MLFlow or Weights \u0026 Biases, use the `report_to` argument, e.g. pass `--report_to wandb` to turn on logging to Weights \u0026 Biases.\n\n### Other models, platforms and LoRA support\nMore instructions to train other models (e.g., FastChat-T5) and use LoRA are in [docs/training.md](docs/training.md).\n\n### Fine-tuning on Any Cloud with SkyPilot\n[SkyPilot](https://github.com/skypilot-org/skypilot) is a framework built by UC Berkeley for easily and cost effectively running ML workloads on any cloud (AWS, GCP, Azure, Lambda, etc.).\nFind SkyPilot documentation [here](https://github.com/skypilot-org/skypilot/tree/master/llm/vicuna) on using managed spot instances to train Vicuna and save on your cloud costs.\n\n## Citation\nThe code (training, serving, and evaluation) in this repository is mostly developed for or derived from the paper below.\nPlease cite it if you find the repository helpful.\n\n```\n@misc{zheng2023judging,\n      title={Judging LLM-as-a-judge with MT-Bench and Chatbot Arena},\n      author={Lianmin Zheng and Wei-Lin Chiang and Ying Sheng and Siyuan Zhuang and Zhanghao Wu and Yonghao Zhuang and Zi Lin and Zhuohan Li and Dacheng Li and Eric. P Xing and Hao Zhang and Joseph E. Gonzalez and Ion Stoica},\n      year={2023},\n      eprint={2306.05685},\n      archivePrefix={arXiv},\n      primaryClass={cs.CL}\n}\n```\n\nWe are also planning to add more of our research to this repository.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Flm-sys%2FFastChat","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Flm-sys%2FFastChat","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Flm-sys%2FFastChat/lists"}