{"id":14835715,"url":"https://github.com/lmarena/arena-hard-auto","last_synced_at":"2025-12-29T22:22:56.639Z","repository":{"id":230775598,"uuid":"714151130","full_name":"lmarena/arena-hard-auto","owner":"lmarena","description":"Arena-Hard-Auto: An automatic LLM benchmark. ","archived":false,"fork":false,"pushed_at":"2024-12-29T23:11:59.000Z","size":4851,"stargazers_count":699,"open_issues_count":5,"forks_count":84,"subscribers_count":8,"default_branch":"main","last_synced_at":"2025-01-09T01:24:45.569Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":"","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/lmarena.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2023-11-04T03:54:08.000Z","updated_at":"2025-01-08T23:40:07.000Z","dependencies_parsed_at":"2024-04-15T01:45:45.838Z","dependency_job_id":"91e3c356-6f11-4371-b62b-f8135ebd0236","html_url":"https://github.com/lmarena/arena-hard-auto","commit_stats":null,"previous_names":["lm-sys/arena-hard","lm-sys/arena-hard-auto","lmarena/arena-hard-auto"],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/lmarena%2Farena-hard-auto","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/lmarena%2Farena-hard-auto/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/lmarena%2Farena-hard-auto/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/lmarena%2Farena-hard-auto/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/lmarena","download_url":"https://codeload.github.com/lmarena/arena-hard-auto/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":233410962,"owners_count":18672293,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-09-19T05:01:03.831Z","updated_at":"2025-12-29T22:22:56.618Z","avatar_url":"https://github.com/lmarena.png","language":"Python","funding_links":[],"categories":["A01_文本生成_文本对话","Python","Benchmark"],"sub_categories":["大语言对话模型及数据","Open ended"],"readme":"\u003cdiv align=\"center\"\u003e\n\n# Arena-Hard-Auto\n\n[![Github](https://img.shields.io/badge/Arena--Hard-black?logo=github\u0026logoColor=white\u0026labelColor=black\u0026color=black)](https://github.com/lmarena/arena-hard-auto) [![arXiv](https://img.shields.io/badge/arXiv-Arena--Hard-b31b1b.svg)](https://arxiv.org/abs/2406.11939) [![Hugging Face Collection](https://img.shields.io/badge/Arena--Hard-fcd022?logo=huggingface\u0026logoColor=000\u0026labelColor)](https://huggingface.co/collections/lmarena-ai/arena-hard-auto-680998796296d1462c729b6c) [![Twitter](https://img.shields.io/badge/LMArena--ai-white?logo=X\u0026logoColor=000\u0026color=000\u0026labelColor=white)](https://x.com/lmarena_ai)\n\n\n\u003cdiv align=\"center\" style=\"font-family: Arial, sans-serif;\"\u003e\n  \u003cp\u003e\n    \u003ca href=\"#news\" style=\"text-decoration: none; font-weight: bold;\"\u003eNews\u003c/a\u003e •\n    \u003ca href=\"#leaderboard\" style=\"text-decoration: none; font-weight: bold;\"\u003eLeaderboard\u003c/a\u003e •\n    \u003ca href=\"#install-dependencies\" style=\"text-decoration: none; font-weight: bold;\"\u003eInstall\u003c/a\u003e •\n    \u003ca href=\"#evaluation\" style=\"text-decoration: none; font-weight: bold;\"\u003eEvaluation\u003c/a\u003e •\n    \u003ca href=\"https://huggingface.co/spaces/lmarena-ai/arena-hard-viewer\" style=\"text-decoration: none; font-weight: bold;\"\u003eDemo\u003c/a\u003e •\n    \u003ca href=\"#citation\" style=\"text-decoration: none; font-weight: bold;\"\u003eCitation\u003c/a\u003e\n  \u003c/p\u003e\n\u003c/div\u003e\n\n\u003c/div\u003e\n\n# News\n- **[Apr 23, 2025]** 🎉 **Arena-Hard-v2.0** is finally here! Better judges, new hard prompts, and additional eval for creative writing.\n- **[Oct 14, 2024]** 🎉 **Style Control** is now supported in Arena-Hard-Auto.\n\n## About\n\nArena-Hard-Auto is an automatic evaluation tool for instruction-tuned LLMs. Arena-Hard-Auto has the highest correlation and separability to LMArena (Chatbot Arena) among popular open-ended LLM benchmarks ([See Paper](https://arxiv.org/abs/2406.11939)). If you are curious to see how well your model might perform on LMArena before deploying, we recommend trying Arena-Hard-Auto's newest evaluation set, **Arena-Hard-v2.0-Preview**.\n\nV2.0 contains 500 fresh, challenging real-world user queries (open-ended software engineering problems, math questions, etc) and 250 creative writing queries sourced from Chatbot Arena. We employs automatic judges, GPT-4.1 and Gemini-2.5, as a cheaper and faster approximator to human preference.\n\nAlthough both Arena-Hard-Auto and Chatbot Arena Category Hard ([See Blog](https://lmsys.org/blog/2024-05-17-category-hard/)) employ similar pipeline to select hard prompts, Arena-Hard-Auto employs automatic judge as a cheaper and faster approximator to human preference. Checkout [BenchBuilder](BenchBuilder) folder for code and resources on how we curate Arena-Hard-Auto. In the paper we also purposed metrics, such as model separability and agreement to human preference, for evaluating benchmarks' ability to rank models (See [Evaluate Benchmarks](#evaluate-benchmarks) for more information and code).\n\n\n## Leaderboard\n\n### Arena-Hard-v2.0-Preview\n\nHard Prompt, Style Control, and Gemini-2.5 as Judge **(Official Configuration)**:\n```console\n                                      Model  Scores (%)         CI (%)\n0                             o3-2025-04-16        85.9  (-0.8 / +0.9)\n1                   o4-mini-2025-04-16-high        79.1  (-1.4 / +1.2)\n2                                gemini-2.5        79.0  (-2.1 / +1.8)\n3                        o4-mini-2025-04-16        74.6  (-1.8 / +1.6)\n4                          gemini-2.5-flash        68.6  (-1.6 / +1.6)\n5                   o3-mini-2025-01-31-high        66.1  (-1.5 / +2.1)\n6                        o1-2024-12-17-high        61.0  (-2.0 / +2.1)\n7   claude-3-7-sonnet-20250219-thinking-16k        59.8  (-2.0 / +1.8)\n8                           Qwen3-235B-A22B        58.4  (-1.9 / +2.1)\n9                               deepseek-r1        58.0  (-2.2 / +2.0)\n10                            o1-2024-12-17        55.9  (-2.2 / +1.8)\n11                          gpt-4.5-preview        50.0  (-1.9 / +2.0)\n12                       o3-mini-2025-01-31        50.0  (-0.0 / +0.0)\n13                                  gpt-4.1        50.0  (-1.9 / +1.7)\n14                             gpt-4.1-mini        46.9  (-2.4 / +2.1)\n15                                Qwen3-32B        44.5  (-2.2 / +2.1)\n16                                  QwQ-32B        43.5  (-2.5 / +2.1)\n17                            Qwen3-30B-A3B        33.9  (-1.6 / +1.5)\n18               claude-3-5-sonnet-20241022        33.0  (-2.3 / +1.8)\n19                                 s1.1-32B        22.3  (-1.7 / +1.5)\n20           llama4-maverick-instruct-basic        17.2  (-1.5 / +1.2)\n21                           Athene-V2-Chat        16.4  (-1.4 / +1.4)\n22                           gemma-3-27b-it        15.0  (-1.4 / +1.0)\n23                                 Qwen3-4B        15.0  (-1.1 / +1.5)\n24                             gpt-4.1-nano        13.7  (-1.1 / +1.0)\n25       Llama-3.1-Nemotron-70B-Instruct-HF        10.3  (-0.8 / +1.0)\n26                     Qwen2.5-72B-Instruct        10.1  (-0.9 / +1.3)\n27                         OpenThinker2-32B         3.2  (-0.3 / +0.3)\n```\n\nHard Prompt, Style Control, and GPT-4.1 as Judge **(If prefer OpenAI API)**\n```console\n                                      Model  Scores (%)         CI (%)\n0                             o3-2025-04-16        87.0  (-1.0 / +1.0)\n1                   o4-mini-2025-04-16-high        81.7  (-1.2 / +1.2)\n2                        o4-mini-2025-04-16        78.0  (-1.3 / +1.4)\n3                   o3-mini-2025-01-31-high        64.8  (-2.1 / +1.9)\n4                        o1-2024-12-17-high        58.7  (-2.3 / +2.1)\n5                                   gpt-4.1        58.3  (-2.0 / +2.3)\n6                             o1-2024-12-17        50.2  (-2.2 / +1.8)\n7                        o3-mini-2025-01-31        50.0  (-0.0 / +0.0)\n8                                gemini-2.5        49.1  (-2.5 / +2.4)\n9                              gpt-4.1-mini        48.6  (-2.7 / +1.9)\n10                              deepseek-r1        48.0  (-2.6 / +2.3)\n11  claude-3-7-sonnet-20250219-thinking-16k        47.0  (-1.9 / +2.3)\n12                          Qwen3-235B-A22B        46.7  (-1.9 / +2.4)\n13                         gemini-2.5-flash        45.1  (-2.7 / +2.1)\n14                          gpt-4.5-preview        43.0  (-1.9 / +2.2)\n15                                  QwQ-32B        36.1  (-2.0 / +2.2)\n16                                Qwen3-32B        35.8  (-2.1 / +2.2)\n17                            Qwen3-30B-A3B        28.7  (-1.4 / +2.1)\n18               claude-3-5-sonnet-20241022        25.8  (-1.7 / +1.8)\n19                                 s1.1-32B        18.3  (-2.3 / +2.2)\n20                             gpt-4.1-nano        15.4  (-1.1 / +1.2)\n21                           Athene-V2-Chat        12.6  (-1.2 / +1.3)\n22                                 Qwen3-4B        12.6  (-1.1 / +1.5)\n23           llama4-maverick-instruct-basic        12.0  (-1.0 / +1.2)\n24                           gemma-3-27b-it         9.7  (-0.9 / +1.1)\n25                     Qwen2.5-72B-Instruct         8.0  (-0.7 / +0.9)\n26       Llama-3.1-Nemotron-70B-Instruct-HF         6.8  (-0.6 / +0.8)\n27                         OpenThinker2-32B         2.3  (-0.2 / +0.3)\n```\n\nCreative Writing, Ensemble GPT-4.1 and Gemini 2.5 as Judges **(Best Configuration for Creative Writing)**\n```console\n                                      Model  Scores (%)         CI (%)\n0                                gemini-2.5        90.8  (-1.2 / +1.3)\n1                             o3-2025-04-16        88.8  (-1.1 / +1.0)\n2                          gemini-2.5-flash        83.9  (-1.3 / +1.4)\n3                               deepseek-r1        77.0  (-2.0 / +1.4)\n4                           Qwen3-235B-A22B        73.5  (-1.8 / +1.5)\n5                            gemma-3-27b-it        69.9  (-1.9 / +1.7)\n6   claude-3-7-sonnet-20250219-thinking-16k        63.9  (-1.7 / +1.9)\n7                                   gpt-4.1        61.5  (-1.9 / +1.9)\n8                                   QwQ-32B        60.9  (-2.0 / +1.6)\n9                        o1-2024-12-17-high        59.9  (-2.1 / +1.7)\n10                  o4-mini-2025-04-16-high        58.7  (-1.8 / +1.9)\n11                            o1-2024-12-17        56.6  (-1.8 / +1.8)\n12                       o4-mini-2025-04-16        55.6  (-1.8 / +2.0)\n13                                Qwen3-32B        53.3  (-1.9 / +1.6)\n14                          gpt-4.5-preview        51.4  (-1.9 / +2.0)\n15                     gemini-2.0-flash-001        50.0  (-0.0 / +0.0)\n16                  o3-mini-2025-01-31-high        43.0  (-1.7 / +2.1)\n17                            Qwen3-30B-A3B        34.9  (-2.0 / +1.6)\n18                             gpt-4.1-mini        28.2  (-1.8 / +1.8)\n19       Llama-3.1-Nemotron-70B-Instruct-HF        26.9  (-2.0 / +1.8)\n20               claude-3-5-sonnet-20241022        24.2  (-1.5 / +1.5)\n21                         OpenThinker2-32B        23.6  (-1.5 / +1.3)\n22                           Athene-V2-Chat        18.1  (-1.6 / +1.5)\n23                                 Qwen3-4B        13.2  (-1.2 / +1.2)\n24                             gpt-4.1-nano        10.7  (-1.1 / +1.1)\n25           llama4-maverick-instruct-basic        10.5  (-1.1 / +1.0)\n26                     Qwen2.5-72B-Instruct        10.2  (-1.1 / +1.1)\n27                                 s1.1-32B         8.2  (-0.9 / +0.\n```\n\nFor older leaderboards, such as Arena-Hard-v0.1, see [past-leaderboards](/misc/past_leaderboards.md)\n\n## Install Dependencies\n```\ngit clone https://github.com/lmarena/arena-hard-auto.git\ncd arena-hard\npip install -r requirements.txt\npip install -r requirements-optional.txt  # Optional dependencies (e.g., anthropic sdk)\n```\n\n## Download dataset\nWe have pre-generated many popular models answers and judgments. You can browse them with an online [demo](https://huggingface.co/spaces/lmarena-ai/arena-hard-viewer) or download them (with [`git-lfs`](https://git-lfs.com) installed) by\n```console\n\u003e git lfs install\n\u003e git clone git@hf.co:datasets/lmarena-ai/arena-hard-auto arena-hard-data\n// copy answers/judgments to the data directory\n\u003e cp -r arena-hard-data/data . \n```\n\nThen run\n```console\n\u003e python show_result.py\n                                      Model  Scores (%)         CI (%)\n0                             o3-2025-04-16        87.6  (-0.8 / +1.0)\n1                   o4-mini-2025-04-16-high        82.7  (-1.4 / +1.3)\n2                        o4-mini-2025-04-16        78.9  (-1.6 / +1.6)\n```\n\n## Evaluate\n\n### Step 1. Set up the endpoint config to your model\n\nFill in your API endpoint in `config/api_config.yaml`. We support OpenAI compatible API server, Anthropic, Vertex AI, and more. You will find examples in `config/api_config.yaml`.\n\nYou may use inference engine such as [vLLM](https://docs.vllm.ai/en/latest/serving/openai_compatible_server.html) or [SGLang](https://github.com/sgl-project/sglang?tab=readme-ov-file#using-local-models) to host your model with an OpenAI compatible API server.\n\nWe also include support for fast built-in inference with SGLang, see examples in `config/api_config.yaml` and implementaton in `utils/completion.py`. See `misc/sglang_setup.bash` for environment setup.\n\n### Step 2. Generate Model Answers\n\nIn `config/gen_answer_config.yaml`, add your model name in `model_list`.\n\nRun the command to generate answers:\n```console\n\u003e python gen_answer.py\n```\n\nCaching feature is implemented. The code will skip generating an answer when there is already an existing answer/judgment to the same prompt (this feature is not supported for built-in SGLang server).\n\n### Step 3. Generate Judgments\n\nIn `config/arena-hard-v2.0.yaml`, add your model name in `model_list`.\n```yaml\n...\n# Add your model below for evaluation\nmodel_list:\n  - deepseek-r1\n  - [YOUR-MODEL-NAME]\n```\n\nWe recommend employing GPT-4.1 as judge for fast, stable judge inference. To use Gemini-2.5, comment out:\n```yaml\njudge_model: gpt-4.1\ntemperature: 0.0\nmax_tokens: 16000\n```\n\nand uncomment:\n```yaml\njudge_model: gemini-2.5\ntemperature: 1.0\nmax_tokens: 32000\n```\n\nRun the command to generate judgments:\n```console\n\u003e python gen_judgment.py\n```\n\nFor Ensemble-as-Judges, we suggest inferencing both judges independently and we will aggregrate the results when displaying leaderboard for you (see step 4).\n\nJudgment caching is also implemented. It will skip generating judgments that has already been generated or lacks one of the model answers.  \n\n### Step 4. Show result\nOutput model win rates for **Arena-Hard-v2.0-Preview (Hard Prompt, Style Control, GPT-4.1 as Judge)**:\n```console\n\u003e python show_result.py --judge-names gpt-4.1 --control-features markdown length\n```\n\nOutput model win rates for **Arena-Hard-v2.0-Preview (Creative Writing, Ensemble GPT-4.1 and Gemini 2.5 as Judges)**:\n```console\n\u003e python show_result.py --judge-names gpt-4.1 gemini-2.5 --category creative_writing\n```\n\n### Step 5. Benchmark Viewer\nYou can review answers and judgment results using our gradio script (`gradio\u003e=5.25.2`).\n```console\n\u003e python qa_browser.py --share\n```\n\n## Style Control\nFollowing the newly introduced Style Control on Chatbot Arena, we release Style Control on Arena Hard Auto! We employ the same Style Control methods as proposed in the [blogpost](https://lmsys.org/blog/2024-08-28-style-control/). Please refer to the blogpost for methodology and technical background.\n\nBefore applying style control, make sure your model answers has proper style attribute generated. Either pull the latest data from [huggingface repo](https://huggingface.co/datasets/lmarena-ai/arena-hard-auto), or run the following script!\n\nTo add style attribute to your model answers, use `add_markdown_info.py`. The following command takes model answers from `--dir`, append style attributes (token length, number of headers, etc), and save the new answers in `--output-dir`.\n\n```console\n\u003e python add_markdown_info.py --dir data/arena-hard-v0.1/model_answer --output-dir data/arena-hard-v0.1/model_answer\n```\n\nTo control for style (token length and markdown elements), use `--control-features` or `-f` when running `show_result.py`.\n\n```console\n\u003e python show_result.py -f markdown length # style control\n\u003e python show_result.py -f markdown # control for markdown density only\n\u003e python show_result.py -f length # length control only\n```\n\n## Evaluate Benchmarks\nWe outline two key properties that the benchmark aiming to approximate human preference should possess to provide meaningful comparisons between models:\n1. Separability: the benchmark should separate models with high confidence.\n2. Alignment with Human Preference: the benchmark should agree with human preference.\n\nWhile previous works have focused on alignment, separability is also a crucial consideration when comparing models of similar quality (e.g., different checkpoints from the same training run). However, achieving high-confidence separability is challenging due to limitations in prompt design and inherent variances in LLM evaluations. Overly simplistic prompts fail to distinguish between models, while the randomness in human and LLM judgments leads to inconsistent predictions. As a result, it is often difficult to confidently determine if a model’s apparent performance reflects a genuine difference in capability or merely noisy observations, highlighting a need for methods to verify whether a benchmark can reliably separate similar models.\n\nStatistical measures like Pearson (Pearson, 1895) and Spearman Correlations (Spearman, 1961), commonly used in benchmarks such as AlpacaEval (Li et al., 2023) to measure correlation to human preference ranking, may fail to adequately address model separability and ranking instability. In addition, these measures only provide a coarse signal of ranking correlation without quantifying the magnitude of performance differences between model pairs. To address these shortcomings, we develop three novel metrics: **Separability with Confidence**, **Agreement with Confidence**, and **Pair Rank Brier Score**.\n\n**Separability with Confidence** quantifies the benchmark’s confidence by measuring its consistency in predicting the winner of a model pair across random seeds through bootstrapping. This is done by calculating the percentage of model pairs that have non-overlapping confidence intervals of their benchmark scores. A higher percentage indicates that the benchmark is more confident in distinguishing between the performance of different models, as the confidence intervals of their scores do not overlap.\n\nFor **Agreement with Confidence**, and **Pair Rank Brier Score**, please refer to section 3 of our [paper](https://arxiv.org/abs/2406.11939). The code for calculating these metrics can be found in this [colab notebook](https://colab.research.google.com/drive/1ar6XLWREN_dXEh404WNOxroFVUe_4njp). \n\n## Integration with Amazon Bedrock API\nWe have now added the capability to benchmark LLMs hosted on **Amazon Bedrock** with arena-hard. Specifically, we added Amazon Bedrock invoke API in `utils/completion.py` which will allow you to use different models hosted on Amazon Bedrock with Arena-Hard.\n\n\nCurrently we support the following models:\n1. Anthropic Models : Claude 3 Haiku, Claude 3 Sonnet, Claude 3.5 Sonnet, Claude 3 Opus, Claude 3.5 Sonnet v2, Claude 3.7 Sonnet\n2. Mistral Models: Mistral 7B Instruct, Mistral 8x7B Instruct, Mistral Large v1, Mistral Large v2, Mistral Small, Pixtral Large\n3. Meta Llama Models: LLaMA 3 8B Instruct, LLaMA 3 70B Instruct, LLaMA 3.1 8B Instruct, LLaMA 3.1 70B Instruct, LLaMA 3.1 405B Instruct\n   LLaMA 3.2 1B Instruct, LLaMA 3.2 3B Instruct, LLaMA 3.2 11B Instruct, LLaMA 3.2 90B Instruct, LLaMA 2 Chat 13B, LLaMA 2 Chat 70B\n4. Amazon Nova Models: Amazon Nova Lite, Amazon Nova Pro, Amazon Nova Micro, Amazon Nova Premier\n5. DeepSeek-R1\n\nTo **Add a new model hosted on Amazon Bedrock**, you need to update two files: `config/api_config.yaml` and `utils/completion.py`.\n\n\n### 1. Update `config/api_config.yaml`\n\nDefine a new entry for the model with the correct `model_id`, `api_type`, and generation parameters.\n\n**Example:**\n\n```yaml\naws_nova_light_v1:\n  model: aws_nova_light_v1\n  model_id: us.amazon.nova-lite-v1:0\n  endpoints: null\n  api_type: aws_nova\n  parallel: 8\n  max_tokens: 4096\n  temperature: 0.0\n```\n**Key Fields**\n1. model: Internal alias used for referencing this config.\n2. model_id: Bedrock-specific model identifier.\n3. api_type: The api_type should be registered through `utils\\completion.py`\n4. endpoints: Set to null for default Bedrock endpoint, or override with custom endpoint.\n5. parallel: Controls parallel inference calls (adjust for throughput).\n6. max_tokens: Maximum output tokens.\n7. temperature: Controls randomness of generation (0.0 for deterministic).\n\nFind more example in `config/api_config_bedrock_models.yaml`\nRefer to Amazon Bedrock documentation (https://docs.aws.amazon.com/bedrock/latest/userguide/models-supported.html) for model IDs and capabilities.\n\n### 2. Register a Model Handler in `utils/completion.py`\nCreate a new function decorated with `@register_api(\"\u003capi_type\u003e\")` to define how inputs are formatted, sent to Bedrock using boto3, and how the response is parsed.\n\nYou can use existing examples as templates:\n\n    \u003e @register_api(\"aws_llama\") handles LLaMA models\n    \u003e @register_api(\"aws_nova\") handles Nova models\n\nThese functions typically use helpers like `create_llama3_body()` or `create_nova_messages()` and send requests using the Bedrock `invoke_model` API.\n\n**Pay attention to**:\n\n    \u003e The api_type in `api_config.yaml` which must match the name used in the `@register_api(...)` decorator.\n    \u003e Input formatting (e.g., prompt structure, message lists)\n    \u003e Parameter mapping (temperature, max_tokens, model_id)\n    \u003e Response parsing (e.g., generation vs nested output.message.content)\n\nBy following this two-step process, users can easily extend support to any Bedrock-hosted model that follows a compatible invocation structure.\nFor examples, see existing handlers for Claude, LLaMA, and Amazon Nova in the repository.\n\n\n## Community Contribution\n\nFeel free to submit a PR or open up an issue!\n\nIf you want to add your model to the leaderboard, please email me the following:\n1. An OpenAI compatible endpoint to your model.\n2. An OpenAI API key for me to inference judgment.\n\nSorry for the inconvience! Since Arena-Hard-Auto is open data, we want to avoid people cheating on our leaderboard. If we find anything suspicious, we reserve the right to not add your model to our leaderboard.\n\n## Citation\nThe code in this repository is developed from the papers below. Please cite it if you find the repository helpful.\n```\n@article{li2024crowdsourced,\n  title={From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline},\n  author={Li, Tianle and Chiang, Wei-Lin and Frick, Evan and Dunlap, Lisa and Wu, Tianhao and Zhu, Banghua and Gonzalez, Joseph E and Stoica, Ion},\n  journal={arXiv preprint arXiv:2406.11939},\n  year={2024}\n}\n@misc{arenahard2024,\n    title = {From Live Data to High-Quality Benchmarks: The Arena-Hard Pipeline},\n    url = {https://lmsys.org/blog/2024-04-19-arena-hard/},\n    author = {Tianle Li*, Wei-Lin Chiang*, Evan Frick, Lisa Dunlap, Banghua Zhu, Joseph E. Gonzalez, Ion Stoica},\n    month = {April},\n    year = {2024}\n}\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Flmarena%2Farena-hard-auto","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Flmarena%2Farena-hard-auto","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Flmarena%2Farena-hard-auto/lists"}