{"id":33176722,"url":"https://github.com/h2oai/h2o-LLM-eval","last_synced_at":"2026-02-03T18:00:51.076Z","repository":{"id":188767140,"uuid":"670715717","full_name":"h2oai/h2o-LLM-eval","owner":"h2oai","description":"Large-language Model Evaluation framework with Elo Leaderboard and A-B testing","archived":false,"fork":false,"pushed_at":"2024-10-24T17:16:05.000Z","size":8867,"stargazers_count":52,"open_issues_count":5,"forks_count":1,"subscribers_count":37,"default_branch":"main","last_synced_at":"2025-08-21T15:58:58.726Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":"https://evalgpt.ai/","language":"Jupyter Notebook","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/h2oai.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2023-07-25T17:05:36.000Z","updated_at":"2025-04-04T04:12:28.000Z","dependencies_parsed_at":null,"dependency_job_id":"764f9983-968d-44bf-8875-7643d3ea6190","html_url":"https://github.com/h2oai/h2o-LLM-eval","commit_stats":null,"previous_names":["h2oai/h2o-llm-eval"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/h2oai/h2o-LLM-eval","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/h2oai%2Fh2o-LLM-eval","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/h2oai%2Fh2o-LLM-eval/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/h2oai%2Fh2o-LLM-eval/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/h2oai%2Fh2o-LLM-eval/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/h2oai","download_url":"https://codeload.github.com/h2oai/h2o-LLM-eval/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/h2oai%2Fh2o-LLM-eval/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":29051269,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-02-03T15:43:47.601Z","status":"ssl_error","status_checked_at":"2026-02-03T15:43:46.709Z","response_time":96,"last_error":"SSL_read: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2025-11-16T03:00:21.739Z","updated_at":"2026-02-03T18:00:51.070Z","avatar_url":"https://github.com/h2oai.png","language":"Jupyter Notebook","funding_links":[],"categories":[],"sub_categories":[],"readme":"# H2O Large Language Model (LLM) Evaluation\n\nIn an era where Large Language Models (LLMs) are rapidly gaining traction for diverse applications, the need for comprehensive evaluation and comparison of these models has never been more critical.\nThis repository is an effort in that direction, providing an evaluation method and the toolkit for the assessment of Large Language Models.\n\nPlease read the [Blog Post](https://h2o.ai/blog/h2o-llm-evalgpt-a-comprehensive-tool-for-evaluating-large-language-models/) for more context.\n\n- [EvalGPT.ai](#evalgptai)\n    - [Elo Leaderboard](#elo-leaderboard)\n    - [Prompts](#prompts)\n    - [Responses](#responses)\n    - [A/B Tests](#ab-tests)\n- [Docker Compose Setup](#docker-compose-setup)\n- [Local Setup](#local-setup)\n- [Reproducing Leaderboard](#reproducing-leaderboard-results)\n- [Roadmap](#roadmap)\n\n\n## EvalGPT.ai\n\n[evalgpt.ai](https://evalgpt.ai/) hosts the Leaderboard of some of the top LLMs ranked by their Elo scores. The leaderboard is updated frequently and provides a comprehensive and fair assessment of Large Language Models. Different features of the website are described below.\n\n### Elo Leaderboard\n\nThe Elo Leaderboard provides a ranking of the top LLMs based on their Elo scores. The Elo scores are computed from the results of A/B tests, wherein the LLMs are pitted against each other in a series of games. The ranking system employed is based on the [Elo Rating System](https://en.wikipedia.org/wiki/Elo_rating_system). The procedure for Elo score computation closely follows the methodology outlined at [this resource](https://lmsys.org/blog/2023-05-25-leaderboard/).\n\n![Elo Leaderboard](docs/images/leaderboard.png)\n\n### Prompts\n\nPrompts tab has the list of 60 prompts used to evaluate the LLMs. The prompts are categorized into different categories based on the type of task they are designed for.\n\n![Prompts](./docs/images/testset.png)\n\n### Responses\n\nIn the Responses section, you can see the responses generated by the LLMs for the prompts. You can also select the LLMs and prompts to compare the responses.\n\n![Responses](./docs/images/responses.png)\n\nClick on the \"Select Models\" button to select the LLMs to compare. You can also select a different prompt using the \"Previous\" and \"Next\" buttons.\n\n![select models and prompts](./docs/images/evalgpt_responses_toolbar.png)\n\nFor any two selected models and the prompt, you can see the evaluation by GPT4 by clicking on the \"Show GPT Eval\" button on the top right.\n\n![show gpt eval button](./docs/images/evalgpt_gpt_eval_button.png)\n\n![show eval gpt](./docs/images/gpt_eval.png)\n\n### A/B Tests\n\n\"Which is Better: A or B?\" provides the interface to perform human evaluation of the LLMs. Each A/B test consists of a prompt and two responses generated by two different LLMs. The user is asked to select the better response among the two.\n\n![A/B Tests](./docs/images/abtests.png)\n\n## Docker Compose Setup\n\n### 1. Clone the repository\n\n```bash\ngit clone https://github.com/h2oai/h2o-LLM-eval.git\ncd h2o-LLM-eval\n```\n\n### 2. Run Docker Compose\n\n```bash\ndocker compose up -d\n```\n\nNavigate to http://localhost:10101/ in your browser\n\n## Local Setup\n\n### 1. Clone the repository\n\n```bash\ngit clone https://github.com/h2oai/h2o-LLM-eval.git\n```\n\n### 2. Setup Database\n\n#### a. Create a docker volume for the database\n\n```bash\ndocker volume create llm-eval-db-data\n```\n\n#### b. Start PostgreSQL 14 in docker\n\n```bash\ndocker run -d --name=llm-eval-db -p 5432:5432 -v llm-eval-db-data:/var/lib/postgresql/data -e POSTGRES_PASSWORD=pgpassword postgres:14\n```\n\n#### c. Install PostgreSQL client\n\n- On Ubuntu:\n\n```bash\nsudo apt update\nsudo apt install postgresql-client\n```\n\n- On macOS:\n\n```bash\nbrew install libpq\necho 'export PATH=\"/usr/local/opt/libpq/bin:$PATH\"' \u003e\u003e ~/.zshrc\n```\n\n#### d. Load the latest data dump into the database\n\n```bash\nPGPASSWORD=pgpassword psql --host=localhost --port=5432 --username=postgres \u003c data/10_init.sql\n```\n\n### 3. Setup the environment\n\nThe setup is tested on Python 3.10\n\n```bash\npython -m venv .venv\n```\n\n```bash\n. .venv/bin/activate\n```\n\n```bash\npip install --upgrade pip\npip install -r requirements.txt\n```\n\n### 4. Run the App\n\n```bash\nPOSTGRES_HOST=localhost POSTGRES_USER=maker POSTGRES_PASSWORD=makerpassword POSTGRES_DB=llm_eval_db H2O_WAVE_NO_LOG=true wave run llm_eval/app.py\n```\n\nNavigate to http://localhost:10101/ in your browser\n\n## Reproducing Leaderboard Results\n\nWe provide [notebooks](notebooks) to generate leaderboard results and reproduce [evalgpt.ai](https://evalgpt.ai).\n\n1. Run [run_all_evaluations.ipynb](notebooks/run_all_evaluations.ipynb) to evaluate any A/B tests that have not yet been evaluated by a chosen evaluation model and insert the outcomes into the database. An A/B test is considered unevaluated by the given model if no evaluation by the model exists for the given combination of models and prompt. After adding a model, running this evaluates all A/B tests for the model against all other models.\n\n2. Run all cells in [calculate_elo_rating_public_leaderboard.ipynb](notebooks/calculate_elo_rating_public_leaderboard.ipynb) to get the Elo leaderboard and relevant charts given the evaluations in the database.\n\n## Roadmap\n\n### Models\n\n1. Add [FreeWilly2](https://stability.ai/blog/freewilly-large-instruction-fine-tuned-models) to the Leaderboard\n\n### Application\n\n1. v2 architecture\n2. Option for users to submit new models\n\n### Eval\n\n1. More prompts in each category\n2. Document Q/A and Retrieval Category with ground truth\n3. Document Summarization Category\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fh2oai%2Fh2o-LLM-eval","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fh2oai%2Fh2o-LLM-eval","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fh2oai%2Fh2o-LLM-eval/lists"}