{"id":26896792,"url":"https://github.com/AlmogBaku/pytest-evals","last_synced_at":"2025-04-01T04:02:01.468Z","repository":{"id":272349217,"uuid":"916294880","full_name":"AlmogBaku/pytest-evals","owner":"AlmogBaku","description":"A pytest plugin for running and analyzing LLM evaluation tests.","archived":false,"fork":false,"pushed_at":"2025-02-05T13:01:50.000Z","size":361,"stargazers_count":116,"open_issues_count":0,"forks_count":3,"subscribers_count":1,"default_branch":"master","last_synced_at":"2025-03-29T02:01:38.668Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":"Jupyter Notebook","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/AlmogBaku.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2025-01-13T20:26:29.000Z","updated_at":"2025-03-24T04:21:18.000Z","dependencies_parsed_at":"2025-01-13T21:31:43.648Z","dependency_job_id":"3b65d38a-e309-47fb-b5c5-4a34a7c357e9","html_url":"https://github.com/AlmogBaku/pytest-evals","commit_stats":null,"previous_names":["almogbaku/pytest-evals"],"tags_count":14,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/AlmogBaku%2Fpytest-evals","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/AlmogBaku%2Fpytest-evals/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/AlmogBaku%2Fpytest-evals/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/AlmogBaku%2Fpytest-evals/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/AlmogBaku","download_url":"https://codeload.github.com/AlmogBaku/pytest-evals/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":246580463,"owners_count":20800110,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2025-04-01T04:02:00.430Z","updated_at":"2025-04-01T04:02:01.452Z","avatar_url":"https://github.com/AlmogBaku.png","language":"Jupyter Notebook","funding_links":[],"categories":["Jupyter Notebook"],"sub_categories":[],"readme":"\u003cdiv id=\"top\"\u003e\u003c/div\u003e\n\n# `pytest-evals` 🚀\n\nTest your LLM outputs against examples - no more manual checking! A (minimalistic) pytest plugin that helps you to\nevaluate that your LLM is giving good answers.\n\n[![PyPI version](https://img.shields.io/pypi/v/pytest-evals.svg)](https://pypi.org/p/pytest-evals)\n[![License](https://img.shields.io/github/license/AlmogBaku/pytest-evals.svg)](https://github.com/AlmogBaku/pytest-evals/blob/main/LICENSE)\n[![Issues](https://img.shields.io/github/issues/AlmogBaku/pytest-evals.svg)](https://github.com/AlmogBaku/pytest-evals/issues)\n[![Stars](https://img.shields.io/github/stars/AlmogBaku/pytest-evals.svg)](https://github.com/AlmogBaku/pytest-evals/stargazers)\n\n# 🧐 Why pytest-evals?\n\nBuilding LLM applications is exciting, but how do you know they're actually working well? `pytest-evals` helps you:\n\n- 🎯 **Test \u0026 Evaluate:** Run your LLM prompt against many cases\n- 📈 **Track \u0026 Measure:** Collect metrics and analyze the overall performance\n- 🔄 **Integrate Easily:** Works with pytest, Jupyter notebooks, and CI/CD pipelines\n- ✨ **Scale Up:** Run tests in parallel with [`pytest-xdist`](https://pytest-xdist.readthedocs.io/) and\n  asynchronously with [`pytest-asyncio`](https://pytest-asyncio.readthedocs.io/).\n\n# 🚀 Getting Started\n\nTo get started, install `pytest-evals` and write your tests:\n\n```bash\npip install pytest-evals\n```\n\n#### ⚡️ Quick Example\n\nFor example, say you're building a support ticket classifier. You want to test cases like:\n\n| Input Text                                             | Expected Classification |\n|--------------------------------------------------------|-------------------------|\n| My login isn't working and I need to access my account | account_access          |\n| Can I get a refund for my last order?                  | billing                 |\n| How do I change my notification settings?              | settings                |\n\n`pytest-evals` helps you automatically test how your LLM perform against these cases, track accuracy, and ensure it\nkeeps working as expected over time.\n\n```python\n# Predict the LLM performance for each case\n@pytest.mark.eval(name=\"my_classifier\")\n@pytest.mark.parametrize(\"case\", TEST_DATA)\ndef test_classifier(case: dict, eval_bag, classifier):\n    # Run predictions and store results\n    eval_bag.prediction = classifier(case[\"Input Text\"])\n    eval_bag.expected = case[\"Expected Classification\"]\n    eval_bag.accuracy = eval_bag.prediction == eval_bag.expected\n\n\n# Now let's see how our app performing across all cases...\n@pytest.mark.eval_analysis(name=\"my_classifier\")\ndef test_analysis(eval_results):\n    accuracy = sum([result.accuracy for result in eval_results]) / len(eval_results)\n    print(f\"Accuracy: {accuracy:.2%}\")\n    assert accuracy \u003e= 0.7  # Ensure our performance is not degrading 🫢\n```\n\nThen, run your evaluation tests:\n\n```bash\n# Run test cases\npytest --run-eval\n\n# Analyze results\npytest --run-eval-analysis\n```\n\n## 😵‍💫 Why Another Eval Tool?\n\n**Evaluations are just tests.** No need for complex frameworks or DSLs. `pytest-evals` is minimalistic by design:\n\n- Use `pytest` - the tool you already know\n- Keep tests and evaluations together\n- Focus on logic, not infrastructure\n\nIt just collects your results and lets you analyze them as a whole. Nothing more, nothing less.\n\u003cp align=\"right\"\u003e(\u003ca href=\"#top\"\u003eback to top\u003c/a\u003e)\u003c/p\u003e\n\n# 📚 User Guide\n\nCheck out our detailed guides and examples:\n\n- [Basic evaluation](example/example_test.py)\n- [Basic of LLM as a judge evaluation](example/example_judge_test.py)\n- [Notebook example](example/example_notebook.ipynb)\n- [Advanced notebook example](example/example_notebook_advanced.ipynb)\n\n## 🤔 How It Works\n\nBuilt on top of [pytest-harvest](https://smarie.github.io/python-pytest-harvest/), `pytest-evals` splits evaluation into\ntwo phases:\n\n1. **Evaluation Phase**: Run all test cases, collecting results and metrics in `eval_bag`. The results are saved in a\n   temporary file to allow the analysis phase to access them.\n2. **Analysis Phase**: Process all results at once through `eval_results` to calculate final metrics\n\nThis split allows you to:\n\n- Run evaluations in parallel (since the analysis test MUST run after all cases are done, we must run them separately)\n- Make pass/fail decisions on the overall evaluation results instead of individual test failures (by passing the\n  `--supress-failed-exit-code --run-eval` flags)\n- Collect comprehensive metrics\n\n**Note**: When running evaluation tests, the rest of your test suite will not run. This is by design to keep the results\nclean and focused.\n\n## 💾 Saving case results\nBy default, `pytest-evals` saves the results of each case in a json file to allow the analysis phase to access them.\nHowever, this might not be a friendly format for deeper analysis. To save the results in a more friendly format, as a\nCSV file, use the `--save-evals-csv` flag:\n\n```bash\npytest --run-eval --save-evals-csv\n```\n\n## 📝 Working with a notebook\n\nIt's also possible to run evaluations from a notebook. To do that, simply\ninstall [ipytest](https://github.com/chmp/ipytest), and load the extension:\n\n```python\n%load_ext pytest_evals\n```\n\nThen, use the magic commands `%%ipytest_eval` in your cell to run evaluations. This will run the evaluation phase and\nthen the analysis phase. By default, using this magic will run both `--run-eval` and `--run-eval-analysis`, but you can\nspecify your own flags by passing arguments right after the magic command (e.g., `%%ipytest_eval --run-eval`).\n\n```python\n%%ipytest_eval\nimport pytest\n\n\n@pytest.mark.eval(name=\"my_eval\")\ndef test_agent(eval_bag):\n    eval_bag.prediction = agent.run(case[\"input\"])\n\n\n@pytest.mark.eval_analysis(name=\"my_eval\")\ndef test_analysis(eval_results):\n    print(f\"F1 Score: {calculate_f1(eval_results):.2%}\")\n```\n\nYou can see an example of this in the [`example/example_notebook.ipynb`](example/example_notebook.ipynb) notebook. Or\nlook at the [advanced example](example/example_notebook_advanced.ipynb) for a more complex example that tracks multiple\nexperiments.\n\u003cp align=\"right\"\u003e(\u003ca href=\"#top\"\u003eback to top\u003c/a\u003e)\u003c/p\u003e\n\n## 🏗️ Production Use\n\n### 📚 Managing Test Data (Evaluation Set)\n\nIt's recommended to use a CSV file to store test data. This makes it easier to manage large datasets and allows you to\ncommunicate with non-technical stakeholders.\n\nTo do this, you can use `pandas` to read the CSV file and pass the test cases as parameters to your tests using\n`@pytest.mark.parametrize` 🙃 :\n\n```python\nimport pandas as pd\nimport pytest\n\ntest_data = pd.read_csv(\"tests/testdata.csv\")\n\n\n@pytest.mark.eval(name=\"my_eval\")\n@pytest.mark.parametrize(\"case\", test_data.to_dict(orient=\"records\"))\ndef test_agent(case, eval_bag, agent):\n    eval_bag.prediction = agent.run(case[\"input\"])\n```\n\nIn case you need to select a subset of the test data (e.g., a golden set), you can simply define an environment variable\nto indicate that, and filter the data with `pandas`.\n\n### 🔀 CI Integration\n\nRun tests and analysis as separate steps:\n\n```yaml\nevaluate:\n  steps:\n    - run: pytest --run-eval -n auto --supress-failed-exit-code  # Run cases in parallel\n    - run: pytest --run-eval-analysis  # Analyze results\n```\n\nUse `--supress-failed-exit-code` with `--run-eval` - let the analysis phase determine success/failure. **If all your\ncases pass, your evaluation set is probably too small!**\n\n### ⚡️ Parallel Testing\n\nAs your evaluation set grows, you may want to run your test cases in parallel. To do this, install\n[`pytest-xdist`](https://pytest-xdist.readthedocs.io/). `pytest-evals` will support that out of the box 🚀.\n\n```bash\nrun: pytest --run-eval -n auto\n```\n\n\u003cp align=\"right\"\u003e(\u003ca href=\"#top\"\u003eback to top\u003c/a\u003e)\u003c/p\u003e\n\n# 👷 Contributing\n\nContributions make the open-source community a fantastic place to learn, inspire, and create. Any contributions you make\nare **greatly appreciated** (not only code! but also documenting, blogging, or giving us feedback) 😍.\n\nPlease fork the repo and create a pull request if you have a suggestion. You can also simply open an issue to give us\nsome feedback.\n\n**Don't forget to give the project [a star](#top)! ⭐️**\n\nFor more information about contributing code to the project, read the [CONTRIBUTING.md](CONTRIBUTING.md) guide.\n\n# 📃 License\n\nThis project is licensed under the MIT License - see the [LICENSE](LICENSE) file for details.\n\u003cp align=\"right\"\u003e(\u003ca href=\"#top\"\u003eback to top\u003c/a\u003e)\u003c/p\u003e","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FAlmogBaku%2Fpytest-evals","html_url":"https://awesome.ecosyste.ms/projects/github.com%2FAlmogBaku%2Fpytest-evals","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FAlmogBaku%2Fpytest-evals/lists"}