{"id":13624255,"url":"https://github.com/langchain-ai/langchain-benchmarks","last_synced_at":"2025-04-15T03:50:25.596Z","repository":{"id":188310814,"uuid":"677148333","full_name":"langchain-ai/langchain-benchmarks","owner":"langchain-ai","description":"🦜💯 Flex those feathers!","archived":false,"fork":false,"pushed_at":"2024-10-21T20:47:16.000Z","size":17036,"stargazers_count":244,"open_issues_count":19,"forks_count":51,"subscribers_count":8,"default_branch":"main","last_synced_at":"2025-04-07T16:16:58.935Z","etag":null,"topics":["benchmark-framework","benchmarking","langchain","langchain-python","llm","llms"],"latest_commit_sha":null,"homepage":"https://langchain-ai.github.io/langchain-benchmarks/","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/langchain-ai.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":"security.md","support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2023-08-10T21:31:11.000Z","updated_at":"2025-04-01T07:48:36.000Z","dependencies_parsed_at":null,"dependency_job_id":"28e1c480-42c8-43c5-96d0-15efb167c354","html_url":"https://github.com/langchain-ai/langchain-benchmarks","commit_stats":{"total_commits":208,"total_committers":12,"mean_commits":"17.333333333333332","dds":0.4326923076923077,"last_synced_commit":"99cf03a50a76f0cd341cf69b166404b34839078f"},"previous_names":["langchain-ai/langchain-benchmarks"],"tags_count":14,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/langchain-ai%2Flangchain-benchmarks","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/langchain-ai%2Flangchain-benchmarks/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/langchain-ai%2Flangchain-benchmarks/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/langchain-ai%2Flangchain-benchmarks/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/langchain-ai","download_url":"https://codeload.github.com/langchain-ai/langchain-benchmarks/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":249003944,"owners_count":21196794,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["benchmark-framework","benchmarking","langchain","langchain-python","llm","llms"],"created_at":"2024-08-01T21:01:40.637Z","updated_at":"2025-04-15T03:50:25.579Z","avatar_url":"https://github.com/langchain-ai.png","language":"Python","funding_links":[],"categories":["A01_文本生成_文本对话","Python"],"sub_categories":["大语言对话模型及数据"],"readme":"# 🦜💯 LangChain Benchmarks\n\n[![Release Notes](https://img.shields.io/github/release/langchain-ai/langchain-benchmarks)](https://github.com/langchain-ai/langchain-benchmarks/releases)\n[![CI](https://github.com/langchain-ai/langchain-benchmarks/actions/workflows/ci.yml/badge.svg)](https://github.com/langchain-ai/langchain-benchmarks/actions/workflows/ci.yml)\n[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)\n[![Twitter](https://img.shields.io/twitter/url/https/twitter.com/langchainai.svg?style=social\u0026label=Follow%20%40LangChainAI)](https://twitter.com/langchainai)\n[![](https://dcbadge.vercel.app/api/server/6adMQxSpJS?compact=true\u0026style=flat)](https://discord.gg/6adMQxSpJS)\n[![Open Issues](https://img.shields.io/github/issues-raw/langchain-ai/langchain-benchmarks)](https://github.com/langchain-ai/langchain-benchmarks/issues)\n\n\n[📖 Documentation](https://langchain-ai.github.io/langchain-benchmarks/index.html)\n\nA package to help benchmark various LLM related tasks.\n\nThe benchmarks are organized by end-to-end use cases, and\nutilize [LangSmith](https://smith.langchain.com/) heavily.\n\nWe have several goals in open sourcing this:\n\n- Showing how we collect our benchmark datasets for each task\n- Showing what the benchmark datasets we use for each task is\n- Showing how we evaluate each task\n- Encouraging others to benchmark their solutions on these tasks (we are always looking for better ways of doing things!)\n\n## Benchmarking Results\n\nRead some of the articles about benchmarking results on our blog.\n\n* [Agent Tool Use](https://blog.langchain.dev/benchmarking-agent-tool-use/)\n* [Query Analysis in High Cardinality Situations](https://blog.langchain.dev/high-cardinality/)\n* [RAG on Tables](https://blog.langchain.dev/benchmarking-rag-on-tables/)\n* [Q\u0026A over CSV data](https://blog.langchain.dev/benchmarking-question-answering-over-csv-data/)\n\n\n### Tool Usage (2024-04-18)\n\nSee [tool usage docs](https://langchain-ai.github.io/langchain-benchmarks/notebooks/tool_usage/benchmark_all_tasks.html) to recreate!\n\n![download](https://github.com/langchain-ai/langchain-benchmarks/assets/3205522/0da33de8-e03f-49cf-bd48-e9ff945828a9)\n\nExplore Agent Traces on LangSmith:\n\n* [Relational Data](https://smith.langchain.com/public/22721064-dcf6-4e42-be65-e7c46e6835e7/d)\n* [Tool Usage (1-tool)](https://smith.langchain.com/public/ac23cb40-e392-471f-b129-a893a77b6f62/d)\n* [Tool Usage (26-tools)](https://smith.langchain.com/public/366bddca-62b3-4b6e-849b-a478abab73db/d)\n* [Multiverse Math](https://smith.langchain.com/public/983faff2-54b9-4875-9bf2-c16913e7d489/d)\n\n## Installation\n\nTo install the packages, run the following command:\n\n```bash\npip install -U langchain-benchmarks\n```\n\nAll the benchmarks come with an associated benchmark dataset stored in [LangSmith](https://smith.langchain.com). To take advantage of the eval and debugging experience, [sign up](https://smith.langchain.com), and set your API key in your environment:\n\n```bash\nexport LANGCHAIN_API_KEY=ls-...\n```\n\n## Repo Structure\n\nThe package is located within [langchain_benchmarks](./langchain_benchmarks/). Check out the [docs](https://langchain-ai.github.io/langchain-benchmarks/index.html) for information on how to get starte.\n\nThe other directories are legacy and may be moved in the future.\n\n\n## Archived\n\nBelow are archived benchmarks that require cloning this repo to run.\n\n- [CSV Question Answering](https://github.com/langchain-ai/langchain-benchmarks/tree/main/archived/csv-qa)\n- [Extraction](https://github.com/langchain-ai/langchain-benchmarks/tree/main/archived/extraction)\n- [Q\u0026A over the LangChain docs](https://github.com/langchain-ai/langchain-benchmarks/tree/main/archived/langchain-docs-benchmarking)\n- [Meta-evaluation of 'correctness' evaluators](https://github.com/langchain-ai/langchain-benchmarks/tree/main/archived/meta-evals)\n\n\n## Related\n\n- For cookbooks on other ways to test, debug, monitor, and improve your LLM applications, check out the [LangSmith docs](https://docs.smith.langchain.com/)\n- For information on building with LangChain, check out the [python documentation](https://python.langchain.com/docs/get_started/introduction) or [JS documentation](https://js.langchain.com/docs/get_started/introduction)\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Flangchain-ai%2Flangchain-benchmarks","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Flangchain-ai%2Flangchain-benchmarks","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Flangchain-ai%2Flangchain-benchmarks/lists"}