{"id":50551875,"url":"https://github.com/vectifyai/mafin2.5-financebench","last_synced_at":"2026-06-04T04:02:48.468Z","repository":{"id":278938348,"uuid":"935556061","full_name":"VectifyAI/Mafin2.5-FinanceBench","owner":"VectifyAI","description":"📈 FinanceBench evaluation of Mafin 2.5","archived":false,"fork":false,"pushed_at":"2025-10-20T05:43:30.000Z","size":188,"stargazers_count":14,"open_issues_count":0,"forks_count":3,"subscribers_count":2,"default_branch":"main","last_synced_at":"2025-10-20T08:28:48.495Z","etag":null,"topics":["financebench"],"latest_commit_sha":null,"homepage":"https://vectify.ai/blog/Mafin2.5","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/VectifyAI.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2025-02-19T16:30:17.000Z","updated_at":"2025-10-20T05:43:33.000Z","dependencies_parsed_at":"2025-02-22T17:40:02.873Z","dependency_job_id":null,"html_url":"https://github.com/VectifyAI/Mafin2.5-FinanceBench","commit_stats":null,"previous_names":["vectifyai/mafin2.5-financebench"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/VectifyAI/Mafin2.5-FinanceBench","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/VectifyAI%2FMafin2.5-FinanceBench","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/VectifyAI%2FMafin2.5-FinanceBench/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/VectifyAI%2FMafin2.5-FinanceBench/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/VectifyAI%2FMafin2.5-FinanceBench/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/VectifyAI","download_url":"https://codeload.github.com/VectifyAI/Mafin2.5-FinanceBench/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/VectifyAI%2FMafin2.5-FinanceBench/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":33888302,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-06-04T02:00:06.755Z","response_time":64,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["financebench"],"created_at":"2026-06-04T04:02:47.762Z","updated_at":"2026-06-04T04:02:48.463Z","avatar_url":"https://github.com/VectifyAI.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# :rocket: Mafin2.5-FinanceBench\n\nThis repository contains the results of our finance benchmark evaluations using our Mafin2.5 system. These evaluations are based on the FinanceBench benchmark as introduced in the paper \n📄 [FinanceBench: A New Benchmark for Financial Question Answering](https://arxiv.org/pdf/2311.11944).\n\n### Mafin2.5 Introduction\n\n**Mafin2.5** is our latest RAG model on Financial reports, built on **PageIndex** -- a vectorless, reasoning-based RAG framework. See [this page](http://pageindex.ai/) for more details.\n\n\n\n### Benchmark Overview\n\n[FinanceBench](https://arxiv.org/pdf/2311.11944) is a pioneering test suite designed to evaluate the performance of large language models (LLMs) on open-book financial question answering (QA). It includes questions about publicly traded companies, each accompanied by corresponding answers and evidence strings. \nIt has the following key features:\n- **Ecologically valid questions:** Covers a diverse set of scenarios relevant to publicly traded companies.\n- **Model evaluation:** Includes assessment of 16 state-of-the-art model configurations such as GPT-4.\n- **Limitations identified:** Highlights the limitations of current LLMs for financial QA, including hallucinations and refusal to answer.\n\n### Evaluation Protocol  \n\nWe follow a **realistic and practical evaluation setup**, where all documents are stored in a **single database**, and Mafin2.5 is tested on the [FinanceBench public set](https://github.com/patronus-ai/financebench). This approach ensures that our model is evaluated under conditions that closely resemble real-world financial applications.  For transparency, we have **open-sourced our evaluation code** in **[Evaluation Code](https://github.com/VectifyAI/Mafin2.5-FinanceBench/blob/main/eval.py)**.  \n\nIn cases where questions are **ambiguous, have multiple valid answers, or are deemed invalid**, we rely on **expert human annotations** to ensure fair and accurate evaluation. For more details, see **[Human Evaluation](https://github.com/VectifyAI/Mafin2.5-FinanceBench/tree/main/human_evaluations)**.\n\n\n## Results\n### 1. Mafin Model Evolution\n\u003c!-- ![evolution](https://github.com/user-attachments/assets/a7f78677-64ac-4bc6-9311-6dfc8bc19033) --\u003e\n\n\u003cp align=\"center\"\u003e\n  \u003cimg src=\"https://github.com/user-attachments/assets/a7f78677-64ac-4bc6-9311-6dfc8bc19033\" alt=\"evolution\" width=\"80%\"\u003e\n\u003c/p\u003e\n\nThis figure showcases the progression of Mafin models, highlighting the significant improvement in accuracy from **Mafin 1 to Mafin 2.5**. The latest iteration, **Mafin 2.5**, achieves a remarkable accuracy of **98.7%**, demonstrating major advancements in reasoning and retrieval capabilities.\n\n\n### 2. Mafin2.5 Performance Across Base Models\n\u003c!-- \u003cimg width=\"1020\" alt=\"comparison\" src=\"https://github.com/user-attachments/assets/f98957f4-dcdd-45fb-807e-198cb73b9452\" /\u003e --\u003e\n\n\u003cp align=\"center\"\u003e\n\u003cimg width=\"80%\" alt=\"414673269-f98957f4-dcdd-45fb-807e-198cb73b9452\" src=\"https://github.com/user-attachments/assets/ca3e510b-b89f-4cff-aba8-24fb1a1beeb3\" /\u003e\n\u003c/p\u003e\n\n\nAs a RAG 3.0 model, Mafin 2.5 is capable of leveraging different base models while maintaining **consistent high performance (98.7%)**. The above figure illustrates its effectiveness across **ChatGPT 4o** and **Deepseek v3**, indicating that its strong performance is **independent of the underlying LLM**. Notably, **Deepseek v3 is a privately deployable model**, offering an alternative for organizations requiring on-premise or self-hosted AI solutions.\n\n\n### 3. Comparison with Market Players\n\u003c!-- ![sandian8](https://github.com/user-attachments/assets/469e2df8-64b5-4a15-9427-be83719dac2d) --\u003e\n\n\u003cp align=\"center\"\u003e\n\u003cimg width=\"80%\" alt=\"414810759-469e2df8-64b5-4a15-9427-be83719dac2d\" src=\"https://github.com/user-attachments/assets/5e54b7ba-405a-46cd-b9dd-7de11642c308\" /\u003e\n\u003c/p\u003e\n\n\u003cdiv align=\"center\"\u003e\n\n| Method               | Accuracy (%) | Full Benchmark? (Coverage) |  Results Public? | Source |\n|----------------------|-------------|----------------------------|-----------------|--------|\n| Mafin2.5           | **98.7**      | **Yes (100%)**             | **Yes**         | [link](https://github.com/VectifyAI/Mafin2.5-FinanceBench) |\n| Quantly             | 94           | **Yes (100%)**             |  No              | [link](https://quantly.substack.com/p/evaluation-of-quantly-on-financebench) |\n| Fintool             | 98           | No (66.7%)                 |  No              | [link](https://fintool.com/benchmark/chatgpt-versus-fintool) |\n| ChatGPT 4o + Search | 31           | No (66.7%)                 |  No              | [link](https://fintool.com/benchmark/chatgpt-versus-fintool) |\n| Perplexity          | 45           | No (66.7%)                 |  No              | [link](https://fintool.com/benchmark/perplexity-versus-fintool) |\n\n\u003c/div\u003e\n\nThis benchmark comparison demonstrates **Mafin 2.5's superiority** over competitors, achieving the **highest accuracy (98.7%)** while covering the **full benchmark (100%)**. Unlike some competitors that only evaluate on **partial benchmarks**, Mafin 2.5 provides a **comprehensive** and **rigorous** assessment.\n\n\n\n\n\n### Key Takeaways\n1. **Mafin 2.5 demonstrates massive improvements over previous versions**, significantly increasing accuracy from **Mafin 1 (38.0%) to Mafin 2.5 (98.7%)**, showcasing strong advancements in financial AI reasoning.\n2. **Mafin 2.5 is highly adaptable across different base models**, achieving **identical high performance (98.7%)** on both **ChatGPT 4o (public cloud) and Deepseek v3 (private deployable)**, making it flexible for various deployment needs.\n3. **Mafin 2.5 outperforms market competitors while covering the full benchmark (100%)**, ensuring a **more comprehensive and fair evaluation** compared to models that only test on 66.7% of the dataset.\n\n\n\n## Limitations of the Current Benchmark  \n\n1. **Errors and Ambiguities in Evaluation**  \n   The current benchmark may contain inconsistencies, ambiguities, or errors in ground truth answers, which can lead to misleading performance evaluations. These issues must be addressed to ensure a fair and reliable assessment of AI capabilities. Establishing a more rigorous annotation and validation process is essential for improving benchmark accuracy.\n\n2. **Lack of Multi-Document Reasoning Tasks**  \n   The current benchmark primarily focuses on simple retrieval tasks based on a single document. However, real-world financial applications require more advanced reasoning capabilities, including multi-step retrieval across multiple documents. To improve the benchmark, we call for the inclusion of complex reasoning tasks that better reflect real-world decision-making and analysis.\n\n\n## Contact\nIf you have questions about these results or want to try our model, email us at contact@vectify.ai.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fvectifyai%2Fmafin2.5-financebench","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fvectifyai%2Fmafin2.5-financebench","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fvectifyai%2Fmafin2.5-financebench/lists"}