{"id":28825647,"url":"https://github.com/mbzuai-oryx/arb","last_synced_at":"2026-03-09T19:13:38.534Z","repository":{"id":294842835,"uuid":"987519639","full_name":"mbzuai-oryx/ARB","owner":"mbzuai-oryx","description":"ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark","archived":false,"fork":false,"pushed_at":"2025-05-22T13:59:36.000Z","size":30300,"stargazers_count":2,"open_issues_count":0,"forks_count":0,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-05-22T15:31:24.311Z","etag":null,"topics":["arabic","benchmark","cot","lmm","reasoning"],"latest_commit_sha":null,"homepage":"https://slmlah.github.io/ARB/","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/mbzuai-oryx.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2025-05-21T07:38:29.000Z","updated_at":"2025-05-22T14:07:09.000Z","dependencies_parsed_at":"2025-05-22T15:41:34.451Z","dependency_job_id":null,"html_url":"https://github.com/mbzuai-oryx/ARB","commit_stats":null,"previous_names":["slmlah/arb","mbzuai-oryx/arb"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/mbzuai-oryx/ARB","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/mbzuai-oryx%2FARB","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/mbzuai-oryx%2FARB/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/mbzuai-oryx%2FARB/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/mbzuai-oryx%2FARB/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/mbzuai-oryx","download_url":"https://codeload.github.com/mbzuai-oryx/ARB/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/mbzuai-oryx%2FARB/sbom","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":260669572,"owners_count":23044299,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["arabic","benchmark","cot","lmm","reasoning"],"created_at":"2025-06-19T02:05:08.328Z","updated_at":"2026-03-09T19:13:38.528Z","avatar_url":"https://github.com/mbzuai-oryx.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"\n\n\u003cdiv align=\"center\"\u003e\n  \u003cimg src=\"assets/arab_logo.png\" width=\"11%\" align=\"left\"/\u003e\n\u003c/div\u003e\n\n\u003cdiv style=\"margin-top:50px;\"\u003e\n  \u003ch1 style=\"font-size: 30px; margin: 0;\"\u003e  ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark\u003c/h1\u003e\n\u003c/div\u003e\n\n \u003cdiv  align=\"center\" style=\"margin-top:10px;\"\u003e \n    \n  [Sara Ghaboura](https://huggingface.co/SLMLAH) \u003csup\u003e * \u003c/sup\u003e \u0026nbsp;\n  [Ketan More](https://github.com/ketanmore2002) \u003csup\u003e * \u003c/sup\u003e \u0026nbsp;\n  [Wafa Alghallabi](https://huggingface.co/SLMLAH) \u0026nbsp;\n  [Omkar Thawakar](https://omkarthawakar.github.io)  \u0026nbsp;\n  \u003cbr\u003e\n  [Jorma Laaksonen](https://scholar.google.com/citations?user=qQP6WXIAAAAJ\u0026hl=en) \u0026nbsp;\n  [Hisham Cholakkal](https://scholar.google.com/citations?hl=en\u0026user=bZ3YBRcAAAAJ) \u0026nbsp;\n  [Salman Khan](https://scholar.google.com/citations?hl=en\u0026user=M59O9lkAAAAJ) \u0026nbsp;\n  [Rao M. Anwer](https://scholar.google.com/citations?hl=en\u0026user=_KlvMVoAAAAJ)\u003cbr\u003e\n  \u003cem\u003e \u003csup\u003e *Equal Contribution  \u003c/sup\u003e \u003c/em\u003e\n  \u003cbr\u003e\n  \u003cbr\u003e  \n  [![arXiv](https://img.shields.io/badge/arXiv-2505.17021-C4EAE5)](https://arxiv.org/abs/2505.17021)\n  [![Our Page](https://img.shields.io/badge/Visit-Our%20Page-C5D9D9?style=flat)](https://mbzuai-oryx.github.io/ARB/)\n  [![GitHub issues](https://img.shields.io/github/issues/mbzuai-oryx/Camel-Bench?color=D8EADC\u0026label=issues\u0026style=flat)](https://github.com/mbzuai-oryx/ARB/issues)\n  [![GitHub stars](https://img.shields.io/github/stars/mbzuai-oryx/TimeTravel?color=C7D7E3\u0026style=flat)](https://github.com/mbzuai-oryx/ARB/stargazers)\n  [![GitHub license](https://img.shields.io/github/license/mbzuai-oryx/Camel-Bench?color=C8B9A7)](https://github.com/mbzuai-oryx/ARB/blob/main/LICENSE)\n  \u003cbr\u003e\n  \u003cem\u003e \u003csup\u003e *Equal Contribution  \u003c/sup\u003e \u003c/em\u003e\n  \u003cbr\u003e\n  \u003cbr\u003e\n  \n\n  \n\u003cp align=\"center\"\u003e\n    \u003cimg src=\"assets/line.png\"  height=\"9px\"\u003e\n\u003c/p\u003e \n\n\u003c/div\u003e\n\u003cdiv align=\"center\"\u003e\n \u003cb\u003e If you like our project, please give us a star ⭐ on GitHub for the latest update. \u003c/b\u003e\u003cbr\u003e\n\u003c/div\u003e\n\u003cbr\u003e\n\u003cp align=\"center\"\u003e\n    \u003cimg src=\"assets/line.png\" height=\"9px\"\u003e\n\u003c/p\u003e \n\u003cbr\u003e\n\u003cbr\u003e\n\u003c/p\u003e \n\n##  \u003cimg src=\"https://github.com/user-attachments/assets/1abcf195-ad44-4500-a14b-f1a4bef9b748\" width=\"40\" height=\"40\" /\u003eLatest Updates\n 🔥  **[22 May 2025]** ARB is **1st** Arabic multimodal benchmark focused on step-by-step reasoning is released.\u003cbr\u003e\n 🤗  **[22 May 2025]** ARB dataset available on [HuggingFace](https://huggingface.co/datasets/MBZUAI/ARB).\u003cbr\u003e\n\n\u003cbr\u003e\n\u003cbr\u003e\n\n  \n\u003cdiv align=\"left\"\u003e\n  \n## \u003cimg src=\"https://github.com/user-attachments/assets/4e69bf65-7b6e-4dd4-8eda-d38b9f1049fc\" width=\"40\" height=\"40\" /\u003e ARB Scope and Diversity\n \n\nARB  is the first benchmark focused on  step-by-step reasoning in Arabic cross both textual and visual modalities, covering 11 diverse domains spanning science, culture, OCR, and historical interpretation.\n\u003cbr\u003e\n\u003c/p\u003e\n\u003cp align=\"center\"\u003e\n   \u003cimg src=\"assets/arb_sample_intro.png\" width=\"750px\" height=\"500px\" alt=\"Figure: ARB Dataset Coverage\"/\u003e\n\u003c/p\u003e\n\u003c/div\u003e\n\u003c/p\u003e\n\n## 🌟 Key Features\n\n- **1,356** multimodal samples, each with an image, Arabic question, and reasoning-based answer.\n- **5,119** curated reasoning steps reflecting human logic\n- **11 diverse domains**, from visual reasoning to historical and scientific analysis.\n- **Native Arabic speakers** and **domain experts** verified.\n- **Hybrid sources**: original Arabic data, high-quality translations, and synthetic samples.\n- **Robust evaluation framework** for final answer accuracy and reasoning quality\n- Fully **open-source dataset** and toolkit to support research in **Arabic reasoning and multimodal AI**.\n\n\u003cbr\u003e\n\n## 🏗️ ARB Construction Pipeline\n\n\u003cp align=\"center\"\u003e\n   \u003cimg src=\"assets/arb_pipeline.png\" width=\"750px\" height=\"180px\" alt=\"Figure: ARB Pipeline Overview\"/\u003e\n\u003c/p\u003e\n\n\u003cbr\u003e\n\n## \u003cimg src=\"https://github.com/user-attachments/assets/626aedcd-fdb0-4a1f-87de-bad0574900e5\" width=\"30\" height=\"30\" /\u003e ARB Collection\n\n\u003cp align=\"center\"\u003e\n   \u003cimg src=\"assets/arb_collection.png\" width=\"600px\" height=\"300px\" alt=\"Figure: ARB Collection\"/\u003e\n\n\n\u003c/p\u003e\n\u003cbr\u003e\n\n## \u003cimg src=\"https://github.com/user-attachments/assets/bdd52751-8503-4b30-9a85-e4e17f69242c\" width=\"30\" height=\"30\" /\u003e ARB Data Distribution over Domains\n\n\u003cp align=\"center\"\u003e\n   \u003cimg src=\"assets/arb_dist.png\" width=\"400px\" height=\"350px\" alt=\"Figure: ARB dist\"/\u003e\n\u003c/p\u003e\n\n\u003cdiv align=\"center\"\u003e\n  \n### Source Types Across Domains\n\n| **Domain**                 | **English Bench** | **Arabic Bench** | **Human-Created** | **Synthetic** |\n|---------------------------|:-----------------:|:----------------:|:-----------------:|:-------------:|\n| Visual Reasoning          | ✅                | –                | –                 | –             |\n| OCR \u0026 Document Analysis   | –                 | –                | ✅                | ✅            |\n| Chart \u0026 Data Table (CDT)  | ✅                | ✅               | ✅                | ✅            |\n| Math \u0026 Logic              | ✅                | –                | –                 | –             |\n| Social \u0026 Cultural         | ✅                | –                | –                 | –             |\n| Computer Vision Perception| ✅                | –                | –                 | –             |\n| Medical Image Analysis    | ✅                | ✅               | –                 | –             |\n| Scientific Reasoning      | ✅                | –                | –                 | –             |\n| Agricultural Interpretation | ✅              | –                | ✅                | ✅            |\n| Remote Sensing Understanding | –             | ✅               | –                 | –             |\n| Historical \u0026 Anthropological | ✅            | –                | ✅                | ✅            |\n\n\u003c/div\u003e\n\u003cbr\u003e\n\n## \u003cimg src=\"https://github.com/user-attachments/assets/2e19e70d-4f4d-4a98-a200-854a80de3cb9\" width=\"30\" height=\"30\" /\u003e  Download\n\n```bash\nfrom datasets import load_dataset\n\n# Login using e.g. `huggingface-cli login` to access this dataset\nds = load_dataset(\"MBZUAI/ARB\")\n```\n\u003cbr\u003e\n\n## \u003cimg src=\"https://github.com/user-attachments/assets/b9105099-639c-4411-8949-9fc4b7f6b86a\" width=\"30\" height=\"30\" /\u003e  Evaluation Protocol\n\u003cdiv\u003e\n\u003cp align=\"left\"\u003e\n  \nWe evaluated 12 open- and closed-source LMMs using:\u003c/p\u003e\n- **Lexical and Semantic Similarity Scoes**: BLEU, ROUGE, BERTScore.\u003c/p\u003e\n- **Cross-lingual semantic alignment**: LaBSE\u003c/p\u003e\n- **Custom Rubric (Arabic):**: Our curated metric rebric includes 10 factors like faithfulness, interpretive depth, coherence, hallucination, and more.\u003c/p\u003e\n\n### \u003cimg src=\"https://github.com/user-attachments/assets/9a10bc39-78ff-41d8-97f8-7cef32df518e\" width=\"30\" height=\"30\" /\u003e LLM-as-Judge (Arabic prompt-based)\n\nWe evaluate models using:\n\n- Step-by-step reasoning quality (coherence, informativeness, commonsense)\n- Final answer accuracy\n- Agreement with human raters (Krippendorff’s Alpha \u003e 87%)\n\u003c/p\u003e\n\u003c/div\u003e\n\u003cbr\u003e\n\n## \u003cimg src=\"https://github.com/user-attachments/assets/508bcdab-08e2-48e0-b89f-94b76190fdc0\" width=\"30\" height=\"30\"\u003e Stepwise Evaluation Results \nFor Closed-Source Models:\u003cbr\u003e\n|                     |   GPT-4o |   GPT-4o-mini |   GPT-4.1 |   o4-mini |   Gemini 1.5 Pro |   Gemini 2.0 Flash |\n|:--------------------|---------:|--------------:|----------:|----------:|-----------------:|-------------------:|\n| Final Answer (%)    |    60.22 |         52.22 |     59.43 |     58.93 |            56.7  |              57.8  |\n| Reasoning Steps (%) |    64.29 |         61.02 |     80.41 |     80.75 |            64.34 |              64.09 |\n\u003cbr\u003e\n\nFor Open-Source Models:\u003cbr\u003e\n|                     |   Qwen2.5-VL-7B |   Llama-3.2-11B |   AIN |   Llama-4 Scout |   Aya-Vision-8B |   InternVL3-8B |\n|:--------------------|----------------:|----------------:|------:|----------------:|----------------:|---------------:|\n| Final Answer (%)    |           37.02 |           25.58 | 27.35 |           48.52 |           28.81 |          31.04 |\n| Reasoning Steps (%) |           64.03 |           53.2  | 52.77 |           77.7  |           63.64 |          54.5  |\n\n\u003cbr\u003e\n\n## 📂 Dataset Structure\n\u003cdiv\u003e\n\u003cp align=\"left\"\u003e\n\nEach sample includes:\n- `image`: Visual input\n- `question`: Arabic reasoning prompt\n- `choices`: The choices for MCQ\n- `steps`: Ordered reasoning chain\n- `answer`: Final solution (Arabic)\n- `domain`: One of 11 categories (e.g., OCR, Scientific, Visual, Math)\n- `curriculum`: One of the 4 curricula followed by the prompt for steps generation (Computational, Sci/Med, Textual/Partial, and General)\n\u003c/p\u003e\n\n\u003c/div\u003e\n\n\n\u003cbr\u003e\n\u003cdiv align=\"left\"\u003e\n\n\n## \u003cimg src=\"https://github.com/user-attachments/assets/08e47f66-e0aa-49b5-b886-ad65ae7a6faa\" width=\"30\" height=\"30\" /\u003e Citation\nIf you use ARB dataset in your research, please consider citing:\n\n```bibtex\n@misc{ghaboura2025arbcomprehensivearabicmultimodal,\n      title={ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark}, \n      author={Sara Ghaboura and Ketan More and Wafa Alghallabi and Omkar Thawakar and Jorma Laaksonen and Hisham Cholakkal and Salman Khan and Rao Muhammad Anwer},\n      year={2025},\n      eprint={2505.17021},\n      archivePrefix={arXiv},\n      primaryClass={cs.CV},\n      url={https://arxiv.org/abs/2505.17021}, \n}\n```\n\n\u003c/div\u003e\n\n\n\n---\n\n\n\u003cp align=\"center\"\u003e\n   \u003cimg src=\"assets/IVAL_logo.png\" width=\"18%\" style=\"display: inline-block; margin: 0 10px;\" /\u003e\n   \u003cimg src=\"assets/Oryx_logo.jpeg\" width=\"12%\" style=\"display: inline-block; margin: 0 10px;\" /\u003e\n   \u003cimg src=\"assets/MBZUAI_logo.png\" width=\"40%\" style=\"display: inline-block; margin: 0 10px;\" /\u003e\n\u003c/p\u003e\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fmbzuai-oryx%2Farb","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fmbzuai-oryx%2Farb","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fmbzuai-oryx%2Farb/lists"}