{"id":31132903,"url":"https://github.com/chirindaopensource/search_benford_law_compatibility","last_synced_at":"2026-04-12T01:36:07.373Z","repository":{"id":314634010,"uuid":"1056233857","full_name":"chirindaopensource/search_benford_law_compatibility","owner":"chirindaopensource","description":"End-to-End Python scalable forensic accounting toolkit implementing Benford's Law analysis for FTSE financial data. Delivers automated anomaly detection with Chi-Squared/MAD testing, comprehensive validation pipelines, and risk-based prioritization of investigative resources. Replicates Ausloos et al.'s (2025) methodology with full reproducibility.","archived":false,"fork":false,"pushed_at":"2025-09-13T17:10:49.000Z","size":83,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":0,"default_branch":"main","last_synced_at":"2025-09-13T19:08:40.536Z","etag":null,"topics":["academic-research","anomaly-detection","benfords-law","chi-squared-test","data-validation","econometrics","financial-analysis","financial-data","forensic-accounting","fraud-detection","ftse","goodness-of-fit","jupyter-notebook","numpy","pandas","python","reproducible-research","risk-management","scipy","statistical-testing"],"latest_commit_sha":null,"homepage":"","language":"Jupyter Notebook","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/chirindaopensource.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2025-09-13T16:59:18.000Z","updated_at":"2025-09-13T17:23:25.000Z","dependencies_parsed_at":"2025-09-13T19:08:42.284Z","dependency_job_id":"30cdf7aa-d015-4223-941d-734511bde374","html_url":"https://github.com/chirindaopensource/search_benford_law_compatibility","commit_stats":null,"previous_names":["chirindaopensource/search_benford_law_compatibility"],"tags_count":null,"template":false,"template_full_name":null,"purl":"pkg:github/chirindaopensource/search_benford_law_compatibility","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/chirindaopensource%2Fsearch_benford_law_compatibility","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/chirindaopensource%2Fsearch_benford_law_compatibility/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/chirindaopensource%2Fsearch_benford_law_compatibility/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/chirindaopensource%2Fsearch_benford_law_compatibility/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/chirindaopensource","download_url":"https://codeload.github.com/chirindaopensource/search_benford_law_compatibility/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/chirindaopensource%2Fsearch_benford_law_compatibility/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":280761857,"owners_count":26386245,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-10-24T02:00:06.418Z","response_time":73,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["academic-research","anomaly-detection","benfords-law","chi-squared-test","data-validation","econometrics","financial-analysis","financial-data","forensic-accounting","fraud-detection","ftse","goodness-of-fit","jupyter-notebook","numpy","pandas","python","reproducible-research","risk-management","scipy","statistical-testing"],"created_at":"2025-09-18T05:02:07.695Z","updated_at":"2025-10-24T07:50:07.539Z","avatar_url":"https://github.com/chirindaopensource.png","language":"Jupyter Notebook","funding_links":[],"categories":[],"sub_categories":[],"readme":"# `README.md`\n\n# A Scalable Framework for Benford's Law Conformity Testing of Financial Data\n\n\u003c!-- PROJECT SHIELDS --\u003e\n[![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](https://opensource.org/licenses/MIT)\n[![Python Version](https://img.shields.io/badge/python-3.9%2B-blue.svg)](https://www.python.org/downloads/)\n[![Code style: black](https://img.shields.io/badge/code%20style-black-000000.svg)](https://github.com/psf/black)\n[![Imports: isort](https://img.shields.io/badge/%20imports-isort-%231674b1?style=flat\u0026labelColor=ef8336)](https://pycqa.github.io/isort/)\n[![Type Checking: mypy](https://img.shields.io/badge/type_checking-mypy-blue)](http://mypy-lang.org/)\n[![Jupyter](https://img.shields.io/badge/Jupyter-%23F37626.svg?style=flat\u0026logo=Jupyter\u0026logoColor=white)](https://jupyter.org/)\n[![arXiv](https://img.shields.io/badge/arXiv-2509.09415-b31b1b.svg)](https://arxiv.org/abs/2509.09415)\n[![Year](https://img.shields.io/badge/Year-2025-purple)](https://github.com/chirindaopensource/search_benford_law_compatibility)\n[![Discipline](https://img.shields.io/badge/Discipline-Forensic%20Accounting%20%26%20Econometrics-blue)](https://github.com/chirindaopensource/search_benford_law_compatibility)\n[![Methodology](https://img.shields.io/badge/Methodology-Benford's%20Law%20%7C%20Goodness--of--Fit-orange)](https://github.com/chirindaopensource/search_benford_law_compatibility)\n[![Data Source](https://img.shields.io/badge/Data-Refinitiv%20EIKON-lightgrey)](https://www.refinitiv.com/en/products/eikon-trading-software)\n[![Pandas](https://img.shields.io/badge/pandas-%23150458.svg?style=flat\u0026logo=pandas\u0026logoColor=white)](https://pandas.pydata.org/)\n[![NumPy](https://img.shields.io/badge/numpy-%23013243.svg?style=flat\u0026logo=numpy\u0026logoColor=white)](https://numpy.org/)\n[![SciPy](https://img.shields.io/badge/SciPy-%23025596?style=flat\u0026logo=scipy\u0026logoColor=white)](https://scipy.org/)\n[![PyYAML](https://img.shields.io/badge/PyYAML-4B5F6E.svg?style=flat)](https://pyyaml.org/)\n\n--\n\n**Repository:** `https://github.com/chirindaopensource/search_benford_law_compatibility`\n\n**Owner:** 2025 Craig Chirinda (Open Source Projects)\n\nThis repository contains an **independent**, professional-grade Python implementation of the research methodology from the 2025 paper entitled **\"Note on pre-taxation reported data by UK FTSE-listed companies. A search for Benford's laws compatibility\"** by:\n\n*   Marcel Ausloos\n*   Probowo Erawan Sastroredjo\n*   Polina Khrennikova\n\nThe project provides a complete, end-to-end computational framework for testing the conformity of financial datasets with Benford's Law. It delivers a modular, auditable, and extensible pipeline that replicates the paper's entire workflow: from rigorous data and configuration validation, through robust, mathematically-grounded digit extraction, to the precise calculation of Chi-Squared and Mean Absolute Deviation (MAD) test statistics and the final generation of publication-quality results tables.\n\n## Table of Contents\n\n- [Introduction](#introduction)\n- [Theoretical Background](#theoretical-background)\n- [Features](#features)\n- [Methodology Implemented](#methodology-implemented)\n- [Core Components (Notebook Structure)](#core-components-notebook-structure)\n- [Key Callable: execute_full_project_workflow](#key-callable-execute_full_project_workflow)\n- [Prerequisites](#prerequisites)\n- [Installation](#installation)\n- [Input Data Structure](#input-data-structure)\n- [Usage](#usage)\n- [Output Structure](#output-structure)\n- [Project Structure](#project-structure)\n- [Customization](#customization)\n- [Contributing](#contributing)\n- [Recommended Extensions](#recommended-extensions)\n- [License](#license)\n- [Citation](#citation)\n- [Acknowledgments](#acknowledgments)\n\n## Introduction\n\nThis project provides a Python implementation of the methodologies presented in the 2025 paper \"Note on pre-taxation reported data by UK FTSE-listed companies. A search for Benford's laws compatibility.\" The core of this repository is the iPython Notebook `search_benford_law_compatibility_draft.ipynb`, which contains a comprehensive suite of functions to replicate the paper's findings, from initial data validation to the final generation and analysis of conformity test results.\n\nThe paper addresses a classic problem in forensic accounting and auditing: using statistical laws to identify potential anomalies in reported financial data. This codebase operationalizes the paper's approach, allowing users to:\n-   Rigorously validate input data and methodological parameters against a predefined schema.\n-   Perform a detailed data quality audit, identifying outliers and missing value patterns without altering the source data.\n-   Create derived analytical variables (e.g., PI/TA ratio) and segmented datasets (e.g., profitable vs. unprofitable firms).\n-   Apply robust, mathematically-grounded algorithms to extract the first, second, and first-two significant digits from numerical data.\n-   Calculate the theoretical probability distributions for Benford's Law (BL1, BL2, BL12) with high precision.\n-   Execute Pearson's Chi-Squared (χ²) and Mean Absolute Deviation (MAD) goodness-of-fit tests to quantify deviations from the theoretical benchmarks.\n-   Generate a full suite of publication-quality tables that precisely replicate the findings in the source paper.\n-   Conduct a comprehensive series of robustness checks, including bootstrap resampling, parameter sensitivity analysis, and temporal stability analysis.\n\n## Theoretical Background\n\nThe implemented methods are grounded in Probability Theory, Statistics, and Forensic Accounting.\n\n**1. Benford's Law (The First-Digit Law):**\nBenford's Law states that in many naturally occurring sets of numbers, the leading significant digit is more likely to be small. The probability of a first digit `d` is given by:\n$$ P(d_1 = i) = \\log_{10}\\left(1 + \\frac{1}{i}\\right) \\quad \\text{for } i \\in \\{1, ..., 9\\} $$\nThis principle extends to second and subsequent digits, as well as to blocks of digits, with specific logarithmic formulas. Deviations from these expected frequencies can suggest that a dataset was not naturally generated and may be the result of manipulation, fabrication, or systemic error.\n\n**2. Chi-Squared (χ²) Goodness-of-Fit Test:**\nThis is a classical statistical test used to determine if there is a significant difference between observed and expected frequencies. The test statistic is calculated as:\n$$ \\chi^2 = \\sum_{i=1}^{K} \\frac{(O_i - E_i)^2}{E_i} $$\nwhere `O` are the observed counts and `E` are the expected counts across `K` bins. A large χ² statistic suggests that the observed data does not fit the theoretical distribution.\n\n**3. Mean Absolute Deviation (MAD) Test:**\nThe MAD test provides a direct measure of the magnitude of the deviation between the observed proportions (`fₒ`) and the expected proportions (`fₑ`). The paper uses the sum of absolute deviations:\n$$ \\text{MAD} = \\sum_{i=1}^{K} |f_{o,i} - f_{e,i}| $$\nThis statistic is less sensitive to sample size than the χ² test and provides an intuitive measure of non-conformity, which is then compared against established thresholds.\n\n## Features\n\nThe provided iPython Notebook (`search_benford_law_compatibility_draft.ipynb`) implements the full research pipeline, including:\n\n-   **Modular, Task-Based Architecture:** The entire pipeline is broken down into 18 distinct, modular tasks, from data validation to final packaging.\n-   **Configuration-Driven Design:** All methodological parameters are managed in an external `config.yaml` file, allowing for easy customization without code changes.\n-   **Professional-Grade Data Validation:** A comprehensive validation suite ensures all inputs (data and configurations) conform to the required schema before execution.\n-   **Robust, Mathematical Digit Extraction:** Numerically stable, from-scratch implementations of logarithmic algorithms for extracting significant digits, avoiding common pitfalls of string-based methods.\n-   **High-Fidelity Statistical Testing:** Precise, vectorized implementations of the Chi-Squared and Mean Absolute Deviation tests.\n-   **Automated Report Generation:** Programmatic generation of all 6 analytical tables (Tables 3-8) from the paper with high fidelity to the original formatting.\n-   **Advanced Robustness Toolkit:**\n    -   A framework for conducting **bootstrap resampling** to assess the stability of test statistics.\n    -   A framework for **parameter sensitivity analysis** to test the impact of varying alpha levels and MAD thresholds.\n    -   A framework for **temporal stability analysis** to check for consistency of results over time.\n-   **Automated Replication Validation:** A final quality assurance step that programmatically compares the generated results against the paper's published statistics and produces a certification report.\n\n## Methodology Implemented\n\nThe core analytical steps directly implement the methodology from the paper:\n\n1.  **Validation (Tasks 1-3):** Ingests and rigorously validates the raw data and `config.yaml` file, and performs a detailed data quality audit.\n2.  **Preprocessing (Tasks 4-6):** Prepares the data by flagging outliers, creating the PI/TA ratio, segmenting the data by profitability, and generating the summary statistics table (Table 1).\n3.  **Digit Extraction (Tasks 7-9):** Applies robust mathematical algorithms to extract the first, second, and first-two significant digits for all relevant variables.\n4.  **Theoretical Distribution Generation (Tasks 10-12):** Calculates the high-precision theoretical probabilities and expected frequencies for BL1, BL2, and BL12.\n5.  **Statistical Testing \u0026 Reporting (Tasks 13-15):** Executes the Chi-Squared and MAD tests for all variable-digit combinations and compiles the results into replications of Tables 3-8.\n6.  **Robustness \u0026 Packaging (Tasks 16-18):** Orchestrates the entire pipeline, runs the optional robustness checks, and performs the final validation and packaging of all outputs.\n\n## Core Components (Notebook Structure)\n\nThe `search_benford_law_compatibility_draft.ipynb` notebook is structured as a logical pipeline with modular orchestrator functions for each of the major tasks. All functions are self-contained, fully documented with type hints and docstrings, and designed for professional-grade execution.\n\n## Key Callable: execute_full_project_workflow\n\nThe central function in this project is `execute_full_project_workflow`. It orchestrates the entire analytical workflow, providing a single entry point for running the baseline study replication and the advanced robustness checks.\n\n```python\ndef execute_full_project_workflow(\n    raw_financial_data: pd.DataFrame,\n    study_configuration: Dict[str, Any],\n    output_directory: str,\n    run_robustness_checks: bool = True\n) -\u003e Dict[str, Any]:\n    \"\"\"\n    Executes the entire, end-to-end research project workflow.\n    \"\"\"\n    # ... (implementation is in the notebook)\n```\n\n## Prerequisites\n\n-   Python 3.9+\n-   Core dependencies: `pandas`, `numpy`, `scipy`, `pyyaml`, `tqdm`.\n\n## Installation\n\n1.  **Clone the repository:**\n    ```sh\n    git clone https://github.com/chirindaopensource/search_benford_law_compatibility.git\n    cd search_benford_law_compatibility\n    ```\n\n2.  **Create and activate a virtual environment (recommended):**\n    ```sh\n    python -m venv venv\n    source venv/bin/activate  # On Windows, use `venv\\Scripts\\activate`\n    ```\n\n3.  **Install Python dependencies:**\n    ```sh\n    pip install pandas numpy scipy pyyaml tqdm\n    ```\n\n## Input Data Structure\n\nThe pipeline requires two primary inputs:\n1.  **`raw_financial_data`:** A `pandas.DataFrame` containing the panel data. It **must** have a `MultiIndex` with the levels `['CompanyID', 'Year']` and the columns `['CompanyName', 'PreTaxIncome_GBP', 'TotalAssets_GBP']`. Financial columns must be `float64`.\n2.  **`study_configuration`:** A Python dictionary loaded from the `config.yaml` file, which controls all methodological parameters.\n\nA mock data generation function is provided in the main notebook to create a valid example DataFrame for testing the pipeline.\n\n## Usage\n\nThe `search_benford_law_compatibility_draft.ipynb` notebook provides a complete, step-by-step guide. The core workflow is:\n\n1.  **Prepare Inputs:** Load your `raw_financial_data` DataFrame. Ensure the `config.yaml` file is present in the same directory.\n2.  **Execute Pipeline:** Call the grand orchestrator function.\n\n    ```python\n    # This single call runs the entire project.\n    final_project_outputs = execute_full_project_workflow(\n        raw_financial_data=my_company_data_df,\n        study_configuration=my_config_dict,\n        output_directory=\"research_outputs\",\n        run_robustness_checks=False  # Set to True for the full analysis\n    )\n    ```\n3.  **Inspect Outputs:** All results are saved to the specified output directory. You can also programmatically access any result from the returned dictionary.\n\n## Output Structure\n\nThe `execute_full_project_workflow` function returns a single, comprehensive dictionary containing all generated artifacts. Additionally, the `output_directory` will be populated with:\n-   `replication_validation_report.json`: A report certifying the accuracy of the replication.\n-   `data_quality_report.json`: A detailed audit of the input data quality.\n-   `table_1_summary_statistics.csv`: The replicated descriptive statistics table.\n-   `table_3_body.csv`, `table_3_stats.csv`, etc.: Pairs of CSV files for each analytical table (3-8).\n-   `raw_test_results.json`: A comprehensive file with the detailed numerical outputs of all statistical tests.\n-   `robustness_analysis_report.json`: (If run) A file with the results of all robustness checks.\n\n## Project Structure\n\n```\nsearch_benford_law_compatibility/\n│\n├── search_benford_law_compatibility_draft.ipynb   # Main implementation notebook\n├── config.yaml                                    # Master configuration file\n├── requirements.txt                               # Python package dependencies\n├── LICENSE                                        # MIT license file\n└── README.md                                      # This documentation file\n```\n\n## Customization\n\nThe pipeline is highly customizable via the `config.yaml` file. Users can easily modify all methodological parameters, such as statistical thresholds, critical values, and bin definitions, without altering the core Python code.\n\n## Contributing\n\nContributions are welcome. Please fork the repository, create a feature branch, and submit a pull request with a clear description of your changes. Adherence to PEP 8, type hinting, and comprehensive docstrings is required.\n\n## Recommended Extensions\n\nFuture extensions could include:\n-   **Additional Goodness-of-Fit Tests:** Implementing other statistical tests for Benford's Law conformity, such as the Kolmogorov-Smirnov test or Kuiper's test.\n-   **Visualization Module:** Creating a function that takes the final results and generates plots of the observed vs. expected distributions, similar to Figures 2-4 in the paper.\n-   **Automated Reporting:** Building a module that uses the generated tables and plots to automatically create a full PDF or HTML summary report of the findings.\n-   **Integration with Database:** Developing a data ingestion module to pull data directly from a SQL database instead of a CSV file.\n\n## License\n\nThis project is licensed under the MIT License. See the `LICENSE` file for details.\n\n## Citation\n\nIf you use this code or the methodology in your research, please cite the original paper:\n\n```bibtex\n@article{ausloos2025note,\n  title={{Note on pre-taxation reported data by UK FTSE-listed companies. A search for Benford's laws compatibility}},\n  author={Ausloos, Marcel and Sastroredjo, Probowo Erawan and Khrennikova, Polina},\n  journal={arXiv preprint arXiv:2509.09415},\n  year={2025}\n}\n```\n\nFor the implementation itself, you may cite this repository:\n```\nChirinda, C. (2025). A Python Implementation for Benford's Law Conformity Testing of Financial Data.\nGitHub repository: https://github.com/chirindaopensource/search_benford_law_compatibility\n```\n\n## Acknowledgments\n\n-   Credit to **Marcel Ausloos, Probowo Erawan Sastroredjo, and Polina Khrennikova** for their foundational research, which forms the entire basis for this computational replication.\n-   This project is built upon the exceptional tools provided by the open-source community. Sincere thanks to the developers of the scientific Python ecosystem, including **Pandas, NumPy, SciPy, and PyYAML**, whose work makes complex computational analysis accessible and robust.\n\n--\n\n*This README was generated based on the structure and content of `search_benford_law_compatibility_draft.ipynb` and follows best practices for research software documentation.*\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fchirindaopensource%2Fsearch_benford_law_compatibility","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fchirindaopensource%2Fsearch_benford_law_compatibility","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fchirindaopensource%2Fsearch_benford_law_compatibility/lists"}