{"id":49537882,"url":"https://github.com/sapsan14/water-quality-ee","last_synced_at":"2026-05-02T12:31:35.671Z","repository":{"id":351037811,"uuid":"1209217702","full_name":"sapsan14/water-quality-ee","owner":"sapsan14","description":"Estonian water quality ML — binary classification of Terviseamet open data, Jupyter + scikit-learn.","archived":false,"fork":false,"pushed_at":"2026-04-27T16:09:38.000Z","size":37946,"stargazers_count":1,"open_issues_count":1,"forks_count":0,"subscribers_count":0,"default_branch":"main","last_synced_at":"2026-04-27T16:25:14.338Z","etag":null,"topics":["classification","estonia","jupyter","ml","open-data","scikit-learn"],"latest_commit_sha":null,"homepage":null,"language":"Jupyter Notebook","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/sapsan14.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":"CITATION.cff","codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2026-04-13T07:56:15.000Z","updated_at":"2026-04-27T16:10:28.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/sapsan14/water-quality-ee","commit_stats":null,"previous_names":["sapsan14/water-quality-ee"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/sapsan14/water-quality-ee","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/sapsan14%2Fwater-quality-ee","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/sapsan14%2Fwater-quality-ee/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/sapsan14%2Fwater-quality-ee/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/sapsan14%2Fwater-quality-ee/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/sapsan14","download_url":"https://codeload.github.com/sapsan14/water-quality-ee/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/sapsan14%2Fwater-quality-ee/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":32534964,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-02T12:25:33.646Z","status":"ssl_error","status_checked_at":"2026-05-02T12:24:51.733Z","response_time":132,"last_error":"SSL_read: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["classification","estonia","jupyter","ml","open-data","scikit-learn"],"created_at":"2026-05-02T12:31:34.941Z","updated_at":"2026-05-02T12:31:35.658Z","avatar_url":"https://github.com/sapsan14.png","language":"Jupyter Notebook","funding_links":[],"categories":[],"sub_categories":[],"readme":"\u003cdiv align=\"center\"\u003e\n  \u003cimg src=\"frontend/public/logo.svg\" alt=\"H2O Atlas\" width=\"96\" height=\"96\"\u003e\n\n  # Water Quality Estonia\n\n  **Probabilistic risk estimator for water quality compliance**\\\n  *TalTech Machine Learning course project (Masinope, spring 2026)*\n\n  [![Tests](https://img.shields.io/github/actions/workflow/status/sapsan14/water-quality-ee/tests.yml?branch=main\u0026style=for-the-badge\u0026logo=githubactions\u0026logoColor=white\u0026label=tests)](https://github.com/sapsan14/water-quality-ee/actions/workflows/tests.yml)\n  [![Frontend CI](https://img.shields.io/github/actions/workflow/status/sapsan14/water-quality-ee/frontend-ci.yml?branch=main\u0026style=for-the-badge\u0026logo=githubactions\u0026logoColor=white\u0026label=frontend)](https://github.com/sapsan14/water-quality-ee/actions/workflows/frontend-ci.yml)\n  [![Python](https://img.shields.io/badge/python-3.11+-3776AB?style=for-the-badge\u0026logo=python\u0026logoColor=white)](https://www.python.org/downloads/)\n  [![License](https://img.shields.io/github/license/sapsan14/water-quality-ee?style=for-the-badge)](LICENSE)\n  [![Ruff](https://img.shields.io/badge/code_style-ruff-D7FF64?style=for-the-badge\u0026logo=ruff\u0026logoColor=white)](https://docs.astral.sh/ruff/)\n\n  [![scikit-learn](https://img.shields.io/badge/scikit--learn-F7931E?style=for-the-badge\u0026logo=scikitlearn\u0026logoColor=white)](https://scikit-learn.org/)\n  [![LightGBM](https://img.shields.io/badge/LightGBM-02569B?style=for-the-badge\u0026logo=microsoft\u0026logoColor=white)](https://lightgbm.readthedocs.io/)\n  [![Next.js](https://img.shields.io/badge/Next.js_16-000000?style=for-the-badge\u0026logo=nextdotjs\u0026logoColor=white)](https://nextjs.org/)\n  [![Jupyter](https://img.shields.io/badge/Jupyter-F37626?style=for-the-badge\u0026logo=jupyter\u0026logoColor=white)](https://jupyter.org/)\n\n  [![Data source](https://img.shields.io/badge/Data-Terviseamet_Open_Data-0063AF?style=for-the-badge)](https://vtiav.sm.ee/index.php/?active_tab_id=A)\n  [![Live](https://img.shields.io/badge/live-h2oatlas.ee-17b0ff?style=for-the-badge\u0026logo=globe\u0026logoColor=white)](https://h2oatlas.ee)\n  [![Colab](https://img.shields.io/badge/Open_in_Colab-F9AB00?style=for-the-badge\u0026logo=googlecolab\u0026logoColor=white)](https://colab.research.google.com/github/sapsan14/water-quality-ee/blob/main/notebooks/colab_quickstart.ipynb)\n\n  **[English]** | [Русский](README.ru.md)\n\n\u003c/div\u003e\n\n---\n\n## Overview\n\nEstonia has thousands of monitored water sites: coastal and inland swimming locations, public pools and SPAs, drinking water networks, and natural water sources. The national Health Board ([Terviseamet](https://vtiav.sm.ee)) publishes laboratory analysis results as open data.\n\nThis project builds a **binary classification model** that estimates **P(violation)** -- the probability that a water sample violates Estonian health norms -- based on 15 chemical and biological parameters plus engineered features. The model is trained on **69,536 samples** across four water domains (2021-2026).\n\n\u003e **What the model predicts:** probability that a sample's measurement profile matches historical violation patterns.\\\n\u003e **What it does NOT predict:** unmeasured contaminants, future water quality, causal reasons for contamination, or safety beyond the measured parameters.\\\n\u003e Full analysis: [`docs/ml_framing.md`](docs/ml_framing.md)\n\n### Key Results\n\n| Model | Recall (violations) | Precision | F1 | ROC-AUC |\n|-------|--------------------:|----------:|---:|--------:|\n| Logistic Regression | 0.827 | 0.454 | 0.586 | 0.936 |\n| Random Forest | 0.949 | 0.919 | 0.934 | 0.992 |\n| Gradient Boosting | 0.954 | 0.946 | 0.950 | 0.994 |\n| **LightGBM (temporal)** | **0.956** | **0.881** | **0.917** | **0.988** |\n\n**Priority metric: Recall on violations** -- a False Negative means predicting water is safe when it contains E. coli. Threshold is optimized via `best_threshold_max_recall_at_precision()` for decision support. Full report: [`docs/report.md`](docs/report.md)\n\n### Live Demo\n\n**[h2oatlas.ee](https://h2oatlas.ee)** -- interactive map of per-location water quality with two layers: **official Terviseamet status** and **ML risk assessment** (P(violation) from 4 models). Supports three languages (RU/ET/EN), dark mode, and mobile-first responsive design.\n\n---\n\n## Table of Contents\n\n- [Quick Start](#quick-start)\n- [Project Structure](#project-structure)\n- [Data](#data)\n- [Notebooks](#notebooks)\n- [Models \u0026 Evaluation](#models--evaluation)\n- [Citizen Service \u0026 Frontend](#citizen-service--frontend)\n- [Architecture](#architecture)\n- [Documentation](#documentation)\n- [Google Colab](#google-colab)\n- [Tests](#tests)\n- [Course Requirements](#course-requirements)\n- [License](#license)\n- [Citation](#citation)\n- [Acknowledgments](#acknowledgments)\n\n---\n\n## Quick Start\n\n```bash\n# 1. Install\npip install -r requirements.txt\npip install -e .                  # editable install: imports work from any cwd\n\n# 2. Download \u0026 parse open data\npython src/data_loader.py         # downloads XML, caches to data/raw/, prints sample\n\n# 3. Run notebooks\njupyter notebook                  # open notebooks/01_eda_supluskoha.ipynb\n```\n\n**Requirements:** Python 3.10+ (3.11+ recommended). For LightGBM + SHAP: `pip install lightgbm shap`.\n\n---\n\n## Project Structure\n\n```\nwater-quality-ee/\n├── src/                           # Core Python modules\n│   ├── data_loader.py             #   XML download, parsing, domain loaders\n│   ├── features.py                #   Feature engineering, ratio-to-norm, imputation\n│   ├── evaluate.py                #   Metrics, ROC, threshold optimization, SHAP\n│   ├── county_infer.py            #   Location -\u003e county inference + geocoding\n│   ├── terviseamet_reference_coords.py  # Official coordinate mappings\n│   └── audit/                     #   Data quality validation modules\n│\n├── notebooks/                     # Jupyter notebooks (run 01 -\u003e 07 in order)\n│   ├── colab_quickstart.ipynb     #   Google Colab setup\n│   ├── 01_eda_supluskoha.ipynb    #   EDA: swimming locations\n│   ├── 02_eda_full.ipynb          #   EDA: all four domains\n│   ├── 03_preprocessing.ipynb     #   Feature engineering, train/test split\n│   ├── 04_models.ipynb            #   LR + RF + GB + GridSearchCV\n│   ├── 05_evaluation.ipynb        #   Confusion matrix, ROC, feature importance\n│   ├── 06_advanced_models.ipynb   #   LightGBM, temporal split, calibration, SHAP\n│   └── 07_data_gaps_audit.ipynb   #   Label-vs-norms divergence analysis\n│\n├── citizen-service/               # Data pipeline (h2oatlas.ee backend)\n│   └── scripts/                   #   Snapshot builder, coordinate enrichment\n│\n├── frontend/                      # Next.js 16 + React 19 + TypeScript\n│   ├── app/                       #   App Router components\n│   └── public/                    #   Logo, favicon, snapshot data\n│\n├── tests/                         # pytest test suite (10 modules)\n├── docs/                          # Project documentation (19 files)\n├── data/                          # Local data (raw XML, processed, reference)\n│\n├── pyproject.toml                 # Package metadata\n├── requirements.txt               # Python dependencies\n├── CITATION.cff                   # Academic citation\n├── LICENSE                        # MIT License\n└── DATA_SOURCES.md                # Comprehensive data source catalog\n```\n\n---\n\n## Data\n\n**Source:** [Terviseamet](https://vtiav.sm.ee/index.php/?active_tab_id=A) (Estonian Health Board) open data -- XML format, 2021-2026.\n\n| Domain | Estonian | Description | Samples |\n|--------|---------|-------------|--------:|\n| `supluskoha` | Supluskohad | Swimming locations (sea, lakes) | ~4,000 |\n| `veevark` | Veevargid | Drinking water networks | ~45,000 |\n| `basseinid` | Basseinid | Swimming pools \u0026 SPAs | ~10,000 |\n| `joogivesi` | Joogiveeallikad | Drinking water sources | ~10,000 |\n\n**15 measured parameters:** E. coli, enterococci, coliforms, pH, turbidity, color, iron, manganese, nitrates, nitrites, ammonium, fluoride, free chlorine, combined chlorine, pseudomonas. Full descriptions: [`docs/parametry.md`](docs/parametry.md)\n\n**Key conventions:**\n- `compliant` = 1 (passes norms) or 0 (violation); derived from Terviseamet `hinnang` field\n- `location_key` -- normalized location identifier; always use instead of raw `location` (handles inter-year renaming). See [`src/data_loader.py`](src/data_loader.py)\n- Estonian number format (comma as decimal separator) handled automatically in the parser\n\nFull data documentation: [`DATA_SOURCES.md`](DATA_SOURCES.md)\n\n---\n\n## Notebooks\n\nNotebooks are numbered sequentially and should be run in order:\n\n| # | Notebook | Purpose |\n|---|----------|---------|\n| 00 | [`polnoye_rukovodstvo`](notebooks/00_polnoye_rukovodstvo.ipynb) | End-to-end walkthrough (optional) |\n| 01 | [`eda_supluskoha`](notebooks/01_eda_supluskoha.ipynb) | EDA for swimming locations |\n| 02 | [`eda_full`](notebooks/02_eda_full.ipynb) | Full EDA across all four domains |\n| 03 | [`preprocessing`](notebooks/03_preprocessing.ipynb) | `build_dataset`, train/test split, impute/scale |\n| 04 | [`models`](notebooks/04_models.ipynb) | LR + RF + GradientBoosting + GridSearchCV |\n| 05 | [`evaluation`](notebooks/05_evaluation.ipynb) | Confusion matrix, ROC, feature importance |\n| 06 | [`advanced_models`](notebooks/06_advanced_models.ipynb) | LightGBM + temporal split + calibration + SHAP |\n| 07 | [`data_gaps_audit`](notebooks/07_data_gaps_audit.ipynb) | Label-vs-norms divergence analysis |\n\n---\n\n## Models \u0026 Evaluation\n\n**Task:** Binary classification -- `compliant` (1 = pass, 0 = violation)\\\n**Class imbalance:** ~12% violations across the full corpus\\\n**Priority:** Minimize False Negatives (FN = \"predicted safe, actually contaminated\")\n\nFour models are trained and compared:\n\n1. **Logistic Regression** -- interpretable baseline\n2. **Random Forest** -- primary model, feature importance\n3. **Gradient Boosting** (sklearn) -- high-accuracy ensemble\n4. **LightGBM** -- best model: native missing-value handling, temporal split validation, SHAP explanations\n\n**Evaluation framework (4 levels):**\n\n| Level | Question | Metric |\n|:-----:|----------|--------|\n| 1 | Does the model separate classes? | ROC-AUC |\n| 2 | What types of errors? | Precision / Recall |\n| 3 | Are probabilities calibrated? | Calibration curve |\n| 4 | Why this prediction? | SHAP values |\n\nFull metrics guide: [`docs/ml_metrics_guide.md`](docs/ml_metrics_guide.md)\n\n---\n\n## Citizen Service \u0026 Frontend\n\nThe project includes a public citizen-facing service deployed at **[h2oatlas.ee](https://h2oatlas.ee)**:\n\n- **Two information layers:** official Terviseamet compliance status + ML risk assessment (P(violation) from 4 models)\n- **Interactive map** with Leaflet clustering, domain-specific icons, and county boundaries\n- **Per-location detail:** latest sample, measurements, history, parameter explanations, SHAP-based risk factors\n- **Three languages:** Russian, Estonian, English (auto-detected + user toggle)\n- **Dark mode** with system preference detection and localStorage persistence\n- **Mobile-first** responsive design (bottom sheet, safe-area insets, touch-optimized)\n\n| Component | Stack | Deployment |\n|-----------|-------|------------|\n| Data pipeline | Python + GitHub Actions (scheduled) | CI/CD |\n| Web frontend | Next.js 16 + React 19 + TypeScript | Cloudflare Pages |\n\nDocumentation: [`citizen-service/README.md`](citizen-service/README.md) | [`frontend/README.md`](frontend/README.md)\n\n---\n\n## Architecture\n\n```mermaid\nflowchart LR\n    subgraph Data[\"Data Layer\"]\n        XML[\"Terviseamet XML\\n(open data)\"]\n        DL[\"data_loader.py\\ndownload + parse\"]\n        XML --\u003e DL\n    end\n\n    subgraph ML[\"ML Pipeline\"]\n        FE[\"features.py\\nengineer features\"]\n        MOD[\"Models\\nLR / RF / GB / LightGBM\"]\n        EV[\"evaluate.py\\nmetrics + threshold\"]\n        DL --\u003e FE --\u003e MOD --\u003e EV\n    end\n\n    subgraph Citizen[\"Citizen Service\"]\n        SNAP[\"build_citizen_snapshot.py\\nsnapshot.json\"]\n        FE2[\"Next.js Frontend\\nh2oatlas.ee\"]\n        EV --\u003e SNAP\n        SNAP --\u003e FE2\n    end\n\n    subgraph CI[\"CI/CD\"]\n        GH[\"GitHub Actions\\ntests + snapshot + coordinates\"]\n        GH -.-\u003e DL\n        GH -.-\u003e SNAP\n    end\n```\n\n---\n\n## Documentation\n\n| Document | Description |\n|----------|-------------|\n| [`docs/report.md`](docs/report.md) | Final course report: EDA, methodology, results, limitations |\n| [`docs/ml_framing.md`](docs/ml_framing.md) | What the model predicts vs. what it cannot |\n| [`docs/ml_metrics_guide.md`](docs/ml_metrics_guide.md) | 4-level metrics guide: ROC-AUC, Precision/Recall, Calibration, SHAP |\n| [`docs/model_card.md`](docs/model_card.md) | Model Card (Mitchell et al. 2019) — intended use, metrics, caveats |\n| [`docs/datasheet.md`](docs/datasheet.md) | Datasheet (Gebru et al. 2021) — data provenance and maintenance |\n| [`docs/ai_act_self_assessment.md`](docs/ai_act_self_assessment.md) | EU AI Act voluntary self-assessment (risk tier + triggers) |\n| [`docs/parametry.md`](docs/parametry.md) | Water parameter descriptions, health effects, norms |\n| [`docs/normy.md`](docs/normy.md) | Regulatory thresholds by parameter and domain |\n| [`docs/glosarij.md`](docs/glosarij.md) | Terminology glossary (RU / ET / EN) |\n| [`docs/learning_journey.md`](docs/learning_journey.md) | Project learning narrative and discoveries |\n| [`docs/phase_10_findings.md`](docs/phase_10_findings.md) | Data-quality audit results (69,536 samples) |\n| [`docs/data_gaps.md`](docs/data_gaps.md) | Label-vs-norms divergence analysis |\n| [`docs/terviseamet_inquiry.md`](docs/terviseamet_inquiry.md) | Draft inquiry to Terviseamet (with audit numbers) |\n| [`DATA_SOURCES.md`](DATA_SOURCES.md) | Comprehensive data source catalog |\n\n### Regulatory posture\n\nh2oatlas.ee is a public visualisation layered on top of already-public Terviseamet data. In its current configuration it is **not** a high-risk AI system under EU AI Act Annex III — it is neither a safety component nor integrated into any operational or regulatory decision loop. We fulfil Art 50 (transparency) through UI disclaimers and adopt the high-risk documentation stack voluntarily: Model Card, Datasheet, per-domain metrics, drift monitor (Phase 2), human-oversight tree (Phase 2), FRIA-light (Phase 2), and cryptographically signed snapshot provenance (Phase 3). Full reasoning and the list of triggers that would move the system into high-risk: [`docs/ai_act_self_assessment.md`](docs/ai_act_self_assessment.md).\n\n---\n\n## Google Colab\n\nUse **[`notebooks/colab_quickstart.ipynb`](https://colab.research.google.com/github/sapsan14/water-quality-ee/blob/main/notebooks/colab_quickstart.ipynb)** to run the project in the cloud:\n\n1. Set `REPO_URL`, run clone + `pip install -r requirements.txt` + `pip install -e .`\n2. Open notebooks `01` through `07` from the file browser\n3. For notebook `06`: additionally install `lightgbm` and `shap`\n\nNote: scikit-learn models do **not** use GPU. Colab T4 runtime won't accelerate LR/RF/GB.\n\n---\n\n## Tests\n\n```bash\npip install -e . pytest \u0026\u0026 pytest tests/\n```\n\n10 test modules covering XML parsing, feature engineering, threshold optimization, geocoding, coordinate resolution, and data quality audits. CI runs automatically on push/PR via [`.github/workflows/tests.yml`](.github/workflows/tests.yml).\n\n---\n\n## Course Requirements\n\nTalTech Masinope (Machine Learning) -- all required components:\n\n- [x] Problem statement and dataset rationale\n- [x] Exploratory data analysis with visualization\n- [x] Data preprocessing and feature engineering\n- [x] Training of 2+ models with hyperparameter tuning\n- [x] Metric comparison and model selection\n- [x] Result interpretation (SHAP, feature importance)\n- [x] Final report and presentation\n\n---\n\n## License\n\nThis project is licensed under the [MIT License](LICENSE).\n\n---\n\n## Citation\n\nIf you use this software or data pipeline in your research, please cite:\n\n```bibtex\n@software{sokolov2026waterquality,\n  author    = {Sokolov, Anton},\n  title     = {Water Quality Estonia: Probabilistic Risk Estimator},\n  year      = {2026},\n  url       = {https://github.com/sapsan14/water-quality-ee},\n  version   = {0.1.0},\n  note      = {TalTech Masinope course project}\n}\n```\n\nSee [`CITATION.cff`](CITATION.cff) for machine-readable citation metadata.\n\n---\n\n## Acknowledgments\n\n- **[Terviseamet](https://vtiav.sm.ee)** (Estonian Health Board) -- open data provider\n- **[Tallinn University of Technology](https://taltech.ee)** -- Masinope (Machine Learning) course, spring 2026\n- **Estonian open data initiative** -- making government data accessible for research and public benefit\n\n---\n\n\u003cdiv align=\"center\"\u003e\n  \u003csub\u003eBuilt with care for Estonian water safety\u003c/sub\u003e\n\u003c/div\u003e\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fsapsan14%2Fwater-quality-ee","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fsapsan14%2Fwater-quality-ee","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fsapsan14%2Fwater-quality-ee/lists"}