{"id":47320119,"url":"https://github.com/martanto/eruption-forecast","last_synced_at":"2026-06-27T12:00:53.301Z","repository":{"id":343197395,"uuid":"1072602598","full_name":"martanto/eruption-forecast","owner":"martanto","description":"Volcanic eruption forecasting using seismic data. Train machine learning models and predict probability of eruptions","archived":false,"fork":false,"pushed_at":"2026-06-16T07:15:47.000Z","size":5206,"stargazers_count":0,"open_issues_count":3,"forks_count":0,"subscribers_count":0,"default_branch":"master","last_synced_at":"2026-06-16T09:15:02.236Z","etag":null,"topics":["eruption","forecasting","forecasting-models","machine-learning","seismic","seismic-processing","volcano","volcano-seismology","volcanoes","volcanology"],"latest_commit_sha":null,"homepage":"https://pypi.org/project/eruption-forecast/","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"other","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/martanto.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2025-10-09T00:41:21.000Z","updated_at":"2026-06-12T12:58:38.000Z","dependencies_parsed_at":null,"dependency_job_id":null,"html_url":"https://github.com/martanto/eruption-forecast","commit_stats":null,"previous_names":["martanto/eruption-forecast"],"tags_count":9,"template":false,"template_full_name":null,"purl":"pkg:github/martanto/eruption-forecast","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/martanto%2Feruption-forecast","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/martanto%2Feruption-forecast/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/martanto%2Feruption-forecast/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/martanto%2Feruption-forecast/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/martanto","download_url":"https://codeload.github.com/martanto/eruption-forecast/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/martanto%2Feruption-forecast/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":34852282,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-05-26T15:22:16.424Z","status":"online","status_checked_at":"2026-06-27T02:00:06.362Z","response_time":126,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["eruption","forecasting","forecasting-models","machine-learning","seismic","seismic-processing","volcano","volcano-seismology","volcanoes","volcanology"],"created_at":"2026-03-17T17:21:05.739Z","updated_at":"2026-06-27T12:00:53.289Z","avatar_url":"https://github.com/martanto.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"# eruption-forecast\n\n[![Version](https://img.shields.io/pypi/v/eruption-forecast?label=version)](https://pypi.org/project/eruption-forecast/)\n[![Python](https://img.shields.io/pypi/pyversions/eruption-forecast?label=python)](https://pypi.org/project/eruption-forecast/)\n[![License](https://img.shields.io/pypi/l/eruption-forecast?label=license)](https://pypi.org/project/eruption-forecast/)\n[![Status](https://img.shields.io/badge/status-active%20development-orange)](https://github.com/martanto/eruption-forecast)\n[![Downloads](https://static.pepy.tech/personalized-badge/eruption-forecast?period=total\u0026units=INTERNATIONAL_SYSTEM\u0026left_color=BLUE\u0026right_color=GREEN\u0026left_text=downloads)](https://pepy.tech/projects/eruption-forecast)\n\nProcess raw seismic tremor, extract time-series features, train multi-seed classifier ensembles, and produce probabilistic volcanic eruption forecasts. Forked from [ddempsey/whakaari](https://github.com/ddempsey/whakaari) and substantially extended.\n\n![Forecast example — Scenario 8, rolling 6h window, prediction 2025-07-27 to 2025-08-22](https://raw.githubusercontent.com/martanto/eruption-forecast/master/assets/forecast_2025-07-27_2025-08-22.png)\n\n## References and Acknowledgments\n\u003e Dempsey, D. E., Cronin, S. J., Mei, S., \u0026 Kempa-Liehr, A. W. (2020). Automatic precursor recognition and real-time forecasting\n\u003e of sudden explosive volcanic eruptions at Whakaari, New Zealand. Nature Communications, 11(1), 1–8. https://doi.org/10.1038/s41467-020-17375-2\n\u003e This model implements a time series feature engineering and classification workflow that issues eruption alerts\n\u003e based on real-time tremor data. https://github.com/ddempsey/whakaari\n\n\u003e Ardid, A., Dempsey, D., Caudron, C., Cronin, S., Kennedy, B., Girona, T., Roman, D., Miller, C., Potter,\n\u003e S., Lamb, O. D., Martanto, A., Cubuk-Sabuncu, Y., Cabrera, L., Ruiz, S., Contreras, R., Pacheco, J., Mora,\n\u003e M. M., \u0026 De Angelis, S. (2025). Ergodic seismic precursors and transfer learning for short term eruption\n\u003e forecasting at data scarce volcanoes. Nature Communications , 16(1), 1–12. https://doi.org/10.1038/s41467-025-56689-x\n\n\u003e Ardid, A., Dempsey, D., Caudron, C., \u0026 Cronin, S. (2022). Seismic precursors to the Whakaari 2019 phreatic eruption\n\u003e are transferable to other eruptions and volcanoes. Nature Communications, 13(1), 2002. https://doi.org/10.1038/s41467-022-29681-y\n\n\u003e Endo, E. T., \u0026 Murray, T. L. (1991). Real-time Seismic Amplitude Measurement (RSAM):\n\u003e a volcano monitoring and prediction tool. Bulletin of Volcanology, 53, 533–545.\n\n\u003e Caudron, C., et al., 2019, Change in seismic attenuation as a long-term precursor of\n\u003e gas-driven eruptions: Geology, https://doi.org/10.1130/G46107.1\n\n\u003e Rey-Devesa, P., Prudencio, J., Benítez, C., Bretón, M., Plasencia, I., León, Z., Ortigosa,\n\u003e F., Gutiérrez, L., Arámbula-Mendoza, R., \u0026 Ibáñez, J. M. (2023).\n\u003e Tracking volcanic explosions using Shannon entropy at Volcán de Colima.\n\u003e Scientific Reports, 13(1), 1–11. https://doi.org/10.1038/s41598-023-36964-x\n\n\u003e Christ, M., Braun, N., Neuffer, J., \u0026 Kempa-Liehr, A. W. (2018). Time Series FeatuRe Extraction on basis of\n\u003e Scalable Hypothesis tests (tsfresh – A Python package). Neurocomputing, 307, 72–77. https://doi.org/10.1016/j.neucom.2018.03.067\n\n\u003e Lei, Y., \u0026 Wu, Z. (2020). Time series classification based on statistical features.\n\u003e Eurasip Journal on Wireless Communications and Networking, 2020(1). https://doi.org/10.1186/s13638-020-1661-4\n\n\u003e Chardot, L., Jolly, A. D., Kennedy, B. M., Fournier, N., \u0026 Sherburn, S. (2015). Using\n\u003e volcanic tremor for eruption forecasting at White Island volcano (Whakaari), New Zealand.\n\u003e Journal of Volcanology and Geothermal Research, 302, 11–23.\n\u003e https://doi.org/10.1016/j.jvolgeores.2015.06.001\n\n\u003e Time-series feature analysis and eruption forecasting for volcano data. Successor package to Whakaari.\n\u003e This model implements a time series feature engineering and classification workflow that issues eruption\n\u003e alerts based on real-time tremor data. https://github.com/ddempsey/puia\n\n## Important Disclaimers\n\n**This software is intended for research purposes only.**\n\n1. **Probabilistic Predictions**: This eruption forecast model provides probabilistic predictions of future volcanic activity, NOT deterministic guarantees. Predictions should be interpreted as likelihood estimates based on historical seismic patterns.\n\n2. **No Guarantee of Accuracy**: This model is **not guaranteed to predict every future eruption**. Volcanic systems are complex and can exhibit unexpected behavior. False negatives (missed eruptions) and false positives (false alarms) are possible.\n\n3. **Software Limitations**: This software is **not guaranteed to be free of bugs or errors**. Users should validate results independently and use this tool as one component of a comprehensive volcano monitoring strategy.\n\n4. **Not for Operational Use**: This package is a research tool and should not be used as the sole basis for public safety decisions, evacuation orders, or emergency response without expert volcanological assessment.\n\n5. **Expert Interpretation Required**: Results should always be interpreted by qualified volcanologists familiar with the specific volcano being monitored.\n\n**Always consult with local volcano observatories and follow official warnings from government agencies.**\n\n---\n\n## Table of Contents\n- [References and Acknowledgments](#references-and-acknowledgments)\n- [Features](#features)\n- [Package Architecture](#package-architecture)\n- [Pipeline Overview](#pipeline-overview)\n- [Installation](#installation)\n- [Data Sources](#data-sources)\n- [Quick Start](#quick-start)\n- [Common Patterns](#common-patterns)\n- [Supported Classifiers](#supported-classifiers)\n- [Cross-Validation Strategies](#cross-validation-strategies)\n- [Output Directory Structure](#output-directory-structure)\n- [Requirements](#requirements)\n- [Development](#development)\n- [Contributing](#contributing)\n- [License](#license)\n\n**Detailed documentation** — the wiki at https://github.com/martanto/eruption-forecast/wiki is the single source of truth (also mirrored under [`wiki/`](wiki/)):\n\n- [Getting Started](wiki/Getting-Started.md) — Prerequisites, install, dev commands\n- [Pipeline Walkthrough](wiki/Pipeline-Walkthrough.md) — Research (`main.py`) + Scenarios (`scenarios.py`) workflows\n- [Training Workflow](wiki/Training-Workflow.md) — Classifiers, CV strategies, imbalance handling\n- [Prediction Workflow](wiki/Prediction-Workflow.md) — Forecast outputs and consensus probabilities\n- [Evaluation Workflow](wiki/Evaluation-Workflow.md) — `MetricsEnsemble`, `ClassifierComparator`\n- [Explanation Workflow](wiki/Explanation-Workflow.md) — `ExplanationModel`, `ExplainerEnsemble`, per-seed SHAP\n- [Configuration](wiki/Configuration.md) — YAML save/replay, Telegram notifications, logging\n- [Visualization](wiki/Visualization.md) — Plot catalog and output paths\n- [Output Structure](wiki/Output-Structure.md) — Full directory tree\n- [Architecture](wiki/Architecture.md) — Package layout and class relationships\n- [API Reference](wiki/API-Reference.md) — Every public class with parameter tables\n\n---\n\n## Features\n\n- **Tremor Calculation** — RSAM, DSAR, and Shannon Entropy across configurable frequency bands, from SDS archives or FDSN web services (with transparent local caching).\n- **Label Building** — Standard sliding-window (`LabelBuilder`) or per-eruption (`DynamicLabelBuilder`) generation from known eruption dates.\n- **Feature Extraction** — tsfresh feature engineering on windowed tremor matrices, with FDR-controlled selection (tsfresh statistical filter; RandomForest permutation importance available as an alternative).\n- **Multi-seed Training** — 11 classifier families (`rf`, `gb`, `xgb`, `svm`, `lr`, `nn`, `dt`, `knn`, `nb`, `voting`, `lite-rf`), three CV strategies, automatic imbalance handling, and per-seed `GridSearchCV`.\n- **Ensemble Packaging** — `SeedEnsemble` bundles every seed for one classifier; `ClassifierEnsemble` bundles multiple classifiers; both implement the sklearn `BaseEstimator + ClassifierMixin` interface.\n- **Probabilistic Forecasting** — `PredictionModel` produces per-seed, per-classifier, and consensus probabilities with uncertainty bands over an unlabelled window grid.\n- **Evaluation + Comparison** — `EvaluationModel` runs metrics over a training or prediction reuse mode; `MetricsEnsemble` persists `(n_samples, n_seeds)` `y_proba` / `y_pred` matrices and keeps per-seed metric tables in memory; `ClassifierComparator` ranks classifiers head-to-head.\n- **Model Explanation** — `ExplanationModel` produces per-seed SHAP explanations over the fitted ensemble via `ExplainerEnsemble` (tree classifiers only — RF / `lite-rf` / GB / XGB). Outputs include per-classifier `ClassifierExplanation_*.pkl`, per-seed bar / beeswarm plots, and per-eruption highest-probability waterfall plots.\n- **Content-Addressable Caching** — `TrainingModel` and `PredictionModel` cache their fitted state under `{output_dir}/cache/` so repeated runs with identical kwargs short-circuit.\n- **Config Round-Trip** — `fm.save_config()` → YAML → `ForecastModel.from_config(path).run()` replays a full pipeline. Every stage model (`TrainingModel`, `PredictionModel`, `EvaluationModel`, `ExplanationModel`) also auto-saves its own per-stage `*.config.yaml` at the end of `fit()` / `forecast()` / `evaluate()` / `explain()`.\n- **Telegram Notifications** — `@notify` decorator + `send_telegram_notification()` for start/finish/error messages and file attachments.\n- **Multi-processing** — `n_jobs` (outer seed workers) × `n_grids` (inner `GridSearchCV` / `FeatureSelector` workers) parallelism, clamped to `total_cpu - 2` automatically.\n\n## Package Architecture\n\n```\nsrc/eruption_forecast/\n├── __init__.py, logger.py, data_container.py\n├── config/        forecast_config, training_config, prediction_config,\n│                  evaluation_config, explanation_config, constants\n├── dataclass/     station_data, classifier_ensemble_summary,\n│                  classifier_explanation\n├── decorators/    notify, decorator_class\n├── ensemble/      base_ensemble, seed_ensemble, classifier_ensemble,\n│                  metrics_ensemble, explainer_ensemble\n├── features/      tremor_matrix_builder, features_builder, feature_selector\n├── label/         label_builder, dynamic_label_builder, label_data, label_plots\n├── model/         base_model, cache_model, forecast_model,\n│                  training_model, prediction_model, evaluation_model,\n│                  explanation_model, classifier_model, classifier_comparator\n├── plots/         styles, tremor_plots, feature_plots, forecast_plots,\n│                  evaluation_plots, explanation_plots\n├── sources/       base, sds, fdsn\n├── tremor/        calculate_tremor, rsam, dsar, shannon_entropy, tremor_data\n└── utils/         array, dataframe, date_utils, formatting, ml,\n                   pathutils, validation, window\n```\n\n\u003e Full directory tree, class relationships, and per-component details: [wiki/Architecture.md](wiki/Architecture.md)\n\n## Pipeline Overview\n\n```\n        ┌──────────────┐     ┌────────────────────┐    ┌─────────────────┐\n        │  Seismic     │     │  CalculateTremor   │    │   TremorData    │\n        │  archive     │ ──► │  rsam/dsar/        │ ─► │  (CSV wrapper)  │\n        │  (SDS|FDSN)  │     │  entropy + bands   │    │                 │\n        └──────────────┘     └────────────────────┘    └────────┬────────┘\n                                                                │\n        ┌─────────── feature pipeline ──────────────────────────┴─────┐\n        │   LabelBuilder / DynamicLabelBuilder                        │\n        │   → TremorMatrixBuilder → FeaturesBuilder (tsfresh)         │\n        │   → FeatureSelector (tsfresh FDR + RF importance)           │\n        └────────────────────────────┬────────────────────────────────┘\n                                     ▼\n                       ┌────────────────────────┐\n                       │     TrainingModel      │\n                       │   build_label →        │\n                       │   extract_features →   │\n                       │   fit (N seeds × M cv) │\n                       └──┬───────────────────┬─┘\n                          │                   │\n                          ▼                   ▼\n                ┌────────────────┐    ┌────────────────────────┐\n                │ SeedEnsemble × │ ─► │   ClassifierEnsemble   │\n                │ N classifiers  │    │  (all SeedEnsembles)   │\n                └────────────────┘    └──────────┬─────────────┘\n                                                 │\n                                                 ▼\n                                  ┌──────────────────────────────┐\n                                  │      PredictionModel         │\n                                  │  build_label →               │\n                                  │  extract_features →          │\n                                  │  forecast (per-seed proba)   │\n                                  └──────────┬───────────────────┘\n                                             │\n            ┌────────────────────────────────┴──────────────────────┐\n            ▼                                                       ▼\n   ┌──────────────────────┐                          ┌────────────────────────┐\n   │   EvaluationModel    │                          │  forecast-results_     │\n   │  training | predict  │  ── MetricsEnsemble ──►  │  *.csv + forecast      │\n   │                      │                          │  PNG/PDF               │\n   └──────────┬───────────┘                          └────────────────────────┘\n              │ writes (n_samples × n_seeds) y_proba / y_pred CSVs\n              │\n              ▼\n   ┌──────────────────────┐\n   │ ClassifierComparator │   ranking_*.csv + comparison figures\n   └──────────────────────┘\n\n   ┌──────────────────────────────────────────────────────────────┐\n   │     ExplanationModel    (BaseModel + CacheModel)             │\n   │     ExplainerEnsemble                                        │\n   │       ── per-seed shap.TreeExplainer (RF / lite-rf / GB /XGB)│\n   │       ── ClassifierExplanation.pkl per classifier            │\n   │       ── per-seed bar + beeswarm, per-eruption waterfall     │\n   └──────────────────────────────────────────────────────────────┘\n```\n\n`ForecastModel.calculate() → train() → predict() → evaluate() → explain()` is the fluent entry point. Each stage caches and persists, so a repeated run with identical kwargs short-circuits via the on-disk cache.\n\n## Installation\n\nThis project uses [uv](https://docs.astral.sh/uv/) as the package manager.\n\n```bash\n# Clone the repository\ngit clone https://github.com/martanto/eruption-forecast.git\ncd eruption-forecast\n\n# Install dependencies\nuv sync\n\n# Install with dev dependencies (ruff, ty, pytest)\nuv sync --group dev\n```\n\n## Data Sources\n\nThe package reads seismic data from two sources, both routed through `CalculateTremor`.\n\n### SDS — SeisComP Data Structure\n\nSDS is the layout used by [SeisComP](https://www.seiscomp.de/) to store waveform data portably. See the [official specification](https://www.seiscomp.de/seiscomp3/doc/applications/slarchive/SDS.html) for full details.\n\n**Directory layout:**\n\n```\n\u003csds_dir\u003e/\n└── YEAR/\n    └── NET/\n        └── STA/\n            └── CHAN.TYPE/\n                └── NET.STA.LOC.CHAN.TYPE.YEAR.DAY\n```\n\n**Example** for network `VG`, station `OJN`, channel `EHZ`, day 075 of 2025:\n\n```\n/data/\n└── 2025/\n    └── VG/\n        └── OJN/\n            └── EHZ.D/\n                └── VG.OJN.00.EHZ.D.2025.075\n```\n\n| Field | Description | Example |\n|-------|-------------|---------|\n| `YEAR` | Four-digit year | `2025` |\n| `NET` | Network code | `VG` |\n| `STA` | Station code | `OJN` |\n| `CHAN` | Channel code | `EHZ` |\n| `LOC` | Location code (may be empty) | `00` |\n| `TYPE` | Data type (`D` = waveform data) | `D` |\n| `DAY` | Three-digit day-of-year | `075` |\n\nFiles are miniSEED format.\n\n### FDSN — Web Service\n\nFDSN downloads waveform data from any FDSN-compatible web service (IRIS, GEOFON, etc.) and caches it locally as SDS miniSEED, so subsequent runs skip the network.\n\n```python\nfrom eruption_forecast import ForecastModel\n\nfm = ForecastModel(station=\"OJN\", channel=\"EHZ\", network=\"VG\", location=\"00\").calculate(\n    start_date=\"2025-01-01\", end_date=\"2025-01-31\",\n    source=\"fdsn\", client_url=\"https://service.iris.edu\",\n)\n```\n\n\u003e SDS vs FDSN read-path diagram and adapter internals: [wiki/Data-Sources.md](wiki/Data-Sources.md)\n\n---\n\n## Quick Start\n\nEnd-to-end pipeline from raw seismic data to forecast — train on Jan–Jul, predict Jul–Aug, evaluate against known eruption dates.\n\n```python\nfrom eruption_forecast import ForecastModel\n\nfm = ForecastModel(\n    station=\"OJN\",\n    channel=\"EHZ\",\n    network=\"VG\",\n    location=\"00\",\n    day_to_forecast=2,           # look-ahead window in days\n    root_dir=\"/path/to/project\",\n    n_jobs=4,\n    verbose=True,\n)\n\n(\n    fm.calculate(\n        start_date=\"2025-01-01\", end_date=\"2025-08-31\",\n        source=\"sds\", sds_dir=\"/path/to/sds\",\n        methods=[\"rsam\", \"dsar\", \"entropy\"],\n        plot_daily=True, save_plot=True,\n    )\n    .train(\n        start_date=\"2025-01-01\", end_date=\"2025-07-26\",\n        eruption_dates=[\n            \"2025-03-20\",\n            \"2025-04-22\",\n            \"2025-05-18\",\n            \"2025-06-17\",\n            \"2025-07-07\",\n        ],\n        window_step=6, window_step_unit=\"hours\",\n        classifiers=[\"lite-rf\", \"rf\", \"gb\", \"xgb\"],\n        cv_strategy=\"shuffle-stratified\",\n        cv_splits=5,\n        seeds=25,\n        top_n_features=20,\n        select_tremor_columns=[\"rsam_f2\", \"rsam_f3\", \"rsam_f4\", \"dsar_f3-f4\", \"entropy\"],\n        resample_method=\"auto\",\n        use_cache=True,\n    )\n    .predict(\n        start_date=\"2025-07-27\", end_date=\"2025-08-22\",\n        window_step=10, window_step_unit=\"minutes\",\n        plot_threshold=0.7,\n        plot_pdf=True,\n    )\n    .evaluate(\n        model=\"prediction\",\n        eruption_dates=[\"2025-08-02\"],   # falls back to train() dates when omitted\n        plot_aggregate=True,\n    )\n    .explain(\n        model=\"prediction\",\n        eruption_dates=[\"2025-08-02\"],   # falls back to train() dates when omitted\n        plot_per_seed=True,\n        plot_aggregate=True,             # per-classifier aggregate bar + beeswarm\n        max_display=20,\n    )\n)\n\n# Pipeline outputs\nprint(fm.results)                         # forecast DataFrame\nprint(fm.evaluation_results)              # per-classifier per-seed metrics\nprint(fm.ExplanationModel.explanations)   # list[ClassifierExplanation]\nprint(fm.TrainingModel.classifier_ensemble_path)\n```\n\n**What this pipeline does:**\n\n1. **Calculate tremor** — RSAM, DSAR, and Shannon Entropy from raw seismic, with outlier removal and daily plots.\n2. **Train** — build labels around the 5 eruption dates, extract tsfresh features over a 6-hour window grid, fit 25 seeds × 4 classifiers under stratified-shuffle CV with `\"auto\"` imbalance handling.\n3. **Predict** — apply the bundled `ClassifierEnsemble` to a fresh 10-minute window grid over the forecast period and emit per-classifier + consensus probabilities.\n4. **Evaluate** — score the forecast against the held-out eruption date by writing `(n_samples, n_seeds)` `y_proba` / `y_pred` matrices per classifier and aggregate metric plots; cross-classifier ranking via `ClassifierComparator`.\n5. **Explain** — produce per-seed SHAP explanations for the tree classifiers in the ensemble, bundled into `ClassifierExplanation_*.pkl` per classifier and rendered as per-seed bar / beeswarm, per-classifier aggregate bar / beeswarm (NaN-padded union feature space), and per-eruption waterfall plots.\n\nSee [`main.py`](main.py) for the full working example and [`scenarios.py`](scenarios.py) for the multi-scenario variant.\n\n\u003e Full per-stage guide: [wiki/Pipeline-Walkthrough.md](wiki/Pipeline-Walkthrough.md)\n\n---\n\n## Common Patterns\n\n### Use FDSN instead of a local SDS archive\n\n```python\nfm.calculate(\n    start_date=\"2025-01-01\", end_date=\"2025-01-31\",\n    source=\"fdsn\",\n    client_url=\"https://service.iris.edu\",\n)\n# Downloaded miniSEED is cached locally as SDS so subsequent runs skip the network.\n```\n\n### Skip tremor calculation when the CSV already exists\n\n`TrainingModel` and `PredictionModel` both accept either a `pd.DataFrame` or a tremor CSV path directly — handy when iterating on training kwargs without re-running `calculate()`.\n\n```python\nfrom eruption_forecast import TrainingModel\n\ntm = (\n    TrainingModel(\n        tremor_data=\"output/VG.OJN.00.EHZ/tremor/VG.OJN.00.EHZ_2025-01-01_2025-08-31.csv\",\n        start_date=\"2025-01-01\", end_date=\"2025-07-26\",\n        classifiers=[\"rf\", \"xgb\"],\n        eruption_dates=[\"2025-03-20\"],\n        window_size=2,\n        top_n_features=20,\n    )\n    .build_label(window_step=6, window_step_unit=\"hours\")\n    .extract_features(select_tremor_columns=[\"rsam_f2\", \"rsam_f3\", \"dsar_f3-f4\", \"entropy\"])\n    .fit(seeds=25, resample_method=\"auto\", plot_features=True)\n)\n```\n\n### Reuse curated features (skip full tsfresh re-extraction)\n\nOnce a `TrainingModel` run has persisted its `features-matrix_*.csv` and `top_{N}_features.csv`, two shortcuts let later runs reuse that work:\n\n```python\n# Reuse features (fastest — tremor matrix unchanged)\n(\n    TrainingModel(...)\n    .load_features(\n        select_features=\"output/.../training/features/.../top_20_features.csv\",\n    )\n    .fit(seeds=25)\n)\n\n# Or re-run tsfresh on the curated columns only\n(\n    TrainingModel(...)\n    .build_label(window_step=6, window_step_unit=\"hours\")\n    .extract_features(\n        select_features=\"output/.../training/features/.../top_20_features.csv\",\n    )\n    .fit(seeds=25)\n)\n```\n\n`select_features` accepts a CSV path or an explicit `list[str]` of fully-qualified tsfresh feature names. Per-seed selection inside `fit()` still runs; when the curated list is already ≤ `top_n_features`, it's effectively a no-op.\n\n### Forecast from a saved `ClassifierEnsemble`\n\n```python\nfrom eruption_forecast import PredictionModel\n\npm = (\n    PredictionModel(\n        model=\"output/.../training/classifiers/ClassifierEnsemble_stratified-shuffle-split.pkl\",\n        tremor_data=\"output/.../tremor/VG.OJN.00.EHZ_2025-01-01_2025-08-31.csv\",\n        start_date=\"2025-07-27\", end_date=\"2025-08-22\",\n        window_size=2,\n    )\n    .build_label(window_step=10, window_step_unit=\"minutes\")\n    .extract_features()\n)\nresults = pm.forecast(plot_threshold=0.7, plot_pdf=True)\n```\n\n`PredictionModel(model=...)` also accepts a live `ClassifierEnsemble` / `SeedEnsemble`, a `ClassifierEnsemble.json`, a `SeedEnsemble_*.pkl`, or a trained-model registry CSV — resolved via `ClassifierEnsemble.from_any(...)`.\n\n### Evaluate from a saved `.pkl`\n\n```python\nfrom eruption_forecast import EvaluationModel\n\nem = EvaluationModel.from_file(\n    \"output/.../PredictionModel_2025-07-27_2025-08-22.pkl\",\n    eruption_dates=[\"2025-08-02\"],   # required for prediction reuse\n)\nmetrics = em.evaluate(plot_aggregate=True)\ncomparator = em.compare()\nprint(comparator.get_ranking())\n```\n\n### Save and replay the pipeline configuration\n\n`fm.evaluate(...)` auto-calls `save_config()`. Call it manually at earlier stages to checkpoint partial runs.\n\n```python\nfm.save_config()                       # → {station_dir}/forecast.config.yaml\nfm.save_config(fmt=\"json\")             # → forecast.config.json\n\n# Replay\nfm2 = ForecastModel.from_config(\"output/VG.OJN.00.EHZ/forecast.config.yaml\")\nfm2.run()                              # replays every captured non-None stage\n```\n\n### Persist stage outputs explicitly\n\n```python\nfm.TrainingModel.save()       # → {output_dir}/TrainingModel_{basename}.pkl\nfm.PredictionModel.save()     # → {output_dir}/PredictionModel_{basename}.pkl\nfm.EvaluationModel.save()     # → {output_dir}/EvaluationModel_{basename}.pkl\n```\n\n### Per-stage config snapshots\n\nEach stage model also auto-saves its own YAML at the end of its main run\nmethod — `tm.save_config()`, `pm.save_config()`, `em.save_config()`,\n`xm.save_config()` all default to a stage-namespaced path under the station\ndir:\n\n```\n{station_dir}/training/training.config.yaml         # auto at end of fit()\n{station_dir}/prediction/prediction.config.yaml     # auto at end of forecast()\n{station_dir}/evaluation/{kind}/evaluation.config.yaml  # auto at end of evaluate()\n{station_dir}/explanation/{kind}/explanation.config.yaml # auto at end of explain()\n```\n\n### Silence logging during batch jobs\n\n```python\nfrom eruption_forecast import disable_logging, enable_logging\nfrom eruption_forecast.logger import set_log_level, set_log_directory\n\ndisable_logging()\nfm.calculate(...).train(...).predict(...)       # silent\nenable_logging()\n\nset_log_level(\"WARNING\")                        # console-only level\nset_log_directory(\"logs/2026-06-10\")            # move file handler\n```\n\n\u003e Telegram + logging full reference: [wiki/Configuration.md](wiki/Configuration.md)\n\n---\n\n## Supported Classifiers\n\n| Key | sklearn class | Imbalance handling | Notes |\n|-----|---------------|---------------------|-------|\n| `rf` | `RandomForestClassifier` | `class_weight=\"balanced\"` | Default; robust baseline |\n| `lite-rf` | `RandomForestClassifier` | `class_weight=\"balanced\"` | Smaller grid for faster training |\n| `gb` | `GradientBoostingClassifier` | natural | — |\n| `xgb` | `XGBClassifier` | `scale_pos_weight` grid | — |\n| `svm` | `SVC` | `class_weight=\"balanced\"` | — |\n| `lr` | `LogisticRegression` | `class_weight=\"balanced\"` | Fast, interpretable |\n| `nn` | `MLPClassifier` | none | — |\n| `dt` | `DecisionTreeClassifier` | `class_weight=\"balanced\"` | Interpretable baseline |\n| `knn` | `KNeighborsClassifier` | none | — |\n| `nb` | `GaussianNB` | none | Fast baseline |\n| `voting` | `VotingClassifier` (RF + XGBoost soft vote) | combined | — |\n\n`classifiers=` accepts `str` or `list[str]`. One `SeedEnsemble` is built per classifier and bundled into a single `ClassifierEnsemble` for consensus forecasting.\n\n\u003e Hyperparameter grids and tuning: [wiki/Training-Workflow.md](wiki/Training-Workflow.md)\n\n## Cross-Validation Strategies\n\n| `cv_strategy` | sklearn class | Best for |\n|---------------|---------------|----------|\n| `shuffle-stratified` (default) | `StratifiedShuffleSplit` | Random splits with stratification |\n| `stratified` | `StratifiedKFold` | Strict k-fold with class-distribution preservation |\n| `shuffle` | `ShuffleSplit` | Random splits without stratification |\n\n`timeseries` (`TimeSeriesSplit`) is available via the lower-level `ClassifierModel` directly but is not exposed on `fm.train(...)`.\n\n## Imbalance Handling\n\n| `resample_method` | Behaviour |\n|-------------------|-----------|\n| `\"auto\"` (default) | Apply `\"under\"` (`RandomUnderSampler`) when the minority-class share is below `minority_threshold` (default `0.15`), otherwise skip |\n| `\"under\"` | Always apply `RandomUnderSampler` |\n| `\"over\"` | Always apply `RandomOverSampler` |\n| `None` | Skip resampling entirely |\n\n`sampling_strategy=0.75` (default) is the target ratio passed to the resampler.\n\n---\n\n## Output Directory Structure\n\nAll outputs land under `{output_dir}/{network}.{station}.{location}.{channel}/` (e.g., `output/VG.OJN.00.EHZ/`).\n\n```\n{station_dir}/\n├── tremor/                                  # CalculateTremor\n│   ├── daily/                               # per-day CSVs (removed when cleanup_daily_dir=True)\n│   ├── figures/                             # daily plots (plot_daily=True)\n│   └── {nslc}_{start}_{end}.csv             # merged tremor CSV\n│\n├── training/                                # TrainingModel\n│   ├── training.config.yaml                 # tm.save_config() — auto at end of fit()\n│   ├── features/{cv-slug}/                  # tsfresh matrix, per-seed CSVs, top-N features\n│   └── classifiers/\n│       ├── ClassifierEnsemble_{cv}.{pkl,json}\n│       └── {clf-slug}/{cv-slug}/\n│           ├── models/{seed:05d}.pkl\n│           ├── trained-model__{suffix}.json\n│           └── SeedEnsemble_{suffix}.pkl\n│\n├── prediction/                              # PredictionModel\n│   ├── prediction.config.yaml               # pm.save_config() — auto at end of forecast()\n│   ├── features/                            # forecast-grid features\n│   ├── results/{clf-slug}/{seed:05d}.csv    # per-seed probabilities (save_seed_result=True)\n│   └── figures/forecast_{basename}.{png,pdf}\n│\n├── evaluation/                              # EvaluationModel\n│   ├── training/                            # when model=\"training\"\n│   └── prediction/                          # when model=\"prediction\"\n│       ├── evaluation.config.yaml                  # em.save_config() — auto at end of evaluate()\n│       ├── classifiers/{ClfName}/\n│       │   ├── predictions/{y_proba,y_pred}.csv   # (n_samples × n_seeds)\n│       │   └── figures/\n│       │       ├── aggregate/{plot_name}.{png,csv}\n│       │       └── {plot_name}/{seed:05d}.png      # plot_per_seed=True\n│       ├── labels/y_true.csv                       # prediction reuse only\n│       └── comparison/                             # ClassifierComparator\n│\n├── explanation/                             # ExplanationModel\n│   ├── training/                            # when upstream model.kind==\"training\"\n│   └── prediction/                          # when upstream model.kind==\"prediction\"\n│       ├── explanation.config.yaml                 # xm.save_config() — auto at end of explain()\n│       ├── classifiers/{ClfName}/\n│       │   ├── ClassifierExplanation_{ClfName}.pkl\n│       │   ├── shap_values/{seed:05d}.pkl\n│       │   └── figures/\n│       │       ├── {bar,beeswarm}/{seed:05d}.png            # plot_per_seed=True\n│       │       └── aggregate/{bar,beeswarm}.{png,csv}       # plot_aggregate=True\n│       └── eruptions/{YYYY-MM-DD}/\n│           └── {ClfName}_{datetime}_seed=_index=.png\n│\n├── cache/                                   # CacheModel\n│   ├── TrainingModel/{hash}.pkl + {hash}.params.json\n│   ├── PredictionModel/{hash}.pkl + {hash}.params.json\n│   └── ExplanationModel/{hash}.pkl + {hash}.params.json\n│\n├── forecast.config.yaml                     # fm.save_config()\n├── forecast-results_{basename}.csv\n└── {Training,Prediction,Evaluation}Model_*.pkl   # optional, via .save()\n```\n\n\u003e Full tree with slug tables and filename conventions: [wiki/Output-Structure.md](wiki/Output-Structure.md)\n\n---\n\n## Requirements\n\n### Core dependencies\n\n- Python ≥ 3.11\n- pandas ≥ 3.0.0, numpy, scipy\n- obspy (seismic data processing)\n- tsfresh (time-series feature extraction)\n- scikit-learn, imbalanced-learn\n- xgboost ≥ 3.x\n- shap ≥ 0.46\n- joblib\n- matplotlib, seaborn\n- loguru\n- python-dotenv (Telegram credentials)\n\n### Development dependencies\n\n- [`ruff`](https://docs.astral.sh/ruff/) — linting + auto-fix\n- [`ty`](https://github.com/astral-sh/ty) — type checking\n- `pytest` — testing\n\n---\n\n## Development\n\n```bash\n# Lint and auto-fix\nuv run ruff check --fix src/\n\n# Type check\nuvx ty check src/\n\n# Run tests\nuv run pytest tests/\n\n# Circular-import check (run after any module move)\nuv run pytest tests/test_imports.py -v\n\n# Run the end-to-end pipeline\nuv run python main.py\n```\n\nProject rules are documented in [`CLAUDE.md`](CLAUDE.md) — including the new-branch-before-any-commit convention, the comprehensive doc-update rule, and the `config.example.yaml` sync requirement.\n\n---\n\n## Contributing\n\n1. Fork the repository.\n2. Create a feature branch from the latest target:\n   - `fix/\u003cname\u003e` — bug fixes\n   - `ft/\u003cname\u003e` — new features\n   - `dev/\u003cname\u003e` — refactors, docs, tooling\n3. Make changes with tests and updated wiki pages.\n4. Pass linting + type checks: `uv run ruff check --fix src/ \u0026\u0026 uvx ty check src/`.\n5. Run the circular-import test: `uv run pytest tests/test_imports.py -v`.\n6. Submit a pull request against the configured base branch.\n\n**Code style:** PEP 8, Google-style docstrings with explicit types, type hints on all public functions, no inline `import`s.\n\n---\n\n## License\n\nMIT License — see [LICENSE](LICENSE) for details.\n\n**Disclaimer of Liability**: This software is provided \"as is\" without warranty of any kind, express or implied. The authors and contributors shall not be liable for any damages or losses arising from the use of this software. Volcanic eruption forecasting is inherently uncertain, and this software should be used only as a research tool, not for operational volcano monitoring or public safety decisions.\n\n---\n\n**Version:** 0.3.3\n**Status:** Active Development\n**Last Updated:** 2026-06-27\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fmartanto%2Feruption-forecast","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fmartanto%2Feruption-forecast","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fmartanto%2Feruption-forecast/lists"}