{"id":20685207,"url":"https://github.com/marpogaus/dvc-stage","last_synced_at":"2026-02-01T03:34:25.975Z","repository":{"id":158369916,"uuid":"578559968","full_name":"MArpogaus/dvc-stage","owner":"MArpogaus","description":"Stop programming common dvc stages. Configure them.","archived":false,"fork":false,"pushed_at":"2025-06-20T13:30:17.000Z","size":161,"stargazers_count":0,"open_issues_count":0,"forks_count":0,"subscribers_count":1,"default_branch":"main","last_synced_at":"2025-11-28T10:16:40.688Z","etag":null,"topics":["ai","data-science","data-version-control","developer-tools","dvc","git","machine-learning","python","reproducibility"],"latest_commit_sha":null,"homepage":"https://marpogaus.github.io/dvc-stage/","language":"Python","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"gpl-3.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/MArpogaus.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":null,"funding":null,"license":"COPYING","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2022-12-15T10:49:02.000Z","updated_at":"2025-06-20T13:30:05.000Z","dependencies_parsed_at":null,"dependency_job_id":"d94f62c1-8f1e-41f8-9df1-52b0739d4676","html_url":"https://github.com/MArpogaus/dvc-stage","commit_stats":null,"previous_names":[],"tags_count":13,"template":false,"template_full_name":null,"purl":"pkg:github/MArpogaus/dvc-stage","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/MArpogaus%2Fdvc-stage","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/MArpogaus%2Fdvc-stage/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/MArpogaus%2Fdvc-stage/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/MArpogaus%2Fdvc-stage/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/MArpogaus","download_url":"https://codeload.github.com/MArpogaus/dvc-stage/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/MArpogaus%2Fdvc-stage/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":28966627,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-02-01T02:14:24.993Z","status":"ssl_error","status_checked_at":"2026-02-01T02:13:55.706Z","response_time":56,"last_error":"SSL_read: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["ai","data-science","data-version-control","developer-tools","dvc","git","machine-learning","python","reproducibility"],"created_at":"2024-11-16T22:26:23.241Z","updated_at":"2026-02-01T03:34:25.968Z","avatar_url":"https://github.com/MArpogaus.png","language":"Python","funding_links":[],"categories":[],"sub_categories":[],"readme":"[![img](https://img.shields.io/github/contributors/MArpogaus/dvc-stage.svg?style=flat-square)](https://github.com/MArpogaus/dvc-stage/graphs/contributors) [![img](https://img.shields.io/github/forks/MArpogaus/dvc-stage.svg?style=flat-square)](https://github.com/MArpogaus/dvc-stage/network/members) [![img](https://img.shields.io/github/stars/MArpogaus/dvc-stage.svg?style=flat-square)](https://github.com/MArpogaus/dvc-stage/stargazers) [![img](https://img.shields.io/github/issues/MArpogaus/dvc-stage.svg?style=flat-square)](https://github.com/MArpogaus/dvc-stage/issues) [![img](https://img.shields.io/github/license/MArpogaus/dvc-stage.svg?style=flat-square)](https://github.com/MArpogaus/dvc-stage/blob/main/LICENSE) [![img](https://img.shields.io/github/actions/workflow/status/MArpogaus/dvc-stage/run_demo.yaml.svg?label=test\u0026style=flat-square)](https://github.com/MArpogaus/dvc-stage/actions/workflows/run_demo.yaml) [![img](https://img.shields.io/github/actions/workflow/status/MArpogaus/dvc-stage/release.yaml.svg?label=release\u0026style=flat-square)](https://github.com/MArpogaus/dvc-stage/actions/workflows/release.yaml) [![img](https://img.shields.io/badge/pre--commit-enabled-brightgreen.svg?logo=pre-commit\u0026style=flat-square)](https://github.com/MArpogaus/dvc-stage/blob/main/.pre-commit-config.yaml) [![img](https://img.shields.io/badge/-LinkedIn-black.svg?style=flat-square\u0026logo=linkedin\u0026colorB=555)](https://linkedin.com/in/MArpogaus)\n\n[![img](https://img.shields.io/pypi/v/dvc-stage.svg?style=flat-square)](https://pypi.org/project/dvc-stage)\n\n\n# DVC-Stage\n\n1.  [About The Project](#org0b9f792)\n2.  [Getting Started](#org2a5d3ea)\n    1.  [Prerequisites](#org4190b97)\n    2.  [Installation](#orge4cf093)\n3.  [Usage](#org57e65e6)\n    1.  [Basic Stage Structure](#org425891e)\n    2.  [Examples](#orgdca970f)\n    3.  [Built-in Transformations](#org32fae61)\n    4.  [Built-in Validations](#orgb49453a)\n    5.  [Using Custom Functions](#org391efe7)\n4.  [Contributing](#org82ce8b3)\n5.  [License](#org1dfb3ef)\n6.  [Contact](#org80d0f10)\n7.  [Acknowledgments](#org70aa47e)\n\n\n\u003ca id=\"org0b9f792\"\u003e\u003c/a\u003e\n\n## About The Project\n\nThis python script provides a easy and parameterizeable way of defining typical dvc (sub-)stages for:\n\n-   data prepossessing\n-   data transformation\n-   data splitting\n-   data validation\n\n\n\u003ca id=\"org2a5d3ea\"\u003e\u003c/a\u003e\n\n## Getting Started\n\nThis is an example of how you may give instructions on setting up your project locally. To get a local copy up and running follow these simple example steps.\n\n\n\u003ca id=\"org4190b97\"\u003e\u003c/a\u003e\n\n### Prerequisites\n\n-   `pandas\u003e=0.20.*`\n-   `dvc\u003e=2.12.*`\n-   `pyyaml\u003e=5`\n\n\n\u003ca id=\"orge4cf093\"\u003e\u003c/a\u003e\n\n### Installation\n\nThis package is available on [PyPI](https://pypi.org/project/dvc-stage/). You install it and all of its dependencies using pip:\n\n```bash\npip install dvc-stage\n```\n\n\n\u003ca id=\"org57e65e6\"\u003e\u003c/a\u003e\n\n## Usage\n\nDVC-Stage works on top of two files: `dvc.yaml` and `params.yaml`. They are expected to be at the root of an initialized [dvc project](https://dvc.org/). From there you can execute `dvc-stage -h` to see available commands or `dvc-stage get-config STAGE` to generate the dvc stages from the `params.yaml` file. The tool then generates the respective yaml which you can then manually paste into the `dvc.yaml` file. Existing stages can then be updated inplace using `dvc-stage update-stage STAGE`.\n\n\n\u003ca id=\"org425891e\"\u003e\u003c/a\u003e\n\n### Basic Stage Structure\n\nStages are defined inside `params.yaml` in the following schema:\n\n```yaml\nSTAGE_NAME:\n  load: {}\n  transformations: []\n  validations: []\n  write: {}\n```\n\nThe `load` and `write` sections both require the yaml-keys `path` and `format` to read and save data respectively.\n\nThe `transformations` and `validations` sections require a sequence of functions to apply, where `transformations` return data and `validations` return a truth value (derived from data). Functions are defined by the key `id` and can be either:\n\n-   Methods defined on Pandas DataFrames, e.g.\n\n    ```yaml\n    transformations:\n      - id: transpose\n    ```\n\n-   Imported from any python module, e.g.\n\n    ```yaml\n    transformations:\n      - id: custom\n        description: duplicate rows\n        import_from: demo.duplicate\n    ```\n\n-   Predefined by DVC-Stage, e.g.\n\n    ```yaml\n    validations:\n      - id: validate_pandera_schema\n        schema:\n          import_from: demo.get_schema\n    ```\n\nWhen writing a custom function, you need to make sure the function gracefully handles data being `None`, which is required for type inference. Data is passed as first argument. Further arguments can be provided as additional keys, as shown above for `validate_pandera_schema`, where schema is passed as second argument to the function.\n\n\n\u003ca id=\"orgdca970f\"\u003e\u003c/a\u003e\n\n### Examples\n\nThe `examples` directory contains a complete working demonstration:\n\n1.  **Setup**: Navigate to the examples directory\n2.  **Data**: Sample data files are provided in `data`\n3.  **Configuration**: `params.yaml` contains all pipeline definitions\n4.  **Custom functions**: `src/demo.py` contains example custom functions\n5.  **DVC configuration**: `dvc.yaml` contains the generated DVC stages\n\nTo run all examples:\n\n```bash\ncd examples\n\n# Update all stage deffinitions\ndvc-stage update-all -y\n\n# Reproduce pipeline\ndvc repro\n```\n\n1.  Example 1: Basic Demo Pipeline\n\n    The simplest example demonstrates basic data loading, transformation, validation, and writing:\n\n    ```yaml\n    demo_pipeline:\n      dvc_stage_args:\n        log-level: ${log_level}\n        log-file: ${log_file}\n      load:\n        path: load.csv\n        format: csv\n      transformations:\n      - id: custom\n        description: duplicate rows\n        import_from: demo.duplicate\n      - id: transpose\n      - id: rename\n        columns:\n          0.0: O1\n          1.0: O2\n          2.0: D1\n          3.0: D2\n      validations:\n      - id: custom\n        description: check none\n        import_from: demo.isNotNone\n      - id: isnull\n        reduction: any\n        expected: false\n      - id: validate_pandera_schema\n        schema:\n          import_from: demo.get_schema\n      write:\n        format: csv\n    ```\n\n    **What this pipeline does:**\n\n    1.  **Load**: Reads data from `load.csv`\n    2.  **Transform**:\n        -   Duplicates all rows using a custom function\n        -   Transposes the DataFrame\n        -   Renames columns from numeric to meaningful names\n    3.  **Validate**:\n        -   Checks that data is not None\n        -   Ensures no null values exist\n        -   Validates against a Pandera schema\n    4.  **Write**: Saves the result to `outdir/out.csv`\n\n    **Run with:**\n\n    ```bash\n    cd examples\n    dvc repro demo_pipeline\n    ```\n\n2.  Example 2: Foreach Pipeline\n\n    Process multiple datasets with the same pipeline using foreach stages:\n\n    ```yaml\n    foreach_pipeline:\n      dvc_stage_args:\n        log-level: ${log_level}\n        log-file: ${log_file}\n      foreach: [dataset_a, dataset_b, dataset_c]\n      load:\n        path: data/${item}/input.csv\n        format: csv\n      transformations:\n      - id: fillna\n        value: 0\n      - id: custom\n        description: normalize data\n        import_from: demo.normalize_data\n        columns: [value1, value2]\n      validations:\n      - id: validate_pandera_schema\n        schema:\n          import_from: demo.get_foreach_schema\n      - id: custom\n        description: check data quality\n        import_from: demo.check_data_quality\n        min_rows: 5\n      write:\n        path: outdir/${item}_${key}_processed.csv\n    ```\n\n    **What this pipeline does:**\n\n    1.  **Foreach**: Processes three datasets (dataset\u003csub\u003ea\u003c/sub\u003e, dataset\u003csub\u003eb\u003c/sub\u003e, dataset\u003csub\u003ec\u003c/sub\u003e)\n    2.  **Load**: Reads from `data/${item}/input.csv` where `${item}` is replaced with each dataset name\n    3.  **Transform**:\n        -   Fills missing values with 0\n        -   Normalizes specified columns using min-max scaling\n    4.  **Validate**:\n        -   Validates against a pandera schema\n        -   Checks data quality (minimum row count)\n    5.  **Write**: Saves each processed dataset to `outdir/${item}_${key}_processed.csv`\n\n    **Run with:**\n\n    ```bash\n    cd examples\n    dvc repro foreach_pipeline\n    ```\n\n3.  Example 3: Advanced Multi-Input Pipeline\n\n    Handle multiple input files with data splitting:\n\n    ```yaml\n    advanced_pipeline:\n      dvc_stage_args:\n        log-level: ${log_level}\n        log-file: ${log_file}\n      load:\n        path:\n        - data/features.csv\n        - data/labels.csv\n        format: csv\n        key_map:\n          features: data/features.csv\n          labels: data/labels.csv\n      transformations:\n      - id: split\n        include: [features]\n        by: id\n        id_col: category\n        left_split_key: train\n        right_split_key: test\n        size: 0.5\n        seed: 42\n      - id: combine\n        include: [train, test]\n        new_key: combined_data\n      validations:\n      - id: validate_pandera_schema\n        schema:\n          import_from: demo.get_advanced_schema\n        include: [combined]\n      write:\n        path: outdir/${key}.csv\n    ```\n\n    **What this pipeline does:**\n\n    1.  **Load**: Reads multiple files and maps them to keys (features, labels)\n    2.  **Transform**:\n        -   The features table is spitted along the categories in two data frames containing each 50% of the data\n        -   The spitted data is again combined into a single table\n    3.  **Validate**: Validates both train and test sets against a schema\n    4.  **Write**: Saves train.csv and test.csv to the output directory\n\n    **Run with:**\n\n    ```bash\n    cd examples\n    dvc repro advanced_pipeline\n    ```\n\n4.  Example 4: Time Series Pipeline\n\n    Process time series data with date-based splitting:\n\n    ```yaml\n    timeseries_pipeline:\n      dvc_stage_args:\n        log-level: ${log_level}\n        log-file: ${log_file}\n      load:\n        path: data/timeseries.csv\n        format: csv\n        parse_dates: [timestamp]\n        index_col: timestamp\n      transformations:\n      - id: reset_index\n      - id: add_date_offset_to_column\n        column: timestamp\n        days: 1\n      - id: split\n        by: date_time\n        left_split_key: train\n        right_split_key: test\n        size: 0.8\n        freq: D\n        date_time_col: timestamp\n      - id: set_index\n        keys: timestamp\n      validations:\n      - id: validate_pandera_schema\n        schema:\n          import_from: demo.get_timeseries_schema\n      - id: custom\n        description: validate split ratio\n        pass_dict_to_fn: true\n        import_from: demo.validate_split_ratio\n        reduction: none\n        expected_ratio: 0.8\n        tolerance: 0.05\n      write:\n        path: outdir/timeseries_${key}.csv\n    ```\n\n    **What this pipeline does:**\n\n    1.  **Load**: Reads time series data with proper datetime parsing\n    2.  **Transform**:\n        -   Reset pandas index\n        -   Adds a date offset to the timestamps\n        -   Splits data chronologically (80% train, 20% test) by date\n        -   Set timestamp as index\n    3.  **Validate**:\n        -   Validates against a time series specific schema\n        -   Validate the split ratio\n    4.  **Write**: Saves timeseries\u003csub\u003etrain.csv\u003c/sub\u003e and timeseries\u003csub\u003etest.csv\u003c/sub\u003e\n\n    **Run with:**\n\n    ```bash\n    cd examples\n    dvc repro timeseries_pipeline\n    ```\n\n\n\u003ca id=\"org32fae61\"\u003e\u003c/a\u003e\n\n### Built-in Transformations\n\nDVC-Stage provides several built-in transformations:\n\n-   **split**: Split data (random, date\u003csub\u003etime\u003c/sub\u003e, or id-based)\n-   **combine**: Combine multiple DataFrames\n-   **column\u003csub\u003etransformer\u003c/sub\u003e\u003csub\u003efit\u003c/sub\u003e**: Fit sklearn column transformers\n-   **column\u003csub\u003etransformer\u003c/sub\u003e\u003csub\u003etransform\u003c/sub\u003e**: Apply fitted transformers\n-   **add\u003csub\u003edate\u003c/sub\u003e\u003csub\u003eoffset\u003c/sub\u003e\u003csub\u003eto\u003c/sub\u003e\u003csub\u003ecolumn\u003c/sub\u003e**: Add time offsets to date columns\n\n    Additionally all pandas DataFrame methods can be used, e.g.:\n\n-   **fillna**: Fill missing values\n-   **dropna**: Drop rows with missing values\n-   **transpose**: Transpose the DataFrame\n-   **rename**: Rename columns\n\n\n\u003ca id=\"orgb49453a\"\u003e\u003c/a\u003e\n\n### Built-in Validations\n\nDVC-Stage provides several built-in validations:\n\n-   **validate\u003csub\u003epandera\u003c/sub\u003e\u003csub\u003eschema\u003c/sub\u003e**: Validate against Pandera schemas\n-   **Custom validations**: Import your own validation functions\n\nAdditionally all pandas DataFrame methods can be used, e.g.:\n\n-   **isnull**: Check for null values\n\n\n\u003ca id=\"org391efe7\"\u003e\u003c/a\u003e\n\n### Using Custom Functions\n\nWhen creating custom functions for transformations or validations:\n\n1.  **Handle None gracefully**: Your function should return appropriate values when data is None\n2.  **First argument is data**: The DataFrame or data structure is always the first parameter\n3.  **Additional parameters**: Pass extra arguments as YAML keys in your stage definition\n4.  **Return appropriate types**: Transformations return data, validations return boolean values\n\nExample custom function:\n\n```python\ndef normalize_data(data: pd.DataFrame, columns: List[str]) -\u003e pd.DataFrame:\n    \"\"\"Normalize specified columns using min-max scaling.\"\"\"\n    if data is None:\n        return None\n\n    result = data.copy()\n    for col in columns:\n        if col in result.columns:\n            min_val = result[col].min()\n            max_val = result[col].max()\n            if max_val \u003e min_val:\n                result[col] = (result[col] - min_val) / (max_val - min_val)\n    return result\n```\n\n\n\u003ca id=\"org82ce8b3\"\u003e\u003c/a\u003e\n\n## Contributing\n\nAny Contributions are greatly appreciated! If you have a question, an issue or would like to contribute, please read our [contributing guidelines](CONTRIBUTING.md).\n\n\n\u003ca id=\"org1dfb3ef\"\u003e\u003c/a\u003e\n\n## License\n\nDistributed under the [GNU General Public License v3](COPYING)\n\n\n\u003ca id=\"org80d0f10\"\u003e\u003c/a\u003e\n\n## Contact\n\n[Marcel Arpogaus](https://github.com/MArpogaus/) - [znepry.necbtnhf@tznvy.pbz](mailto:znepry.necbtnhf@tznvy.pbz) (encrypted with [ROT13](\u003chttps://rot13.com/\u003e))\n\nProject Link: \u003chttps://github.com/MArpogaus/dvc-stage\u003e\n\n\n\u003ca id=\"org70aa47e\"\u003e\u003c/a\u003e\n\n## Acknowledgments\n\nParts of this work have been funded by the Federal Ministry for the Environment, Nature Conservation and Nuclear Safety due to a decision of the German Federal Parliament (AI4Grids: 67KI2012A).\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fmarpogaus%2Fdvc-stage","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fmarpogaus%2Fdvc-stage","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fmarpogaus%2Fdvc-stage/lists"}