{"id":51859212,"url":"https://github.com/PawanSikawat/faucet-stream","last_synced_at":"2026-07-24T08:00:42.859Z","repository":{"id":346293674,"uuid":"1189252746","full_name":"PawanSikawat/faucet-stream","owner":"PawanSikawat","description":"The fast, config-driven way to move data in Rust — a native CLI and an embeddable Rust ETL library","archived":false,"fork":false,"pushed_at":"2026-07-20T05:22:41.000Z","size":8861,"stargazers_count":7,"open_issues_count":19,"forks_count":4,"subscribers_count":4,"default_branch":"main","last_synced_at":"2026-07-20T07:12:43.616Z","etag":null,"topics":["cdc","cli","connectors","data-engineering","data-integration","data-pipeline","elt","etl","kafka","parquet","rust","singer","streaming"],"latest_commit_sha":null,"homepage":"https://pawansikawat.github.io/faucet-stream/","language":"Rust","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"other","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/PawanSikawat.png","metadata":{"files":{"readme":"README.md","changelog":"CHANGELOG.md","contributing":"CONTRIBUTING.md","funding":null,"license":"LICENSE-APACHE","code_of_conduct":"CODE_OF_CONDUCT.md","threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":"SECURITY.md","support":null,"governance":null,"roadmap":"docs/roadmap.md","authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null,"notice":null,"maintainers":null,"copyright":null,"agents":null,"dco":null,"cla":null}},"created_at":"2026-03-23T06:05:04.000Z","updated_at":"2026-07-19T18:43:12.000Z","dependencies_parsed_at":null,"dependency_job_id":"8238517e-8936-43fc-8d6f-e29b66d05f3a","html_url":"https://github.com/PawanSikawat/faucet-stream","commit_stats":null,"previous_names":["pawansikawat/faucet-stream"],"tags_count":386,"template":false,"template_full_name":null,"purl":"pkg:github/PawanSikawat/faucet-stream","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/PawanSikawat%2Ffaucet-stream","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/PawanSikawat%2Ffaucet-stream/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/PawanSikawat%2Ffaucet-stream/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/PawanSikawat%2Ffaucet-stream/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/PawanSikawat","download_url":"https://codeload.github.com/PawanSikawat/faucet-stream/tar.gz/refs/heads/main","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/PawanSikawat%2Ffaucet-stream/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":35832970,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-07-20T02:08:10.276Z","status":"online","status_checked_at":"2026-07-24T02:00:07.870Z","response_time":62,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["cdc","cli","connectors","data-engineering","data-integration","data-pipeline","elt","etl","kafka","parquet","rust","singer","streaming"],"created_at":"2026-07-24T04:00:40.324Z","updated_at":"2026-07-24T08:00:42.845Z","avatar_url":"https://github.com/PawanSikawat.png","language":"Rust","funding_links":[],"categories":["Data Ingestion","Table of Contents"],"sub_categories":["Data Pipeline"],"readme":"\u003cp align=\"center\"\u003e\n  \u003cimg src=\"https://raw.githubusercontent.com/PawanSikawat/faucet-stream/main/.github/assets/social-banner.png\" alt=\"faucet-stream\" width=\"640\"\u003e\n\u003c/p\u003e\n\n# faucet-stream\n\n[![Crates.io](https://img.shields.io/crates/v/faucet-stream.svg)](https://crates.io/crates/faucet-stream)\n[![Docs.rs](https://docs.rs/faucet-stream/badge.svg)](https://docs.rs/faucet-stream)\n[![Guide](https://img.shields.io/badge/guide-pawansikawat.github.io-1f6feb)](https://pawansikawat.github.io/faucet-stream/)\n[![CI](https://github.com/PawanSikawat/faucet-stream/actions/workflows/ci.yml/badge.svg)](https://github.com/PawanSikawat/faucet-stream/actions/workflows/ci.yml)\n[![Coverage](https://codecov.io/gh/PawanSikawat/faucet-stream/branch/main/graph/badge.svg)](https://codecov.io/gh/PawanSikawat/faucet-stream)\n[![Downloads](https://img.shields.io/crates/d/faucet-stream.svg)](https://crates.io/crates/faucet-stream)\n[![MSRV](https://img.shields.io/crates/msrv/faucet-stream.svg)](rust-toolchain.toml)\n[![Dependencies](https://img.shields.io/badge/deps-cargo--deny-blue)](deny.toml)\n[![License](https://img.shields.io/crates/l/faucet-stream.svg)](#license)\n[![Changelog](https://img.shields.io/badge/changelog-keep%20a%20changelog-orange)](CHANGELOG.md)\n\n**The fast, config-driven way to move data in Rust.**\n\n**Move data at Rust speed, govern it in flight, and ship it as a single binary.** On a\n1M-row CSV→JSONL move, faucet sustains **712k rows/s in 11.8 MiB of RAM** — **~96× faster and\n~62× less memory than Meltano**, output identical row-for-row ([see the benchmarks](BENCHMARKS.md)).\nNo Python runtime, no platform to stand up, no daemon to babysit.\n\nfaucet-stream is a **data-movement platform** for Rust — with governance built in: **33 source**\nand **25 sink** connectors (**58 in total**) plus in-flight transforms, including a page-level\nembedded-DuckDB `sql` transform — wired by a single `faucet` binary that runs pipelines\ndeclaratively from YAML/JSON (no Rust code required), or embedded in your own service through\nthe typed `Source` / `Sink` traits. One platform, whether you want a CLI you can drop on any\nbox or a library you compile in.\n\n📖 **[Guide](https://pawansikawat.github.io/faucet-stream/)** · 📊 **[Benchmarks](BENCHMARKS.md)** · 📜 **[Connector spec (FCP v0)](docs/spec/faucet-connector-spec-v0.md)**\n\n```bash\nbrew install PawanSikawat/faucet-stream/faucet-cli   # the CLI — prebuilt, no Rust needed\n# — or —\ncurl -LsSf https://github.com/PawanSikawat/faucet-stream/releases/latest/download/faucet-cli-installer.sh | sh\n# — or —\ncargo install faucet-cli          # build the CLI from source\n# — or —\ncargo add faucet-stream           # the library\n```\n\n### Why faucet-stream\n\n- **🚀 Built for throughput** — native streaming with bounded memory, connection\n  pooling, multi-row inserts, bulk APIs, and parallel I/O. Throughput is a first-class\n  design goal for every connector. On single-machine batch throughput, faucet runs\n  **~1–2 orders of magnitude faster than a Python Singer runtime**: a reproducible\n  1M-row CSV→JSONL move (a best case that maximally exposes Python's per-row\n  overhead) hit **712k rows/s in 11.8 MiB** vs Meltano's 7.4k rows/s / 724 MiB\n  (~96× faster, ~62× less memory, exact row parity); sink-bound moves like\n  Postgres→Postgres narrow the gap. See [`BENCHMARKS.md`](BENCHMARKS.md) for the\n  methodology, the sink-bound scenario, and honest caveats.\n- **🧩 Config-driven _or_ embeddable** — run `faucet run pipeline.yaml`, or call\n  `Pipeline::new(\u0026source, \u0026sink).run().await?` from Rust. Same orchestration either way.\n- **⚙️ A runtime, not just connectors** — incremental + resumable replication, change-data-capture,\n  effectively-once delivery (idempotent dedup-on-resume), upsert/delete write modes, dead-letter\n  queues, automatic retries, adaptive batch sizing, secrets-manager interpolation, cron scheduling,\n  and an HTTP control plane with event-driven triggers — plus built-in Prometheus metrics +\n  `tracing` spans, all with zero per-connector code.\n- **🛡️ Governance in the movement path** — the guardrails most pipelines bolt on downstream, native\n  and zero-config: data-quality checks, versioned data contracts, PII masking (applied *before* any\n  sink sees a row), schema-drift detection \u0026 policy, column-level lineage (OpenLineage) + a\n  data-movement catalog, and freshness/volume SLA monitoring.\n- **📦 Pay only for what you use** — every connector is a Cargo feature, so a slim build can\n  be just REST + JSONL, or pull in all 58 connectors with `--features full`.\n\n**Documentation:** the [faucet-stream guide](https://pawansikawat.github.io/faucet-stream/)\n(getting started, tutorials, cookbook, operations) · API reference on\n[docs.rs](https://docs.rs/faucet-stream) · [`cli/README.md`](cli/README.md) for the full config grammar.\n\n---\n\n## Table of contents\n\n- [Quickstart — the CLI](#quickstart--the-cli)\n- [Quickstart — the library](#quickstart--the-library)\n- [What's in the box](#whats-in-the-box)\n- [Connectors](#connectors)\n- [How it compares](#how-it-compares)\n- [When to use faucet-stream](#when-to-use-faucet-stream)\n- [Architecture](#architecture)\n- [Performance](#performance)\n- [Observability](#observability)\n- [Feature flags](#feature-flags)\n- [Using faucet-stream as a Rust library](#using-faucet-stream-as-a-rust-library)\n- [Building custom connectors](#building-custom-connectors)\n- [Project structure](#project-structure)\n- [Contributing](#contributing)\n- [License](#license)\n\n---\n\n## Quickstart — the CLI\n\n\u003e **Just want to poke at it?** From a clone, run `./scripts/try-local.sh` — it\n\u003e builds a light feature set, runs a no-infrastructure demo (transforms,\n\u003e quality, contracts, masking, lineage, catalog, DLQ replay) against generated\n\u003e data, then leaves the [web console](https://pawansikawat.github.io/faucet-stream/getting-started/try-it-locally.html)\n\u003e running so you can browse Runs, Datasets, and Lineage. No Docker or cloud\n\u003e accounts needed.\n\nMove data without writing any Rust:\n\n```bash\ncargo install faucet-cli\nfaucet init my_pipeline --source postgres --sink bigquery   # scaffold pipeline.yaml from schemas\nfaucet validate pipeline.yaml                               # parse + resolve secrets, no run\nfaucet doctor pipeline.yaml                                 # preflight: probe auth/network/permissions\nfaucet test tests/*.yaml                                    # offline fixture tests for pipeline logic\nfaucet run pipeline.yaml                                    # one-shot run to completion\nfaucet discover conn.yaml -o pipeline.yaml                  # introspect a database and generate a config\nfaucet backfill pipeline.yaml --from 2026-06-01 --to 2026-07-01 --window 1d   # resumable historical replay\nfaucet schedule pipeline.yaml                               # run on a cron schedule (add a schedule: block)\nfaucet serve --no-auth                                      # HTTP control plane: submit/poll/cancel runs over REST\n```\n\nA minimal config — fetch open GitHub issues and write them to JSON Lines:\n\n```yaml\n# faucet.yaml — `faucet run` auto-discovers this file (and a sibling `.env`) in cwd\nversion: 1\npipeline:\n  source:\n    type: rest\n    config:\n      base_url: https://api.github.com\n      path: /repos/PawanSikawat/faucet-stream/issues\n      method: GET\n      auth: { type: api_key, config: { header: Authorization, value: \"Bearer ${env:GITHUB_TOKEN}\" } }\n      query_params: { state: open }\n      pagination: { type: LinkHeader }\n      max_retries: 3\n  transforms:\n    - type: keys_case          # re-case every key: snake / camel / pascal / kebab / screaming_snake\n      config: { mode: snake }\n  sink:\n    type: jsonl\n    config:\n      path: ./out/issues.jsonl\n```\n\nRun many invocations from one config with a `matrix:` block (independent fan-out, a\nparent/child DAG, **or** `depends_on:` completion ordering between rows), and bound\nconcurrency with `execution:`. See [`cli/README.md`](cli/README.md)\nfor the full grammar, [`cli/examples/rest_to_bigquery_matrix.yaml`](cli/examples/rest_to_bigquery_matrix.yaml)\nfor matrix fan-out, and [`cli/examples/rest_users_posts_dag.yaml`](cli/examples/rest_users_posts_dag.yaml)\nfor the DAG pattern. The [`cli/examples/`](cli/examples) directory has runnable configs for\nevery common source→sink combination.\n\n## Quickstart — the library\n\nEmbed the same engine in a Rust service:\n\n```bash\ncargo add faucet-stream --features sink-jsonl   # default already has the REST source\ncargo add tokio --features full\n```\n\n```rust\nuse faucet_stream::{Pipeline, RestStream, RestStreamConfig, PaginationStyle};\nuse faucet_stream::sink::jsonl::{JsonlSink, JsonlSinkConfig};\n\n#[tokio::main]\nasync fn main() -\u003e Result\u003c(), Box\u003cdyn std::error::Error\u003e\u003e {\n    let source = RestStream::new(\n        RestStreamConfig::new(\"https://api.example.com\", \"/v1/users\")\n            .records_path(\"$.data[*]\")\n            .pagination(PaginationStyle::Cursor {\n                next_token_path: \"$.meta.next_cursor\".into(),\n                param_name: \"cursor\".into(),\n            }),\n    )?;\n    let sink = JsonlSink::new(JsonlSinkConfig::new(\"./users.jsonl\"));\n\n    let result = Pipeline::new(\u0026source, \u0026sink).run().await?;\n    println!(\"Wrote {} records\", result.records_written);\n    Ok(())\n}\n```\n\nMore library recipes — pagination styles, OAuth2, streaming, incremental replication,\npartitions, transforms, and custom connectors — are in\n[Using faucet-stream as a Rust library](#using-faucet-stream-as-a-rust-library) below and the\n[library tutorial](https://pawansikawat.github.io/faucet-stream/tutorials/library.html).\n\n## What's in the box\n\nfaucet-stream is a full data-movement **runtime**, not just a bag of connectors. Every\ncapability below works across all connectors with zero per-connector code, and each is a\none-block addition to your YAML:\n\n| Capability | What it does | Learn more |\n|---|---|---|\n| **Streaming, bounded memory** | Sources stream page-by-page; sinks write each page as it arrives — memory stays at one `batch_size` regardless of total volume. | [concepts](https://pawansikawat.github.io/faucet-stream/getting-started/concepts.html) |\n| **Incremental + resumable** | Bookmark-based replication: only fetch what changed, resume mid-run from a durable state store (file / Redis / Postgres). | [state](https://pawansikawat.github.io/faucet-stream/cookbook/state.html) |\n| **Change data capture** | Streaming row-level CDC for **PostgreSQL** (logical replication), **MySQL** (binlog), **MongoDB** (change streams), and **SQL Server** (CDC change tables) — resumable. | [CDC guide](https://pawansikawat.github.io/faucet-stream/reference/connectors.html) |\n| **Effectively-once delivery** | Monotonic per-page commit tokens committed atomically with the data (SQL sinks, Iceberg, BigQuery), so a resumed run re-delivers no duplicates. This is idempotent at-least-once (dedup on resume), not distributed-consensus exactly-once. | [state](https://pawansikawat.github.io/faucet-stream/cookbook/state.html) |\n| **Upsert / delete write modes** | `write_mode: upsert \\| delete` with a `key` + `delete_marker` — merge by key on Postgres / MySQL / SQL Server / SQLite / Mongo / Elasticsearch. | [upsert](https://pawansikawat.github.io/faucet-stream/cookbook/upsert.html) |\n| **Data-quality checks** | 13 per-record and per-batch assertions (not-null, regex, ranges, uniqueness, row-count, JSON Schema, …) with quarantine routing or abort policies. | [quality](https://pawansikawat.github.io/faucet-stream/cookbook/quality.html) |\n| **Data contracts** | A versioned promise about the output shape (types, nullability, enums, patterns, bounds) enforced per page — breaches fail, quarantine, or warn; export as JSON Schema / OpenLineage via `faucet contract`. | [contracts](https://pawansikawat.github.io/faucet-stream/cookbook/contracts.html) |\n| **SLA monitoring** | Declared freshness (`max_staleness_secs`) + volume floors and learned-baseline anomaly detection (z-score / IQR) per pipeline — violations emit metrics + warnings and surface in `faucet doctor`, never failing the run. | [SLA](https://pawansikawat.github.io/faucet-stream/cookbook/sla.html) |\n| **Dead-letter queue** | Route failed rows to any sink instead of aborting the run, with a fixed envelope and reason. | [DLQ](https://pawansikawat.github.io/faucet-stream/cookbook/dlq.html) |\n| **Adaptive batch sizing** | Opt-in AIMD controller that tunes write batch size from observed sink latency and error rate. | [tuning](https://pawansikawat.github.io/faucet-stream/) |\n| **Secrets-manager interpolation** | `${vault:…}`, `${aws-sm:…}`, `${gcp-sm:…}`, `${azure-kv:…}` resolved at load time, redacted from logs. | [secrets](https://pawansikawat.github.io/faucet-stream/cookbook/secrets.html) |\n| **Cron scheduling** | `faucet schedule` — DST-correct cron, overlap policies, run timeouts, graceful drain. | [scheduling](https://pawansikawat.github.io/faucet-stream/) |\n| **HTTP control plane** | `faucet serve` — submit / poll / cancel runs over REST, idempotency keys, run history, optional embedded web console (`serve-ui`), clustered execution. | [serve](https://pawansikawat.github.io/faucet-stream/) |\n| **Event-driven triggers** | `faucet serve --triggers` — auto-enqueue runs on object-arrival (S3/GCS), webhook, or queue-depth (Redis/Kafka). | [triggers](https://pawansikawat.github.io/faucet-stream/reference/triggers.html) |\n| **OpenLineage emission** | Emit START/RUNNING/COMPLETE/FAIL events with schema facets and column-level lineage over HTTP / file / Kafka. | [lineage](https://pawansikawat.github.io/faucet-stream/cookbook/lineage.html) |\n| **Observability** | Automatic Prometheus metrics + `tracing` spans for every source, sink, transform, and state op — labelled by pipeline / row / connector. | [observability](cli/README.md#observability-prometheus--tracing) |\n| **Transforms** | `flatten`, `rename_keys`, `keys_case`, `select`, `drop`, `set`, `rename_field`, `cast`, `redact`, `value_case`, `spell_symbols`, `cdc_unwrap`, and `sql` (embedded DuckDB, page-level). | [transforms](https://pawansikawat.github.io/faucet-stream/cookbook/transforms.html) |\n\n## Connectors\n\nAll connector crates depend only on `faucet-core`, so any source pairs with any sink. See\nthe [connector capability matrix](https://pawansikawat.github.io/faucet-stream/reference/connectors.html)\n(streaming, resumable state, compression, auth per connector) and the\n[choosing-a-connector guide](https://pawansikawat.github.io/faucet-stream/reference/choosing.html)\nfor help picking between overlapping connectors (Postgres query vs CDC, S3 vs Parquet, Redis vs Kafka, …).\n\n\u003e **Support tiers** (the **Tier** column below). A connector is **Tier-1 ✅** when\n\u003e it invokes and passes the [`faucet-conformance`](crates/conformance) battery in\n\u003e CI against the connector's real backend — valid config schema, bounded-memory\n\u003e streaming, bookmark round-trip, idempotent replay, truthful capabilities,\n\u003e errors-not-panics (see the [Faucet Connector Protocol spec](docs/spec/faucet-connector-spec-v0.md)).\n\u003e **That battery *is* the tiering mechanism** — there is no separate scheme, and\n\u003e the Tier-1 set grows as more connectors wire it in. **ᵐ** marks a connector\n\u003e whose battery runs against a **wiremock HTTP mock** in CI rather than a live\n\u003e service (the `rest`, `graphql`, `xml`, `elasticsearch`, `bigquery`, `snowflake`\n\u003e sources and the `http` sink): the mock faithfully drives the paging / schema /\n\u003e error paths the checks assert, but it is not an end-to-end test against the real\n\u003e system. **ᵉ** marks the Cloud Spanner pair, whose battery runs against Google's\n\u003e official **Spanner emulator** (Docker) — a real gRPC Spanner implementation.\n\u003e **Tier-2** connectors are **not conformance-certified in CI** — either their\n\u003e full battery can't run without a live cloud backend (the BigQuery / Snowflake /\n\u003e Elasticsearch sinks are tested against wiremock, which can't prove real\n\u003e idempotent dedup; the GCS source/sink need a real gRPC backend the emulator\n\u003e doesn't provide), or the connector's shape doesn't fit a check (the webhook\n\u003e source is buffer-shaped; the append-only Iceberg sink's terminal `flush`\n\u003e doesn't fit the effectively-once replay check). They still have their own\n\u003e extensive integration tests (wiremock / testcontainers) and are used in\n\u003e production — Tier-2 means \"not certified,\" **not** \"low quality.\" The Singer\n\u003e bridge is additionally **experimental (v0, single-stream)** ⚠️.\n\n### Sources (33)\n\n`Tier`: **T1 ✅** = passes the `faucet-conformance` battery in CI; **T2** = not yet\nwired into the battery (see the support-tiers note above).\n\n| Crate | Tier | Description |\n|-------|------|-------------|\n| [`faucet-source-rest`](crates/source/rest) | **T1 ✅ᵐ** | REST API — auth, pagination, extraction, schema inference |\n| [`faucet-source-graphql`](crates/source/graphql) | T1 ✅ᵐ | GraphQL API — cursor-based pagination, variable injection |\n| [`faucet-source-xml`](crates/source/xml) | T1 ✅ᵐ | XML/SOAP API — XML-to-JSON conversion, dot-path extraction |\n| [`faucet-source-grpc`](crates/source/grpc) | T1 ✅ | gRPC — dynamic protobuf via `prost-reflect`, unary + server-streaming |\n| [`faucet-source-postgres`](crates/source/postgres) | **T1 ✅** | PostgreSQL — run SQL queries, return rows as JSON |\n| [`faucet-source-postgres-cdc`](crates/source/postgres-cdc) | T1 ✅ | PostgreSQL CDC — logical replication via pgoutput, resumable |\n| [`faucet-source-mysql`](crates/source/mysql) | T1 ✅ | MySQL — run SQL queries, return rows as JSON |\n| [`faucet-source-mysql-cdc`](crates/source/mysql-cdc) | T1 ✅ | MySQL CDC — binlog row events, resumable via file/pos or GTID |\n| [`faucet-source-mssql`](crates/source/mssql) | T1 ✅ | Microsoft SQL Server — streaming queries, incremental replication |\n| [`faucet-source-mssql-cdc`](crates/source/mssql-cdc) | T1 ✅ | Microsoft SQL Server CDC — change tables (`fn_cdc_get_all_changes`), LSN bookmarks, resumable |\n| [`faucet-source-sqlite`](crates/source/sqlite) | **T1 ✅** | SQLite — run SQL queries, return rows as JSON |\n| [`faucet-source-mongodb`](crates/source/mongodb) | T1 ✅ | MongoDB — find() with filter, projection, sort |\n| [`faucet-source-mongodb-cdc`](crates/source/mongodb-cdc) | T1 ✅ | MongoDB CDC — Change Streams, resumable via resumeToken |\n| [`faucet-source-redis`](crates/source/redis) | T1 ✅ | Redis — read from streams, lists, or key patterns |\n| [`faucet-source-kafka`](crates/source/kafka) | T1 ✅ | Apache Kafka — consumer with idle/max-messages termination |\n| [`faucet-source-kinesis`](crates/source/kinesis) | T1 ✅ | AWS Kinesis Data Streams — sharded consumer with resumable sequence checkpoints |\n| [`faucet-source-pubsub`](crates/source/pubsub) | T2 | Google Cloud Pub/Sub — streaming pull; per-message records + attributes, resumable |\n| [`faucet-source-s3`](crates/source/s3) | T1 ✅ | AWS S3 — read objects as JSONL, JSON array, or raw text |\n| [`faucet-source-gcs`](crates/source/gcs) | T2 | Google Cloud Storage — read objects as JSONL, JSON array, or raw text |\n| [`faucet-source-azure-blob`](crates/source/azure-blob) | T1 ✅ | Azure Blob / ADLS Gen2 — read objects as JSONL, JSON array, or raw text |\n| [`faucet-source-parquet`](crates/source/parquet) | T1 ✅ | Apache Parquet — local file, glob, or S3; vectorized Arrow reader, projection |\n| [`faucet-source-delta`](crates/source/delta) | T2 | Apache Delta Lake — local FS or S3/Azure/GCS; time travel, projection pushdown |\n| [`faucet-source-databricks`](crates/source/databricks) | T3 | Databricks SQL query source (Statement Execution API) — typed rows, chunk pagination, incremental |\n| [`faucet-source-redshift`](crates/source/redshift) | T1 ✅ | Amazon Redshift — SQL query over the PostgreSQL wire, incremental replication |\n| [`faucet-source-clickhouse`](crates/source/clickhouse) | T2 | ClickHouse — HTTP interface, `FORMAT JSONEachRow` streaming, incremental replication |\n| [`faucet-source-elasticsearch`](crates/source/elasticsearch) | T1 ✅ᵐ | Elasticsearch — search/scroll API |\n| [`faucet-source-bigquery`](crates/source/bigquery) | T1 ✅ᵐ | Google BigQuery — `jobs.query` + `getQueryResults`, type-aware decoding |\n| [`faucet-source-snowflake`](crates/source/snowflake) | T1 ✅ᵐ | Snowflake — SQL REST API, server-side partition pagination, JWT / OAuth |\n| [`faucet-source-spanner`](crates/source/spanner) | T1 ✅ᵉ | Google Cloud Spanner — streaming SQL over gRPC, incremental replication, stale reads, PK-range sharding |\n| [`faucet-source-webhook`](crates/source/webhook) | T2 | Webhook — temporary HTTP server collecting POST payloads |\n| [`faucet-source-websocket`](crates/source/websocket) | T1 ✅ | WebSocket — live streaming feed; subscribe frames, reconnect, keepalive |\n| [`faucet-source-csv`](crates/source/csv) | **T1 ✅** | CSV — read CSV files as JSON objects |\n| [`faucet-source-singer`](crates/source/singer) | T2 ⚠️ | **Singer tap bridge** — run any Singer tap and adapt its output. Passes the battery, but **experimental (v0, single-stream)** |\n\n### Sinks (25)\n\n| Crate | Tier | Description |\n|-------|------|-------------|\n| [`faucet-sink-bigquery`](crates/sink/bigquery) | T2 | Google BigQuery — streaming inserts; effectively-once via MERGE |\n| [`faucet-sink-iceberg`](crates/sink/iceberg) | T2 | Apache Iceberg — append snapshots via REST/Glue/SQL/HMS catalogs |\n| [`faucet-sink-postgres`](crates/sink/postgres) | T1 ✅ | PostgreSQL — JSONB or auto-mapped columns; upsert/delete |\n| [`faucet-sink-mysql`](crates/sink/mysql) | T1 ✅ | MySQL — JSON column or auto-mapped columns; upsert/delete |\n| [`faucet-sink-mssql`](crates/sink/mssql) | T1 ✅ | Microsoft SQL Server — JSON or auto-mapped columns, 2100-param split |\n| [`faucet-sink-sqlite`](crates/sink/sqlite) | **T1 ✅** | SQLite — JSON column or auto-mapped columns; upsert/delete; effectively-once |\n| [`faucet-sink-snowflake`](crates/sink/snowflake) | T2 | Snowflake — SQL REST API with JWT/OAuth |\n| [`faucet-sink-redshift`](crates/sink/redshift) | T1 ✅ | Amazon Redshift — COPY-from-S3 (staged) or multi-row `INSERT`; append-only |\n| [`faucet-sink-clickhouse`](crates/sink/clickhouse) | T2 | ClickHouse — `INSERT … FORMAT JSONEachRow`; optional `async_insert`; append-only |\n| [`faucet-sink-mongodb`](crates/sink/mongodb) | T1 ✅ | MongoDB — insert_many; upsert/delete by key |\n| [`faucet-sink-redis`](crates/sink/redis) | T1 ✅ | Redis — write to streams, lists, or key-value |\n| [`faucet-sink-kafka`](crates/sink/kafka) | T1 ✅ | Apache Kafka — producer with batching, multi-topic routing |\n| [`faucet-sink-kinesis`](crates/sink/kinesis) | T1 ✅ | AWS Kinesis Data Streams — batched PutRecords with partition-key routing |\n| [`faucet-sink-pubsub`](crates/sink/pubsub) | T2 | Google Cloud Pub/Sub — batched publish; optional ordering key, per-entry retry |\n| [`faucet-sink-spanner`](crates/sink/spanner) | T1 ✅ᵉ | Google Cloud Spanner — batched mutations; upsert/delete, effectively-once commit tokens, schema evolution |\n| [`faucet-sink-elasticsearch`](crates/sink/elasticsearch) | T2 | Elasticsearch — bulk index API; upsert/delete by `_id` |\n| [`faucet-sink-s3`](crates/sink/s3) | T1 ✅ | AWS S3 — write JSONL files to bucket |\n| [`faucet-sink-gcs`](crates/sink/gcs) | T2 | Google Cloud Storage — write JSONL files to bucket |\n| [`faucet-sink-azure-blob`](crates/sink/azure-blob) | T2 | Azure Blob / ADLS Gen2 — write JSONL blobs; batch/byte rollover |\n| [`faucet-sink-parquet`](crates/sink/parquet) | T1 ✅ | Apache Parquet — local file or S3; schema inference, row/byte rollover |\n| [`faucet-sink-delta`](crates/sink/delta) | T2 | Apache Delta Lake — append-only; local FS or S3/Azure/GCS; one commit per flush |\n| [`faucet-sink-jsonl`](crates/sink/jsonl) | **T1 ✅** | JSON Lines — file output with append/truncate |\n| [`faucet-sink-csv`](crates/sink/csv) | T1 ✅ | CSV — write JSON records as CSV rows |\n| [`faucet-sink-http`](crates/sink/http) | T1 ✅ᵐ | HTTP — POST records to any endpoint |\n| [`faucet-sink-stdout`](crates/sink/stdout) | T1 ✅ | Stdout/stderr — JSON Lines, pretty JSON, or TSV |\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003eSupporting crates\u003c/b\u003e — core, shared connector libraries, state stores, lineage, transforms, umbrella, CLI\u003c/summary\u003e\n\n| Crate | Description |\n|-------|-------------|\n| [`faucet-core`](crates/core) | Shared types, traits (`Source`, `Sink`, `AuthProvider`), pipeline orchestration, transforms, error types |\n| [`faucet-auth`](crates/auth) | Shared single-flight auth providers (OAuth2, token-endpoint) for `auth: { ref }` |\n| [`faucet-common-bigquery`](crates/common/bigquery) | Shared BigQuery types — `BigQueryCredentials` enum and `build_client` helper |\n| [`faucet-common-elasticsearch`](crates/common/elasticsearch) | Shared `ElasticsearchAuth` enum for Elasticsearch source/sink |\n| [`faucet-common-gcs`](crates/common/gcs) | Shared GCS types — credentials enum, Storage/StorageControl client builders |\n| [`faucet-common-kafka`](crates/common/kafka) | Shared Kafka types — auth, value formats, Schema Registry client |\n| [`faucet-common-snowflake`](crates/common/snowflake) | Shared Snowflake types — `SnowflakeAuth` enum + auth header helpers |\n| [`faucet-common-spanner`](crates/common/spanner) | Shared Cloud Spanner types — credentials enum, connection/client builders, value conversion |\n| [`faucet-common-mssql`](crates/common/mssql) | Shared MSSQL types — connection/TLS config, `tiberius`+`bb8` pool builder, identifier quoting |\n| [`faucet-state-redis`](crates/state/redis) | Redis-backed `StateStore` for persistent bookmarks |\n| [`faucet-state-postgres`](crates/state/postgres) | PostgreSQL-backed `StateStore` for persistent bookmarks |\n| [`faucet-lineage`](crates/lineage) | OpenLineage event emission — HTTP/file/Kafka transports, schema facets, column-lineage analysis |\n| [`faucet-transform-sql`](crates/transform-sql) | Embedded DuckDB SQL transform — run DuckDB SQL over each page (`batch` relation) |\n| [`faucet-stream`](faucet-stream) | Umbrella crate — feature-gated re-exports of all connectors and state backends |\n| [`faucet-cli`](cli) | `faucet` binary — YAML/JSON config-driven pipeline runner (`run`, `validate`, `schema`, `list`, `preview`, `init`, `doctor`, `test`, `schedule`, `serve`) |\n\n\u003c/details\u003e\n\n### Install\n\n**CLI — prebuilt binaries** (macOS arm64/x86_64, Linux x86_64/aarch64; no Rust toolchain\nrequired). Includes every first-party connector plus `serve` (web console), `schedule`,\nand `lineage`:\n\n```bash\n# Homebrew\nbrew install PawanSikawat/faucet-stream/faucet-cli\n\n# Shell installer (installs to ~/.cargo/bin by default, checksummed)\ncurl -LsSf https://github.com/PawanSikawat/faucet-stream/releases/latest/download/faucet-cli-installer.sh | sh\n\n# Or download an archive + SHA256 checksum from the latest faucet-cli GitHub Release\n```\n\n**CLI — from source** (full/custom feature sets, e.g. DuckDB SQL transforms, OTLP export):\n\n```bash\ncargo install faucet-cli                   # default feature set\ncargo install faucet-cli --features full   # everything\n```\n\n**Library:**\n\n```bash\n# Everything (default includes the REST source)\ncargo add faucet-stream\n\n# All sources / all sinks / all connectors\ncargo add faucet-stream --features source\ncargo add faucet-stream --features sink\ncargo add faucet-stream --features full\n\n# Pick individual connectors\ncargo add faucet-stream --features source-rest,sink-postgres,sink-s3\n\n# Or depend on individual connector crates directly\ncargo add faucet-source-rest\ncargo add faucet-source-mongodb\n```\n\n## How it compares\n\nThere are many great data-movement tools. faucet-stream's niche is being **a single fast\nnative binary _and_ an embeddable Rust library** — config-driven, with no Python runtime, no\nplatform to operate, and a typed library API when you want to compile pipelines into your\nown service.\n\n| | **faucet-stream** | Meltano (Singer) | Airbyte | Benthos / Redpanda Connect | Vector | Fivetran |\n|---|---|---|---|---|---|---|\n| Runtime | Rust, native binary | Python | Java/Python on Docker | Go, native binary | Rust, native binary | Hosted SaaS |\n| Single static binary | ✓ | ✗ | ✗ | ✓ | ✓ | n/a |\n| Config-driven (YAML/JSON) | ✓ | ✓ | via UI/API | ✓ | ✓ | via UI |\n| Embeddable as a library | ✓ (Rust) | ✗ | ✗ | ✓ (Go) | ✗ | ✗ |\n| Connector count | 58, growing | 600+ taps | 350+ | dozens | dozens | 500+ |\n| Change data capture | ✓ Postgres / MySQL / Mongo / SQL Server | partial¹ | ✓ | partial | ✗ | ✓ |\n| Incremental + resumable state | ✓ | ✓ | ✓ | partial | n/a | ✓ |\n| Effectively-once delivery³ | ✓ (11 sinks incl. Kafka, Iceberg, BigQuery) | ✗ | partial | ✗ | ✗ | ✓ |\n| Built-in data-quality checks | ✓ native | ✗ | paywalled add-on | ✗ | ✗ | paywalled add-on |\n| Built-in metrics + tracing | ✓ Prometheus + OTLP + `tracing` | partial | ✓ (platform) | ✓ | ✓ | ✓ (hosted) |\n| Self-hosted, no daemon | ✓ run-to-completion | ✓ | ✗ needs platform | usually a service | agent | ✗ SaaS |\n| License | MIT / Apache-2.0 | MIT | ELv2 + MIT | Apache-2.0 / source-available² | MPL-2.0 | Proprietary |\n\n¹ Singer CDC depends on the individual tap. ² The original Benthos is Apache-2.0; Redpanda Connect's maintained build is source-available. ³ \"Effectively-once\" = idempotent at-least-once: per-page commit tokens are committed atomically with the data so a resumed run drops duplicates — not distributed-consensus exactly-once (see [delivery guarantees](https://pawansikawat.github.io/faucet-stream/cookbook/state.html)). *Comparison reflects the general shape of each tool as of 2026-05 — check each project for current details.*\n\nFor reference: **[Singer](https://www.singer.io/)** is a connector spec and **[Meltano](https://meltano.com/)**\nis its most common runtime; both appear above. faucet-stream is a full **ETL** tool — it\ntransforms data *in flight* as it moves (11 record transforms plus filter,\nexplode, and CDC-unwrap stages and a page-level embedded-DuckDB `sql` transform\nfor aggregation, joins, and filtering), not just extract-and-load. **[dbt](https://www.getdbt.com/) is complementary, not a competitor:**\nit models transformations *in the warehouse* on data already loaded (the \"T\" of ELT, at\nwarehouse scale) — pair the two when you need heavy in-warehouse modeling on top of what\nfaucet extracts, transforms, and loads.\n\n**Deep dives** (detailed, and honest about where each tool wins): [vs. Meltano](https://pawansikawat.github.io/faucet-stream/comparison/meltano.html) · [vs. Airbyte](https://pawansikawat.github.io/faucet-stream/comparison/airbyte.html) · [vs. Singer](https://pawansikawat.github.io/faucet-stream/comparison/singer.html).\n\n## When to use faucet-stream\n\n**Reach for it when:**\n\n- You want **one fast static binary** (or a Rust library) to move data between APIs, databases, object stores, and warehouses — without standing up a platform, scheduler, or Python environment.\n- You want **version-controlled, config-driven pipelines** you can run anywhere: locally, in CI, behind cron, or inside another service.\n- You need **streaming with bounded memory, incremental/resumable replication, CDC, effectively-once delivery, data-quality assertions, retries, dead-letter queues, and metrics** without hand-writing that plumbing.\n- You're **already in Rust** and want typed `Source`/`Sink` traits you can embed and extend.\n\n**Look elsewhere (for now) when:**\n\n- You need a connector faucet-stream **doesn't ship yet and can't write** — [Meltano](https://meltano.com/) (600+ Singer taps) and [Airbyte](https://airbyte.com/) (350+) have far broader catalogs today.\n- You want a **fully-managed, hosted service** with a UI and a team operating it — Fivetran or Airbyte Cloud.\n- Your job is **heavy in-warehouse transformation modeling** — a tested, version-controlled SQL DAG built on data already in the warehouse, at warehouse scale. That's dbt's domain; pair it with faucet-stream, which extracts, transforms in-flight, and loads.\n- You need a **continuous record-by-record streaming processor** — [Benthos / Redpanda Connect](https://www.redpanda.com/connect) and [Vector](https://vector.dev/) are purpose-built for that. faucet-stream runs discrete pipelines to completion; even the long-running modes (`faucet schedule`, `faucet serve`) orchestrate complete runs rather than a never-ending stream.\n\n## Architecture\n\nA `Source` streams batches of records, optional `Transform`s reshape them, and the\n`Pipeline` writes each batch to a `Sink` — bounding memory at one batch on both sides\nregardless of total volume. The pipeline also drives the cross-cutting runtime (bookmarks,\ndead-letter routing, quality checks, metrics) so connectors stay simple:\n\n```mermaid\n%%{init: {'theme':'base','flowchart':{'curve':'basis','nodeSpacing':50,'rankSpacing':72,'padding':14},'themeVariables':{'fontFamily':'-apple-system,BlinkMacSystemFont,Segoe UI,sans-serif','fontSize':'14px','lineColor':'#a5b4c4','clusterBkg':'#f8fafc','clusterBorder':'#e2e8f0'}}}%%\nflowchart LR\n    S[\"\u003cb\u003eSource\u003c/b\u003e\u003cbr/\u003eREST · DB · CDC\u003cbr/\u003eKafka · S3 · Parquet\"]\n    T[\"\u003cb\u003eTransforms\u003c/b\u003e\u003cbr/\u003eflatten · rename · keys_case\u003cbr/\u003eselect · drop · set · cast\u003cbr/\u003eredact · value_case · spell_symbols\u003cbr/\u003esql (DuckDB, page-level)\"]\n    P{{\"\u003cb\u003ePipeline\u003c/b\u003e\"}}\n    K[\"\u003cb\u003eSink\u003c/b\u003e\u003cbr/\u003eBigQuery · Postgres\u003cbr/\u003eParquet · Kafka · ...\"]\n    ST[(\"State store\u003cbr/\u003efile · Redis · Postgres\")]\n    D[(\"Dead-letter\u003cbr/\u003equeue\")]\n    O([\"Prometheus\u003cbr/\u003e+ tracing\"])\n\n    S --\u003e|StreamPage batches| T --\u003e P --\u003e|write_batch| K\n    P -.-\u003e|bookmark per page| ST\n    ST -.-\u003e|resume from bookmark| S\n    P -.-\u003e|failed rows| D\n    P -.-\u003e|metrics + spans| O\n    classDef src fill:#e0f2f1,stroke:#26a69a,stroke-width:1.5px,color:#00695c\n    classDef proc fill:#eceff8,stroke:#7986cb,stroke-width:1.5px,color:#303f9f\n    classDef dec fill:#fff3e0,stroke:#ffa726,stroke-width:1.5px,color:#e65100\n    classDef bad fill:#fdecec,stroke:#ef9a9a,stroke-width:1.5px,color:#c62828\n    classDef store fill:#f3e5f5,stroke:#ab47bc,stroke-width:1.5px,color:#6a1b9a\n    classDef sink fill:#e3f2fd,stroke:#42a5f5,stroke-width:1.5px,color:#1565c0\n    class S src\n    class T,O proc\n    class P dec\n    class D bad\n    class ST store\n    class K sink\n```\n\nfaucet-stream is a Cargo workspace with **80 crates** — 33 sources, 25 sinks, 13 shared\nconnector libraries, the shared auth-provider library, 2 state-store backends, the lineage\ncrate, the SQL transform crate, the conformance test battery, the shared core, the umbrella\ncrate, and the CLI binary. See\nthe [Connectors](#connectors) table above and the\n[architecture guide](https://pawansikawat.github.io/faucet-stream/getting-started/concepts.html).\n\n## Performance\n\n\u003e **Performance and reliability are why this library exists.** Every connector is optimised\n\u003e for throughput out of the box — there are no \"slow defaults\" to tune away.\n\n| Technique | Where |\n|-----------|-------|\n| **Parallel I/O** | S3/GCS read/write objects concurrently (configurable `concurrency`); HTTP sink sends requests in parallel; REST source processes partitions concurrently |\n| **Multi-row INSERT** | PostgreSQL, MySQL, SQLite, and SQL Server sinks batch records into single INSERT statements instead of one per row (MSSQL auto-splits at the 2100-parameter limit) |\n| **Transaction wrapping** | SQLite sink wraps batches in `BEGIN`/`COMMIT` for a large write speedup |\n| **Connection pooling** | All database connectors use connection pools with configurable `max_connections` |\n| **Connection reuse** | S3, MongoDB, Redis, Elasticsearch, and HTTP connectors create clients once and reuse them across all operations |\n| **Redis pipelining** | Redis sink batches commands with `pipe()`; Redis source uses `MGET` for bulk key reads |\n| **Bulk APIs** | Elasticsearch uses the bulk NDJSON API; BigQuery uses `insertAll`; MongoDB uses `insert_many` |\n| **Buffered I/O** | JSONL sink uses `BufWriter`; CSV uses buffered readers/writers in blocking threads |\n| **Streaming pagination** | Sources stream pages one at a time via `stream_pages()` to bound memory |\n| **Adaptive batch sizing** | Opt-in AIMD controller tunes the effective write batch size from observed sink latency and error rate |\n\n## Observability\n\nEvery pipeline emits OTel-compatible `tracing` spans and Prometheus metrics automatically —\nlabelled by `pipeline`, `row` (matrix row id), and `connector`, with **zero per-connector\ncode**. The CLI exposes a `/metrics` endpoint via the optional `observability:` block in\n`faucet.yaml`. See the [CLI README](cli/README.md#observability-prometheus--tracing) for the\nYAML grammar and the OpenTelemetry bridge snippet.\n\n## Feature flags\n\nDefault features: `source-rest`, `transform-flatten`, `transform-rename-keys`,\n`transform-keys-case`. Every connector and capability is an opt-in Cargo feature.\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003eFull feature-flag table (umbrella crate)\u003c/b\u003e\u003c/summary\u003e\n\n| Feature | Default | Description |\n|---------|---------|-------------|\n| `source-rest` | yes | REST API source |\n| `source-graphql` | no | GraphQL API source |\n| `source-xml` | no | XML/SOAP API source |\n| `source-grpc` | no | gRPC source |\n| `source-postgres` | no | PostgreSQL query source |\n| `source-postgres-cdc` | no | PostgreSQL CDC source (logical replication) |\n| `source-mysql` | no | MySQL query source |\n| `source-mysql-cdc` | no | MySQL CDC source (binlog replication) |\n| `source-mssql` | no | Microsoft SQL Server query source |\n| `source-mssql-cdc` | no | Microsoft SQL Server CDC source (change tables) |\n| `source-sqlite` | no | SQLite query source |\n| `source-mongodb` | no | MongoDB query source |\n| `source-mongodb-cdc` | no | MongoDB CDC source (Change Streams) |\n| `source-redis` | no | Redis source |\n| `source-kafka` | no | Apache Kafka consumer source |\n| `source-kinesis` | no | AWS Kinesis Data Streams source |\n| `source-pubsub` | no | Google Cloud Pub/Sub source |\n| `source-s3` | no | AWS S3 file source |\n| `source-gcs` | no | Google Cloud Storage file source |\n| `source-azure-blob` | no | Azure Blob / ADLS Gen2 file source |\n| `source-parquet` | no | Apache Parquet file source (local, glob, S3) |\n| `source-delta` | no | Apache Delta Lake source (local FS; S3/Azure/GCS via `delta-s3`/`delta-azure`/`delta-gcs`) |\n| `source-databricks` | no | Databricks SQL query source (Statement Execution API) |\n| `sink-delta` | no | Apache Delta Lake sink (local FS; S3/Azure/GCS via `delta-s3`/`delta-azure`/`delta-gcs`) |\n| `source-elasticsearch` | no | Elasticsearch source |\n| `source-bigquery` | no | Google BigQuery query source |\n| `source-snowflake` | no | Snowflake query source |\n| `source-redshift` | no | Amazon Redshift query source (PostgreSQL wire) |\n| `source-clickhouse` | no | ClickHouse query source (HTTP interface) |\n| `source-spanner` | no | Google Cloud Spanner query source |\n| `source-webhook` | no | Webhook HTTP receiver |\n| `source-websocket` | no | WebSocket live streaming source |\n| `source-csv` | no | CSV file source |\n| `sink-bigquery` | no | Google BigQuery sink |\n| `sink-iceberg` | no | Apache Iceberg sink (append, REST/Glue/SQL/HMS catalogs) |\n| `sink-postgres` | no | PostgreSQL sink |\n| `sink-mysql` | no | MySQL sink |\n| `sink-mssql` | no | Microsoft SQL Server sink |\n| `sink-sqlite` | no | SQLite sink |\n| `sink-snowflake` | no | Snowflake sink |\n| `sink-redshift` | no | Amazon Redshift sink (COPY-from-S3 or multi-row INSERT) |\n| `sink-clickhouse` | no | ClickHouse sink (INSERT … FORMAT JSONEachRow) |\n| `sink-mongodb` | no | MongoDB sink |\n| `sink-redis` | no | Redis sink |\n| `sink-kafka` | no | Apache Kafka producer sink |\n| `sink-kinesis` | no | AWS Kinesis Data Streams sink |\n| `sink-pubsub` | no | Google Cloud Pub/Sub sink |\n| `sink-spanner` | no | Google Cloud Spanner sink |\n| `sink-elasticsearch` | no | Elasticsearch bulk index sink |\n| `sink-s3` | no | AWS S3 file sink |\n| `sink-gcs` | no | Google Cloud Storage file sink |\n| `sink-azure-blob` | no | Azure Blob / ADLS Gen2 file sink |\n| `sink-parquet` | no | Apache Parquet file sink (local, S3) |\n| `sink-jsonl` | no | JSON Lines file sink |\n| `sink-csv` | no | CSV file sink |\n| `sink-http` | no | HTTP POST sink |\n| `sink-stdout` | no | Stdout/stderr sink (JSON Lines, pretty JSON, TSV) |\n| `kafka-schema-registry` | no | Confluent Schema Registry support for the Kafka pair (Avro, Protobuf, JSON Schema) |\n| `state-redis` | no | Redis-backed `StateStore` backend |\n| `state-postgres` | no | PostgreSQL-backed `StateStore` backend |\n| `source` / `sink` / `state` | no | All sources / all sinks / all state backends |\n| `auth` | no | Shared OAuth2 / token-endpoint auth providers |\n| `full` | no | Every connector, state backend, and capability |\n| `transform-flatten` | yes | Flatten nested objects |\n| `transform-rename-keys` | yes | Regex key renaming |\n| `transform-keys-case` | yes | Re-case every key — snake / camel / pascal / kebab / screaming_snake |\n| `transform-select` / `transform-drop` / `transform-set` | no | Keep / remove / add top-level fields |\n| `transform-rename-field` | no | Exact-name field rename (single or batch) |\n| `transform-cast` | no | Per-field type coercion with `on_error` policy |\n| `transform-redact` | no | Replace listed field values with a mask |\n| `transform-value-case` | no | Lowercase / uppercase / trim string field values |\n| `transform-spell-symbols` | no | Spell out symbols in keys (`%` → `percent`, `#` → `number`, …) |\n| `transform-cdc-unwrap` | no | Normalize a CDC envelope into a flat row + `__op` marker (pairs with upsert sinks) |\n| `transforms` | no | All built-in transforms above |\n| `transform-sql` | no | Embedded DuckDB SQL transform — run DuckDB SQL over each page (`batch` relation; `batch_size: 0` for global aggregation) |\n| `compression` | no | gzip / zstd read+write on JSONL/CSV/S3/GCS source and sink connectors |\n| `encryption` | no | Encryption at rest (AES-256-GCM) for `file` state-store bookmarks and per-line JSONL/DLQ output |\n| `cli-tui` | no | Live terminal UI for `faucet run --tui` — per-invocation throughput, errors, DLQ, bookmark age |\n\n`RecordTransform::Custom` is always available regardless of feature flags. CLI-only features\n(`schedule`, `serve`, `serve-ui`, `serve-history-*`, `triggers*`, `lineage`, `quality`,\n`contract`, `secrets-*`) live in `faucet-cli`, not the umbrella crate — see\n[`cli/README.md`](cli/README.md).\n\n\u003c/details\u003e\n\n## Using faucet-stream as a Rust library\n\nThe CLI is just one consumer of the engine. Everything it does is available as a typed API.\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003ePagination styles\u003c/b\u003e — cursor, page-number, offset, Link header, next-link-in-body\u003c/summary\u003e\n\n```rust\nuse faucet_stream::{RestStream, RestStreamConfig, Auth, PaginationStyle};\n\n// Cursor-based pagination with Bearer auth\nlet stream = RestStream::new(\n    RestStreamConfig::new(\"https://api.example.com\", \"/v1/users\")\n        .auth(Auth::Bearer { token: \"my-token\".into() })\n        .records_path(\"$.data[*]\")\n        .pagination(PaginationStyle::Cursor {\n            next_token_path: \"$.meta.next_cursor\".into(),\n            param_name: \"cursor\".into(),\n        })\n        .max_pages(50),\n)?;\nlet users: Vec\u003cserde_json::Value\u003e = stream.fetch_all().await?;\n\n// Page-number pagination with an API key\nlet stream = RestStream::new(\n    RestStreamConfig::new(\"https://api.example.com\", \"/v2/orders\")\n        .auth(Auth::ApiKey { header: \"X-Api-Key\".into(), value: \"secret\".into() })\n        .records_path(\"$.results[*]\")\n        .pagination(PaginationStyle::PageNumber {\n            param_name: \"page\".into(),\n            start_page: 1,\n            page_size: Some(100),\n            page_size_param: Some(\"per_page\".into()),\n        }),\n)?;\n```\n\n| Style | Use when |\n|-------|----------|\n| `Cursor` | API returns a next-page token in the response body |\n| `PageNumber` | API uses `?page=1\u0026per_page=100` style |\n| `Offset` | API uses `?offset=0\u0026limit=50` style |\n| `LinkHeader` | API returns pagination in the `Link` HTTP header (GitHub-style) |\n| `NextLinkInBody` | API returns the full next-page URL in the response body |\n\nEvery pagination style has a termination/loop guard. `Cursor`, `LinkHeader`, and\n`NextLinkInBody` stop when the same token/link repeats; `PageNumber` stops on a zero-record\npage or a repeated page body (content-fingerprint detection); `Offset` stops when the offset\nreaches `total` or a page returns fewer records than the limit. `max_pages` is a hard cap\nacross all styles.\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003eAuthentication\u003c/b\u003e — Bearer, Basic, API key, OAuth2, token endpoint, custom\u003c/summary\u003e\n\n| Method | Description |\n|--------|-------------|\n| `Bearer` | `Authorization: Bearer \u003ctoken\u003e` header |\n| `Basic` | `Authorization: Basic \u003cbase64\u003e` header |\n| `ApiKey` | Custom header (e.g. `X-Api-Key: secret`) |\n| `ApiKeyQuery` | API key as a query parameter (e.g. `?api_key=secret`) |\n| `OAuth2` | Client-credentials flow with automatic token caching and refresh |\n| `TokenEndpoint` | Fetch a token from any HTTP API via JSONPath, with caching and refresh |\n| `Custom` | Arbitrary headers |\n\n```rust\nuse faucet_stream::{Auth, fetch_oauth2_token};\n\n// OAuth2 client credentials\nlet token = fetch_oauth2_token(\n    \"https://auth.example.com/oauth/token\",\n    \"client-id\",\n    \"client-secret\",\n    \u0026[\"read:data\".into()],\n).await?;\nlet config = RestStreamConfig::new(\"https://api.example.com\", \"/data\")\n    .auth(Auth::Bearer { token });\n```\n\n`Auth::TokenEndpoint` fetches a token from an external API (a login endpoint, secrets\nmanager, or custom auth service) via a JSONPath, caches it across pages, and refreshes it at\n`expiry_ratio` of the reported lifetime (default 90%). See the\n[auth cookbook](https://pawansikawat.github.io/faucet-stream/cookbook/auth.html) for the full\nshape and `ResponseValidator` customization.\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003eStreaming, incremental replication, partitions, and typed deserialization\u003c/b\u003e\u003c/summary\u003e\n\n```rust\nuse faucet_stream::{RestStream, RestStreamConfig, ReplicationMethod, PaginationStyle};\nuse futures::StreamExt;\nuse serde::Deserialize;\nuse serde_json::json;\n\n// Stream page-by-page (bounded memory)\nlet mut pages = stream.stream_pages();\nwhile let Some(result) = pages.next().await {\n    let records = result?;\n    println!(\"processing page of {} records\", records.len());\n}\n\n// Incremental replication — only fetch records newer than a stored bookmark\nlet stream = RestStream::new(\n    RestStreamConfig::new(\"https://api.example.com\", \"/events\")\n        .records_path(\"$.data[*]\")\n        .replication_method(ReplicationMethod::Incremental)\n        .replication_key(\"updated_at\")\n        .start_replication_value(json!(\"2024-06-01T00:00:00Z\")),\n)?;\nlet (records, bookmark) = stream.fetch_all_incremental().await?;  // persist `bookmark` for next run\n\n// Typed deserialization straight into your structs\n#[derive(Debug, Deserialize)]\nstruct User { id: u64, name: String, email: String }\nlet users: Vec\u003cUser\u003e = stream.fetch_all_as::\u003cUser\u003e().await?;\n```\n\n`add_partition(...)` runs the same stream config across multiple contexts (e.g. per-org,\nper-repo) and concatenates the results.\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003eTransforms\u003c/b\u003e — reshape records as they're extracted\u003c/summary\u003e\n\nWrap any `Source` with `TransformingSource`. Built-in transforms are feature-gated (the\nthree case/flatten/rename ones are on by default):\n\n```rust\nuse faucet_stream::{\n    KeyCaseMode, Labels, RecordTransform, RestStream, RestStreamConfig, Source, TransformingSource,\n};\n\nlet inner = RestStream::new(\n    RestStreamConfig::new(\"https://api.example.com\", \"/data\").records_path(\"$.results[*]\"),\n)?;\nlet stream = TransformingSource::new(\n    Box::new(inner) as Box\u003cdyn Source\u003e,\n    vec![\n        RecordTransform::Flatten { separator: \"__\".into() },         // {\"user\":{\"id\":1}} -\u003e {\"user__id\":1}\n        RecordTransform::KeysCase { mode: KeyCaseMode::Snake },        // re-case every key\n        RecordTransform::RenameKeys { pattern: r\"^_sdc_\".into(), replacement: \"\".into() },\n        RecordTransform::custom(|mut record| {                         // arbitrary closure\n            if let serde_json::Value::Object(ref mut map) = record {\n                map.insert(\"_source\".to_string(), serde_json::json!(\"my-api\"));\n            }\n            record\n        }),\n    ],\n    Labels::for_named(\"rest\"),\n)?;\n```\n\n\u003c/details\u003e\n\n\u003cdetails\u003e\n\u003csummary\u003e\u003cb\u003eConfig loading and schema introspection\u003c/b\u003e\u003c/summary\u003e\n\nAll connector configs load from JSON files, environment variables, or `.env` files, and can\ndescribe themselves:\n\n```rust\nuse faucet_core::config::{load_json, load_env, load_env_file};\nuse faucet_source_rest::RestStreamConfig;\n\nlet source: RestStreamConfig = load_json(\"source_config.json\")?;       // from a JSON file\nlet source: RestStreamConfig = load_env(\"REST\")?;                       // from REST_* env vars\nlet source: RestStreamConfig = load_env_file(\".env\", \"REST\")?;          // from a .env file\n\n// Every source/sink can print a full JSON Schema of its config — always in sync with the code\nlet schema = RestStream::new(source)?.config_schema();\nprintln!(\"{}\", serde_json::to_string_pretty(\u0026schema)?);\n```\n\n\u003c/details\u003e\n\n## Building custom connectors\n\nfaucet-stream is a **marketplace ecosystem** — third-party developers can publish their own\n`faucet-source-*` / `faucet-sink-*` crates with minimal friction. The only dependency you\nneed is `faucet-core`; it re-exports everything required (`async_trait`, `serde_json`,\n`Value`, `json!`, `JsonSchema`, `schema_for!`).\n\n```rust\nuse faucet_core::{async_trait, FaucetError, Source, Sink, Value, json, JsonSchema, schema_for};\n\n#[async_trait]\nimpl Source for MySource {\n    async fn fetch_all(\u0026self) -\u003e Result\u003cVec\u003cValue\u003e, FaucetError\u003e {\n        Ok(vec![json!({\"id\": 1, \"name\": \"example\"})])\n    }\n    fn config_schema(\u0026self) -\u003e Value {\n        serde_json::to_value(schema_for!(MySourceConfig)).expect(\"schema serialization\")\n    }\n}\n\n#[async_trait]\nimpl Sink for MySink {\n    async fn write_batch(\u0026self, records: \u0026[Value]) -\u003e Result\u003cusize, FaucetError\u003e {\n        Ok(records.len())\n    }\n    fn config_schema(\u0026self) -\u003e Value {\n        serde_json::to_value(schema_for!(MySinkConfig)).expect(\"schema serialization\")\n    }\n}\n```\n\nAny `Source` works with any `Sink` via `Pipeline::new(\u0026source, \u0026sink).run().await?`. Map your\nown failures to `FaucetError` variants (`Source` / `Sink` / `Config`, or `Custom(boxed_err)`\nto wrap any `std::error::Error` without losing the chain). Publish under the naming\nconvention `faucet-source-\u003cname\u003e` / `faucet-sink-\u003cname\u003e`. Full walkthrough:\n[authoring connectors](https://pawansikawat.github.io/faucet-stream/extending/authoring-connectors.html).\n\nTo use a third-party connector **from a `faucet.yaml` config** (not just from Rust), build a\ncustom `faucet` binary that registers it via `PluginRegistry` and `faucet_cli::run_main` — the\nconnector then works as `type: \u003cname\u003e` across every CLI command with zero runtime overhead. See\n[Custom binaries with third-party connectors](cli/README.md#custom-binaries-with-third-party-connectors)\nand the runnable [`cli/examples/custom-cli/`](cli/examples/custom-cli/main.rs).\n\n## Project structure\n\n```\nCargo.toml                    — workspace manifest (80 crates)\ncrates/\n  core/                       — faucet-core: shared types, traits, pipeline, transforms, config\n  auth/                       — faucet-auth: shared OAuth2 / token-endpoint providers\n  source/                     — 33 source connectors (rest, graphql, xml, grpc, *-cdc, kafka, s3, azure-blob, redshift, clickhouse, pubsub, delta, databricks, singer, …)\n  sink/                       — 25 sink connectors (bigquery, iceberg, delta, postgres, parquet, kafka, redshift, clickhouse, pubsub, azure-blob, …)\n  common/                     — 13 shared connector libraries (bigquery, elasticsearch, gcs, kafka, snowflake, mssql, kinesis, spanner, delta, redshift, pubsub, clickhouse, azure)\n  state/                      — Redis- and Postgres-backed StateStore backends\n  lineage/                    — faucet-lineage: OpenLineage event emission\n  transform-sql/              — faucet-transform-sql: embedded DuckDB SQL transform\nfaucet-stream/                — umbrella crate with feature-gated re-exports\ncli/                          — faucet-cli: `faucet` binary, YAML/JSON pipeline runner\n  examples/                   — ready-to-run pipeline YAMLs\n  tests/                      — assert_cmd + wiremock integration tests\nexamples/                     — repo-level examples: docker-compose infra stack + run index\n  orchestration/              — ELT recipe: faucet (EL) + dbt (T) + Airflow/Dagster\nscripts/                      — helper scripts (try-local.sh interactive demo, cleanup-artifacts.sh)\ndocs/book/                    — mdBook documentation site (source under docs/book/src)\n.github/workflows/            — ci.yml, release-plz.yml, docs.yml (mdBook → GitHub Pages)\n.github/assets/               — brand assets: logo, wordmark, social-preview banner, favicon\n```\n\n## Star history\n\n\u003ca href=\"https://www.star-history.com/?repos=PawanSikawat%2Ffaucet-stream\u0026type=date\u0026legend=top-left\"\u003e\n \u003cpicture\u003e\n   \u003csource media=\"(prefers-color-scheme: dark)\" srcset=\"https://api.star-history.com/chart?repos=PawanSikawat/faucet-stream\u0026type=date\u0026theme=dark\u0026legend=top-left\u0026sealed_token=9G0SLiIBTDIhFTaBki1bRfURPsuWihZWbMamJ6mn6RIc_wQxXlvqeoPe6MLo5Zx2jx5i4s8VgB-Yu-MSjq8bMIzXcDn7hf0oe5BOU7H6RRz2w3Qnia4vog1vgULXzVCY0zPe10ztVjfTTzu2va4KvyZxVZR-8ZfNK3ku3XuVeoFCrIQgo3dtkQXJCmR8\" /\u003e\n   \u003csource media=\"(prefers-color-scheme: light)\" srcset=\"https://api.star-history.com/chart?repos=PawanSikawat/faucet-stream\u0026type=date\u0026legend=top-left\u0026sealed_token=9G0SLiIBTDIhFTaBki1bRfURPsuWihZWbMamJ6mn6RIc_wQxXlvqeoPe6MLo5Zx2jx5i4s8VgB-Yu-MSjq8bMIzXcDn7hf0oe5BOU7H6RRz2w3Qnia4vog1vgULXzVCY0zPe10ztVjfTTzu2va4KvyZxVZR-8ZfNK3ku3XuVeoFCrIQgo3dtkQXJCmR8\" /\u003e\n   \u003cimg alt=\"Star History Chart\" src=\"https://api.star-history.com/chart?repos=PawanSikawat/faucet-stream\u0026type=date\u0026legend=top-left\u0026sealed_token=9G0SLiIBTDIhFTaBki1bRfURPsuWihZWbMamJ6mn6RIc_wQxXlvqeoPe6MLo5Zx2jx5i4s8VgB-Yu-MSjq8bMIzXcDn7hf0oe5BOU7H6RRz2w3Qnia4vog1vgULXzVCY0zPe10ztVjfTTzu2va4KvyZxVZR-8ZfNK3ku3XuVeoFCrIQgo3dtkQXJCmR8\" /\u003e\n \u003c/picture\u003e\n\u003c/a\u003e\n\n**Using faucet-stream in production?** Open a PR to add your team here — real adopters are the\nbest signal for the next person deciding whether to bet a pipeline on it.\n\n## Contributing\n\nContributions — core changes and third-party connectors alike — are welcome. See\n[CONTRIBUTING.md](CONTRIBUTING.md) for setup, the checks CI runs, and the add-a-connector\nchecklist, and the\n[authoring guide](https://pawansikawat.github.io/faucet-stream/extending/authoring-connectors.html)\nfor building your own `faucet-source-*` / `faucet-sink-*` crate. Please review our\n[Code of Conduct](CODE_OF_CONDUCT.md). To report a vulnerability, see [SECURITY.md](SECURITY.md).\n\n## License\n\nLicensed under either of\n\n- Apache License, Version 2.0 ([LICENSE-APACHE](LICENSE-APACHE) or \u003chttp://www.apache.org/licenses/LICENSE-2.0\u003e)\n- MIT license ([LICENSE-MIT](LICENSE-MIT) or \u003chttp://opensource.org/licenses/MIT\u003e)\n\nat your option.\n\nUnless you explicitly state otherwise, any contribution intentionally submitted for inclusion\nin the work by you, as defined in the Apache-2.0 license, shall be dual licensed as above,\nwithout any additional terms or conditions.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FPawanSikawat%2Ffaucet-stream","html_url":"https://awesome.ecosyste.ms/projects/github.com%2FPawanSikawat%2Ffaucet-stream","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FPawanSikawat%2Ffaucet-stream/lists"}