{"id":13565353,"url":"https://github.com/hpgrahsl/kafka-connect-mongodb","last_synced_at":"2026-02-28T13:11:28.009Z","repository":{"id":54427172,"uuid":"80776285","full_name":"hpgrahsl/kafka-connect-mongodb","owner":"hpgrahsl","description":"**Unofficial / Community** Kafka Connect MongoDB Sink Connector -\u003e integrated 2019 into the official MongoDB Kafka Connector here: https://www.mongodb.com/kafka-connector","archived":false,"fork":false,"pushed_at":"2023-04-02T20:02:28.000Z","size":475,"stargazers_count":151,"open_issues_count":8,"forks_count":60,"subscribers_count":19,"default_branch":"master","last_synced_at":"2025-07-12T19:49:38.219Z","etag":null,"topics":["avro","azure-cosmosdb","bson","cdc","change-data-capture","confluent-hub","connector","cosmosdb","debezium","json","kafka","kafka-connect","mongodb","sink","sink-connector"],"latest_commit_sha":null,"homepage":"","language":"Java","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/hpgrahsl.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE.txt","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null}},"created_at":"2017-02-02T22:49:07.000Z","updated_at":"2025-05-13T23:57:06.000Z","dependencies_parsed_at":"2024-01-14T04:12:32.221Z","dependency_job_id":null,"html_url":"https://github.com/hpgrahsl/kafka-connect-mongodb","commit_stats":null,"previous_names":[],"tags_count":6,"template":false,"template_full_name":null,"purl":"pkg:github/hpgrahsl/kafka-connect-mongodb","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hpgrahsl%2Fkafka-connect-mongodb","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hpgrahsl%2Fkafka-connect-mongodb/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hpgrahsl%2Fkafka-connect-mongodb/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hpgrahsl%2Fkafka-connect-mongodb/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/hpgrahsl","download_url":"https://codeload.github.com/hpgrahsl/kafka-connect-mongodb/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hpgrahsl%2Fkafka-connect-mongodb/sbom","scorecard":{"id":470149,"data":{"date":"2025-08-11","repo":{"name":"github.com/hpgrahsl/kafka-connect-mongodb","commit":"772f543b363bf708fa4b2af44b148ac2474d492d"},"scorecard":{"version":"v5.2.1-40-gf6ed084d","commit":"f6ed084d17c9236477efd66e5b258b9d4cc7b389"},"score":3.2,"checks":[{"name":"Token-Permissions","score":-1,"reason":"No tokens found","details":null,"documentation":{"short":"Determines if the project's workflows follow the principle of least privilege.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#token-permissions"}},{"name":"Packaging","score":-1,"reason":"packaging workflow not detected","details":["Warn: no GitHub/GitLab publishing workflow detected."],"documentation":{"short":"Determines if the project is published as a package that others can easily download, install, easily update, and uninstall.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#packaging"}},{"name":"Code-Review","score":1,"reason":"Found 3/30 approved changesets -- score normalized to 1","details":null,"documentation":{"short":"Determines if the project requires human code review before pull requests (aka merge requests) are merged.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#code-review"}},{"name":"Maintained","score":0,"reason":"0 commit(s) and 0 issue activity found in the last 90 days -- score normalized to 0","details":null,"documentation":{"short":"Determines if the project is \"actively maintained\".","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#maintained"}},{"name":"Dangerous-Workflow","score":-1,"reason":"no workflows found","details":null,"documentation":{"short":"Determines if the project's GitHub Action workflows avoid dangerous patterns.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#dangerous-workflow"}},{"name":"Binary-Artifacts","score":10,"reason":"no binaries found in the repo","details":null,"documentation":{"short":"Determines if the project has generated executable (binary) artifacts in the source repository.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#binary-artifacts"}},{"name":"Pinned-Dependencies","score":-1,"reason":"no dependencies found","details":null,"documentation":{"short":"Determines if the project has declared and pinned the dependencies of its build process.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#pinned-dependencies"}},{"name":"CII-Best-Practices","score":0,"reason":"no effort to earn an OpenSSF best practices badge detected","details":null,"documentation":{"short":"Determines if the project has an OpenSSF (formerly CII) Best Practices Badge.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#cii-best-practices"}},{"name":"Security-Policy","score":0,"reason":"security policy file not detected","details":["Warn: no security policy file detected","Warn: no security file to analyze","Warn: no security file to analyze","Warn: no security file to analyze"],"documentation":{"short":"Determines if the project has published a security policy.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#security-policy"}},{"name":"Fuzzing","score":0,"reason":"project is not fuzzed","details":["Warn: no fuzzer integrations found"],"documentation":{"short":"Determines if the project uses fuzzing.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#fuzzing"}},{"name":"License","score":10,"reason":"license file detected","details":["Info: project has a license file: LICENSE.txt:0","Info: FSF or OSI recognized license: Apache License 2.0: LICENSE.txt:0"],"documentation":{"short":"Determines if the project has defined a license.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#license"}},{"name":"Signed-Releases","score":-1,"reason":"no releases found","details":null,"documentation":{"short":"Determines if the project cryptographically signs release artifacts.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#signed-releases"}},{"name":"Branch-Protection","score":0,"reason":"branch protection not enabled on development/release branches","details":["Warn: branch protection not enabled for branch 'master'"],"documentation":{"short":"Determines if the default and release branches are protected with GitHub's branch protection settings.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#branch-protection"}},{"name":"SAST","score":0,"reason":"SAST tool is not run on all commits -- score normalized to 0","details":["Warn: 0 commits out of 28 are checked with a SAST tool"],"documentation":{"short":"Determines if the project uses static code analysis.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#sast"}},{"name":"Vulnerabilities","score":10,"reason":"0 existing vulnerabilities detected","details":null,"documentation":{"short":"Determines if the project has open, known unfixed vulnerabilities.","url":"https://github.com/ossf/scorecard/blob/f6ed084d17c9236477efd66e5b258b9d4cc7b389/docs/checks.md#vulnerabilities"}}]},"last_synced_at":"2025-08-19T13:41:35.641Z","repository_id":54427172,"created_at":"2025-08-19T13:41:35.641Z","updated_at":"2025-08-19T13:41:35.641Z"},"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":29935078,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-02-28T13:00:17.143Z","status":"ssl_error","status_checked_at":"2026-02-28T12:59:13.669Z","response_time":90,"last_error":"SSL_read: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["avro","azure-cosmosdb","bson","cdc","change-data-capture","confluent-hub","connector","cosmosdb","debezium","json","kafka","kafka-connect","mongodb","sink","sink-connector"],"created_at":"2024-08-01T13:01:45.278Z","updated_at":"2026-02-28T13:11:27.964Z","avatar_url":"https://github.com/hpgrahsl.png","language":"Java","funding_links":["https://www.paypal.com/cgi-bin/webscr?cmd=_s-xclick\u0026hosted_button_id=E3P9D3REZXTJS"],"categories":["大数据","Databases","Java"],"sub_categories":["MongoDB"],"readme":"# Kafka Connect MongoDB\n\n[![Build Status](https://travis-ci.org/hpgrahsl/kafka-connect-mongodb.svg?branch=master)](https://travis-ci.org/hpgrahsl/kafka-connect-mongodb) [![Codacy Badge](https://api.codacy.com/project/badge/Grade/9ce80f1868154f02ad839eb76521d582)](https://www.codacy.com/app/hpgrahsl/kafka-connect-mongodb?utm_source=github.com\u0026amp;utm_medium=referral\u0026amp;utm_content=hpgrahsl/kafka-connect-mongodb\u0026amp;utm_campaign=Badge_Grade) [![Codacy Badge](https://api.codacy.com/project/badge/Coverage/9ce80f1868154f02ad839eb76521d582)](https://www.codacy.com/app/hpgrahsl/kafka-connect-mongodb?utm_source=github.com\u0026amp;utm_medium=referral\u0026amp;utm_content=hpgrahsl/kafka-connect-mongodb\u0026amp;utm_campaign=Badge_Coverage)\n[![Maven Central](https://maven-badges.herokuapp.com/maven-central/at.grahsl.kafka.connect/kafka-connect-mongodb/badge.svg)](https://maven-badges.herokuapp.com/maven-central/at.grahsl.kafka.connect/kafka-connect-mongodb)\n[![Donate](https://img.shields.io/badge/Donate-PayPal-green.svg)](https://www.paypal.com/cgi-bin/webscr?cmd=_s-xclick\u0026hosted_button_id=E3P9D3REZXTJS)\n\nIt's a basic [Apache Kafka](https://kafka.apache.org/) [Connect SinkConnector](https://kafka.apache.org/documentation/#connect) for [MongoDB](https://www.mongodb.com/).\nThe connector uses the official MongoDB [Java Driver](http://mongodb.github.io/mongo-java-driver/3.10/).\nFuture releases might additionally support the [asynchronous driver](http://mongodb.github.io/mongo-java-driver/3.10/driver-async/).\n\n### Users / Testimonials\n\n| Company |   |\n|---------|---|\n| [![QUDOSOFT](docs/logos/qudosoft.png)](http://www.qudosoft.de/) | \"As a subsidiary of a well-established major german retailer,\u003cbr/\u003eQudosoft is challenged by incorporating innovative and\u003cbr/\u003eperformant concepts into existing workflows. At the core of\u003cbr/\u003ea novel event-driven architecture, Kafka has been in\u003cbr/\u003eexperimental use since 2016, followed by Connect in 2017.\u003cbr/\u003e\u003cbr/\u003eSince MongoDB is one of our databases of choice, we were\u003cbr/\u003eglad to discover a production-ready sink connector for it.\u003cbr/\u003e We use it, e.g. to persist customer contact events, making\u003cbr/\u003ethem available to applications that aren't integrated into our\u003cbr/\u003eKafka environment. Currently, this MongoDB sink connector\u003cbr/\u003eruns on five workers consuming approx. 50 - 200k AVRO\u003cbr/\u003emessages per day, which are written to a replica set.\" |\n| [![RUNTITLE](docs/logos/runtitle.png)](https://www.runtitle.com/) | \"RunTitle.com is a data-driven start-up in the Oil \u0026 Gas space.\u003cbr/\u003eWe curate mineral ownership data from millions of county\u003cbr/\u003erecords and help facilitate deals between mineral owners\u003cbr/\u003eand buyers. We use Kafka to create an eco-system of loosely\u003cbr/\u003ecoupled, specialized applications that share information.\u003cbr/\u003e\u003cbr/\u003eWe have identified the mongodb-sink-connector to be central\u003cbr/\u003eto our plans and were excited that it was readily enhanced to\u003cbr/\u003e[support our particular use-case.](https://github.com/hpgrahsl/kafka-connect-mongodb#custom-write-models) The connector documentation\u003cbr/\u003eand code are very clean and thorough. We are extremely positive\u003cbr/\u003eabout relying on OSS backed by such responsive curators.\" |\n\n### Supported Sink Record Structure\nCurrently the connector is able to process Kafka Connect SinkRecords with\nsupport for the following schema types [Schema.Type](https://kafka.apache.org/22/javadoc/org/apache/kafka/connect/data/Schema.Type.html):\n*INT8, INT16, INT32, INT64, FLOAT32, FLOAT64, BOOLEAN, STRING, BYTES, ARRAY, MAP, STRUCT*.\n\nThe conversion is able to generically deal with nested key or value structures - based on the supported types above - like the following example which is based on [AVRO](https://avro.apache.org/)\n\n```json\n{\"type\": \"record\",\n  \"name\": \"Customer\",\n  \"namespace\": \"at.grahsl.data.kafka.avro\",\n  \"fields\": [\n    {\"name\": \"name\", \"type\": \"string\"},\n    {\"name\": \"age\", \"type\": \"int\"},\n    {\"name\": \"active\", \"type\": \"boolean\"},\n    {\"name\": \"address\", \"type\":\n    {\"type\": \"record\",\n      \"name\": \"AddressRecord\",\n      \"fields\": [\n        {\"name\": \"city\", \"type\": \"string\"},\n        {\"name\": \"country\", \"type\": \"string\"}\n      ]}\n    },\n    {\"name\": \"food\", \"type\": {\"type\": \"array\", \"items\": \"string\"}},\n    {\"name\": \"data\", \"type\": {\"type\": \"array\", \"items\":\n    {\"type\": \"record\",\n      \"name\": \"Whatever\",\n      \"fields\": [\n        {\"name\": \"k\", \"type\": \"string\"},\n        {\"name\": \"v\", \"type\": \"int\"}\n      ]}\n    }},\n    {\"name\": \"lut\", \"type\": {\"type\": \"map\", \"values\": \"double\"}},\n    {\"name\": \"flags\",\n        \"type\": [ \"null\", \n                 {\"type\": \"map\", \"values\": {\"type\": \"array\", \"items\": \"string\"} } \n                ],\n        \"default\": null },\n    {\"name\": \"raw\", \"type\": \"bytes\"}\n  ]\n}\n```\n\n##### Logical Types\nBesides the standard types it is possible to use [AVRO logical types](http://avro.apache.org/docs/1.8.2/spec.html#Logical+Types) in order to have field type support for\n\n* **Decimal**\n* **Date**\n* **Time** (millis/micros)\n* **Timestamp** (millis/micros)\n\nThe following AVRO schema snippet based on exemplary logical type definitions should make this clearer:\n\n```json\n{\n  \"type\": \"record\",\n  \"name\": \"MyLogicalTypesRecord\",\n  \"namespace\": \"at.grahsl.data.kafka.avro\",\n  \"fields\": [\n    {\n      \"name\": \"myDecimalField\",\n      \"type\": {\n        \"type\": \"bytes\",\n        \"logicalType\": \"decimal\",\n        \"connect.parameters\": {\n          \"scale\": \"2\"\n        }\n      }\n    },\n    {\n      \"name\": \"myDateField\",\n      \"type\": {\n        \"type\": \"int\",\n        \"logicalType\": \"date\"\n      }\n    },\n    {\n      \"name\": \"myTimeMillisField\",\n      \"type\": {\n        \"type\": \"int\",\n        \"logicalType\": \"time-millis\"\n      }\n    },\n    {\n      \"name\": \"myTimeMicrosField\",\n      \"type\": {\n        \"type\": \"long\",\n        \"logicalType\": \"time-micros\"\n      }\n    },\n    {\n      \"name\": \"myTimestampMillisField\",\n      \"type\": {\n        \"type\": \"long\",\n        \"logicalType\": \"timestamp-millis\"\n      }\n    },\n    {\n      \"name\": \"myTimestampMicrosField\",\n      \"type\": {\n        \"type\": \"long\",\n        \"logicalType\": \"timestamp-micros\"\n      }\n    }\n  ]\n}\n```\n\nNote that if you are using AVRO code generation for logical types in order to use them from a Java-based producer app you end-up with the following Java type mappings:\n\n* org.joda.time.LocalDate myDateField;\n* org.joda.time.LocalTime mytimeMillisField;\n* long myTimeMicrosField;\n* org.joda.time.DateTime myTimestampMillisField;\n* long myTimestampMicrosField;\n\nSee [this discussion](https://github.com/hpgrahsl/kafka-connect-mongodb/issues/5) if you are interested in some more details.\n\nFor obvious reasons, logical types can only be supported for **AVRO** and **JSON + Schema** data (see section below).\n\n### Supported Data Formats\nThe sink connector implementation is configurable in order to support\n\n* **AVRO** (makes use of Confluent's Kafka Schema Registry and is the recommended format)\n* **JSON with Schema** (offers JSON record structure with explicit schema information)\n* **JSON plain** (offers JSON record structure without any attached schema)\n* **RAW JSON** (string only - JSON structure not managed by Kafka connect)\n\nSince key and value settings can be independently configured, it is possible to work with different data formats for records' keys and values respectively.\n\n_NOTE: Even when using RAW JSON mode i.e. with [StringConverter](https://kafka.apache.org/21/javadoc/index.html?org/apache/kafka/connect/storage/StringConverter.html) the expected Strings have to be valid and parsable JSON._\n\n##### Configuration example for AVRO\n```properties\nkey.converter=io.confluent.connect.avro.AvroConverter\nkey.converter.schema.registry.url=http://localhost:8081\n\nvalue.converter=io.confluent.connect.avro.AvroConverter\nvalue.converter.schema.registry.url=http://localhost:8081\n```\n\n##### Configuration example for JSON with Schema\n```properties\nkey.converter=org.apache.kafka.connect.json.JsonConverter\nkey.converter.schemas.enable=true\n\nvalue.converter=org.apache.kafka.connect.json.JsonConverter\nvalue.converter.schemas.enable=true\n```\n\n### Post Processors\nRight after the conversion, the BSON documents undergo a **chain of post processors**. There are the following 4 processors to choose from:\n\n* **DocumentIdAdder** (mandatory): uses the configured _strategy_ (explained below) to insert an **_id field**\n* **BlacklistProjector** (optional): applicable for _key_ + _value_ structure\n* **WhitelistProjector** (optional): applicable for _key_ + _value_ structure\n* **FieldRenamer** (optional): applicable for _key_ + _value_ structure\n\nFurther post processors can be easily implemented based on the provided abstract base class [PostProcessor](https://github.com/hpgrahsl/kafka-connect-mongodb/blob/master/src/main/java/at/grahsl/kafka/connect/mongodb/processor/PostProcessor.java), e.g.\n\n* remove fields with null values\n* redact fields containing sensitive information\n* etc.\n\nThere is a configuration property which allows to customize the post processor chain applied to the converted records before they are written to the sink. Just specify a comma separated list of fully qualified class names which provide the post processor implementations, either existing ones or new/customized ones, like so:\n\n```properties\nmongodb.post.processor.chain=at.grahsl.kafka.connect.mongodb.processor.field.renaming.RenameByMapping\n```\n\nThe DocumentIdAdder is automatically added at the very first position in the chain in case it is not present. Other than that, the chain can be built more or less arbitrarily. However, currently each post processor can only be specified once.\n\nFind below some documentation how to configure the available ones:\n\n##### DocumentIdAdder (mandatory)\nThe sink connector is able to process both, the key and value parts of kafka records. After the conversion to MongoDB BSON documents, an *_id* field is automatically added to value documents which are finally persisted in a MongoDB collection. The *_id* itself is filled by the **configured document id generation strategy**, which can be one of the following:\n\n* a MongoDB **BSON ObjectId** (default)\n* a Java **UUID**\n* **Kafka meta-data** comprised of the string concatenation based on [topic-partition-offset] information\n* **full key** using the sink record's complete key structure\n* **provided in key** expects the sink record's key to contain an *_id* field which is used as is (error if not present or null)\n* **provided in value** expects the sink record's value to contain an *_id* field which is used as is (error if not present or null)\n* **partial key** using parts of the sink record's key structure \n* **partial value** using parts of the sink record's value structure\n\n_Note: the latter two of which can be configured to use the blacklist/whitelist field projection mechanisms described below._\n\nThe strategy is set by means of the following property:\n\n```properties\nmongodb.document.id.strategy=at.grahsl.kafka.connect.mongodb.processor.id.strategy.BsonOidStrategy\n```\n\nThere is a configuration property which allows to customize the applied id generation strategy. Thus, if none of the available strategies fits your needs, further strategies can be easily implemented based on the interface [IdStrategy](https://github.com/hpgrahsl/kafka-connect-mongodb/blob/master/src/main/java/at/grahsl/kafka/connect/mongodb/processor/id/strategy/IdStrategy.java)\n\nAll custom strategies that should be available to the connector can be registered by specifying a list of fully qualified class names for the following configuration property:\n\n```properties\nmongodb.document.id.strategies=...\n```\n\n**It's important to keep in mind that the chosen / implemented id strategy has direct implications on the possible delivery semantics.** Obviously, if it's set to BSON ObjectId or UUID respectively, it can only ever guarantee at-least-once delivery of records, since new ids will result due to the re-processing on retries after failures. The other strategies permit exactly-once semantics iff the respective fields forming the document *_id* are guaranteed to be unique in the first place.\n\n##### Blacklist-/WhitelistProjector (optional)\nBy default the current implementation converts and persists the full value structure of the sink records.\nKey and/or value handling can be configured by using either a **blacklist or whitelist** approach in order to remove/keep fields\nfrom the record's structure. By using the \".\" notation to access sub documents it's also supported to do \nredaction of nested fields. It is also possible to refer to fields of documents found within arrays by the same notation. See two concrete examples below about the behaviour of these two projection strategies\n\nGiven the following fictional data record:\n\n```json\n{ \"name\": \"Anonymous\", \n  \"age\": 42,\n  \"active\": true, \n  \"address\": {\"city\": \"Unknown\", \"country\": \"NoWhereLand\"},\n  \"food\": [\"Austrian\", \"Italian\"],\n  \"data\": [{\"k\": \"foo\", \"v\": 1}],\n  \"lut\": {\"key1\": 12.34, \"key2\": 23.45}\n}\n```\n\n##### Example blacklist projection:\n```properties\nmongodb.[key|value].projection.type=blacklist\nmongodb.[key|value].projection.list=age,address.city,lut.key2,data.v\n```\n\nwill result in:\n\n```json\n{ \"name\": \"Anonymous\", \n  \"active\": true, \n  \"address\": {\"country\": \"NoWhereLand\"},\n  \"food\": [\"Austrian\", \"Italian\"],\n  \"data\": [{\"k\": \"foo\"}],\n  \"lut\": {\"key1\": 12.34}\n}\n```\n\n##### Example whitelist projection:\n```properties\nmongodb.[key|value].projection.type=whitelist\nmongodb.[key|value].projection.list=age,address.city,lut.key2,data.v\n```\n\nwill result in:\n\n```json\n{ \"age\": 42, \n  \"address\": {\"city\": \"Unknown\"},\n  \"data\": [{\"v\": 1}],\n  \"lut\": {\"key2\": 23.45}\n}\n```\n\nTo have more flexibility in this regard there might be future support for:\n\n* explicit null handling: the option to preserve / ignore fields with null values\n* investigate if it makes sense to support array element access for field projections based on an index or a given value to project simple/primitive type elements\n\n##### How wildcard pattern matching works:\nThe configuration supports wildcard matching using a __'\\*'__ character notation. A wildcard\nis supported on any level in the document structure in order to include (whitelist) or\nexclude (blacklist) any fieldname at the corresponding level. A part from that there is support\nfor __'\\*\\*'__ which can be used at any level to include/exclude the full sub structure\n(i.e. all nested levels further down in the hierarchy).\n\n_NOTE: A bunch of more concrete examples of field projections including wildcard pattern matching can be found in a corresponding [test class](https://github.com/hpgrahsl/kafka-connect-mongodb/blob/master/src/test/java/at/grahsl/kafka/connect/mongodb/processor/field/projection/FieldProjectorTest.java)._\n\n##### Whitelist examples:\n\nExample 1: \n\n```properties\nmongodb.[key|value].projection.type=whitelist\nmongodb.[key|value].projection.list=age,lut.*\n```\n\n-\u003e will include: the *age* field, the *lut* field and all its immediate subfiels (i.e. one level down)\n\nExample 2: \n\n```properties\nmongodb.[key|value].projection.type=whitelist\nmongodb.[key|value].projection.list=active,address.**\n```\n\n-\u003e will include: the *active* field, the *address* field and its full sub structure (all available nested levels)\n\nExample 3:\n\n```properties\nmongodb.[key|value].projection.type=whitelist\nmongodb.[key|value].projection.list=*.*\n```\n\n-\u003e will include: all fields on the 1st and 2nd level\n\n##### Blacklist examples:\nExample 1:\n\n```properties\nmongodb.[key|value].projection.type=blacklist\nmongodb.[key|value].projection.list=age,lut.*\n```\n\n-\u003e will exclude: the *age* field, the *lut* field and all its immediate subfields (i.e. one level down)\n\nExample 2: \n\n```properties\nmongodb.[key|value].projection.type=blacklist\nmongodb.[key|value].projection.list=active,address.**\n```\n\n-\u003e will exclude: the *active* field, the *address* field and its full sub structure (all available nested levels)\n\nExample 3:\n\n```properties\nmongodb.[key|value].projection.type=blacklist\nmongodb.[key|value].projection.list=*.*\n```\n\n-\u003e will exclude: all fields on the 1st and 2nd level\n\n##### FieldRenamer (optional)\nThere are two different options to rename any fields in the record, namely a simple and rigid 1:1 field name mapping or a more flexible approach using regexp. Both config options are defined by inline JSON arrays containing objects which describe the renaming.\n\nExample 1:\n\n```properties\nmongodb.field.renamer.mapping=[{\"oldName\":\"key.fieldA\",\"newName\":\"field1\"},{\"oldName\":\"value.xyz\",\"newName\":\"abc\"}]\n```\n\nThese settings cause:\n\n1) a field named _fieldA_ to be renamed to _field1_ in the **key document structure**\n2) a field named _xyz_ to be renamed to _abc_ in the **value document structure**\n\nExample 2:\n\n```properties\nmongodb.field.renamer.mapping=[{\"regexp\":\"^key\\\\..*my.*$\",\"pattern\":\"my\",\"replace\":\"\"},{\"regexp\":\"^value\\\\..*-.+$\",\"pattern\":\"-\",\"replace\":\"_\"}]\n```\n\nThese settings cause:\n\n1) **all field names of the key structure containing 'my'** to be renamed so that **'my' is removed**\n2) **all field names of the value structure containing a '-'** to be renamed by replacing **'-' with '_'**\n\nNote the use of the **\".\" character** as navigational operator in both examples. It's used in order to refer to nested fields in sub documents of the record structure. The prefix at the very beginning is used as a simple convention to distinguish between the _key_ and _value_ structure of a document.\n\n### Custom Write Models\nThe default behaviour for the connector whenever documents are written to MongoDB collections is to make use of a proper [ReplaceOneModel](http://mongodb.github.io/mongo-java-driver/3.10/javadoc/com/mongodb/client/model/ReplaceOneModel.html) with [upsert mode](http://mongodb.github.io/mongo-java-driver/3.10/javadoc/com/mongodb/client/model/ReplaceOneModel.html) and **create the filter document based on the _id field** which results from applying the configured DocumentIdAdder in the value structure of the sink document.\n\nHowever, there are other use cases which need different approaches and the **customization option for generating custom write models** can support these. The configuration entry (_mongodb.writemodel.strategy_) allows for such customizations. Currently, the following strategies are implemented:\n\n* **default behaviour** at.grahsl.kafka.connect.mongodb.writemodel.strategy.**ReplaceOneDefaultStrategy**\n* **business key** (-\u003e see [use case 1](https://github.com/hpgrahsl/kafka-connect-mongodb#use-case-1-employing-business-keys)) at.grahsl.kafka.connect.mongodb.writemodel.strategy.**ReplaceOneBusinessKeyStrategy**\n* **add inserted/modified timestamps** (-\u003e see [use case 2](https://github.com/hpgrahsl/kafka-connect-mongodb#use-case-2-add-inserted-and-modified-timestamps)) at.grahsl.kafka.connect.mongodb.writemodel.strategy.**UpdateOneTimestampsStrategy**\n* **monotonic write behaviour** (-\u003e see [use case 3](https://github.com/hpgrahsl/kafka-connect-mongodb#use-case-3-prevent-updates-for-stale-data)) at.grahsl.kafka.connect.mongodb.writemodel.strategy.**MonotonicWritesDefaultStrategy**\n* **delete on null values** at.grahsl.kafka.connect.mongodb.writemodel.strategy.**DeleteOneDefaultStrategy** implicitly used when config option _mongodb.delete.on.null.values=true_ for [convention-based deletion](https://github.com/hpgrahsl/kafka-connect-mongodb#convention-based-deletion-on-null-values)\n\n_NOTE:_ Future versions will allow to make use of arbitrary, individual strategies that can be registered and easily used as _mongodb.writemodel.strategy_ configuration setting.\n\n##### Use Case 1: Employing Business Keys\nLet's say you want to re-use a unique business key found in your sink records while at the same time have _BSON ObjectIds_ created for the resulting MongoDB documents.\nTo achieve this a few simple configuration steps are necessary:\n\n1) make sure to **create a unique key constraint** for the business key of your target MongoDB collection\n2) use the **PartialValueStrategy** as the DocumentIdAdder's strategy in order to let the connector know which fields belong to the business key\n3) use the **ReplaceOneBusinessKeyStrategy** instead of the default behaviour\n\nThese configuration settings then allow to have **filter documents based on the original business key but still have _BSON ObjectIds_ created for the _id field** during the first upsert into your target MongoDB target collection. Find below how such a setup might look like:\n \nGiven the following fictional Kafka record\n\n```json\n{ \n  \"fieldA\": \"Anonymous\", \n  \"fieldB\": 42,\n  \"active\": true, \n  \"values\": [12.34, 23.45, 34.56, 45.67]\n}\n```\n\ntogether with the sink connector config below \n\n```json\n{\n  \"name\": \"mdb-sink\",\n  \"config\": {\n    ...\n    \"mongodb.document.id.strategy\": \"at.grahsl.kafka.connect.mongodb.processor.id.strategy.PartialValueStrategy\",\n    \"mongodb.key.projection.list\": \"fieldA,fieldB\",\n    \"mongodb.key.projection.type\": \"whitelist\",\n    \"mongodb.writemodel.strategy\": \"at.grahsl.kafka.connect.mongodb.writemodel.strategy.ReplaceOneBusinessKeyStrategy\"\n  }\n}\n```\n\nwill eventually result in a MongoDB document looking like:\n\n```json\n{ \n  \"_id\": ObjectId(\"5abf52cc97e51aae0679d237\"),\n  \"fieldA\": \"Anonymous\", \n  \"fieldB\": 42,\n  \"active\": true, \n  \"values\": [12.34, 23.45, 34.56, 45.67]\n}\n```\n\nAll upsert operations are done based on the unique business key which for this example is a compound one that consists of the two fields _(fieldA,fieldB)_.\n\n##### Use Case 2: Add Inserted and Modified Timestamps\nLet's say you want to attach timestamps to the resulting MongoDB documents such that you can store the point in time of the document insertion and at the same time maintain a second timestamp reflecting when a document was modified.\n\nAll that needs to be done is use the **UpdateOneTimestampsStrategy** instead of the default behaviour. What results from this is that\nthe custom write model will take care of attaching two timestamps to MongoDB documents:\n\n1) **_insertedTS**: will only be set once in case the upsert operation results in a new MongoDB document being inserted into the corresponding collection\n2) **_modifiedTS**: will be set each time the upsert operation\nresults in an existing MongoDB document being updated in the corresponding collection\n\nGiven the following fictional Kafka record\n\n```json\n{ \n  \"_id\": \"ABCD-1234\",\n  \"fieldA\": \"Anonymous\", \n  \"fieldB\": 42,\n  \"active\": true, \n  \"values\": [12.34, 23.45, 34.56, 45.67]\n}\n```\n\ntogether with the sink connector config below \n\n```json\n{\n  \"name\": \"mdb-sink\",\n  \"config\": {\n    ...\n    \"mongodb.document.id.strategy\": \"at.grahsl.kafka.connect.mongodb.processor.id.strategy.ProvidedInValueStrategy\",\n    \"mongodb.writemodel.strategy\": \"at.grahsl.kafka.connect.mongodb.writemodel.strategy.UpdateOneTimestampsStrategy\"\n  }\n}\n```\n\nwill result in a new MongoDB document looking like:\n\n```json\n{ \n  \"_id\": \"ABCD-1234\",\n  \"_insertedTS\": ISODate(\"2018-07-22T09:19:000Z\"),\n  \"_modifiedTS\": ISODate(\"2018-07-22T09:19:000Z\"),\n  \"fieldA\": \"Anonymous\",\n  \"fieldB\": 42,\n  \"active\": true, \n  \"values\": [12.34, 23.45, 34.56, 45.67]\n}\n```\n\nIf at some point in time later there is a Kafka record referring to the same _id but containing updated data\n\n```json\n{ \n  \"_id\": \"ABCD-1234\",\n  \"fieldA\": \"anonymous\", \n  \"fieldB\": -23,\n  \"active\": false, \n  \"values\": [12.34, 23.45]\n}\n```\n\nthen the existing MongoDB document will get updated together with a fresh timestamp for the **_modifiedTS** value:\n\n```json\n{ \n  \"_id\": \"ABCD-1234\",\n  \"_insertedTS\": ISODate(\"2018-07-22T09:19:000Z\"),\n  \"_modifiedTS\": ISODate(\"2018-07-31T19:09:000Z\"),\n  \"fieldA\": \"anonymous\",\n  \"fieldB\": -23,\n  \"active\": false, \n  \"values\": [12.34, 23.45]\n}\n```\n\n##### Use Case 3: Prevent updates for stale data\nWhen the sink connector processes data it can happen that the same records might get reprocessed. For instance, this affects any records for which the corresponding offsets haven't been successfully committed for whatever reason. Ofentimes, the reprocessing as such isn't a big deal especially if write models use upsert semantics. **However, for certain scenarios it might be unacceptable that \"older\" records can lead to the overwriting of \"newer\" documents which are already present in the sink.** Given this behaviour, it might happen for queries against the sink to temporarily see stale data until the reprocessing has caught up.\n\nThe **MonotonicWritesDefaultStrategy allows to prevent any such writes based on stale data** against the sink. It adds the Kafka coordinates of processed records to the actual SinkDocument as _meta-data_ before it gets written to the MongoDB collection. The pre-defined and currently not(!) configurable data format for this is using a sub-document with the following structure, field names and value \u0026lt;PLACEHOLDERS\u0026gt;\n\n```json\n {\n    ...,\n    \"_kafkaCoords\":{\n        \"_topic\": \"\u003cTOPIC_NAME\u003e\",\n        \"_partition\": \u003cPARTITION_NUMBER\u003e,\n        \"_offset\": \u003cOFFSET_NUMBER\u003e;\n    },\n    ...\n }\n```\n\nThis _meta-data_ is used to perform the actual staleness check, namely, that upsert operations based on the corresponding document's **_id** field will get suppressed, in case newer data has already been written to the sink in the past. **Newer data means that a document exhibiting a greater than or equal offset for the same Kafka topic and partition is already present in the corresponding MongoDB collection.**\n\n_! IMPORTANT NOTE !_\nThis WriteModelStrategy needs **MongoDB version 4.2+** and **Java Driver 3.11+** since lower versions of either lack the support for leveraging update pipeline syntax which is needed to perform the conditional checks during write operations.   \n\n### Change Data Capture Mode\nThe sink connector can also be used in a different operation mode in order to handle change data capture (CDC) events. Currently, the following CDC events from [Debezium](http://debezium.io/) can be processed:\n\n* [MongoDB](http://debezium.io/docs/connectors/mongodb/) \n* [MySQL](http://debezium.io/docs/connectors/mysql/)\n* [PostgreSQL](http://debezium.io/docs/connectors/postgresql/)\n* **Oracle** ([incubating at Debezium Project](http://debezium.io/docs/connectors/oracle/))\n* **SQL Server** ([incubating at Debezium Project](http://debezium.io/docs/connectors/sqlserver/))\n\nThis effectively allows to replicate all state changes within the source databases into MongoDB collections. Debezium produces very similar CDC events for MySQL and PostgreSQL. The so far addressed use cases worked fine based on the same code which is why there is only one _RdbmsHandler_ implementation to support them both at the moment. Compatibility with Debezium's Oracle \u0026 SQL Server CDC format will be addressed in a future release.\n \nAlso note that **both serialization formats (JSON+Schema \u0026 AVRO) can be used** depending on which configuration is a better fit for your use case.\n\n##### CDC Handler Configuration\nThe sink connector configuration offers a property called *mongodb.change.data.capture.handler* which is set to the fully qualified class name of the respective CDC format handler class. These classes must extend from the provided abstract class *[CdcHandler](https://github.com/hpgrahsl/kafka-connect-mongodb/blob/master/src/main/java/at/grahsl/kafka/connect/mongodb/cdc/CdcHandler.java)*. As soon as this configuration property is set the connector runs in **CDC operation mode**. Find below a JSON based configuration sample for the sink connector which uses the current default implementation that is capable to process Debezium CDC MongoDB events. This config can be posted to the [Kafka connect REST endpoint](https://docs.confluent.io/current/connect/references/restapi.html) in order to run the sink connector.\n\n```json\n{\n  \"name\": \"mdb-sink-debezium-cdc\",\n  \"config\": {\n    \"key.converter\":\"io.confluent.connect.avro.AvroConverter\",\n    \"key.converter.schema.registry.url\":\"http://localhost:8081\",\n    \"value.converter\":\"io.confluent.connect.avro.AvroConverter\",\n    \"value.converter.schema.registry.url\":\"http://localhost:8081\",\n  \t\"connector.class\": \"at.grahsl.kafka.connect.mongodb.MongoDbSinkConnector\",\n    \"topics\": \"myreplset.kafkaconnect.mongosrc\",\n    \"mongodb.connection.uri\": \"mongodb://mongodb:27017/kafkaconnect?w=1\u0026journal=true\",\n    \"mongodb.change.data.capture.handler\": \"at.grahsl.kafka.connect.mongodb.cdc.debezium.mongodb.MongoDbHandler\",\n    \"mongodb.collection\": \"mongosink\"\n  }\n}\n```\n\n##### Convention-based deletion on null values\nThere are scenarios in which there is no CDC enabled source connector in place. However, it might be required to still be able to handle record deletions. For these cases the sink connector can be configured to delete records in MongoDB whenever it encounters sink records which exhibit _null_ values. This is a simple convention that can be activated by setting the following configuration option:\n\n```properties\nmongodb.delete.on.null.values=true\n```\n\nBased on this setting the sink connector tries to delete a MongoDB document from the corresponding collection based on the sink record's key or actually the resulting *_id* value thereof, which is generated according to the specified [DocumentIdAdder](https://github.com/hpgrahsl/kafka-connect-mongodb#documentidadder-mandatory).  \n\n### MongoDB Persistence\nThe sink records are converted to BSON documents which are in turn inserted into the corresponding MongoDB target collection. The implementation uses unorderd bulk writes. According to the chosen write model strategy either a [ReplaceOneModel](http://mongodb.github.io/mongo-java-driver/3.10/javadoc/com/mongodb/client/model/ReplaceOneModel.html) or an [UpdateOneModel](http://mongodb.github.io/mongo-java-driver/3.10/javadoc/com/mongodb/client/model/ReplaceOneModel.html) - both of which are run in [upsert mode](http://mongodb.github.io/mongo-java-driver/3.6/javadoc/com/mongodb/client/model/UpdateOptions.html) - is used whenever inserts or updates are handled. If the connector is configured to process convention-based deletes when _null_ values of sink records are discovered then it uses a [DeleteOneModel](http://mongodb.github.io/mongo-java-driver/3.10/javadoc/com/mongodb/client/model/ReplaceOneModel.html) respectively.\n\nData is written using acknowledged writes and the configured write concern level of the connection as specified in the connection URI. If the bulk write fails (totally or partially) errors are logged and a simple retry logic is in place. More robust/sophisticated failure mode handling has yet to be implemented.\n \n### Sink Connector Configuration Properties \n\nAt the moment the following settings can be configured by means of the *connector.properties* file. For a config file containing default settings see [this example](https://github.com/hpgrahsl/kafka-connect-mongodb/blob/master/config/MongoDbSinkConnector.properties).\n\n| Name                                           | Description                                                                                          | Type    | Default                                                                       | Valid Values                                                                                                     | Importance |\n|------------------------------------------------|------------------------------------------------------------------------------------------------------|---------|-------------------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------------|------------|\n| mongodb.collection                             | single sink collection name to write to                                                              | string  | \"\"                                                                            |                                                                                                                  | high       |\n| mongodb.connection.uri                         | the monogdb connection URI as supported by the offical drivers                                       | string  | mongodb://localhost:27017/kafkaconnect?w=1\u0026journal=true                       |                                                                                                                  | high       |\n| mongodb.document.id.strategy                   | class name of strategy to use for generating a unique document id (_id)                              | string  | at.grahsl.kafka.connect.mongodb.processor.id.strategy.BsonOidStrategy         |                                                                                                                  | high       |\n| mongodb.collections                            | names of sink collections to write to for which there can be topic-level specific properties defined | string  | \"\"                                                                            |                                                                                                                  | medium     |\n| mongodb.delete.on.null.values                  | whether or not the connector tries to delete documents based on key when value is null               | boolean | false                                                                         |                                                                                                                  | medium     |\n| mongodb.max.batch.size                         | maximum number of sink records to possibly batch together for processing                             | int     | 0                                                                             | [0,...]                                                                                                          | medium     |\n| mongodb.max.num.retries                        | how often a retry should be done on write errors                                                     | int     | 3                                                                             | [0,...]                                                                                                          | medium     |\n| mongodb.retries.defer.timeout                  | how long in ms a retry should get deferred                                                           | int     | 5000                                                                          | [0,...]                                                                                                          | medium     |\n| mongodb.change.data.capture.handler            | class name of CDC handler to use for processing                                                      | string  | \"\"                                                                            |                                                                                                                  | low        |\n| mongodb.change.data.capture.handler.operations | comma separated list of CDC operations that should be performed (not listed ones get suppressed)     | string  | \"c,r,u,d\"                                                                     | any string based on subset of [c,r,u,d]                                                                          | low        |\n| mongodb.document.id.strategies                 | comma separated list of custom strategy classes to register for usage                                | string  | \"\"                                                                            |                                                                                                                  | low        |\n| mongodb.field.renamer.mapping                  | inline JSON array with objects describing field name mappings                                        | string  | []                                                                            |                                                                                                                  | low        |\n| mongodb.field.renamer.regexp                   | inline JSON array with objects describing regexp settings                                            | string  | []                                                                            |                                                                                                                  | low        |\n| mongodb.key.projection.list                    | comma separated list of field names for key projection                                               | string  | \"\"                                                                            |                                                                                                                  | low        |\n| mongodb.key.projection.type                    | whether or not and which key projection to use                                                       | string  | none                                                                          | [none, blacklist, whitelist]                                                                                     | low        |\n| mongodb.post.processor.chain                   | comma separated list of post processor classes to build the chain with                               | string  | at.grahsl.kafka.connect.mongodb.processor.DocumentIdAdder                     |                                                                                                                  | low        |\n| mongodb.rate.limiting.every.n                  | after how many processed batches the rate limit should trigger (NO rate limiting if n=0)             | int     | 0                                                                             | [0,...]                                                                                                          | low        |\n| mongodb.rate.limiting.timeout                  | how long in ms processing should wait before continue processing                                     | int     | 0                                                                             | [0,...]                                                                                                          | low        |\n| mongodb.value.projection.list                  | comma separated list of field names for value projection                                             | string  | \"\"                                                                            |                                                                                                                  | low        |\n| mongodb.value.projection.type                  | whether or not and which value projection to use                                                     | string  | none                                                                          | [none, blacklist, whitelist]                                                                                     | low        |\n| mongodb.writemodel.strategy                    | how to build the write models for the sink documents                                                 | string  | at.grahsl.kafka.connect.mongodb.writemodel.strategy.ReplaceOneDefaultStrategy |                                                                                                                  | low        |\n\nThe above listed *connector.properties* are the 'original' (still valid / supported) way to configure the sink connector. The main drawback with it is that only one MongoDB collection could be used so far to sink data from either a single or multiple Kafka topic(s).\n\n### Collection-aware Configuration Settings\n\nIn the past several sink connector instances had to be configured and run separately, one for each topic / collection which needed to have individual processing settings applied. Starting with version 1.2.0 it is possible to **configure multiple Kafka topic \u003c-\u003e MongoDB collection mappings.** This allows for a lot more flexibility and **supports complex data processing needs within one and the same sink connector instance.**\n\nEssentially all relevant *connector.properties* can now be defined individually for each topic / collection.\n \n##### Topic \u003c-\u003e Collection Mappings\n\nThe most important change in configuration options is about defining the named-relation between configured Kafka topics and MongoDB collections like so:\n\n```properties\n\n#Kafka topics to consume from\ntopics=foo-t,blah-t\n\n#MongoDB collections to write to\nmongodb.collections=foo-c,blah-c\n\n#Named topic \u003c-\u003e collection mappings\nmongodb.collection.foo-t=foo-c\nmongodb.collection.blah-t=blah-c\n\n```\n\n**NOTE:** In case there is no explicit mapping between Kafka topic names and MongoDB collection names the following convention applies:\n\n* if the configuration property for ```mongodb.collection``` is set to any non-empty string this MongoDB collection name will be taken for any Kafka topic for which there is no defined mapping\n* if no _default name_ is configured with the above configuration property the connector falls back to using the original Kafka topic name as MongoDB collection name   \n\n##### Individual Settings for each Collection\n\nConfiguration properties can then be defined specifically for any of the collections for which there is a named mapping defined. The following configuration fragments show how to apply different settings for *foo-c* and *blah-c* MongoDB sink collections.\n\n```properties\n\n#specific processing settings for topic 'foo-t' -\u003e collection 'foo-c'\n\nmongodb.document.id.strategy.foo-c=at.grahsl.kafka.connect.mongodb.processor.id.strategy.UuidStrategy\nmongodb.post.processor.chain.foo-c=at.grahsl.kafka.connect.mongodb.processor.DocumentIdAdder,at.grahsl.kafka.connect.mongodb.processor.BlacklistValueProjector\nmongodb.value.projection.type.foo-c=blacklist\nmongodb.value.projection.list.foo-c=k2,k4 \nmongodb.max.batch.size.foo-c=100\n\n```\n\nThese properties result in the following actions for messages originating form Kafka topic 'foo-t':\n\n* document identity (*_id* field) will be given by a generated UUID\n* value projection will be done using a blacklist approach in order to remove fields *k2* and *k4*\n* at most 100 documents will be written to the MongoDB collection 'foo-c' in one bulk write operation\n\nThen there are also individual settings for collection 'blah-c':\n\n```properties\n\n#specific processing settings for topic 'blah-t' -\u003e collection 'blah-c'\n\nmongodb.document.id.strategy.blah-c=at.grahsl.kafka.connect.mongodb.processor.id.strategy.ProvidedInValueStrategy\nmongodb.post.processor.chain.blah-c=at.grahsl.kafka.connect.mongodb.processor.WhitelistValueProjector\nmongodb.value.projection.type.blah-c=whitelist\nmongodb.value.projection.list.blah-c=k3,k5 \nmongodb.writemodel.strategy.blah-c=at.grahsl.kafka.connect.mongodb.writemodel.strategy.UpdateOneTimestampsStrategy\n\n```\n\nThese settings result in the following actions for messages originating from Kafka topic 'blah-t':\n\n* document identity (*_id* field) will be taken from the value structure of the message\n* value projection will be done using a whitelist approach to remove only retain *k3* and *k5*\n* the chosen write model strategy will keep track of inserted and modified timestamps for each written document\n\n##### Fallback to Defaults\n\nWhenever the sink connector tries to apply collection specific settings where no such settings are in place, it automatically falls back to either:\n\n* what was explicitly configured for the same collection-agnostic property\n\nor\n\n* what is implicitly defined for the same collection-agnostic property\n\nFor instance, given the following configuration fragment:\n\n```properties\n\n#explicitly defined fallback for document identity\nmongodb.document.id.strategy=at.grahsl.kafka.connect.mongodb.processor.id.strategy.FullKeyStrategy\n\n#collections specific overriding for document identity\nmongodb.document.id.strategy.foo-c=at.grahsl.kafka.connect.mongodb.processor.id.strategy.UuidStrategy\n\n#collections specific overriding for write model\nmongodb.writemodel.strategy.blah-c=at.grahsl.kafka.connect.mongodb.writemodel.strategy.UpdateOneTimestampsStrategy\n\n```\n\nmeans that:\n\n* **document identity would fallback to the explicitly given default which is the *FullKeyStrategy*** for all collections other than 'foo-c' for which it uses the specified *UuidStrategy*\n* **write model strategy would fallback to the implicitly defined *ReplaceOneDefaultStrategy*** for all collections other than 'blah-c' for which it uses the specified *UpdateOneTimestampsStrategy*\n\n### Running in development\n\n```\nmvn clean package\nexport CLASSPATH=\"$(find target/ -type f -name '*.jar'| grep '\\-package' | tr '\\n' ':')\"\n$CONFLUENT_HOME/bin/connect-standalone $CONFLUENT_HOME/etc/schema-registry/connect-avro-standalone.properties config/MongoDbSinkConnector.properties\n```\n\n### Donate\nIf you like this project and want to support its further development and maintanance we are happy about your [PayPal donation](https://www.paypal.com/cgi-bin/webscr?cmd=_s-xclick\u0026hosted_button_id=E3P9D3REZXTJS)\n\n### License Information\nThis project is licensed according to [Apache License Version 2.0](https://www.apache.org/licenses/LICENSE-2.0)\n\n```\nCopyright (c) 2019. Hans-Peter Grahsl (grahslhp@gmail.com)\n\nLicensed under the Apache License, Version 2.0 (the \"License\");\nyou may not use this file except in compliance with the License.\nYou may obtain a copy of the License at\n\n    http://www.apache.org/licenses/LICENSE-2.0\n\nUnless required by applicable law or agreed to in writing, software\ndistributed under the License is distributed on an \"AS IS\" BASIS,\nWITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.\nSee the License for the specific language governing permissions and\nlimitations under the License.\n```\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fhpgrahsl%2Fkafka-connect-mongodb","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fhpgrahsl%2Fkafka-connect-mongodb","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fhpgrahsl%2Fkafka-connect-mongodb/lists"}