{"id":20228447,"url":"https://github.com/hortonworks-spark/cloud-integration","last_synced_at":"2025-10-14T15:10:45.223Z","repository":{"id":49965752,"uuid":"91702365","full_name":"hortonworks-spark/cloud-integration","owner":"hortonworks-spark","description":"Spark cloud integration: tests, cloud committers and more","archived":false,"fork":false,"pushed_at":"2025-01-30T18:19:24.000Z","size":911,"stargazers_count":19,"open_issues_count":4,"forks_count":9,"subscribers_count":3,"default_branch":"master","last_synced_at":"2025-04-10T17:46:44.238Z","etag":null,"topics":["apache-spark","aws-s3","azure","gcs","spark"],"latest_commit_sha":null,"homepage":"","language":"Scala","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/hortonworks-spark.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null,"zenodo":null}},"created_at":"2017-05-18T14:21:40.000Z","updated_at":"2025-01-30T18:19:28.000Z","dependencies_parsed_at":"2025-04-10T17:32:21.569Z","dependency_job_id":"b5a82d54-5eef-432a-bcd2-77b747eb2730","html_url":"https://github.com/hortonworks-spark/cloud-integration","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/hortonworks-spark/cloud-integration","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hortonworks-spark%2Fcloud-integration","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hortonworks-spark%2Fcloud-integration/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hortonworks-spark%2Fcloud-integration/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hortonworks-spark%2Fcloud-integration/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/hortonworks-spark","download_url":"https://codeload.github.com/hortonworks-spark/cloud-integration/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/hortonworks-spark%2Fcloud-integration/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":279019298,"owners_count":26086709,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","status":"online","status_checked_at":"2025-10-14T02:00:06.444Z","response_time":60,"last_error":null,"robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":true,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["apache-spark","aws-s3","azure","gcs","spark"],"created_at":"2024-11-14T07:30:21.706Z","updated_at":"2025-10-14T15:10:45.177Z","avatar_url":"https://github.com/hortonworks-spark.png","language":"Scala","funding_links":[],"categories":[],"sub_categories":[],"readme":"# Cloud Integration for Apache Spark\n\nThe [cloud-integration](https://github.com/hortonworks-spark/cloud-integration)\nrepository provides modules to improve Apache Spark's integration with cloud infrastructures.\n\n\n\n## Module `spark-cloud-integration`\n\nClasses and Tools to make Spark work better in-cloud\n\n* Committer integration with the s3a committers.\n* Proof of concept cloud-first distcp replacement.\n* Serialization for Hadoop `Configuration`: class `ConfigSerDeser`. Use this\nto get a configuration into an RDD method\n* Trait `HConf` to manipulate the hadoop options in a spark config.\n* Anything else which turns out to be useful.\n* Variant of `FileInputStream` for cloud storage, `org.apache.spark.streaming.cloudera.CloudInputDStream`\n\nSee [Spark Cloud Integration](spark-cloud-integration/src/main/site/markdown/index.md)\n\n\n\n## Module `cloud-examples`\n\nThis does the packaging/integration tests for Spark and cloud against AWS, Azure and Google GCS.\n\nThese are basic tests of the core functionality of I/O, streaming, and verify that\nthe commmitters work.\n\nAs well as running as unit tests, they have CLI entry points which can be used for scalable functional testing.\n\n\n## Module `minimal-integration-test`\n\nThis is a minimal JAR for integration tests\n\nUsage\n```bash\nspark-submit --class com.cloudera.spark.cloud.integration.Generator \\\n--master yarn \\\n--num-executors 2 \\\n--driver-memory 512m \\\n--executor-memory 512m \\\n--executor-cores 1 \\\nminimal-integration-test-1.0-SNAPSHOT.jar \\\nadl://example.azuredatalakestore.net/output/dest/1 \\\n2 2 15\n```\n\n\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fhortonworks-spark%2Fcloud-integration","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fhortonworks-spark%2Fcloud-integration","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fhortonworks-spark%2Fcloud-integration/lists"}