{"id":23041016,"url":"https://github.com/kwartile/spark-benchmark","last_synced_at":"2026-04-26T23:31:55.797Z","repository":{"id":194672520,"uuid":"92214193","full_name":"kwartile/spark-benchmark","owner":"kwartile","description":"Spark Benchmark suite to evaluate cluster configuration and compare the performance with other big data frameworks.","archived":false,"fork":false,"pushed_at":"2017-05-26T20:34:01.000Z","size":29,"stargazers_count":2,"open_issues_count":0,"forks_count":0,"subscribers_count":3,"default_branch":"master","last_synced_at":"2025-04-03T00:30:26.076Z","etag":null,"topics":["apache-spark","benchmark","benchmarking-suite","cdh","cloudera-hadoop","hadoop","hive","impala","performance","scala","spark"],"latest_commit_sha":null,"homepage":null,"language":"Scala","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":null,"status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/kwartile.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":null,"code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null}},"created_at":"2017-05-23T20:00:03.000Z","updated_at":"2021-03-15T04:26:02.000Z","dependencies_parsed_at":"2023-09-14T16:06:11.292Z","dependency_job_id":null,"html_url":"https://github.com/kwartile/spark-benchmark","commit_stats":null,"previous_names":["kwartile/spark-benchmark"],"tags_count":0,"template":false,"template_full_name":null,"purl":"pkg:github/kwartile/spark-benchmark","repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kwartile%2Fspark-benchmark","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kwartile%2Fspark-benchmark/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kwartile%2Fspark-benchmark/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kwartile%2Fspark-benchmark/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/kwartile","download_url":"https://codeload.github.com/kwartile/spark-benchmark/tar.gz/refs/heads/master","sbom_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/kwartile%2Fspark-benchmark/sbom","scorecard":null,"host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":286080680,"owners_count":32317163,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2026-04-26T23:26:28.701Z","status":"ssl_error","status_checked_at":"2026-04-26T23:26:25.802Z","response_time":129,"last_error":"SSL_read: unexpected eof while reading","robots_txt_status":"success","robots_txt_updated_at":"2025-07-24T06:49:26.215Z","robots_txt_url":"https://github.com/robots.txt","online":false,"can_crawl_api":true,"host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["apache-spark","benchmark","benchmarking-suite","cdh","cloudera-hadoop","hadoop","hive","impala","performance","scala","spark"],"created_at":"2024-12-15T19:28:46.246Z","updated_at":"2026-04-26T23:31:55.773Z","avatar_url":"https://github.com/kwartile.png","language":"Scala","funding_links":[],"categories":[],"sub_categories":[],"readme":"## Spark Benchmark\n\n#### Overview\nSpark Benchmark suite helps you evaluate Spark cluster configuration.  This benchmark can also be used to compare the speed, throughput, and resource usage of Spark jobs with other big data frameworks such as Impala and Hive. It contains a set of Spark RDD based operations that performs map, filter, reduceByKey, and join operations.\n\n#### Data\nThe benchmark uses the dataset used for Impala performance measurement (http://docs.aws.amazon.com/emr/latest/DeveloperGuide/query-impala-generate-data.html).  The dataset consists of three different files:\n* Books\n* Customers\n* Transactions\n\n```\n\u003e head books\n0|5-54687-602-6|FOREIGN-LANGUAGE-STUDY|1989-09-29|Saraiva|83.99\n1|7-20527-497-2|PHILOSOPHY|1999-09-14|Kyowon|40.99\n2|8-98211-350-2|JUVENILE-NONFICTION|1975-01-11|Wolters Kluwer|173.99\n3|6-52228-529-3|MATHEMATICS|2010-06-26|Bungeishunju|24.99\n4|8-98702-825-4|HUMOR|1990-07-15|China Publishing Group Corporate|64.99\n5|3-11023-371-2|LITERARY-CRITICISM|1971-06-04|AST|137.99\n```\n\n```\n\u003e head customers\nCustomers\n0|Sophia PERKINS|1975-11-18|F|OK|sophia.perkins.1975@gmail.com|963-341-4876\n1|Brianna MURRAY|2001-11-02|F|MT|brianna.murray.2001@gmail.com|260-164-6277\n2|James SCOTT|1997-09-17|M|UT|james.scott.1997@gmail.com|920-899-8587\n3|Samuel GREEN|2013-05-22|F|CO|samuel.green.2013@live.com|263-707-8321\n4|Logan COLEMAN|1997-12-10|F|NE|logan.coleman.1997@hotmail.com|333-318-5685\n5|Matthew BENNETT|1975-12-26|M|CO|matthew.bennett.1975@outlook.com|717-808-3733\n6|Jace SPENCER|2013-10-30|M|KS|jace.spencer.2013@live.com|448-105-3939\n```\n\n```\n\u003e head transactions\n0|29948726|124004825|21|2000-10-03 12:08:37\n1|76896577|10225228|17|2001-04-23 15:21:18\n2|77394742|62037151|23|2008-02-22 11:52:36\n3|23558280|21960491|29|2000-06-22 10:14:48\n4|5742930|73207419|15|2004-11-26 00:46:53\n5|101531051|122609274|13|2008-01-14 05:26:46\n```\n#### Benchmark\nThe benchmarks contains four tests:\nRDDScan: RDDScan reads the customer file and performs a filter operation.  It is equivalent to a select-where statement in SQL.\nRDDAggregate: RDDAgregate operation scans the books file and perform reduceByKey to aggregate count of books by category.  It then sorts the results based on the book count.\nRDDTwoWayJoin: This operation performs a join between books and transactions between 2008 and 2010, aggregates the results on book category, and returns sorted results based on the total transaction amount.\nRDDThreeWayJoin: This operation is similar to the above except we perform addition join with the customer table and filter the results on three states\n\n#### Supported Platform\nThe pom file currently includes support for Spark 1.6 on CDH 5.8.  But this can be easily modified to run on Spark 2.x and other version of Cloudera, HortonWorks, and Apache distribution.\n\n#### Build\nUse the standard maven command ```mvn package``` to build.\n\n#### Run\n* Generate Data: Download the DBGen utlity from http://docs.aws.amazon.com/emr/latest/DeveloperGuide/query-impala-generate-data.html.  Follow the instructions to generate data. \n* Copy generated data to HDFS.\n* Run the following command to run the benchmark. You may need to adjust the memory parameters to tune the job.\n```\nspark-submit  --class com.kwartile.benchmark.spark.\u003cRDDScan | RDDAggregate | RDDTwoWayJoin | RDDTHreeWayJoin\u003e --master yarn --executor-memory \u003cmem\u003e --executor-cores \u003cnum\u003e --num-executors \u003cnum\u003e  --conf spark.yarn.executor.memoryOverhead=\u003cmem_in_mb\u003e perf-benchmark-1.0-SNAPSHOT-jar-with-dependencies.jar --input-path \u003chdfs location\u003e\n```\n#### Hive and Impala Query\nYou can use the following equivalent Hive/Impala query to compare the performance.\n\n```\n# Scan Query\nSELECT COUNT(*)\nFROM customers256gb\nWHERE name = 'Scarlett STEVENS';\n```\n\n```\n# Aggregation\nSELECT category, count(*) cnt\nFROM books256gb\nGROUP BY category\nORDER BY cnt DESC LIMIT 10;\n```\n\n```\n# Two Way Join\nSELECT tmp.book_category, ROUND(tmp.revenue, 2) AS revenue\nFROM (\nSELECT books256gb.category AS book_category, SUM(books256gb.price * transactions256gb.quantity) AS revenue\nFROM books256gb JOIN transactions256gb ON (\ntransactions256gb.book_id = books256gb.id\nAND YEAR(transactions256gb.transaction_date) BETWEEN 2008 AND 2010\n)\nGROUP BY books256gb.category\n) tmp\nORDER BY revenue DESC LIMIT 10;\n```\n\n```\n# Three Way Join\nSELECT tmp.book_category, ROUND(tmp.revenue, 2) AS revenue\nFROM (\n  SELECT books256gb.category AS book_category, SUM(books256gb.price * transactions256gb.quantity) AS revenue\n  FROM books256gb\n  JOIN transactions256gb ON (\n    transactions256gb.book_id = books256gb.id\n  )\n  JOIN customers256gb ON (\n    transactions256gb.customer_id = customers256gb.id\n    AND customers256gb.state IN ('WA', 'CA', 'NY')\n  )\n  GROUP BY books256gb.category\n) tmp\nORDER BY revenue DESC LIMIT 10;\n```\n### Questions \u0026 Feedback\nPlease contact us as labs@kwartile.com for any question or enhancement request.\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fkwartile%2Fspark-benchmark","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Fkwartile%2Fspark-benchmark","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Fkwartile%2Fspark-benchmark/lists"}