{"id":14982376,"url":"https://github.com/lynnlangit/learning-hadoop-and-spark","last_synced_at":"2025-05-16T13:05:00.339Z","repository":{"id":41242011,"uuid":"193248399","full_name":"lynnlangit/learning-hadoop-and-spark","owner":"lynnlangit","description":"Companion to Learning Hadoop and Learning Spark courses on Linked In Learning","archived":false,"fork":false,"pushed_at":"2024-12-10T17:52:23.000Z","size":14250,"stargazers_count":195,"open_issues_count":0,"forks_count":167,"subscribers_count":16,"default_branch":"master","last_synced_at":"2025-05-12T07:16:27.422Z","etag":null,"topics":["apache-spark","dataproc","emr","hadoop","learning-hadoop","mapreduce","spark","wordcount"],"latest_commit_sha":null,"homepage":"https://www.linkedin.com/learning/learning-hadoop-2","language":"HTML","has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"apache-2.0","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/lynnlangit.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2019-06-22T15:20:09.000Z","updated_at":"2025-05-11T01:15:16.000Z","dependencies_parsed_at":"2023-02-17T23:15:54.370Z","dependency_job_id":"e0b2f665-2756-4bd8-9e9e-96c5b20c10af","html_url":"https://github.com/lynnlangit/learning-hadoop-and-spark","commit_stats":{"total_commits":216,"total_committers":2,"mean_commits":108.0,"dds":0.00462962962962965,"last_synced_commit":"731825370a67ce10f818e7a27550e5a5a693bdf9"},"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/lynnlangit%2Flearning-hadoop-and-spark","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/lynnlangit%2Flearning-hadoop-and-spark/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/lynnlangit%2Flearning-hadoop-and-spark/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/lynnlangit%2Flearning-hadoop-and-spark/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/lynnlangit","download_url":"https://codeload.github.com/lynnlangit/learning-hadoop-and-spark/tar.gz/refs/heads/master","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":254535827,"owners_count":22087399,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":["apache-spark","dataproc","emr","hadoop","learning-hadoop","mapreduce","spark","wordcount"],"created_at":"2024-09-24T14:05:18.216Z","updated_at":"2025-05-16T13:05:00.270Z","avatar_url":"https://github.com/lynnlangit.png","language":"HTML","funding_links":[],"categories":[],"sub_categories":[],"readme":"# Learning Hadoop and Spark\n\n## Contents\n\nThis is the companion repo to my Linked In Learning Courses on Apache Hadoop and Apache Spark.  \n\n🐘  **1. Learning Hadoop** - [link](https://www.linkedin.com/learning/learning-hadoop-23008320)  \n    - this course demos I use mostly GCP Dataproc   \n    - for running Hadoop \u0026 associated libraries (i.e. Hive, Pig, Spark...) workloads    \n    \n🌩️  **2. Cloud Hadoop: Scaling Apache Spark** - [link](https://www.linkedin.com/learning/cloud-hadoop-scaling-apache-spark) \u0026 [link to content area in this repo](https://github.com/lynnlangit/learning-hadoop-and-spark/tree/master/5-Use-Spark)  \n    - this course demos I use GCP DataProc, AWS EMR --or--   \n    - I use Databricks on AWS or on GCP\n    \n⛈️  **3. Azure Databricks Spark Essential Training** - [link](https://www.linkedin.com/learning/azure-databricks-essential-training) \u0026 [link to content area in this repo](https://github.com/lynnlangit/learning-hadoop-and-spark/tree/master/5-Use-Spark/Jupyter-Notebooks)    \n    - this course demos I use Azure with Databricks  \n    - for scaling Apache Spark workloads  \n\n---\n\n\n## Other LinkedIn Learning Courses on Hadoop or Spark\n\nThere are ~ 10 courses on Hadoop/Spark topics on LinkedIn Learning.  See graphic below  \n![Learning Paths](https://github.com/lynnlangit/learning-hadoop-and-spark/blob/master/images/path.png)\n\n- **Hadoop** for Data Science Tips and Tricks - [link](https://www.linkedin.com/learning/hadoop-for-data-science-tips-tricks-techniques)\n    - Set up Cloudera Enviroment\n    - Working with Files in HDFS\n    - Connecting to Hadoop Hive\n    - Complex Data Structures in Hive\n- **Spark** courses - [link](https://www.linkedin.com/learning/search?entityType=COURSE\u0026keywords=Spark\u0026software=Apache%20Spark~Hadoop)\n    - Various Topics - see screenshot below\n\n![LinkedInLearningSpark](https://github.com/lynnlangit/learning-hadoop-and-spark/blob/master/images/spark-courses.png)\n\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Flynnlangit%2Flearning-hadoop-and-spark","html_url":"https://awesome.ecosyste.ms/projects/github.com%2Flynnlangit%2Flearning-hadoop-and-spark","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2Flynnlangit%2Flearning-hadoop-and-spark/lists"}