Ecosyste.ms: Awesome

An open API service indexing awesome lists of open source software.

Awesome Lists | Featured Topics | Projects

Apache Spark

Apache Spark is an open source distributed general-purpose cluster-computing framework. It provides an interface for programming entire clusters with implicit data parallelism and fault tolerance.

https://github.com/thanaraklee/dataflow-with-gcp

This project demonstrates the workflow of a Data Engineer. It utilizes the Google Cloud Platform and Google Colab as the main tools.

airflow apache-spark data-engineering etl pandas spark

Last synced: 25 Dec 2024

https://github.com/angelcervera/poc-drivingdistance

Proof of concept to implement a service to calculate the driving distance using osm network

akka openstreetmap osm osm4scala scala spark

Last synced: 10 Feb 2025

https://github.com/pedropark99/introd-pyspark

An open and introductory book for the Python API of Apache Spark (pyspark)

book pyspark python spark

Last synced: 14 Oct 2024

https://github.com/pankajsingh09/data_engineering_using_aws

This Repository contains the contents related to Data Engineering Using AWS

aws data-ingestion dataengineering event-bridge lambda-functions pipeline pycharm-ide pyspark python s3 spark

Last synced: 12 Feb 2025

https://github.com/lmouhib/auto-register-spark-ui-k8s

A lightweight operator to automatically expose Spark UI manage its ingress when running Spark on Kubernetes

spark spark-kubernetes spark-sql spark-streaming spark-ui

Last synced: 10 Feb 2025

https://github.com/kemalcanbora/ba_bigdata_docker

Docker containers provide a way to package applications with everything needed to run them, including base operating system images, databases, libraries, and binaries.

bigdata hadoop hue kafka spark

Last synced: 24 Jan 2025

https://github.com/bluejoe2008/hippo-rpc

Hippo Transport Library enhances spark-commons with easy stream management & handling

kraps rpc spark stream

Last synced: 10 Feb 2025

https://github.com/pomadchin/vlm-performance

GeoTrellis RasterSources Ingest benchmark

aws emr geotrellis gis raster spark

Last synced: 17 Jan 2025

https://github.com/multivacplatform/multivac-ml

Pre-trained ML models for Apache Spark

machine-learning nlp spark spark-ml

Last synced: 12 Jan 2025

https://github.com/piotr-kalanski/spark-local

API enabling switching between Spark execution engine and local fast implementation based on Scala collections.

scala spark unit-testing

Last synced: 13 Feb 2025

https://github.com/maxinexiong/item-based-collaborative-filtering

This project utilizes PySpark DataFrames and PySpark RDD to implement item-based collaborative filtering. By calculating cosine similarity scores or identifying movies with the highest number of shared viewers, the system recommends 10 similar movies for a given target movie that aligns users’ preferences.

apache-spark collaborative-filtering movie-recommendation pyspark python spark spark-dataframes spark-rdd

Last synced: 21 Dec 2024

https://github.com/bedrockstreaming/sparktest

A testing tool for Scala and Spark developers

scala spark

Last synced: 31 Dec 2024

https://github.com/jabhij/crimerate_classification

Developing a system that could classify crime descriptions into different categories which would help the authorities to assign officers to crimes based on the report.

classification crime-analysis crime-classification crime-rates machine-learning mllib pyspark python spark tensorflow

Last synced: 17 Jan 2025

https://github.com/comcast/spark-util

little spark util library

spark

Last synced: 14 Nov 2024

https://github.com/simplexspatial/osm-facts

Proofs and checks about osm pbf format and data content facts

osm osm4scala scala spark

Last synced: 15 Jan 2025

https://github.com/maxinexiong/degrees-of-separation-with-breadth-first-search

This project utilizes PySpark RDD and the Breadth-first Search (BFS) algorithm to find the shortest path and degrees of separation between two given Marvel superheroes based on based on their appearances together in the same comic books, empowering users to discover connections between their favourite superheroes in the Marvel universe.

apache-spark bfs-algorithm breadth-first-search degrees-of-separation marvel-characters pyspark python spark spark-rdd

Last synced: 21 Dec 2024

https://github.com/superruzafa/scala-spark-big-data

My solutions to the Coursera's Big Data Analysis with Scala and Spark course

big-data coursera scala spark

Last synced: 30 Dec 2024

https://github.com/r13i/spark-record-deduplicating

Data cleansing problem statement: Data in a record are often duplicated. How do we find the duplicate probability ? [Work In Progress]

big-data deduplication record-linkage records-management scala spark

Last synced: 02 Feb 2025

https://github.com/renardeinside/terrametria

Source code 3D population density map of Germany, with ETL and app logic on top the Databricks Platform.

databricks deckgl python react spark

Last synced: 03 Dec 2024

https://github.com/gmartinezramirez-old/data-science-portafolio

:notebook: [Active] Portafolio of data science projects. Using: Python, PyTorch, Spark, Tensorflow, Scikit, Keras. Includes Classification, Regression, Time series, NLP, Deep learning, among others.

data-science data-science-learning data-science-notebook data-science-portfolio hadoop jupyter-notebook keras notebook pandas pyspark python pytorch r sci-kit spark tensorflow

Last synced: 05 Dec 2024

https://github.com/mtpatter/bilao

Jupyter notebooks for filtering Kafka data with Spark Streaming.

avro docker jupyter-notebook kafka spark spark-streaming

Last synced: 12 Jan 2025

https://github.com/dimajix/docker-spark

Repository for building Docker containers for Spark

cluster docker hadoop spark

Last synced: 05 Jan 2025

https://github.com/figuran04/big-data

📃 Praktikum Big Data

anaconda big data hadoop hive mongodb pig spark

Last synced: 01 Nov 2024

https://github.com/multivacplatform/multivac-wikipedia

Wonderful reusable codes, libraries and scripts to process Wikipedia page views by using Apache Spark.

data-frame multivac-wikipedia spark spark-sql wikipedia

Last synced: 12 Jan 2025

https://github.com/multivacplatform/multivac-fakenews

Detecting users and communities which propagate fake news on Twitter by Apache Spark

deep-learning fakenews machine-learning spark twitter

Last synced: 12 Jan 2025

https://github.com/timvw/adobe-analytics-datafeed-datasource

Apache Spark data source for Adobe Analytics Data Feed

adobe-analytics clickstream python scala spark

Last synced: 08 Nov 2024

https://github.com/hussaintaj-w/spark_submit_project

An easy to use script that automatically adds files to the spark-submit command.

python spark spark-submit

Last synced: 23 Jan 2025

https://github.com/gaelfoppolo/self-service-data-analytics

Data analysis made for business users

aws big-data data-analytics hadoop spark

Last synced: 03 Feb 2025

https://github.com/garystafford/dataproc-java-demo

Demonstration of Google Cloud Dataproc for running Spark jobs with Java

big-data-analytics dataproc gcp google java spark

Last synced: 06 Dec 2024

https://github.com/bataeves/isparkcache

Jupyter модуль для кеширования Spark DataFrame, полученных в результате выполнения ячейки

cache ipython jupyter pyspark spark

Last synced: 06 Feb 2025

https://github.com/boazmohar/pysparkutils

A collection of utilities for handling pySpark's SparkContext

pyspark python spark

Last synced: 09 Feb 2025

https://github.com/renardeinside/dbx-kafka-protobuf-example

Sample code for working with Kafka & Protobuf in Databricks

databricks kafka protobuf scala spark spark-streaming

Last synced: 06 Feb 2025

https://github.com/kmohamedalie/big-data-hadoop-spark-lab

Big Data🛢️ with Hadoop🐘 and Spark⭐ lab🧪🥼

big-data coursera data-engineering docker hadoop ibm kubernetes spark

Last synced: 02 Jan 2025

https://github.com/spratiher9/valido

PySpark ⚡ dataframe workflow ⚒ validator

apache apache-spark bigdata databricks decorators pyspark python3 spark testing

Last synced: 01 Feb 2025

https://github.com/tranthe170/nyc-taxi-pipeline

Building Data Lakehouse by open source technology. Support end to end data pipeline, from source data on AWS S3 to Lakehouse, visualize.

airflow delta-lake hive lakehouse presto python s3 spark superset

Last synced: 17 Jan 2025

https://github.com/jatin-8898/sparkwebsite

A clean and very interesting looking website. :sparkles:

bootstrap4 css html javascript spark typescript

Last synced: 17 Jan 2025

https://github.com/inbravo/spark-movie-lens

Various examples of analytics using Apache Spark

apache-spark scala spark

Last synced: 02 Feb 2025

https://github.com/mangalaman93/dspark

Run spark in docker containers

big-data containers docker microservices spark

Last synced: 18 Jan 2025

https://github.com/tashi-2004/fma-a-dataset-for-music-analysis

🎶 Scripts for music feature analysis, model training, and real-time recommendation using Apache Kafka. Extract features with Librosa 🎹, store them in MongoDB 🗄️, and process the data with Apache Spark ⚡. A 🌐 web interface 💻✨ is also included. Contributors: Tashfeen Abbasi 👤, Laiba Mazhar 👤, and Rafia Khan 👤.

html kafka kafka-consumer kafka-producer kafka-streaming linux mongodb mongodb-compass python3 spark ubuntu web-application

Last synced: 03 Dec 2024

https://github.com/aamend/texata-r2-2017

This project has been created in a 4h time for the purpose of the Texata Big Data world championship.

bigdata gdelt hackathon spark texata

Last synced: 30 Dec 2024

https://github.com/jinsyin/datalink

⚡ 数据集成 | DataLink is a lightweight data integration framework build on top of DataX, Spark and Flink

batch big-data bigdata cdc data data-collection data-exchange data-integration data-pipeline data-synchronization datalink etl flink flink-cdc framework integration pipeline spark streaming

Last synced: 15 Nov 2024

https://github.com/ugurcanerdogan/machine-learning-with-spark

BBM469*ASG3 - Machine Learning with Spark

apache-spark data-science machine-learning spark

Last synced: 12 Feb 2025

https://github.com/jms0522/hadoop_system

✅ hadoop eco system을 구성하고 파이프라인 제작합니다.

hadoop pipeline spark

Last synced: 14 Feb 2025

https://github.com/kwartile/spark-benchmark

Spark Benchmark suite to evaluate cluster configuration and compare the performance with other big data frameworks.

apache-spark benchmark benchmarking-suite cdh cloudera-hadoop hadoop hive impala performance scala spark

Last synced: 08 Feb 2025

https://github.com/policratus/sparkmage

🐘 A tool for blazing fast analysis and clustering of similar images using 🐘 Hadoop and ⚡ Spark.

big-data computer-vision hadoop image-processing spark

Last synced: 02 Nov 2024

https://github.com/cclient/elasticsearch-spark-upsert-from-kafka

elasticsearch-hadoop官方不支持upsert doc,修改源码实现,spark kafka streaming 示例 upsert { "upsert": {}, "doc": {...} }

elasticsearch elasticsearch-hadoop kafka kafka-streams spark upsert upsert-doc

Last synced: 16 Jan 2025

https://github.com/alfex4936/spark-studies

Apache Spark 공부 in Python

apache pyspark python spark

Last synced: 27 Jan 2025

https://github.com/derlin/workshop-data-sciences

A two-days workshop material introducing data sciences for big data

big-data data-science hdfs hive spark workshop workshop-materials zeppelin

Last synced: 20 Jan 2025

https://github.com/vasnake/spark.ml.spatialjointransformer

spark.ml.transformer: join two datasets using spatial relations

geospatial join ml-pipeline python scala spark spark-ml spatial transformer

Last synced: 03 Jan 2025

https://github.com/jldbc/big-data

Coursework from Big Data (CS3390) -- Machine Learning tasks performed using Hadoop, MapReduce, and Spark

big-data hadoop pagerank recommender-system spark

Last synced: 04 Jan 2025

https://github.com/mgrojo/adasearch

Custom search engine for the Ada programming language

ada custom-search-google search-engine spark spark-ada

Last synced: 27 Oct 2024

https://github.com/badoo/hadoop-xargs

Util to run heterogenous applications on Hadoop synchronously

hadoop java spark

Last synced: 12 Nov 2024

https://github.com/hifly81/1brc_streaming

1brc challenge with streaming solutions for Apache Kafka

1brc apache camel-kafka flink kafka kafkastreams ksqldb nifi spark spring-kafka streaming

Last synced: 02 Nov 2024

https://github.com/rpytel1/supercomputing-labs

Fork of the repository for Supercomputing in Big Data class on TU Delft. Scala, Spark and Kafka were used to perform processing and streaming of GDelt data segments.

big-data gdelt-data kafka scala spark

Last synced: 18 Jan 2025

https://github.com/anskarl/auxlib-spark-nlp

NLP utilities for Apache Spark

nlp opennlp scala spark

Last synced: 12 Feb 2025

https://github.com/yucl80/avrodemo

write , append avro to hdfs file

avro hdfs hive java kafka log scala spark sparksql sparkstreaming tomcat-log

Last synced: 27 Jan 2025

https://github.com/chimera-suite/pysparql

This is a simple module that allows developer to query SPARQL endpoints and analyze the results with Apache Spark.

apache apache-spark construct-query dataframe graphframe jena-fuseki spark sparql

Last synced: 01 Dec 2024

https://github.com/dustin-decker/elasticsearchsql

A simple example of using Apache Spark SQL against Elasticsearch 5

elasticsearch spark sql

Last synced: 29 Jan 2025

https://github.com/rezacsedu/Mining-Maximal-Frequent-Pattern-Spark

Implementation of Static mining part of "Mining maximal frequent patterns in transactional databases and dynamic data streams: A spark-based approach" Information Sciences, Volume 432, March 2018, Pages 278-300

data-mining data-stream frequent-pattern-mining java maximal-frequent-pattern spark structured-streaming

Last synced: 30 Oct 2024

https://github.com/makohn/lambda-architecture-poc

♨️ A PoC implementation of the λ-Architecture for collecting and analysing tweets

cassandra kafka lambda-architecture sbt scala spark

Last synced: 12 Feb 2025

https://github.com/jaehyeon-kim/emr-local-dev

Spark Local Development Environment Using Docker (and vscode)

aws docker emr spark vscode

Last synced: 30 Oct 2024

https://github.com/joyceannie/us-immigrations-data-warehouse

A data warehouse to perform analytics on the immigration trends in the US.

airflow data-engineering etl pyspark redshift s3 spark

Last synced: 29 Jan 2025

https://github.com/brooksian/sparkpipeline2mleapbundle

Convert Spark Pipeline Models to MLeap Bundles

mleap-bundle spark

Last synced: 19 Jan 2025

https://github.com/brooksian/epaairnow

Exploring EPA Air Now Time Series Data with Apache Spark and Apache Zeppelin

spark sparksql time-series zeppelin-notebook

Last synced: 19 Jan 2025

https://github.com/kadnan/vagrant-spark2

Vagrant Box with Python 3.6.1, Apache Spark 2.1.1 with Scala 2.11.8 and PySpark (2.1.1).

pyspark python3 spark vagrant vagrant-boxes

Last synced: 20 Jan 2025

https://github.com/brooksian/ds_gtdb

KMeans Clustering on Global Terrorism Database

global-terrorism-database machine-learning spark sparksql zeppelin-notebook

Last synced: 19 Jan 2025

https://github.com/wittline/sparksql-with-python

This repository has some examples of using Spark and SparkSQL with Python through PySpark

flask-api python spark sparksql

Last synced: 29 Jan 2025

https://github.com/brooksian/solrtosparknotebook

Connecting Solr and Spark In An Apache Zeppelin Notebook

solr spark zeppelin-notebook

Last synced: 19 Jan 2025

https://github.com/xpcosmos/injestao-dados-enem-sql

Esse projeto tem o objetivo de estruturar dados do enem em bancos de dados e analisar os dados utilizando métodos estatísticos.

docker docker-compose postgresql pyspark python spark sql statistics

Last synced: 14 Jan 2025

https://github.com/tpvasconcelos/sparkypandy

It's not spark, it's not pandas, it's just awkward...

dataframe pandas pyspark spark

Last synced: 05 Nov 2024

https://github.com/hb-chen/spark-elasticsearch-recommender

Zeppelin-v0.8.0 Notebook演示使用Spark -v2.3.2+ Elasticsearch-v6.3.2构建推荐系统

elasticsearch recommender spark zeppelin

Last synced: 08 Jan 2025

https://github.com/yalishanda42/scala-recsys

Scala(-ble) recommender system architecture using functional programming (PoC)

cats cats-effect functional-programming movielens recommender-system recsys scala spark

Last synced: 28 Dec 2024

https://github.com/angelotc/MacroDAG

A Dockerized Airflow ETL pipeline that processes macroeconomic indicators from the Federal Reserve.

airflow docker spark

Last synced: 06 Nov 2024

https://github.com/fpopic/gg-interview-challenge

(Interview) GG Interview Challenge in Scala/Spark

apache-spark json logstash parsing regex scala spark sparksql

Last synced: 10 Jan 2025

https://github.com/omar-besbes/football-big-data

This is a comprehensive solution for real-time football analytics, leveraging Apache Spark execution on yarn for both streaming and batch processing, Hadoop HDFS for distributed storage, Kafka for real-time data ingestion, RethinkDB for live data updates and Next.js for data visualization as well as a custom built search engine.

batch-processing hadoop kafka nextjs rethinkdb spark streaming t3-stack yarn

Last synced: 20 Jan 2025

https://github.com/nhviet03/is405_bigdata_mapreduce_knn

A MapReduce-Based k-Nearest Neighbor Approach for Big Data Classification on Apache Spark

knn-classification mapreduce pyspark spark

Last synced: 17 Jan 2025

https://github.com/saadsalmanakram/data-processing

This repo is focused on all key frameworks, libraries or tools use for Data Processing

big-data pandas polars spark

Last synced: 14 Feb 2025

https://github.com/gunantos/php-spark

PHP Server Develop

php server serverless-framework spark

Last synced: 13 Jan 2025