awesome-bigdata
A curated list of awesome big data frameworks, ressources and other awesomeness
https://github.com/Anyz01/awesome-bigdata
Last synced: 13 days ago
JSON representation
-
Search engine and framework
- ElasticSearch - Search and analytics engine based on Apache Lucene.
- Enigma.io
- Facebook Unicorn - social graph search platform.
- Lily HBase Indexer - quickly and easily search for any content stored in HBase.
- LinkedIn Galene - search architecture at LinkedIn.
- Sphinx Search Server - fulltext search engine.
- Elassandra - is a fork of Elasticsearch modified to run on top of Apache Cassandra in a scalable and resilient peer-to-peer architecture.
- LinkedIn Bobo - is a Faceted Search implementation written purely in Java, an extension to Apache Lucene.
- LinkedIn Cleo - is a flexible software library for enabling rapid development of partial, out-of-order and real-time typeahead search.
- LinkedIn Zoie - is a realtime search/indexing system written in Java.
- MG4J - MG4J (Managing Gigabytes for Java) is a full-text search engine for large document collections written in Java. It is highly customisable, high-performance and provides state-of-the-art features and new research algorithms.
- Sphinx Search Server - fulltext search engine.
- Vespa - is an engine for low-latency computation over large data sets. It stores and indexes your data such that queries, selection and processing over the data can be performed at serving time.
-
Security
- BDA - The vulnerability detector for Hadoop and Spark
-
Service Programming
- Google Chubby - a lock service for loosely-coupled distributed systems.
- OpenMPI - message passing framework.
- Hydrosphere Mist - a service for exposing Apache Spark analytics jobs and machine learning models as realtime, batch or reactive web services.
- Spotify Luigi - a Python package for building complex pipelines of batch jobs. It handles dependency resolution, workflow management, visualization, handling failures, command line integration, and much more.
- Twitter Elephant Bird - libraries for working with LZOP-compressed data.
- Twitter Finagle - asynchronous network stack for the JVM.
- Serf - decentralized solution for service discovery and orchestration.
- Spring XD - distributed and extensible system for data ingestion, real time analytics, batch processing, and data export.
-
SQL-like processing
- Actian SQL for Hadoop - high performance interactive SQL access to all Hadoop data.
- Apache HCatalog - table and storage management layer for Hadoop.
- Aster Database - SQL-like analytic processing for MapReduce.
- Facebook PrestoDB - distributed SQL query engine.
- Spark Catalyst - is a Query Optimization Framework for Spark and Shark.
- Splice Machine - a full-featured SQL-on-Hadoop RDBMS with ACID transactions.
- Trafodion - enterprise-class SQL-on-HBase solution targeting big data transactional or operational workloads.
- Cloudera Impala - framework for interactive analysis, Inspired by Dremel.
- Concurrent Lingual - SQL-like query language for Cascading.
- Pivotal HDB - SQL-like data warehouse system for Hadoop.
- PipelineDB - an open-source relational database that runs SQL queries continuously on streams, incrementally storing results in tables.
- Actian SQL for Hadoop - high performance interactive SQL access to all Hadoop data.
- Apache Phoenix - SQL skin over HBase.
- RainstorDB - database for storing petabyte-scale volumes of structured and semi-structured data.
-
System Deployment
- Brooklyn - library that simplifies application deployment and management.
- Buildoop - Similar to Apache BigTop based on Groovy language.
- Google Borg - job scheduling and monitoring system.
- Google Omega - job scheduling and monitoring system.
- Kubernetes - a system for automating deployment, scaling, and management of containerized applications.
- Apache Slider - is a YARN application to deploy existing distributed applications on YARN.
- Brooklyn - library that simplifies application deployment and management.
- Buildoop - Similar to Apache BigTop based on Groovy language.
- Marathon - Mesos framework for long-running services.
- Cloudera HUE - web application for interacting with Hadoop.
-
Time-Series Databases
- Axibase Time Series Database - Integrated time series database on top of HBase with built-in visualization, rule-engine and SQL support.
- InfluxDB - distributed time series database.
- Prometheus - a time series database and service monitoring system.
- Rhombus - series object store for Cassandra that handles all the complexity of building wide row indexes.
- Chronix - a time series storage built to store time series highly compressed and for fast access times.
- Cube - uses MongoDB to store time series data.
- Heroic - is a scalable time series database based on Cassandra and Elasticsearch.
- Kairosdb - similar to OpenTSDB but allows for Cassandra.
- Newts - a time series database based on Apache Cassandra.
- Beringei - Facebook's in-memory time-series database.
- Druid
- Akumuli - series database. It can be used to capture, store and process time-series data in real-time. The word "akumuli" can be translated from esperanto as "accumulate".
- Dalmatiner DB
- Blueflood
- Timely
- Thanos - Thanos is a set of components to create a highly available metric system with unlimited storage capacity using multiple (existing) Prometheus deployments.
-
Videos
-
2001 - 2010
- Spark in Motion - Spark in Motion teaches you how to use Spark for batch and streaming data analytics.
-
Programming Languages
Categories
Interesting Papers
47
Distributed Programming
44
Data Visualization
38
Machine Learning
31
Applications
25
Key-value Data Model
23
NewSQL Databases
23
Books
21
Graph Data Model
20
Business Intelligence
18
Data Ingestion
17
Time-Series Databases
16
SQL-like processing
14
Distributed Filesystem
14
Search engine and framework
13
Columnar Databases
10
MySQL forks and evolutions
10
System Deployment
10
Internet of things and sensor data
9
Key Map Data Model
9
Service Programming
8
PostgreSQL forks and evolutions
7
Benchmarking
6
Scheduling
5
Document Data Model
5
RDBMS
4
Embedded Databases
4
Interesting Readings
3
Memcached forks and evolutions
3
Frameworks
3
Security
1
Distributed Index
1
Videos
1
Sub Categories
Keywords
database
14
go
6
graph
5
python
5
analytics
5
mysql
4
visualization
4
data-visualization
4
geospatial
4
spark
4
in-memory
3
sql
3
data
3
nosql
3
metrics
3
big-data
3
time-series
3
graph-database
3
awesome-list
3
awesome
3
java
3
golang
2
key-value
2
aggregation
2
javascript
2
distributed
2
machine-learning
2
thrift
2
serverless
2
scale
2
webgl
2
hadoop
2
timeseries-database
2
timeseries
2
distributed-database
2
rest-api
2
data-science
2
dashboard
2
index
2
data-analysis
2
business-intelligence
2
accumulo
2
kafka
2
hbase
2
c-plus-plus
2
postgresql
2
lists
2
resources
2
bi
2
mysql-compatibility
1