An open API service indexing awesome lists of open source software.

awesome-bigdata

A curated list of awesome big data frameworks, ressources and other awesomeness
https://github.com/Anyz01/awesome-bigdata

Last synced: 13 days ago
JSON representation

  • Embedded Databases

    • HanoiDB - Erlang LSM BTree Storage.
    • LevelDB - a fast key-value storage library written at Google that provides an ordered mapping from string keys to string values.
    • BerkeleyDB - a software library that provides a high-performance embedded database for key/value data.
  • Frameworks

    • IBM Streams - platform for distributed processing and real-time analytics. Integrates with many of the popular technologies in the Big Data ecosystem (Kafka, HDFS, Spark, etc.)
    • Tigon - High Throughput Real-time Stream Processing Framework.
    • Pachyderm - Pachyderm is a data storage platform built on Docker and Kubernetes to provide reproducible data processing and analysis.
  • Graph Data Model

    • MapGraph - Massively Parallel Graph processing on GPUs.
    • Neo4j - graph database written entirely in Java.
    • Titan - distributed graph database, built over Cassandra.
    • NodeXL - A free, open-source template for Microsoft® Excel® 2007, 2010, 2013 and 2016 that makes it easy to explore network graphs.
    • AgensGraph - a new generation multi-model graph database for the modern complex data environment.
    • AgensGraph - a new generation multi-model graph database for the modern complex data environment.
    • DGraph - A scalable, distributed, low latency, high throughput graph database aimed at providing Google production level scale and throughput, with low enough latency to be serving real time user queries, over terabytes of structured data.
    • EliasDB - a lightweight graph based database that does not require any third-party libraries.
    • GCHQ Gaffer - Gaffer by GCHQ is a framework that makes it easy to store large-scale graphs in which the nodes and edges have statistics.
    • Google Cayley - open-source graph database.
    • Gremlin - graph traversal Language.
    • Infovore - RDF-centric Map/Reduce framework.
    • Phoebus - framework for large scale graph processing.
    • Titan - distributed graph database, built over Cassandra.
    • Twitter FlockDB - distributed graph database.
    • Facebook TAO - TAO is the distributed data store that is widely used at facebook to store and serve the social graph.
    • OrientDB - document and graph database.
    • GraphLab PowerGraph - a core C++ GraphLab API and a collection of high-performance machine learning and data mining toolkits built on top of the GraphLab API.
    • Intel GraphBuilder - tools to construct large-scale graphs on top of Hadoop.
    • GraphX - resilient Distributed Graph System on Spark.
  • Interesting Papers

    • 2001 - 2010

      • 2003 - **Google** - The Google File System.
      • 2010 - **Google** - Pregel: A System for Large-Scale Graph Processing.
      • 2010 - **Facebook** - Finding a needle in Haystack: Facebook’s photo storage.
      • 2010 - **AMPLab** - Spark: Cluster Computing with Working Sets.
      • 2010 - **Google** - Large-scale Incremental Processing Using Distributed Transactions and Notifications base of Percolator and Caffeine.
      • 2010 - **Google** - Dremel: Interactive Analysis of Web-Scale Datasets.
      • 2010 - **Yahoo** - S4: Distributed Stream Computing Platform.
      • 2009 - HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads.
      • 2008 - **AMPLab** - Chukwa: A large-scale monitoring system.
      • 2006 - **Google** - The Chubby lock service for loosely-coupled distributed systems.
      • 2004 - **Google** - MapReduce: Simplied Data Processing on Large Clusters.
      • 2010 - **Google** - Large-scale Incremental Processing Using Distributed Transactions and Notifications base of Percolator and Caffeine.
      • 2010 - **Yahoo** - S4: Distributed Stream Computing Platform.
      • 2009 - HadoopDB: An Architectural Hybrid of MapReduce and DBMS Technologies for Analytical Workloads.
      • 2007 - **Amazon** - Dynamo: Amazon’s Highly Available Key-value Store.
    • 2011 - 2012

      • 2012 - **Twitter** - The Unified Logging Infrastructure
      • 2012 - **AMPLab** - Blink and It’s Done: Interactive Queries on Very Large Data.
      • 2012 - **AMPLab** - Fast and Interactive Analytics over Hadoop Data with Spark.
      • 2012 - **AMPLab** - Shark: Fast Data Analysis Using Coarse-grained Distributed Memory.
      • 2012 - **Microsoft** - Paxos Replicated State Machines as the Basis of a High-Performance Data Store.
      • 2012 - **AMPLab** - BlinkDB: Queries with Bounded Errors and Bounded Response Times on Very Large Data.
      • 2012 - **Google** - Processing a trillion cells per mouse click.
      • 2012 - **Google** - Spanner: Google’s Globally-Distributed Database.
      • 2011 - **AMPLab** - Scarlett: Coping with Skewed Popularity Content in MapReduce Clusters.
      • 2011 - **AMPLab** - Mesos: A Platform for Fine-Grained Resource Sharing in the Data Center.
      • 2011 - **Google** - Megastore: Providing Scalable, Highly Available Storage for Interactive Services.
      • 2012 - **Twitter** - The Unified Logging Infrastructure
      • 2012 - **AMPLab** - BlinkDB: Queries with Bounded Errors and Bounded Response Times on Very Large Data.
      • 2012 - **Google** - Processing a trillion cells per mouse click.
      • 2012 - **Microsoft** - Paxos Made Parallel.
    • 2013 - 2014

      • 2014 - **Stanford** - Mining of Massive Datasets.
      • 2013 - **AMPLab** - Presto: Distributed Machine Learning and Graph Processing with Sparse Matrices.
      • 2013 - **AMPLab** - MLbase: A Distributed Machine-learning System.
      • 2013 - **AMPLab** - Shark: SQL and Rich Analytics at Scale.
      • 2013 - **AMPLab** - GraphX: A Resilient Distributed Graph System on Spark.
      • 2013 - **Google** - HyperLogLog in Practice: Algorithmic Engineering of a State of The Art Cardinality Estimation Algorithm.
      • 2013 - **Metamarkets** - Druid: A Real-time Analytical Data Store.
      • 2013 - **Google** - F1: A Distributed SQL Database That Scales.
      • 2013 - **Facebook** - Scaling Memcache at Facebook.
      • 2014 - **Stanford** - Mining of Massive Datasets.
      • 2013 - **Google** - HyperLogLog in Practice: Algorithmic Engineering of a State of The Art Cardinality Estimation Algorithm.
      • 2013 - **Metamarkets** - Druid: A Real-time Analytical Data Store.
      • 2013 - **Google** - F1: A Distributed SQL Database That Scales.
      • 2013 - **Facebook** - Scuba: Diving into Data at Facebook.
      • 2013 - **Google** - Online, Asynchronous Schema Change in F1.
    • 2015 - 2016

      • 2015 - **Facebook** - One Trillion Edges: Graph Processing at Facebook-Scale.
      • 2015 - **Facebook** - One Trillion Edges: Graph Processing at Facebook-Scale.
  • Interesting Readings

  • Internet of things and sensor data

    • Apache Edgent (Incubating) - a programming model and micro-kernel style runtime that can be embedded in gateways and small footprint edge devices enabling local, real-time, analytics on the edge devices.
    • TempoIQ - Cloud-based sensor analytics.
    • Pubnub - Data stream network
    • IFTTT - If this then that
    • Evrything - Making products smart
    • Azure IoT Hub - Cloud-based bi-directional monitoring and messaging hub
    • ThingWorx - Rapid development and connection of intelligent systems
    • NetLytics - Analytics platform to process network data on Spark.
    • NetLytics - Analytics platform to process network data on Spark.
  • Key Map Data Model

  • Key-value Data Model

    • Amazon DynamoDB - distributed key/value store, implementation of Dynamo paper.
    • Ignite - is an in-memory key-value data store providing full SQL-compliant data access that can optionally be backed by disk storage.
    • LinkedIn Krati - is a simple persistent data store with very low latency and high throughput.
    • Linkedin Voldemort - distributed key/value storage system.
    • Redis - in memory key value datastore.
    • Bolt - an embedded key-value database for Go.
    • BTDB - Key Value Database in .Net with Object DB Layer, RPC, dynamic IL and much more
    • BuntDB - a fast, embeddable, in-memory key/value database for Go with custom indexing and geospatial support.
    • Edis - is a protocol-compatible Server replacement for Redis.
    • ElephantDB - Distributed database specialized in exporting data from Hadoop.
    • HyperDex - a scalable, next generation key-value and document store with a wide array of features, including consistency, fault tolerance and high performance.
    • Linkedin Voldemort - distributed key/value storage system.
    • Oracle NoSQL Database - distributed key-value database by Oracle Corporation.
    • Riak - a decentralized datastore.
    • Storehaus - library to work with asynchronous key value stores, by Twitter.
    • SummitDB - an in-memory, NoSQL key/value database, with disk persistance and using the Raft consensus algorithm.
    • Tarantool - an efficient NoSQL database and a Lua application server.
    • TiKV - a distributed key-value database powered by Rust and inspired by Google Spanner and HBase.
    • Tile38 - a geolocation data store, spatial index, and realtime geofence, supporting a variety of object types including latitude/longitude points, bounding boxes, XYZ tiles, Geohashes, and GeoJSON
    • TreodeDB - key-value store that's replicated and sharded and provides atomic multirow writes.
    • Badger - a fast, simple, efficient, and persistent key-value store written natively in Go.
    • EventStore - distributed time series database.
    • GridDB - suitable for sensor data stored in a timeseries.
  • Machine Learning

    • Azure ML Studio - Cloud-based AzureML, R, Python Machine Learning platform
    • DataVec - A vectorization and data preprocessing library for deep learning in Java and Scala. Part of the Deeplearning4j ecosystem.
    • Deeplearning4j - Fast, open deep learning for the JVM (Java, Scala, Clojure). A neural network configuration layer powered by a C++ library. Uses Spark and Hadoop to train nets on multiple GPUs and CPUs.
    • etcML - text classification with machine learning.
    • GraphLab Create - A machine learning platform in Python with a broad collection of ML toolkits, data engineering, and deployment tools.
    • MLbase - distributed machine learning libraries for the BDAS stack.
    • MonkeyLearn - Text mining made easy. Extract and classify data from text.
    • ND4J - A matrix library for the JVM. Numpy for Java.
    • RL4J - Reinforcement learning for Java and Scala. Includes Deep-Q learning and A3C algorithms, and integrates with Open AI's Gym. Runs in the Deeplearning4j ecosystem.
    • Sibyl - System for Large Scale Machine Learning at Google.
    • Theano - A Python-focused machine learning library supported by the University of Montreal.
    • Torch - A deep learning library with a Lua API, supported by NYU and Facebook.
    • WEKA - suite of machine learning software.
    • brain - Neural networks in JavaScript.
    • Concurrent Pattern - machine learning library for Cascading.
    • convnetjs - Deep Learning in Javascript. Train Convolutional Neural Networks (or ordinary ones) in your browser.
    • Decider - Flexible and Extensible Machine Learning in Ruby.
    • Etsy Conjecture - scalable Machine Learning in Scalding.
    • H2O - statistical, machine learning and math runtime with Hadoop. R and Python.
    • MLPNeuralNet - Fast multilayer perceptron neural network library for iOS and Mac OS X.
    • scikit-learn - scikit-learn: machine learning in Python.
    • TensorFlow - Library from Google for machine learning using data flow graphs.
    • Velox - System for serving machine learning predictions.
    • BidMach - CPU and GPU-accelerated Machine Learning Library.
    • WEKA - suite of machine learning software.
    • Vowpal Wabbit - learning system sponsored by Microsoft and Yahoo!.
    • Keras - An intuitive neural net API inspired by Torch that runs atop Theano and Tensorflow.
    • nupic - Numenta Platform for Intelligent Computing: a brain-inspired machine intelligence platform, and biologically accurate neural network based on cortical learning algorithms.
    • Cloudera Oryx - real-time large-scale machine learning.
    • Sibyl - System for Large Scale Machine Learning at Google.
    • MonkeyLearn - Text mining made easy. Extract and classify data from text.
  • Memcached forks and evolutions

  • MySQL forks and evolutions

    • Amazon RDS - MySQL databases in Amazon's cloud.
    • Drizzle - evolution of MySQL 6.0.
    • MariaDB - enhanced, drop-in replacement for MySQL.
    • MySQL Cluster - MySQL implementation using NDB Cluster storage engine.
    • TokuDB - TokuDB is a storage engine for MySQL and MariaDB.
    • WebScaleSQL - is a collaboration among engineers from several companies that face similar challenges in running MySQL at scale.
    • Drizzle - evolution of MySQL 6.0.
    • ProxySQL - High Performance Proxy for MySQL.
    • WebScaleSQL - is a collaboration among engineers from several companies that face similar challenges in running MySQL at scale.
    • Percona Server - enhanced, drop-in replacement for MySQL.
  • NewSQL Databases

    • BayesDB - statistic oriented SQL database.
    • CitusDB - scales out PostgreSQL through sharding and replication.
    • FoundationDB - distributed database, inspired by F1.
    • Google F1 - distributed SQL database built on Spanner.
    • Google Spanner - globally distributed semi-relational database.
    • InfiniSQL - infinity scalable RDBMS.
    • Map-D - GPU in-memory database, big data analysis and visualization platform.
    • Pivotal GemFire XD - Low-latency, in-memory, distributed SQL data store. Provides SQL interface to in-memory table data, persistable in HDFS.
    • SAP HANA - is an in-memory, column-oriented, relational database management system.
    • Sky - database used for flexible, high performance analysis of behavioral data.
    • Sky - database used for flexible, high performance analysis of behavioral data.
    • ActorDB - a distributed SQL database with the scalability of a KV store, while keeping the query capabilities of a relational database.
    • Cockroach - Scalable, Geo-Replicated, Transactional Datastore.
    • Comdb2 - a clustered RDBMS built on optimistic concurrency control techniques.
    • Haeinsa - linearly scalable multi-row, multi-table transaction library for HBase based on Percolator.
    • InfiniSQL - infinity scalable RDBMS.
    • SenseiDB - distributed, realtime, semi-structured database.
    • TiDB - TiDB is a distributed SQL database. Inspired by the design of Google F1.
    • SymmetricDS - open source software for both file and database synchronization.
    • Actian Ingres - commercially supported, open-source SQL relational database management system.
    • NuoDB - SQL/ACID compliant distributed database.
    • H-Store - is an experimental main-memory, parallel database management system that is optimized for on-line transaction processing (OLTP) applications.
    • Oracle TimesTen in-Memory Database - in-memory, relational database management system with persistence and recoverability.
  • PostgreSQL forks and evolutions

    • RecDB - Open Source Recommendation Engine Built Entirely Inside PostgreSQL.
    • Stado - open source MPP database system solely targeted at data warehousing and data mart applications.
    • Yahoo Everest - multi-peta-byte database / MPP derived by PostgreSQL.
    • Postgres-XL - Scalable Open Source PostgreSQL-based Database Cluster.
    • HadoopDB - hybrid of MapReduce and DBMS.
    • IBM Netezza - high-performance data warehouse appliances.
    • TimescaleDB - An open-source time-series database optimized for fast ingest and complex queries
  • RDBMS

  • Scheduling