An open API service indexing awesome lists of open source software.

awesome-bigdata

A curated list of awesome big data frameworks, ressources and other awesomeness
https://github.com/Anyz01/awesome-bigdata

Last synced: 13 days ago
JSON representation

    • ElasticSearch - Search and analytics engine based on Apache Lucene.
    • Enigma.io
    • Facebook Unicorn - social graph search platform.
    • Lily HBase Indexer - quickly and easily search for any content stored in HBase.
    • LinkedIn Galene - search architecture at LinkedIn.
    • Sphinx Search Server - fulltext search engine.
    • Elassandra - is a fork of Elasticsearch modified to run on top of Apache Cassandra in a scalable and resilient peer-to-peer architecture.
    • LinkedIn Bobo - is a Faceted Search implementation written purely in Java, an extension to Apache Lucene.
    • LinkedIn Cleo - is a flexible software library for enabling rapid development of partial, out-of-order and real-time typeahead search.
    • LinkedIn Zoie - is a realtime search/indexing system written in Java.
    • MG4J - MG4J (Managing Gigabytes for Java) is a full-text search engine for large document collections written in Java. It is highly customisable, high-performance and provides state-of-the-art features and new research algorithms.
    • Sphinx Search Server - fulltext search engine.
    • Vespa - is an engine for low-latency computation over large data sets. It stores and indexes your data such that queries, selection and processing over the data can be performed at serving time.
  • Security

    • BDA - The vulnerability detector for Hadoop and Spark
  • Service Programming

    • Google Chubby - a lock service for loosely-coupled distributed systems.
    • OpenMPI - message passing framework.
    • Hydrosphere Mist - a service for exposing Apache Spark analytics jobs and machine learning models as realtime, batch or reactive web services.
    • Spotify Luigi - a Python package for building complex pipelines of batch jobs. It handles dependency resolution, workflow management, visualization, handling failures, command line integration, and much more.
    • Twitter Elephant Bird - libraries for working with LZOP-compressed data.
    • Twitter Finagle - asynchronous network stack for the JVM.
    • Serf - decentralized solution for service discovery and orchestration.
    • Spring XD - distributed and extensible system for data ingestion, real time analytics, batch processing, and data export.
  • SQL-like processing

    • Actian SQL for Hadoop - high performance interactive SQL access to all Hadoop data.
    • Apache HCatalog - table and storage management layer for Hadoop.
    • Aster Database - SQL-like analytic processing for MapReduce.
    • Facebook PrestoDB - distributed SQL query engine.
    • Spark Catalyst - is a Query Optimization Framework for Spark and Shark.
    • Splice Machine - a full-featured SQL-on-Hadoop RDBMS with ACID transactions.
    • Trafodion - enterprise-class SQL-on-HBase solution targeting big data transactional or operational workloads.
    • Cloudera Impala - framework for interactive analysis, Inspired by Dremel.
    • Concurrent Lingual - SQL-like query language for Cascading.
    • Pivotal HDB - SQL-like data warehouse system for Hadoop.
    • PipelineDB - an open-source relational database that runs SQL queries continuously on streams, incrementally storing results in tables.
    • Actian SQL for Hadoop - high performance interactive SQL access to all Hadoop data.
    • Apache Phoenix - SQL skin over HBase.
    • RainstorDB - database for storing petabyte-scale volumes of structured and semi-structured data.
  • System Deployment

    • Brooklyn - library that simplifies application deployment and management.
    • Buildoop - Similar to Apache BigTop based on Groovy language.
    • Google Borg - job scheduling and monitoring system.
    • Google Omega - job scheduling and monitoring system.
    • Kubernetes - a system for automating deployment, scaling, and management of containerized applications.
    • Apache Slider - is a YARN application to deploy existing distributed applications on YARN.
    • Brooklyn - library that simplifies application deployment and management.
    • Buildoop - Similar to Apache BigTop based on Groovy language.
    • Marathon - Mesos framework for long-running services.
    • Cloudera HUE - web application for interacting with Hadoop.
  • Time-Series Databases

    • Axibase Time Series Database - Integrated time series database on top of HBase with built-in visualization, rule-engine and SQL support.
    • InfluxDB - distributed time series database.
    • Prometheus - a time series database and service monitoring system.
    • Rhombus - series object store for Cassandra that handles all the complexity of building wide row indexes.
    • Chronix - a time series storage built to store time series highly compressed and for fast access times.
    • Cube - uses MongoDB to store time series data.
    • Heroic - is a scalable time series database based on Cassandra and Elasticsearch.
    • Kairosdb - similar to OpenTSDB but allows for Cassandra.
    • Newts - a time series database based on Apache Cassandra.
    • Beringei - Facebook's in-memory time-series database.
    • Druid
    • Akumuli - series database. It can be used to capture, store and process time-series data in real-time. The word "akumuli" can be translated from esperanto as "accumulate".
    • Dalmatiner DB
    • Blueflood
    • Timely
    • Thanos - Thanos is a set of components to create a highly available metric system with unlimited storage capacity using multiple (existing) Prometheus deployments.
  • Videos

    • 2001 - 2010

      • Spark in Motion - Spark in Motion teaches you how to use Spark for batch and streaming data analytics.