data-engineering-collection
A collection of awesome software, libraries, Learning Tutorials, documents, books, resources and interesting stuff about Big Data Science & Engineering
https://github.com/exajobs/data-engineering-collection
Last synced: about 9 hours ago
JSON representation
-
System Deployment
- Linkis - Linkis helps easily connect to various back-end computation/storage engines.
- Cloudera HUE - web application for interacting with Hadoop.
- Apache Helix - cluster management framework.
- Apache YARN - Cluster manager.
- Hortonworks HOYA - application that can deploy HBase cluster on YARN.
-
Time-Series Databases
- QuestDB - high-performance, open-source SQL database for applications in financial services, IoT, machine learning, DevOps and observability.
- Rhombus - series object store for Cassandra that handles all the complexity of building wide row indexes.
- Axibase Time Series Database - Integrated time series database on top of HBase with built-in visualization, rule-engine and SQL support.
- InfluxDB - a time series database with optimised IO and queries, supports pgsql and influx wire protocols.
- M3DB - a distributed time series database that can be used for storing realtime metrics at long retention.
- Prometheus - a time series database and service monitoring system.
- Chronix - a time series storage built to store time series highly compressed and for fast access times.
- Cube - uses MongoDB to store time series data.
- Heroic - is a scalable time series database based on Cassandra and Elasticsearch.
- Newts - a time series database based on Apache Cassandra.
- Beringei - Facebook's in-memory time-series database.
- VictoriaMetrics - fast, scalable and resource-effective open-source TSDB compatible with Prometheus. Single-node and cluster versions included
- TDengine - a time series database in C utilizing unique features of IoT to improve read/write throughput and reduce space needed to store data
- Druid
- IronDB - scalable, general-purpose time series database.
- Thanos - Thanos is a set of components to create a highly available metric system with unlimited storage capacity using multiple (existing) Prometheus deployments.
- OpenTSDB - distributed time series database on top of HBase.
- TrailDB - an efficient tool for storing and querying series of events.
-
Videos
-
2001 - 2010
- Spark in Motion - Spark in Motion teaches you how to use Spark for batch and streaming data analytics.
- Machine Learning, Data Science and Deep Learning with Python - LiveVideo tutorial that covers machine learning, Tensorflow, artificial intelligence, and neural networks.
- Data warehouse schema design - dimensional modeling and star schema - Introduction to schema design for data warehouse using the star schema method.
- Elasticsearch 7 and Elastic Stack - LiveVideo tutorial that covers searching, analyzing, and visualizing big data on a cluster with Elasticsearch, Logstash, Beats, Kibana, and more.
-
Programming Languages
Categories
`Distributed Programming `
53
Interesting Papers
46
Data Visualization
44
Databases
44
Machine Learning
42
Books
35
Applications
31
`NewSQL Databases`
28
Data Ingestion
26
SQL-like processing
24
Graph Data Model
24
Key-value Data Model
24
Business Intelligence
20
Time-Series Databases
18
Search engine and framework
17
`Distributed Filesystem `
17
System Deployment
14
MySQL forks and evolutions
13
`Columnar Databases`
13
Service Programming
11
`Key Map Data Model `
11
Scheduling
10
Internet of things and sensor data
10
PostgreSQL forks and evolutions
8
`Frameworks `
6
Benchmarking
6
`Document Data Model `
5
Memcached forks and evolutions
5
Videos
4
Embedded Databases
4
Interesting Readings
4
`RDBMS `
4
Security
3
License
2
`Distributed Index `
1
Sub Categories
Keywords
database
21
machine-learning
13
deep-learning
11
python
10
data-science
9
go
8
sql
7
analytics
7
kafka
6
java
6
time-series
6
data-visualization
6
graph
6
mysql
5
network-embedding
5
visualization
5
awesome-list
5
spark
5
awesome
5
metrics
5
distributed-database
5
golang
5
kubernetes
4
tensorflow
4
network-science
4
classifier
4
graph-embedding
4
random-forest
4
pytorch
4
monitoring
4
geospatial
4
distributed
4
node-embedding
4
tsdb
3
hadoop
3
workflow
3
c-plus-plus
3
graph-database
3
in-memory
3
data-analysis
3
big-data
3
postgresql
3
jupyter
3
gradient-boosting
3
rust
3
stream-processing
3
nosql
3
distributed-systems
3
etl
3
node2vec
3