awesome-apache-iceberg
A curated list of apache iceberg and surrounding ecosystem
https://github.com/zriyansh/awesome-apache-iceberg
Last synced: 10 days ago
JSON representation
-
📢 Acknowledgements
-
Contribution Guidelines
-
-
📂 Additional Sections
-
2. Tutorials and Learning Resources
- Apache Iceberg Documentation
- Getting Started with Apache Iceberg
- Iceberg vs. Delta Lake vs. Hudi
- Databricks Blog on Iceberg
- YouTube - Apache Iceberg Tutorials
- Confluent YouTube Channel
- Simplilearn - Iceberg Tutorials
- edureka! Data Engineering Tutorials
- Udemy: Apache Iceberg Essentials
- Coursera: Data Engineering on Google Cloud
- LinkedIn Learning: Modern Data Warehousing with Apache Iceberg
- Pluralsight: Advanced Data Engineering with Apache Iceberg
- Towards Data Science - Apache Iceberg Articles
- Towards Data Science - Apache Iceberg Articles
-
3. Open-source Projects
- lakeFS - Git-like version control for data lakes.
- Great Expectations - Data validation framework.
- Delta Sharing - Open protocol for secure data sharing.
- Materialize - Streaming SQL database for real-time analytics.
- Airbyte - Data integration platform.
- Amundsen - Data discovery and metadata engine.
- DataHub - Metadata platform for the modern data stack.
-
-
Blog Posts & Articles
-
Introduction & Overview
- Apache Iceberg University - Your guide to learning concepts and practices
- Apache Iceberg: A Beginner's Complete Guide - DataCamp's complete introduction
- Apache Iceberg: A Beginner's Complete Guide - DataCamp's complete introduction
-
Migration Guides
- Your Guide to Migrating to an Apache Iceberg Lakehouse - Comprehensive migration strategies
- Migrating to Apache Iceberg: A Complete Guide - Medium article by Hugo Lu.
- Migrating to Apache Iceberg: A Complete Guide - Medium article by Hugo Lu.
-
Technical Deep Dives
- Iceberg at Netflix: Data Platform Architecture - Netflix Engineering Team
- AutoOptimize: Netflix's Data Layout Optimization - Automatic file optimization at petabyte scale
- FastIngest: Low-latency Gobblin with Apache Iceberg - LinkedIn's 45min to 5min latency reduction
- Iceberg at Netflix: Data Platform Architecture - Netflix Engineering Team
- AutoOptimize: Netflix's Data Layout Optimization - Automatic file optimization at petabyte scale
- Iceberg at Netflix: Data Platform Architecture - Netflix Engineering Team
- AutoOptimize: Netflix's Data Layout Optimization - Automatic file optimization at petabyte scale
- Iceberg at Netflix: Data Platform Architecture - Netflix Engineering Team
- AutoOptimize: Netflix's Data Layout Optimization - Automatic file optimization at petabyte scale
-
-
Books & Courses
-
Books
- Apache Iceberg: The Definitive Guide - O'Reilly Media
-
-
Community
-
Communication Channels
- Slack Workspace - Primary community channel
-
Community Events
- Iceberg Community Events - Events such as conferences and meetups, aimed to educate and inspire Iceberg users.
- Iceberg Dev Events - Events such as the triweekly Iceberg sync, aimed to discuss the project roadmap and how to implement features.
-
Development
- GitHub Issues - Bug reports and feature requests
- Contributing Guide - How to contribute
-
-
Conference Talks
-
Iceberg Summit
- Iceberg Summit 2024 - All 32 recordings from inaugural summit (May 14-15, 2024)
- Iceberg Summit 2025 - April 8 (in-person), April 9 (virtual)
- Iceberg Summit 2024 - All 32 recordings from inaugural summit (May 14-15, 2024)
- Iceberg Summit 2025 - April 8 (in-person), April 9 (virtual)
-
-
🤝 Contributing
-
Contribution Guidelines
-
-
🔑 Core Sections
-
1. Iceberg Fundamentals
- Introducing Apache Iceberg - Official Blog Post
- Iceberg: Table Format for Large Analytics Datasets - SlideShare Presentation
- Academic Papers on Iceberg
-
2. Key Iceberg Technologies
- Apache Iceberg - Core table format for managing large datasets.
- Delta Lake - Complementary storage layer providing ACID transactions.
- Apache Hudi - Another storage layer option with unique features.
- Apache Spark - Unified analytics engine for large-scale data processing.
- Trino - High-performance distributed SQL query engine.
- Presto - Distributed SQL query engine for big data.
- Hive - Data warehouse software for querying and managing large datasets.
- Hive Metastore - Central repository for metadata.
- Amundsen - Data discovery and metadata engine.
- DataHub - Metadata platform for the modern data stack.
- Apache Atlas - Governance and metadata framework.
-
3. ETL/ELT Tools for Iceberg
- Apache Flink - Stream processing framework.
- Airbyte - Open-source data integration platform.
- Meltano - Open-source data integration tool built on Singer.
- dbt (data build tool) - Transform data in your warehouse more effectively.
- Matillion - Data integration and transformation tool.
- Apache Kafka - Distributed event streaming platform.
- Spark Streaming - Scalable stream processing.
-
4. Data Orchestration
- Apache Airflow - Platform to programmatically author, schedule, and monitor workflows.
- Dagster - Data orchestrator for machine learning, analytics, and ETL.
- Prefect - Workflow management system.
-
5. BI and Analytics on Iceberg
-
6. ML and AI Workflows
- MLflow - Open-source platform for managing the ML lifecycle.
- Databricks Machine Learning - Unified environment for ML.
-
7. Monitoring and Observability
- OpenTelemetry - Observability framework for cloud-native software.
- Prometheus - Monitoring system and time series database.
- Grafana - Open-source analytics and monitoring platform.
- Great Expectations - Data validation framework.
-
8. Cost Optimization
- DuckDB - An in-process SQL OLAP Database Management System.
- ClickHouse - Fast open-source column-oriented database management system.
- Materialize - Streaming database for real-time applications.
-
-
Ecosystem Tools
-
Development Tools
- Ducklake - Lightweight data lake with DuckDB (100+ stars)
-
-
🌟 Influential Personalities
-
**1. Ryan Blue**
-
**2. Benjamin Haimowitz**
-
**3. Paige Nord**
-
**4. Robin Moffatt**
-
**5. James Nestor**
-
**6. Maxime Beauchemin**
-
-
Key People
-
Official Resources
-
Core Documentation
- Documentation - Comprehensive docs covering all libraries and integrations
- Specification - Official format specification (stable with new features each version)
- Multi-Engine Support - Compatibility matrix for Spark, Flink, and Hive versions
-
GitHub Repositories
- apache/iceberg - Core Java implementation (6.3k+ stars)
- apache/iceberg-python - PyIceberg - Python implementation (748 stars)
- apache/iceberg-rust - Rust implementation (700+ stars)
- apache/iceberg-go - Go implementation with CLI tools (200+ stars)
- apache/iceberg-cpp - C++ implementation
-
-
Open Source Tools
-
Catalog Implementations
- Project Nessie - Git-like version control for data (1.2k stars)
- Lakekeeper - Secure REST catalog with fine-grained access control (350+ stars)
-
Migration Tools
- Iceberg Catalog Migrator - Bulk migration between catalogs
- Apache XTable - Cross-table converter for Delta, Hudi, Iceberg (800+ stars)
-
REST Catalog Implementations
- Cloudcheflabs Iceberg REST Catalog - Easy-to-deploy server
- Tabular Iceberg REST - Docker image for REST API (200+ stars)
-
-
Vendor Documentation
-
Cloud Providers
- AWS Prescriptive Guidance - Best practices for Iceberg on AWS
- Amazon Athena - Native Iceberg v1.4.2 support
- AWS Glue - ETL jobs and Data Catalog integration
- Amazon EMR - Support from version 6.5.0+
-
Categories
Sub Categories
2. Tutorials and Learning Resources
14
2. Key Iceberg Technologies
11
Technical Deep Dives
9
3. Open-source Projects
7
3. ETL/ELT Tools for Iceberg
7
Original Creators (Netflix)
6
GitHub Repositories
5
Cloud Providers
4
7. Monitoring and Observability
4
Iceberg Summit
4
**3. Paige Nord**
3
1. Iceberg Fundamentals
3
**2. Benjamin Haimowitz**
3
4. Data Orchestration
3
8. Cost Optimization
3
Introduction & Overview
3
**4. Robin Moffatt**
3
**6. Maxime Beauchemin**
3
Core Documentation
3
Migration Guides
3
Contribution Guidelines
3
**5. James Nestor**
3
**1. Ryan Blue**
3
5. BI and Analytics on Iceberg
3
6. ML and AI Workflows
2
Migration Tools
2
Community Events
2
Catalog Implementations
2
Development
2
Active PMC Members & Contributors
2
REST Catalog Implementations
2
Governance & Security
1
Books
1
Communication Channels
1
Development Tools
1
Keywords
iceberg
7
apache
5
data-engineering
3
golang
2
data-quality
2
metadata
2
data-discovery
2
data-catalog
2
spark
2
postgresql
2
pipeline
2
java
2
rust
2
data
2
data-governance
1
datahub
1
lists
1
apache-spark
1
apache-sparksql
1
aws-s3
1
azure-blob-storage
1
azure-storage
1
data-lake
1
awesome-list
1
data-version-control
1
data-versioning
1
datalake
1
datalakes
1
git-for-data
1
go
1
awesome
1
data-pipeline
1
elt
1
etl
1
data-integration
1
mssql
1
mysql
1
data-collection
1
data-analysis
1
change-data-capture
1
bigquery
1
python
1
redshift
1
s3
1
self-hosted
1
snowflake
1
unicorns
1
resources
1
aws-lambda
1
git
1