awesome-bdccai-tools
Werkzeuge für Big Data und Cloud-Computing für KI
https://github.com/cyberlytics/awesome-bdccai-tools
Last synced: 18 days ago
JSON representation
-
Appendix: More Free Student Stuff
-
Big Data
- Apache **Hadoop**
- Key/Value-Stores - Stores](https://db-engines.com/de/ranking/document+store) | [Wide-Column-Stores](https://db-engines.com/de/ranking/wide+column+store)
- Snowflake - native DWaaS
- CockroachDB - cockroachdb-windows)\]: Open-Source NewSQL; **PostgreSQL**\-compatible; built on a transactional and strongly-consistent key-value store
- YugabyteDB - start/docker/)\]: Open-source NewSQL; **PostgreSQL**\-compatible
- PingCAP **TiDB** - source NewSQL (OLTP/HTAP workloads); **MySQL**\-compatible; built on a transactional key-value store
- Twitter **Summingbird**
- Apache **Flink**
- Stratosphere - dr-volker-markl.html))
- **Kafka** Streams
- Apache **Druid** - time (i.e., sub-second) analytics database, with separation of ingest compute and query compute
- **InfluxDB** Open Source - time (i.e., sub-second) analytics of IoT Data; core component of the [TICK stack](https://www.influxdata.com/time-series-platform/)
- **Splunk** Free
- logz.io
- OpenCL - und/oder datenbasierter Parallelität, als HW-Plattform-Abstraktion für CPUs, GPUs, DSPs, FGPAs, usw.
- JOCL - Abstraktion für OpenGL, OpenAL, OpenCL und OpenVR
- Big Data Landscape - data/) (= Machine Learning, AI & Data)
- Oracle Big Data Lite Virtual Machine - data-lite)
- HowTo: Big Data Europe
- HowTo: Hadoop 3.x
- HowTo: Spark as part of jupyter - p 8888:8888 -e JUPYTER_ENABLE_LAB=yes --name pyspark jupyter/pyspark-notebook\] or [HowTo: Spark Cluster](https://medium.com/@MarinAgli1/setting-up-a-spark-standalone-cluster-on-docker-in-layman-terms-8cbdc9fdd14b)
- HowTo: Spark on Hadoop 3.x - data/how-to-install-apache-spark-on-windows)
- choco install mongodb.install \ - compass\] or [Studio 3T](https://studio3t.com/download/) \[choco install studio3t\]
- ScyllaDB - in replacement for Cassandra, providing the same CQL interface and drivers \[docker run --name scylla -d scylladb/scylla\]
- Community Edition (CE) - itd --name couchbase-server -p 8091-8094:8091-8094 -p 11210:11210 couchbase:community\]
- Desktop - Managed » Community](https://neo4j.com/deployment-center/?gdb-selfmanaged)
- Databricks - edition)
- JOCL - Abstraktion für OpenGL, OpenAL, OpenCL und OpenVR
- Awesome Big Data #1 - bigdata)
- Snowflake
- Matano
- Spark Shell - it apache/spark /opt/spark/bin/spark-shell`\]
- HowTo: Spark as part of jupyter - p 8888:8888 -e JUPYTER_ENABLE_LAB=yes --name pyspark jupyter/pyspark-notebook\] or [HowTo: Spark Cluster](https://medium.com/@MarinAgli1/setting-up-a-spark-standalone-cluster-on-docker-in-layman-terms-8cbdc9fdd14b)
- Community Edition (CE) - itd --name couchbase-server -p 8091-8094:8091-8094 -p 11210:11210 couchbase:community\]
- Databricks - edition)
- CockroachDB - cockroachdb-windows)\]: Open-Source NewSQL; **PostgreSQL**-compatible; built on a transactional and strongly-consistent key-value store
- Stratosphere - dr-volker-markl.html))
- **InfluxDB** Open Source - time (i.e., sub-second) analytics of IoT Data; core component of the [TICK stack](https://www.influxdata.com/time-series-platform/)
-
Cloud-Computing
- Micronaut
- Amazon **AWS** - services/), [**Google** Cloud](https://cloud.google.com/free), [**Alibaba** Cloud](https://www.alibabacloud.com/free), [**IBM** Cloud](https://www.ibm.com/cloud/free), [**Tencent** Cloud](https://www.tencentcloud.com/campaign/freetier), [Oracle **OCI**](https://www.oracle.com/cloud/free/), Heroku (no free tier anymore \*sigh\*), **[DigitalOcean](https://www.digitalocean.com/pricing/app-platform)**, [**SAP** BTP](https://www.sap.com/products/technology-platform/trial.html)
- Vercel - service/static/) \[**[Awesome JAMstack](https://github.com/automata/awesome-jamstack)**\]
- podman - Gehversuche](https://github.com/containers/podman/blob/main/docs/tutorials/podman-for-windows.md)) sowie ggf. [buildah](https://buildah.io/) direkt
- KataContainers - microvm.GitHub.io/), google [gVisor](https://gvisor.dev/) (und historisch: CoreOS [rkt](https://github.com/rkt/rkt), später per [CNCF archived](https://www.cncf.io/archived-projects/))
- CRI-O
- Play with Docker - basierte Docker-Umgebung
- hadolint - -rm -i hadolint/hadolint < Dockerfile**\]
- . - ignore-dockerignore-2/)\-Generator!?)
- Divio - basierte Proof-of-Concept WebApps
- Alpine
- Quay
- Play with K8s - basierte Kubernetes-Umgebung
- **OpenShift** Playground - basierte OpenShift-Umgebung
- Rancher Desktop - desktop**\]: Runs Kubernetes and container management on your desktop
- RKE1 - Container-basiert, über [RancherOS](https://github.com/rancher/os), mit Docker als Container Engine) \[**choco install rke**\]
- K8sGPT
- Waypoint
- Resilience4j
- Arkade - of-apps) for Kubernetes
- Portainer
- cAdvisor
- Quarkus
- Val Town
- Apache **OpenWhisk**
- Knative - sponsored Open Source Serverless Cloud Platform
- OpenFaaS - based [Scale to Zero](https://docs.openfaas.com/openfaas-pro/scale-to-zero/)
- Cellery - based [Knative](https://knative.dev/)
- Cold-Start - Zeiten:
- SWE » DevOps
- Kepler - Metriken exportierbar
- Nubenetes - awesome-lists/) | [Awesome Sysadmin](https://github.com/awesome-foss/awesome-sysadmin) | [Awesome Chaos Engineering](https://github.com/dastergon/awesome-chaos-engineering) | [Awesome AWS](https://github.com/donnemartin/awesome-aws) | [Awesome Serverless](https://github.com/anaibol/awesome-serverless) | [Awesome Lambda Essentials](https://github.com/danteata/awesome-aws-lambda) | [Awesome Lambda Layers](https://github.com/mthenw/awesome-layers)
- Kubernetes Failure Stories
- dockur
- docker run … [docker.io/registry:2
- Funky Penguin's Geek Cookbook
- Europäische Alternativen
- Azure for Students
- Alpine
- k3d
- k3s
- RKE2
- K8sGPT
- KEDA - based Event-Driven Autoscaler
-
dApps
- LBRY - based file-sharing, social networks and video platform („open, free, and fair network for digital content“)
- DappRadar
- Awesome dApps - web3)
- Substrate
-
Data Science
- KNIME
- RapidMiner Studio auch Open Source
- **Juyter** Docker Stacks - p 8888:8888 jupyter/scipy-notebook**\]
- binder - Bindeglied zwischen Jupyter und Ihren git-gehosteten Notebook-Dateien (Obacht: Wenn in den USA die Leute aufstehen, dann geht der globally-shared Infrastruktur von binder ggf. die Puste aus, daher ggf. nicht ausreichend zuverlässig für Lehrveranstaltungen oder Konferenz-Demos)
- CoCalc
- SageMath
- SQL Notebook
- Tad
- Google **Collab**
- **kaggle** Datasets - public-datasets)** ⭐
- **100 interesting data sets** for statistics - public-data-sets-data-science-project/)
- Worldbank Open Data
- **Quick Draw!** The Data - ähnlich; bspw. [Ameisen](https://quickdraw.withgoogle.com/data/ant))
- The Stack - Listing](https://www.numbersstation.ai/post/introducing-nsql-open-source-sql-copilot-foundation-models)
- Rucio
- Awesome Data Science - awesome.org/krzjoa/awesome-python-data-science)
- JASP
- Jamovi Desktop
- SOFA Statistics
- Positron
- VersaTiles
- Laion Open Dataset
- R Studio - R](https://tinn-r.org/en/) \[choco install tinn-r\]
- RapidMiner
- DBpedia
-
Datenbanksysteme
- Docker Container
- TPC-H Benchmark - Systeme (suite of business oriented ad-hoc decision support queries and concurrent data modifications)
- DuckDB - Memory-basiertes In-Process-fähiges ACID-konformes RDBMS für analytische Workloads
- Embeddable - reasons-duckdb-slaps/)
- Elastic - Stack \[via **[Docker](https://www.elastic.co/guide/en/elasticsearch/reference/current/docker.html)** oder **choco install elasticsearch** sowie **choco install kibana**\]: im Kern eine verteilte Volltext-Suchmaschine, basierend auf Lucene; aber auch als skalierbares NoSQL-System verwendbar
- MySQL
- MariaDB - d -p 3306:3306 -e MYSQL_ROOT_PASSWORD=geheim mariadb:latest**\] sowie [Maria **Galera**](https://mariadb.com/kb/en/galera-cluster/) = Multi-Master-Cluster
- PostgreSQL - -params '"/Password:geheim /Port:5432"' --params-global**\]
- SQLite - studio.portable**\]
- **Oracle** Database Free - d -p 1521:1521 -e ORACLE_PASSWORD=geheim gvenzl/oracle-free**\]
- Microsoft **SQL Server** Express - server-express** sowie **choco install sql-server-management-studio**\] ggf. noch statischen Port 1433 konfigurieren
- IBM **Db2** Community Edition
- CouchDB
- BaseX
- Neo4j - community**\]: [Cypher](https://neo4j.com/developer/cypher/) Anfragesprache (Further Reading: [Neo4j GraphAcademy](https://neo4j.com/graphacademy/online-training/))
- bit.io - 1.4))
- **OCI Cloud** Free Tier - Instanzen, je 20GB, verschieden Typen, bspw. Exadata oder NoSQL
- dbfiddle - basierter SQL-Datenbank-Playground (diverse Datenbanksysteme)
- MongoDB Atlas - Variante des klassischen NoSQL-Systems (The „M“ in MEAN and MERN) – kostenlos für 512MB
- **CockroachDB** SQL-Playground - Variante des NewSQL-Datenbanksystems (s. unten)
- LeanStore - performance OLTP storage engine optimized for many-core CPUs and NVMe SSDs (Prof. Viktor Leis)
- HyPer - memory-based relational DBMS for mixed OLTP and OLAP workloads (aquired by Tableau)
- Umbra - based system with in-memory performance
- Peloton - driving main-memory-based relational DBMS for mixed OLTP and OLAP workloads
- Oracle Docker Images
- Couchbase Capella
- neo4j AuraDB
- EXASOL - Memory-basiertes MPP-fähiges ACID-konformes RDBMS für analytische Workloads
- Peloton - driving main-memory-based relational DBMS for mixed OLTP and OLAP workloads
- Milvus - U pymilvus**\], [ChromaDB](https://github.com/chroma-core/chroma) \[pip install chromadb\], ...
- KDB
- AWS MemoryDB
- LanceDB
- LanceDB Cloud
- db-engines
- EXASOL - Memory-basiertes MPP-fähiges ACID-konformes RDBMS für analytische Workloads
- **CockroachDB** SQL-Playground - Variante des NewSQL-Datenbanksystems (s. unten)
- HyPer - memory-based relational DBMS for mixed OLTP and OLAP workloads (aquired by Tableau)
- **CockroachDB** SQL-Playground - Variante des NewSQL-Datenbanksystems (s. unten)
- EXASOL - Memory-basiertes MPP-fähiges ACID-konformes RDBMS für analytische Workloads
- Community Edition
- IBM **Db2** Community Edition
- Neo4j - community**\]: [Cypher](https://neo4j.com/developer/cypher/) Anfragesprache (Further Reading: [Neo4j GraphAcademy](https://neo4j.com/graphacademy/online-training/))
- **CockroachDB** SQL-Playground - Variante des NewSQL-Datenbanksystems (s. unten)
-
Datenbankwerkzeuge
- DbVisualizer - visualizer**\]: Advanced SQL Editor
- **DBeaver** Community Edition - Source Hintergrund](https://github.com/dbeaver/dbeaver), **[DbVisualizer](https://www.dbvis.com/pricing/)** \[**choco install db-visualizer**\], etc.
- DataGrip
- DbSchema
- SchemaCrawler - Source Database Schema Discovery and Comprehension Tool; generates Schema Diagrams
- SchemaSpy
- Navicat Data Modeler - data-modeler-essentials);
- Ab Initio - UIW-731/images/Gartner-MQ-2022-Quadrant.png) auftaucht)
- Artikel von Atlan
- Talend Open Studio for Data Quality - US/8.0/studio-user-guide-open-studio-for-data-quality/detecting-anomalies-in-columns-functional-dependency-analysis)
- Data Catalog
- Artikel von Atlan
- Dozenten-Werkzeuge
- SQLFlow
- Artikel von Atlan
- CleverCSV - Code für CSV-Imports, zzgl. Bibliotheksfunktionen für Data Cleaning
- OpenRefine
- DataVault Builder - builder.com/pricing/)
- RelaX
- Functional Dependency Calculator
- Awesome Database Tools - Source Data Engineering](https://github.com/gunnarmorling/awesome-opensource-data-engineering) (ähnliche [Liste mit kommerziellen Optionen](https://github.com/igorbarinov/awesome-data-engineering))
- Beekeeper - studio.portable**\]: Advanced SQL Editor
- Collibra - content/uploads/2022/11/figure1.png), Collibra ist ggf. hier nicht Best-in-Class aber wegen Collibras Überlappung mit DQ und MDM attraktiv)
- Alteryx **Trifacta** - license/)
- Airbyte
- DbGate - Source Cross-Plattform SQL Editor, inkl. NoSQL-Unterstützung und starkem JSON-Viewer
- Talend Open Studio
- DbGate - Source Cross-Plattform SQL Editor, inkl. NoSQL-Unterstützung und starkem JSON-Viewer
- Functional Dependency Calculator
- Flyway
- Talend Open Studio
- Airbyte
- Talend Open Studio for Data Quality - core)
- WinPure Clean & Match
- RelaX
-
Datenvisualisierung
- Tableau - desktop**\]
- kostenlos für Studierende
- Looker
- Qlik Sense
- kostenlos für Studierende
- Eclipse **BIRT** - Installation\]
- **KNOWAGE** Community Edition
Programming Languages
Categories
Security
86
Moderne Web-Anwendungsentwicklung
66
Privacy
64
ML / AI
54
Datenbanksysteme
44
Cloud-Computing
44
Big Data
38
Datenbankwerkzeuge
35
Data Science
25
Datenvisualisierung
22
DevOps
21
Low-Code / No-Code
11
Edge / Fog / IoT
10
Verteilte Systeme
9
Semantic Web / Knowledge Representation
8
Mobile Apps
8
Operations Research / Optimization
5
Footer
4
Operations Research (OR) / Optimization
4
dApps
4
MLOps
3
Appendix: More Free Student Stuff
3
Uncategorized
2
Sub Categories
Keywords
awesome-list
19
awesome
13
security
10
python
10
machine-learning
9
llm
7
database
5
docker
5
privacy
5
ai
5
osint
4
golang
4
kubernetes
4
deep-learning
4
artificial-intelligence
4
pytorch
3
go
3
compliance
3
natural-language-processing
3
devops
3
students
3
free
3
react
3
anonymization
3
javascript
3
data-visualization
3
data-science
3
llmops
3
pentesting
2
aws
2
mysql
2
business-intelligence
2
chart
2
dotnet
2
energy-monitor
2
sql
2
static-analysis
2
electron
2
rag
2
electronjs
2
dockerfile
2
openai
2
distributed-database
2
cloud-native
2
windows
2
vuejs
2
azure
2
macos
2
ccpa
2
linux
2