awesome-sre-tools
A curated list of Site Reliability and Production Engineering Tools
https://github.com/SquadcastHub/awesome-sre-tools
Last synced: 20 days ago
JSON representation
-
AI SRE Tools & SRE Copilots
-
Incident Communication
- Sherlocks.ai
- Resolve.ai
- Deductive.ai
- IncidentFox
- metoro.io
- IncidentFox
- IncidentFox
- Ops AI by Middleware
- tailscale-mcp - MCP server with 52 tools for managing Tailscale tailnets from AI assistants like Claude Code and Cursor.
- KubeStellar Console - AI-powered multi-cluster Kubernetes management console with MCP server (kc-agent) for AI-assisted cluster operations, pod inspection, deployment management, and real-time observability across distributed environments.
- Cynative - Deep research agent for your infra - sandboxed, read-only, covers AWS, GCP, Azure, Kubernetes, GitHub and GitLab.
- Aurora - Open source (Apache 2.0) AI SRE agent that autonomously investigates incidents and performs root cause analysis across AWS, Azure, GCP, and Kubernetes. Self-hosted via Docker Compose or Helm, works with major LLM providers or local models via Ollama.
- Anyshift - AI SRE built on a versioned resource graph of your infrastructure, for root cause analysis and predicting the impact of changes before they ship.
- Radar - Open source Kubernetes visibility tool with a built-in MCP server for AI-assisted cluster operations — topology, service traffic, events, logs, and a 31-check best-practices audit.
- KnoxOps - AI-native ops agent that gives agents production-safe execution with human review and a built-in knowledge graph.
- Hyground - Self-hosted AI SRE agent that goes beyond on-call incident resolution.
-
-
Continuous Delivery
-
Container
-
Container Orchestration
-
Container Registry
-
Deployment
- AWS CodeDeploy
- Octopus Deploy
- IBM UrbanCode
- DeployBot
- Shippable
- Codar Continuous Delivery
- Wercker
- Humanitec
- ArgoCD
- Buddy Works
- werf
- Google Cloud Build
- Codar Continuous Delivery
- Qovery - Enterprise Kubernetes management platform for deploying applications, databases, Helm charts, and Terraform modules on AWS, GCP, Azure, and Scaleway.
-
Infrastructure orchestration
-
-
Continuous Integration
-
Build
-
Integration
-
-
Continuous Monitoring
-
Container Orchestration
- AWS CloudWatch
- DebugBear
- Prometheus
- Sensu
- Kapacitor
- loggly
- NewRelic
- Pingdom
- ServerDensity
- Zabbix
- InsightOps
- Chaos Genius
- Thanos
- Mimir
- Dynatrace
- Datadog
- Elastic APM
- OnlineOrNot - Uptime monitoring for websites, APIs, and cron jobs, with integrated status pages.
- CopperEgg
- Streamdal - Code-Native Data Privacy - embed privacy controls in your application code to detect and monitor PII. [](https://github.com/streamdal/streamdal)
- Sentry
- Logstash
- Papertrail
- VictoriaMetrics
- DoctorGPT - Brings GPT into production for application log error monitoring
- Dash0 - OpenTelemetry Native Observability, built on CNCF Open Standards such as PromQL, Perses and OTLP with full cost control. Supporting Metrics, Traces and Logs with full custom dashboarding and alerting capabilities.
- CICube - AI DevOps monitoring platform by monitoring your CI workflows, detect anomalies, and provide actionable fixes.
- Shipfox - Boost GitHub Actions speed by 2x and cut costs by up to 75%, with smarter caching, deep CI insights, and zero-config setup.
- ReleaseRun Vulnerability Scanner
- Cloud Waste Scanner - Detects cloud waste and helps DevOps/platform teams identify quick cloud cost optimization opportunities.
- SSL Certificate Monitor - Open-source SSL/TLS certificate expiry monitoring tool with email alerts
- DNS Propagation Checker - Open-source DNS propagation monitoring tool with global DNS server coverage
- Ingero - eBPF-based GPU causal observability agent. Traces CUDA APIs and host kernel events to build causal chains explaining GPU latency. Includes MCP server for AI-assisted incident investigation.
- cloud-audit - AWS security auditing CLI that runs 17 checks across IAM, S3, EC2, VPC, and RDS with built-in remediation engine generating AWS CLI commands and Terraform snippets.
- FlareWarden - Uptime, content, and dependency monitoring with multi-region verification, status pages, and incident management.
- API Status Check - Real-time status monitoring dashboard for 250+ developer APIs including AWS, Stripe, GitHub, and OpenAI. Free, no signup required.
- ReleaseRun Vulnerability Scanner
- StackDriver
- Crashlytics
- LynxDB - Lightweight columnar log analytics database for SRE workflows, with a pipe-style query language inspired by SPL for investigating production logs.
- Riftmap - Cross-repo infrastructure dependency discovery and change impact analysis for multi-repo environments using Terraform, Docker, Helm, and more.
- Oack - HTTP monitoring with TCP kernel telemetry, 6-phase latency breakdown, Server-Timing header capture, Cloudflare CDN enrichment, and built-in incident management with on-call scheduling.
- OpenClaw Monitor - Real-time AI agent monitoring dashboard for OpenClaw agents. Track Gateway status, sessions, token usage & trends.
- DevHelm - Developer-first uptime monitoring with HTTP, DNS, TCP, ICMP, and heartbeat checks, dependency intelligence for 80+ providers, hosted status pages, incident management, and a full developer surface (CLI, SDKs, Terraform provider, MCP server).
- agenttrace - TUI observability for AI coding agents. Track cost, tokens, tool failures, latency, anomalies, health, diffs, and CI gates across Claude Code, Codex CLI, Gemini CLI, Aider, and Cursor exports.
- net-benchmark - DNS/HTTP/SSL benchmarking with CSV, Excel, PDF, and JSON exports.
- Prismix - Real-time status dashboard for 75+ AI services (OpenAI, Anthropic, Gemini, Mistral, etc.) with starring, email/webhook alerts, 30-day uptime history, and embeddable SVG badges.
- sunwatch - Crypto-paid uptime monitoring for side projects. Pay per monitor with USDC on Base; webhook alerts on down/up state changes.
- Respan - Observability platform for LLM and AI agent applications, with tracing, evals, prompt management, and a gateway across 250+ models.
- Faultline - Deterministic CI failure analysis CLI that classifies build logs into explainable failure types with evidence and fix steps.
-
-
Continuous Testing
-
Code Editors and IDEs
- JUnit
- TestNG
- NUnit
- TestSigma
- Unified Functional Testing (UFT)
- Tricentis Tosca
- IBM Rational Functional Tester
- TestComplete
- Waitr
- Zephyr
- accelQ
- Apache JMeter
- Appium
- steadybit
- k6
- Gatling
- Cypress
- Waitr
- Waitr
- Waitr
- Waitr
- Waitr
- Unified Functional Testing (UFT)
- TestComplete
- Selenium
- Waitr
- TestRail
- Bencher
- Zephyr
- flakybin
-
-
Development
-
Bug / Defect Tracking Software
-
Code Editors and IDEs
-
Programming Languages
Categories
Sub Categories
Container Orchestration
90
Code Editors and IDEs
47
Incident Communication
28
Integration
25
Project Management & Issue Tracking Software
21
Build
16
Infrastructure orchestration
15
Deployment
14
IT Service Management
13
Container Registry
12
Bug / Defect Tracking Software
9
Source Code Management
8
Container
8
Keywords
devops
10
monitoring
6
sre
6
observability
6
incident-management
4
prometheus
4
python
4
cli
4
grafana
3
alerting
3
slo
3
golang
3
performance
2
cli-tool
2
open-source
2
alerts
2
cloud-security
2
aiops
2
on-call
2
sli
2
service-level-objective
2
service-level-indicator
2
slack
2
metrics
2
continuous-delivery
2
ci
2
continuous-deployment
2
continuous-integration
2
docker
2
continuous-testing
2
automation
1
chatops
1
sre-workbook
1
log-analysis
1
slo-exporter
1
devops-tools
1
incident
1
production
1
exporter
1
cloud-management
1
github-actions
1
chatbot
1
terraform-github-actions
1
infrastructure-as-code
1
infrastructure-automation
1
infrastructure-orchestration
1
terraform
1
tacos
1
ocaml
1
opentofu
1