awesome-sre
A curated list of Site Reliability and Production Engineering resources.
https://github.com/dastergon/awesome-sre
Last synced: 19 days ago
JSON representation
-
Blogs
- Brendan Gregg's Blog - Highly Technical Blog Posts About Systems Internals, Performance and SRE.
- Everything Sysadmin - Blog Posts About SysAdmin/DevOps/SRE by Tom Limoncelli.
- rachelbythebay - Techincal Blog Posts.
- SysAdvent - One article for each day of December, ending on the 25th article.
- Stephen Thorne's Blog - Blog Posts About SRE
- Increment - A digital magazine about how teams build and operate software systems at scale.
- GopherSRE - Blog Posts about Go and SRE.
- Cindy Sridharan - Blog posts about distributed systems and their management.
- Resilience Roundup - Weekly analysis of Resilience Engineering and Human Factors research designed for software systems
- Squadcast Blog - Blog posts about SRE best practices, reliability, on-call and incident management.
- FireHydrant Blog - Posts about complex systems, incident response, and SRE best practices.
- Rootly Blog - Incident management best practices and guides.
- incident.io Blog - Guides, advice and resources on incident management and response.
- Logit.io Blog - Resources on log management, SRE and devOps.
- Susan J. Fowler - Various blog posts about SRE, Software Engineering and Microservices.
- Rootly Blog - Incident management best practices and guides.
- GopherSRE - Blog Posts about Go and SRE.
- Cindy Sridharan - Blog posts about distributed systems and their management.
- Blameless Blog - Blog posts about SRE culture and practices.
- FireHydrant Blog - Posts about complex systems, incident response, and SRE best practices.
- incident.io Blog - Guides, advice and resources on incident management and response.
- Logit.io Blog - Resources on log management, SRE and devOps.
-
Books
- Practical Linux Infrastructure
- Building Secure and Reliable Systems
- Observability Engineering: Achieving Production Excellence
- The Practice Of Cloud System Administration: Designing and Operating Large Distributed Systems
- Web Operations - Keeping the Data On Time
- The Checklist Manifesto: How to Get Things Right
- Microservices in Production - Standard Principles and Requirements
- Production-Ready Microservices - Building Standardized Systems Across an Engineering Organization
- Systems Performance: Enterprise and the Cloud
- Monitoring Distributed Systems: Case Studies from Google's SRE Teams
- Chaos Engineering: Building Confidence in System Behavior through Experiment
- Incident Management for Operations
- Real-World SRE
- Seeking SRE
- What is SRE?
- Engineering Reliable Mobile Applications: Strategies for Developing Resilient Native Mobile Applications
- 97 Things Every SRE Should Know
- Four Steps to Creating Effective Game Day Tests
- The Linux Programming Interface
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- Practical Linux Infrastructure
- The Checklist Manifesto: How to Get Things Right
- Practical Linux Infrastructure
- Systems Performance: Enterprise and the Cloud
- Practical Linux Infrastructure
- Production-Ready Microservices - Building Standardized Systems Across an Engineering Organization
- Web Operations - Keeping the Data On Time
- Incident Management for Operations
- Seeking SRE
- Real-World SRE
-
Capacity Planning
- Capacity Planning
- SouthBay SRE: Cloud Capacity Planning
- Intent-based Capacity Planning and Autoscaling with Kubernetes
- How do you do Capacity Planning
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
- How Back Market SREs prepared for Black Friday
-
Conferences & Meetups
- SRECon Conferences - The Official SRE Conference.
- LISA Conferences - Prominent Conference About SysAdmin/DevOps/SRE.
- South Bay Site Reliability Engineering (Sunnyvale, CA) Meetup - A Group For Individuals Who Tackle Reliability Challenges For Web-Scale Systems.
- San Francisco Reliability Engineering - A Group Of People Who Are Passionate About Reliable, Performant Software Systems.
- Site Reliability Engineering Munich, Germany - SRE Meetup in the greater area of Oktoberfest city.
- ADDO - All Day DevOps - A 24 hour conference that is completely online and free.
- Site Reliability Engineering Paris, France - SRE Meetup in the city of light.
- Site Reliability Engineering India - SRE Meetup India
- SRE Tech Talks - SRE Talks Hosted by Google.
- South Bay Site Reliability Engineering (Sunnyvale, CA) Meetup - A Group For Individuals Who Tackle Reliability Challenges For Web-Scale Systems.
- San Francisco Reliability Engineering - A Group Of People Who Are Passionate About Reliable, Performant Software Systems.
- Site Reliability Engineering Munich, Germany - SRE Meetup in the greater area of Oktoberfest city.
- Site Reliability Engineering Paris, France - SRE Meetup in the city of light.
-
Culture
- What is Site Reliability Engineering?
- Keys To SRE by Ben Treynor
- Google SRE Resources
- Notes from Production Engineering by Pedro Canahuati
- PostOps: Recovery from Operations
- Love DevOps? Wait 'till you meet SRE - k)
- How Google Does Planet-Scale Engineering for Planet-Scale Infra
- Site Reliability Engineering at Facebook
- A History of Site Reliability Engineering at Uber
- Case Study: Adopting SRE Principles at StackOverflow
- Site Reliability Engineering at Dropbox
- Site Reliability Engineers — Keeping Google up and running 24/7
- video
- SRE@Google: Thousands of DevOps Since 2004
- Transactional System Administration Is Killing Us and Must be Stopped
- A hierarchy of SRE needs
- PostOps: A Non-Surgical Tale of Software, Fragility, and Reliability
- SRE: An incomplete guide to cultural Narnia - [[Video]](https://www.youtube.com/watch?v=__wypEhdcrQ&t=0s)
- Putting Together Great SRE Teams
- Work at Google: Meet our Production Engineers for Site Reliability Hangout on Air
- Toil: A Word Every Engineer Should Know
- Engineering Reliability into Web Sites: Google SRE
- DEVOPS & SRE AMA - Building High Performance Organizations
- How SysAdmins Devalue Themselves
- The Softer Side of DevOps
- SRE, noun. See also: confidence, trust.
- Site Reliability Engineering with Stephen Weinberg
- We are the Google Site Reliability team. We make Google’s websites work. Ask us Anything!
- We are the Google Site Reliability Engineering team. Ask us Anything!
- The Ops Identity Crisis
- The Irreproducibility Of Bugs In Large-Scale Production Systems
- SE-Radio Episode 276: Björn Rabenstein on Site Reliability Engineering
- Microservices, DevOps and Production Complexity
- Introducing Google Customer Reliability Engineering
- Evolution or Rebellion? The rise of Site Reliability Engineers (SRE)
- The difference between Site Reliability Engineering, System Administration, and DevOps
- SRE in the Small and in the Large
- SBSRE Meetup: Different SRE roles and challenges(Netflix)
- Panel: Who/What Is SRE?
- Hope Is Not a Strategy
- Tenets of SRE
- Site Reliability Engineering Demystified
- Is Site Reliability Engineering the True ‘Ops’ in DevOps?
- SRE vs. DevOps vs. Cloud Native: The Server Cage Match
- SRE: What’s The Big Idea?
- Building the SRE Culture at LinkedIn
- Podcast #111 – SRE: Occasionally Maintaining Infrastructure That You Hate
- Splicing SRE DNA Sequences in the Biggest Software Company on the Planet
- Why should your app get SRE support? - CRE life lessons
Categories
On-Call
175
Culture
163
Reliability
87
Monitoring & Observability & Alerting
69
Books
69
Post-Mortem
68
Service Level Agreement
56
Capacity Planning
47
Misc Articles
31
Education
25
Blogs
22
Twitter
16
Conferences & Meetups
13
Newsletters
9
Podcasts
7
Hiring
7
Performance
5
Programming
5
SRE Tools
3
Real-time Messaging
3
Sub Categories
Keywords
post-mortem
4
site-reliability-engineering
3
awesome
3
postmortem
3
sre
3
devops
3
reliability
2
awesome-list
2
failures
1
simian-army
1
resilience
1
netflix-chaos-monkey
1
chaos-testing
1
chaos-monkey
1
chaos-engineering
1
chaos-community
1
chaos
1
development
1
developer-tools
1
continuous-integration
1
comparison
1
ci
1
site-reliability
1
postmortem-templates
1
incident-response
1
incident-reports
1
incident-reporting
1
service-level-objective
1
service-level-monitoring
1
service-level-agreement
1
reliability-engineering
1
production
1
monitoring-tools
1
monitoring
1
list
1
incident-responce
1
incident-management
1
devops-tools
1
availability
1
production-engineering
1
kubernetes
1
incidents
1
debugging
1