{"id":13684914,"url":"https://github.com/OpsPAI/awesome-AIOps","last_synced_at":"2025-05-01T00:33:43.974Z","repository":{"id":40561676,"uuid":"378397750","full_name":"OpsPAI/awesome-AIOps","owner":"OpsPAI","description":"A curated list of awesome academic researches and industrial materials about Artificial Intelligence for IT Operations (AIOps).","archived":false,"fork":false,"pushed_at":"2025-02-12T03:18:05.000Z","size":165,"stargazers_count":252,"open_issues_count":3,"forks_count":36,"subscribers_count":9,"default_branch":"main","last_synced_at":"2025-04-22T18:06:01.487Z","etag":null,"topics":[],"latest_commit_sha":null,"homepage":null,"language":null,"has_issues":true,"has_wiki":null,"has_pages":null,"mirror_url":null,"source_name":null,"license":"mit","status":null,"scm":"git","pull_requests_enabled":true,"icon_url":"https://github.com/OpsPAI.png","metadata":{"files":{"readme":"README.md","changelog":null,"contributing":null,"funding":null,"license":"LICENSE","code_of_conduct":null,"threat_model":null,"audit":null,"citation":null,"codeowners":null,"security":null,"support":null,"governance":null,"roadmap":null,"authors":null,"dei":null,"publiccode":null,"codemeta":null}},"created_at":"2021-06-19T11:39:38.000Z","updated_at":"2025-04-20T02:34:39.000Z","dependencies_parsed_at":"2024-06-01T05:12:37.986Z","dependency_job_id":null,"html_url":"https://github.com/OpsPAI/awesome-AIOps","commit_stats":null,"previous_names":[],"tags_count":0,"template":false,"template_full_name":null,"repository_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/OpsPAI%2Fawesome-AIOps","tags_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/OpsPAI%2Fawesome-AIOps/tags","releases_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/OpsPAI%2Fawesome-AIOps/releases","manifests_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories/OpsPAI%2Fawesome-AIOps/manifests","owner_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners/OpsPAI","download_url":"https://codeload.github.com/OpsPAI/awesome-AIOps/tar.gz/refs/heads/main","host":{"name":"GitHub","url":"https://github.com","kind":"github","repositories_count":251425692,"owners_count":21587435,"icon_url":"https://github.com/github.png","version":null,"created_at":"2022-05-30T11:31:42.601Z","updated_at":"2022-07-04T15:15:14.044Z","host_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub","repositories_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repositories","repository_names_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/repository_names","owners_url":"https://repos.ecosyste.ms/api/v1/hosts/GitHub/owners"}},"keywords":[],"created_at":"2024-08-02T14:00:40.565Z","updated_at":"2025-05-01T00:33:43.967Z","avatar_url":"https://github.com/OpsPAI.png","language":null,"funding_links":[],"categories":["Other awesome AI lists","Other Lists","Educational Resources","Community Lists"],"sub_categories":["TeX Lists","Visualization \u0026 Dashboards"],"readme":"# awesome-AIOps\nA curated list of awesome academic researches and industrial materials about Artificial Intelligence for IT Operations (AIOps).\n\n- [Researchers](#researchers)\n- [Industrial Materials](#industrial-materials)\n  - [Competitions](#competitions)\n  - [White Papers](#white-papers)\n  - [Blogs \u0026 Tutorials \u0026 Magazines](#blogs--tutorials--magazines)\n  - [Benchmarks](#benchmarks)\n  - [Tools](#tools)\n  - [Companies](#companies)\n- [Academic Materials](#academic-materials)\n  - [Talks](#talks)\n  - [Workshops](#workshops)\n- [Papers](#papers)\n  - [Survey \u0026 Empirical Study](#survey--empirical-study)\n  - [Benchmarks](#benchmarks)\n  - [(Large) Language Models for IT Operations](#large-language-models-for-it-operations)\n  - [Knowledge Graph for AIOps](#knowledge-graph-for-aiops)\n  - [Microservices and Serverless](#microservices-and-serverless)\n  - [Dependency and Tracing](#dependency-and-tracing)\n  - [Anomaly/Failure Detection](#anomalyfailure-detection)\n  - [Root Cause Analysis](#root-cause-analysis)\n  - [Incident and Alarm Management](#incident-and-alarm-management)\n  - [Node, Disk, and Storage](#node-disk-and-storage)\n  - [VM Analysis and Management](#vm-analysis-and-management)\n  - [Deployment](#deployment)\n- [Datasets](#datasets)\n- [Others](#others)\n  - [Courses](#courses)\n\n## Researchers\n| China (\u0026 HK SAR) | |||\n| :---------| :------ | :------ | :------ |\n| [Michael R. Lyu](http://www.cse.cuhk.edu.hk/lyu/), CUHK | [Dongmei Zhang](https://www.microsoft.com/en-us/research/people/dongmeiz/), Microsoft | [Pengfei Chen](http://sdcs.sysu.edu.cn/content/3747), SYSU | [Dan Pei](https://netman.aiops.org/~peidan/), Tsinghua |\n| [Xin Peng](https://cspengxin.github.io/), Fudan ||||\n| **USA** ||||\n| [Ryan Huang](https://www.cs.jhu.edu/~huang/), JHU | [Yingnong Dang](https://scholar.google.com.hk/citations?user=InqtwxcAAAAJ\u0026hl=en), Microsoft | [Christina Delimitrou](https://www.csl.cornell.edu/~delimitrou/), MIT EECS ||\n| **Europe** |||||\n| [Odej Kao](https://www.cit.tu-berlin.de/kao/), TU Berlin ||||\n| **Australia** ||||\n| [Hongyu Zhang](http://hongyujohn.github.io/), UON ||||\n\n\n## Industrial Materials\n### Competitions\n- [AIOps Challenge] [A series of AIOps competitions hosted by Tsinghua University](https://competition.aiops-challenge.com/home/competition)\n- [PAKDD2020] [Alibaba AIOps Competition](https://tianchi.aliyun.com/competition/entrance/231775/introduction?lang=en-us)\n\n### White Papers\n- [VMware] [Proactive Incident and Problem Management](https://docplayer.net/8854482-Proactive-incident-and-problem-management.html)\n- [GREATOPS 高效运维社区] [《企业级 AIOps 实施建议》白皮书](https://pic.huodongjia.com/ganhuodocs/2018-04-16/1523873064.74.pdf)\n- [Awesome Open Source] [Aiops Handbook](https://awesomeopensource.com/project/chenryn/aiops-handbook)\n\n### Blogs \u0026 Tutorials \u0026 Magazines\n- [Moogsoft] [What is AIOps?](https://www.moogsoft.com/resources/aiops/guide/everything-aiops/)\n- [Tsinghua University] [清华裴丹：AIOps落地的15条原则](https://mp.weixin.qq.com/s/Ov1gQlQ0mRpk58cNL_YlVg)\n- [Tsinghua University] [清华裴丹：AIOps效果落地最后一公里](https://mp.weixin.qq.com/s/VhaRfvjc839bAXBMfzv11g)\n- [Alibaba Cloud] [基于大数据的智能网络分析-齐天](https://developer.aliyun.com/article/590290)\n- [Microsoft] [Advancing Azure service quality with artificial intelligence: AIOps](https://azure.microsoft.com/en-us/blog/advancing-azure-service-quality-with-artificial-intelligence-aiops/)\n- [Grafana] [GrafanaCON: Grafana Observability Conference 2022](https://grafana.com/about/events/observabilitycon/2022/)\n- [InfoQ] [2023，可观测性需求将迎来“爆发之年”？](https://mp.weixin.qq.com/s/6na952N3c5RzcopanZGs6w)\n- [Alibaba] [阿里云张建锋谈新型计算体系：云正在重构硬件、软件和终端世界](https://mp.weixin.qq.com/s/IQvurZ_9Vm0SufV1K0sK1A)\n\n### Benchmarks\n- [Cornell] [DeathStarBench (An open-source benchmark suite for cloud microservices)](https://github.com/delimitrou/DeathStarBench/tree/master)\n- [Google Cloud] [Online Boutique (A microservices demo application)](https://github.com/GoogleCloudPlatform/microservices-demo)\n- [Fudan] [Train Ticket (A benchmark microservice system)](https://github.com/FudanSELab/train-ticket)\n- [Weaveworks] [Sock Shop (A microservices demo application)](https://microservices-demo.github.io/)\n\n### Tools\n- [Log Analytics] [LogPAI](https://github.com/logpai)\n- [AI for Cloud Operation] [OpsPAI](https://github.com/OpsPAI)\n- [Outlier Detection] [PyOD](https://github.com/yzhao062/pyod)\n- [Anomaly Detection] [ADTK](https://github.com/arundo/adtk)\n- [Anomaly Detection] [PySAD](https://github.com/selimfirat/pysad)\n- [Online Machine Learning] [River](https://riverml.xyz/)\n- [Online Machine Learning] [scikit-multiflow](https://scikit-multiflow.readthedocs.io/)\n- [Fault Injection] [Chaos Mesh](https://github.com/chaos-mesh/chaos-mesh)\n- [Fault Injection] [ChaosBlade](https://github.com/chaosblade-io/chaosblade)\n- [Container Monitoring] [cAdvisor](https://github.com/google/cadvisor)\n- [Performance Monitoring] [Netdata](https://www.netdata.cloud/)\n- [Anomaly Detection Labeling Tool] [Microsoft TagAnomaly](https://github.com/Microsoft/TagAnomaly)\n- [Serverless App Dev. Framework] [AWS Serverless Application Model (AWS SAM)](https://github.com/aws/serverless-application-model)\n- [Performance Testing Tool] [Locust](https://locust.io/)\n- [Alibaba Java Diagnostic Tool] [Arthas](https://arthas.aliyun.com/)\n\n### Companies\n- [Datadog](https://www.datadoghq.com/): A monitoring and security platform for cloud applications\n- [必示 bizseer](https://www.bizseer.com/)\n- [日志易](https://www.rizhiyi.com/)\n- [博睿数据](https://www.bonree.com/)\n- [听云 TINGYUN](https://www.tingyun.com/lp.html): 端到端的全平台应用性能管理系统\n- [Loom Systems](https://www.loomsystems.com/)\n- [Keep](https://www.keephq.dev): Open-source alert management and AIOps platform\n\n\n## Academic Materials\n\n### Talks\n- [Michael R. Lyu] [Reliability-Driven AIOps for Cloud Resilience (Keynote talk at ICSE '21)](http://ariselab.cse.cuhk.edu.hk/assets/files/ICSE2021_keynote_lyu.pdf)\n\n### Workshops\n- [ICSE21 Workshop on Cloud Intelligence](http://cloudintelligenceworkshop.org/index.html)\n- [AAAI-20 Workshop on Cloud Intelligence](http://cloudintelligenceworkshop.org/2020/index.html)\n- [AIOPS 2020 (International Workshop on Artificial Intelligence for IT Operations)](https://aiopsworkshop.github.io/)\n\n\n## Papers\n\n### Survey \u0026 Empirical Study\n- [arXiv '24] [A Survey on Failure Analysis and Fault Injection in AI Systems](https://arxiv.org/abs/2407.00125)\n- [arXiv '23] [AI for IT Operations (AIOps) on Cloud Platforms: Reviews, Opportunities and Challenges](https://arxiv.org/abs/2304.04661)\n- [CSUR '22] [Anomaly Detection and Failure Root Cause Analysis in (Micro) Service-Based Cloud Applications: A Survey](https://dl.acm.org/doi/full/10.1145/3501297)\n- [ASE '22] [Going through the Life Cycle of Faults in Clouds: Guidelines on Fault Handling](https://github.com/IntelligentDDS/Post-mortems-Analysis)\n- [arXiv '21] [Experience Report: Deep Learning-based System Log Analysis for Anomaly Detection](https://arxiv.org/abs/2107.05908)\n- [CSUR '21] [A Survey on Automated Log Analysis for Reliability Engineering](https://arxiv.org/abs/2009.07237)\n- [ESEC/FSE '20] [Towards intelligent incident management: why we need it and how we make it](https://dl.acm.org/doi/abs/10.1145/3368089.3417055)\n- [arXiv '20] [A Systematic Mapping Study in AIOps](https://arxiv.org/abs/2012.09108)\n- [ICSE '19] [AIOps: Real-World Challenges and Research Innovations](https://ieeexplore.ieee.org/document/8802836)\n- [HotOS '19] [What bugs cause production cloud incidents?](https://dl.acm.org/doi/10.1145/3317550.3321438)\n- [ISSRE '16] [Experience Report: System Log Analysis for Anomaly Detection](https://ieeexplore.ieee.org/abstract/document/7774521)\n- [ASE '13] [Software analytics for incident management of online services: An experience report](https://ieeexplore.ieee.org/document/6693105)\n\n### Benchmarks\n- [arXiv '22] [Constructing Large-Scale Real-World Benchmark Datasets for AIOps](https://arxiv.org/abs/2208.03938)\n- [ASPLOS '19] [An Open-Source Benchmark Suite for Microservices and Their Hardware-Software Implications for Cloud and Edge Systems](https://dl.acm.org/doi/10.1145/3297858.3304013)\n\n\n### (Large) Language Models for IT Operations\n- [ISSTA '24] [LILAC: Log Parsing using LLMs with Adaptive Parsing Cache](https://arxiv.org/abs/2310.01796)\n- [arXiv '24] [Exploring LLM-based Agents for Root Cause Analysis](https://arxiv.org/abs/2403.04123)\n- [arXiv '24] [Nissist: An Incident Mitigation Copilot based on Troubleshooting Guides](https://arxiv.org/abs/2402.17531)\n- [arXiv '24] [Automated Root Causing of Cloud Incidents using In-Context Learning with GPT-4](https://arxiv.org/abs/2401.13810)\n- [arXiv '23] [Automatic Root Cause Analysis via Large Language Models for Cloud Incidents](https://arxiv.org/abs/2305.15778)\n- [arXiv '23] [OpsEval: A Comprehensive Task-Oriented AIOps Benchmark for Large Language Models](https://arxiv.org/abs/2310.07637)\n- [arXiv '23] [Xpert: Empowering Incident Management with Query Recommendations via Large Language Models](https://arxiv.org/abs/2312.11988)\n- [arXiv '23] [Exploring the Effectiveness of LLMs in Automated Logging Generation: An Empirical Study](https://arxiv.org/abs/2307.05950)\n- [arXiv '23] [Assess and Summarize: Improve Outage Understanding with Large Language Models](https://arxiv.org/pdf/2305.18084.pdf)\n- [arXiv '23] [Empower Large Language Model to Perform Better on Industrial Domain-Specific Question Answering](https://arxiv.org/pdf/2305.11541.pdf)\n- [arXiv '23] [Recommending Root-Cause and Mitigation Steps for Cloud Incidents using Large Language Models](https://arxiv.org/abs/2301.03797)\n- [SoCC '19] [A System-Wide Debugging Assistant Powered by Natural Language Processing](https://dl.acm.org/doi/10.1145/3357223.3362701)\n\n\n### Knowledge Graph for AIOps\n- [ICSE-SEIP '22] [Mining Root Cause Knowledge from Cloud Service Incident Investigations for AIOps](https://arxiv.org/abs/2204.11598)\n- [ICSE-SEIP '21] [Neural knowledge extraction from cloud service incidents](https://dl.acm.org/doi/abs/10.1109/ICSE-SEIP52600.2021.00031)\n- [arXiv '21] [SoftNER: Mining Knowledge Graphs From Cloud Incidents](https://arxiv.org/abs/2101.05961)\n- [APPLSCI '20] [A Causality Mining and Knowledge Graph Based Method of Root Cause Diagnosis for Performance Anomaly in Cloud Applications](https://www.mdpi.com/2076-3417/10/6/2166)\n\n\n### Microservices and Serverless\n- [ASPLOS '21] [Sage: Practical \u0026 Scalable ML-Driven Performance Debugging in Microservices](https://dl.acm.org/doi/abs/10.1145/3445814.3446700)\n- [ICDCS '21] [Defuse: A Dependency-Guided Function Scheduler to Mitigate Cold Starts on FaaS Platforms](https://ieeexplore.ieee.org/document/9546470)\n- [FSE '20] [Graph-based trace analysis for microservice architecture understanding and problem diagnosis](https://dl.acm.org/doi/10.1145/3368089.3417066)\n- [OSDI '20] [FIRM: An Intelligent Fine-grained Resource Management Framework for SLO-Oriented Microservices](https://www.usenix.org/conference/osdi20/presentation/qiu)\n- [ESEC/FSE '19] [Latent Error Prediction and Fault Localization for Microservice Applications by Learning from System Trace Logs](https://dl.acm.org/doi/10.1145/3338906.3338961)\n- [TSE '18] [Fault Analysis and Debugging of Microservice Systems: Industrial Survey, Benchmark System, and Empirical Study](https://ieeexplore.ieee.org/document/8580420/)\n\n\n### Dependency and Tracing\n- [ASE '21] [AID: Efficient Prediction of Aggregated Intensity of Dependency in Large-scale Cloud Systems](https://arxiv.org/abs/2109.04893) [[code](https://github.com/OpsPAI/aid)]\n- [NSDI '07] [X-Trace: A Pervasive Network Tracing Framework](https://www.usenix.org/conference/nsdi-07/x-trace-pervasive-network-tracing-framework)\n- [HotNets '06] [Discovering Dependencies for Network Management](https://www.microsoft.com/en-us/research/wp-content/uploads/2006/11/hotnets06.pdf)\n\n\n### Anomaly/Failure Detection\n- [POMACS '24] [The Tale of Errors in Microservices](https://dl.acm.org/doi/pdf/10.1145/3700436) [[data](https://zenodo.org/records/13947828)]\n- [ICSE '23] [CONAN: Diagnosing Batch Failures for Cloud Systems](http://windows-microsoft-en.com/research/uploads/prod/2022/12/Conan_ICSE23_CR.pdf)\n- [ISSRE '22] [Share or Not Share? Towards the Practicability of Deep Models for Unsupervised Anomaly Detection in Modern Online Systems](https://ieeexplore.ieee.org/document/9978953) [[code](https://github.com/IntelligentDDS/Uni-AD)]\n- [ICSE '22] [Adaptive Performance Anomaly Detection for Online Service Systems via Pattern Sketching](https://arxiv.org/abs/2201.02944) [[code](https://github.com/OpsPAI/ADSketch)]\n- [KDD '19] [Time-Series Anomaly Detection Service at Microsoft](https://dl.acm.org/doi/10.1145/3292500.3330680)\n- [ESEC/FSE '18] [Identifying Impactful Service System Problems via Log Analysis](https://dl.acm.org/doi/10.1145/3236024.3236083)\n- [CCS '17] [DeepLog: Anomaly Detection and Diagnosis from System Logs through Deep Learning](https://dl.acm.org/doi/10.1145/3133956.3134015)\n\n\n### Root Cause Analysis\n- [SIGCOMM '23] [Murphy: Performance Diagnosis of Distributed Cloud Applications](https://dl.acm.org/doi/abs/10.1145/3603269.3604877)\n- [FSE '23] [Nezha: Interpretable Fine-Grained Root Causes Analysis for Microservices on Multi-modal Observability Data](https://dl.acm.org/doi/10.1145/3611643.3616249)\n- [OSDI '18] [Capturing and Enhancing In Situ System Observability for Failure Detection](https://www.usenix.org/conference/osdi18/presentation/huang)\n\n### Incident and Alarm Management\n- [ATC '23] [AutoARTS: Taxonomy, Insights and Tools for Root Cause Labelling of Incidents in Microsoft Azure](https://www.usenix.org/conference/atc23/presentation/dogga)\n- [ICSE '23] [Incident-aware Duplicate Ticket Aggregation for Cloud Systems](https://arxiv.org/abs/2302.09520)\n- [SoCC '22] [How to Fight Production Incidents? An Empirical Study on a Large-scale Cloud Service](https://dl.acm.org/doi/10.1145/3542929.3563482)\n- [DSN '22] [Characterizing and Mitigating Anti-patterns of Alerts in Industrial Cloud Systems](https://arxiv.org/abs/2204.09670)\n- [USENIX ATC '21] [Fighting the Fog of War: Automated Incident Detection for Cloud Systems](https://www.usenix.org/conference/atc21/presentation/li-liqun)\n- [ASE '21] [Graph-based Incident Aggregation for Large-Scale Online Service Systems](https://arxiv.org/abs/2108.12179)\n- [ASE '21] [Groot: An Event-graph-based Approach for Root Cause Analysis in Industrial Settings](https://arxiv.org/abs/2108.00344)\n- [SIGCOMM '20] [Scouts: Improving the Diagnosis Process Through Domain-customized Incident Routing](https://dl.acm.org/doi/10.1145/3387514.3405867)\n- [ASE '20] [How Incidental are the Incidents?: Characterizing and Prioritizing Incidents for Large-Scale Online Service Systems](https://dl.acm.org/doi/10.1145/3324884.3416624)\n- [ESEC/FSE '20] [Identifying linked incidents in large-scale online service systems](https://dl.acm.org/doi/10.1145/3368089.3409768)\n- [ESEC/FSE '20] [Efficient incident identification from multi-dimensional issue reports via meta-heuristic search](https://dl.acm.org/doi/abs/10.1145/3368089.3409741)\n- [ESEC/FSE '20] [Real-time incident prediction for online service systems](https://dl.acm.org/doi/abs/10.1145/3368089.3409672)\n- [ESEC/FSE '20] [How to mitigate the incident? an effective troubleshooting guide recommendation technique for online service systems](https://dl.acm.org/doi/abs/10.1145/3368089.3417054)\n- [ICSE '20] [Understanding and Handling Alert Storm for Online Service Systems](https://dl.acm.org/doi/10.1145/3377813.3381363)\n- [HotOS '19] [What bugs cause production cloud incidents?](https://dl.acm.org/doi/10.1145/3317550.3321438)\n- [ASE '19] [Continuous Incident Triage for Large-Scale Online Service Systems](https://dl.acm.org/doi/10.1109/ASE.2019.00042)\n- [ICSE '19] [An empirical investigation of incident triage for online service systems](https://dl.acm.org/doi/10.1109/ICSE-SEIP.2019.00020)\n- [WWW '19] [Outage Prediction and Diagnosis for Cloud Service Systems](https://dl.acm.org/doi/10.1145/3308558.3313501)\n- [KDD '14] [Correlating Events with Time Series for Incident Diagnosis](https://dl.acm.org/doi/10.1145/2623330.2623374)\n\n\n### Node, Disk, and Storage\n- [FAST '23] [Perseus: A Fail-Slow Detection Framework for Cloud Storage Systems](https://www.usenix.org/conference/fast23/presentation/lu) [[data](https://tianchi.aliyun.com/dataset/144479)]\n- [DSN '21] [General Feature Selection for Failure Prediction in Large-scale SSD Deployment](https://ieeexplore.ieee.org/document/9505157)\n- [TOSEM '20] [Predicting Node Failures in an Ultra-Large-Scale Cloud Computing Platform: An AIOps Solution](https://dl.acm.org/doi/10.1145/3385187)\n- [ICDCS '20] [Toward Adaptive Disk Failure Prediction via Stream Mining](https://ieeexplore.ieee.org/document/9355640)\n- [VLDB '20] [Diagnosing root causes of intermittent slow queries in cloud databases](https://dl.acm.org/doi/abs/10.14778/3389133.3389136)\n- [USENIX ATC '19] [IASO: A Fail-Slow Detection and Mitigation Framework for Distributed Storage Services](https://www.usenix.org/conference/atc19/presentation/panda)\n- [NSDI '18] [Deepview: Virtual Disk Failure Diagnosis and Pattern Detection for Azure](https://www.usenix.org/conference/nsdi18/presentation/zhang-qiao)\n- [ESEC/FSE '18] [Predicting Node Failure in Cloud Service Systems](https://dl.acm.org/doi/10.1145/3236024.3236060)\n- [USENIX ATC '18] [Improving Service Availability of Cloud Systems by Predicting Disk Error](https://www.usenix.org/conference/atc18/presentation/xu-yong)\n\n\n### VM Analysis and Management\n- [NSDI '22] [CloudCluster: Unearthing the Functional Structure of a Cloud Service](https://www.usenix.org/conference/nsdi22/presentation/pang)\n- [OSDI '20] [Predictive and Adaptive Failure Mitigation to Avert Production Cloud VM Interruptions](https://www.usenix.org/conference/osdi20/presentation/levy)\n\n\n### Deployment\n- [SOSP '21] [Understanding and Detecting Software Upgrade Failures in Distributed Systems](https://dl.acm.org/doi/10.1145/3477132.3483577)\n- [NSDI '20] [Gandalf: An Intelligent, End-To-End Analytics Service for Safe Deployment in Large-Scale Cloud Infrastructure](https://www.usenix.org/conference/nsdi20/presentation/li)\n\n\n## Datasets\n- [CUHK] [Loghub](https://github.com/logpai/loghub)\n- [Microsoft Azure] [Azure Public Dataset](https://github.com/Azure/AzurePublicDataset)\n- [Tsinghua] [AIOps Challenge Dataset](http://iops.ai/dataset_list/)\n- [Google] [Cluster Traces](https://github.com/google/cluster-data)\n- [Backblaze] [Hard Drive Dataset](https://www.backblaze.com/b2/hard-drive-test-data.html)\n- [Baidu] [SMART Dataset of PAKDD CUP 2020](https://pan.baidu.com/share/link?shareid=189977\u0026uk=4278294944#list/path=%2FS.M.A.R.T.dataset)\n- [Red Hat] [Ceph Device Telemetry Dataset](https://ceph.io/en/users/telemetry/device-telemetry/)\n- [Alibaba] [SSD SMART logs and failure data](https://github.com/alibaba-edu/dcbrain/tree/master/ssd_open_data)\n- [Alibaba] [Alibaba Cluster Trace Program](https://github.com/alibaba/clusterdata)\n- [CloudWise] [GAIA Dataset](https://github.com/CloudWise-OpenSource/GAIA-DataSet)\n- [Huawei Cloud] [Serverless traces](https://github.com/sir-lab/data-release?tab=readme-ov-file)\n\n\n## Others\n\n### Courses\n- [Coursera] [Cloud-Based Network Design \u0026 Management Techniques](https://www.coursera.org/learn/cloud-based-network-design-and-management)\n- [Tsinghua] [AIOps Course of Tsinghua](https://netman.aiops.org/courses/advanced-network-management-spring2021-course/)\n","project_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FOpsPAI%2Fawesome-AIOps","html_url":"https://awesome.ecosyste.ms/projects/github.com%2FOpsPAI%2Fawesome-AIOps","lists_url":"https://awesome.ecosyste.ms/api/v1/projects/github.com%2FOpsPAI%2Fawesome-AIOps/lists"}