Abstract: Artificial Intelligence for IT Operations (AIOps) has emerged as a critical approach for managing complex IT systems, yet its ad⁃ vancement has been hindered by the lack of comprehensive public datasets for evaluation. This study addresses this challenge by introduc⁃ ing three large-scale, real-world benchmark datasets specifically designed for AIOps research and development. We present datasets cov⁃ ering key AIOps scenarios: key performance indicator (KPI) anomaly detection with diverse patterns from multiple Internet companies, multi-dimensional root cause localization with 400 labeled cases, and failure discovery and diagnosis containing 169 injected failures across various system components. These datasets are distinguished by their scale, authenticity, and detailed ground-truth labels, enabling rigorous evaluation of AIOps algorithms. The practical value of our datasets has been validated through three annual AIOps algorithm com ⁃ petitions (2018‒2020), which attracted over 400 participating teams and led to numerous research publications. By providing these com⁃ prehensive benchmarks, our work establishes a foundation for reproducible AIOps research and facilitates fair comparison of different ap ⁃ proaches, serving a role for AIOps analogous to that of ImageNet in computer vision.
Keywords: AIOps; benchmark datasets; KPI anomaly detection; root cause localization; failure discovery and diagnosis