Process-Aware LLM-Agent Scaffolds for Metric-Based Microservice Root-Cause Analysis with Evidence-Trace Scoring

Authors

  • Matthew Liao Author

DOI:

https://doi.org/10.61424/zngee941

Keywords:

AIOps; microservices; root-cause analysis; LLM agents; metric telemetry; evidence grounding; process evaluation; fault diagnosis

Abstract

Root-cause analysis in microservice systems increasingly relies on agents that retrieve telemetry, form hypotheses, and explain a diagnosis. Final-answer accuracy alone does not reveal whether the cited evidence supports the conclusion. This study evaluated a process-aware diagnostic scaffold for large-language-model agents on the metric-only RE1-OB corpus from RCAEval. The corpus contains 125 fault-injection cases arranged as five root services, five fault types, and five repetitions. Each case was represented by equal 300-second pre-injection and post-injection windows. To prevent acquisition and schema leakage, the primary predictors were restricted to 35 CPU, memory, and workload streams present in every case; timestamps, file duration, column count, and missingness patterns were excluded. Five leave-one-repetition-out folds compared six conventional classifiers with a tri-head tree ensemble that combined joint service–fault prediction with separate localization and fault-identification heads. The process-aware model achieved 0.960 joint accuracy (95% bootstrap confidence interval [0.920, 0.992]), 0.957 macro-F1, 0.968 service accuracy, 0.992 fault accuracy, and 0.992 top-three accuracy. Injection-aligned evidence retrieval reached 0.952 root-service recall at five. Reranking evidence against the predicted service increased recall to 1.000, reason score from 0.653 to 0.719, and evidence-trace score from 0.871 to 0.893. Five errors remained, dominated by loss-fault service confusion. The findings show that strong diagnosis and auditable evidence are separable objectives: a structured retrieval and verification layer added measurable grounding without changing the underlying top-one predictions.

References

Cai, Y., Nie, X., Yin, K., Pei, C., Sun, Y., Zhang, S., Liu, H., Liu, G., Wen, X., Situ, F., & Pei, D. (2026). A multi-dataset benchmark for evaluating LLM agents in microservice failure diagnosis [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2606.29193

Chen, Y., Shetty, M., Somashekar, G., Ma, M., Simmhan, Y., Mace, J., Bansal, C., Wang, R., & Rajmohan, S. (2025). AIOpsLab: A holistic framework to evaluate AI agents for enabling autonomous clouds. Proceedings of Machine Learning and Systems, 7. https://proceedings.mlsys.org/paper_files/paper/2025/hash/d1f9e4a9f109b6e8b75ed362736f22ec-Abstract-Conference.html

Chen, Y., Xie, H., Ma, M., Kang, Y., Gao, X., Shi, L., Cao, Y., Gao, X., Fan, H., Wen, M., Zeng, J., Ghosh, S., Zhang, X., Zhang, C., Lin, Q., Rajmohan, S., Zhang, D., & Xu, T. (2024). Automatic root cause analysis via large language models for cloud incidents. In Proceedings of the Nineteenth European Conference on Computer Systems (pp. 674–688). Association for Computing Machinery. https://doi.org/10.1145/3627703.3629553

Clark, J., Su, Y., Pial, S. M. R., Tian, Y., Gniedziejko, L., Jacobsen, H.-A., Chen, Y., & Xu, T. (2026). SREGym: A live benchmark for AI SRE agents with high-fidelity failure scenarios [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2605.07161

Fang, A., Yang, Y., Shang, J., Lu, Q., Xu, J., Wang, R., Zhang, S., Zhang, Y., Yu, B., & He, P. (2026). OpenRCA 2.0: From outcome labels to causal process supervision [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2606.27154

Ikram, A., Chakraborty, S., Mitra, S., Saini, S., Bagchi, S., & Kocaoglu, M. (2022). Root cause analysis of failures in microservices through causal discovery. Advances in Neural Information Processing Systems, 35, 31158–31170. https://proceedings.neurips.cc/paper_files/paper/2022/hash/c9fcd02e6445c7dfbad6986abee53d0d-Abstract-Conference.html

Jha, S., Arora, R. R., Watanabe, Y., Yanagawa, T., Chen, Y., Clark, J., Bhavya, B., Verma, M., Kumar, H., Kitahara, H., Zheutlin, N., Takano, S., Pathak, D., George, F., Wu, X., Turkkan, B. O., Vanloo, G., Nidd, M., Dai, T., . . . Puri, R. (2025). ITBench: Evaluating AI agents across diverse real-world IT automation tasks. Proceedings of Machine Learning Research, 267, 27134–27197. https://proceedings.mlr.press/v267/jha25a.html

Jin, J. (2025). Evidence-chain reliable RAG: Hallucination detection, source attribution, and deterministic provenance explanations. Journal of Technology Informatics and Engineering, 4(2), 520–533. https://doi.org/10.51903/jtie.v4i2.535

Lee, C., Yang, T., Chen, Z., Su, Y., & Lyu, M. R. (2023). Eadro: An end-to-end troubleshooting framework for microservices on multi-source data. In 2023 IEEE/ACM 45th International Conference on Software Engineering (pp. 1750–1762). IEEE. https://doi.org/10.1109/ICSE48619.2023.00150

Li, Z., Zhao, N., Zhang, S., Sun, Y., Chen, P., Wen, X., Ma, M., & Pei, D. (2022). Constructing large-scale real-world benchmark datasets for AIOps [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2208.03938

Lin, J., Chen, P., & Zheng, Z. (2018). Microscope: Pinpoint performance issues with causal graphs in micro-service environments. In Service-Oriented Computing (pp. 3–20). Springer. https://doi.org/10.1007/978-3-030-03596-9_1

Liu, Y., Pei, C., Xu, L., Chen, B., Sun, M., Zhang, Z., Sun, Y., Zhang, S., Wang, K., Zhang, H., Li, J., Xie, G., Wen, X., Nie, X., Ma, M., & Pei, D. (2023). OpsEval: A comprehensive IT operations benchmark suite for large language models [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2310.07637

Naakka, A., Wang, Y., & Mäntylä, M. V. (2026). LATS-RCA: Language agent tree search for root cause analysis in microservices [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2605.03505

Pei, C., Wang, Z., Liu, F., Li, Z., Liu, Y., He, X., Kang, R., Zhang, T., Chen, J., Li, J., Xie, G., & Pei, D. (2025). Flow-of-Action: SOP enhanced LLM-based multi-agent system for root cause analysis. In Companion Proceedings of the ACM Web Conference 2025 (pp. 422–431). Association for Computing Machinery. https://doi.org/10.1145/3701716.3715225

Pham, L. (2025). RCAEval: A benchmark for root cause analysis of microservice systems [Data set]. Zenodo. https://doi.org/10.5281/zenodo.14590730

Pham, L., Ha, H., & Zhang, H. (2024). BARO: Robust root cause analysis for microservices via multivariate Bayesian online change point detection. Proceedings of the ACM on Software Engineering, 1(FSE), 2214–2237. https://doi.org/10.1145/3660805

Pham, L., Zhang, H., Ha, H., Salim, F., & Zhang, X. (2025). RCAEval: A benchmark for root cause analysis of microservice systems with telemetry data. In Companion Proceedings of the ACM Web Conference 2025 (pp. 777–780). Association for Computing Machinery. https://doi.org/10.1145/3701716.3715290

Qin, Y., Liang, S., Ye, Y., Zhu, K., Yan, L., Lu, Y., Lin, Y., Cong, X., Tang, X., Qian, B., Zhao, S., Hong, L., Tian, R., Xie, R., Zhou, J., Gerstein, M., Li, D., Liu, Z., & Sun, M. (2024). ToolLLM: Facilitating large language models to master 16000+ real-world APIs. In The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=dHng2O0Jjr

Roy, D., Zhang, X., Bhave, R., Bansal, C., Las-Casas, P., Fonseca, R., & Rajmohan, S. (2024). Exploring LLM-based agents for root cause analysis. In Companion Proceedings of the 32nd ACM International Conference on the Foundations of Software Engineering (pp. 208–219). Association for Computing Machinery. https://doi.org/10.1145/3663529.3663841

Wang, Z., Liu, Z., Zhang, Y., Zhong, A., Wang, J., Yin, F., Fan, L., Wu, L., & Wen, Q. (2024). RCAgent: Cloud root cause analysis by autonomous agents with tool-augmented large language models. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management (pp. 4966–4974). Association for Computing Machinery. https://doi.org/10.1145/3627673.3680016

Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q. V., & Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35, 24824–24837. https://proceedings.neurips.cc/paper/2022/hash/9d5609613524ecf4f15af0f7b31abca4-Abstract-Conference.html

Wu, L., Tordsson, J., Elmroth, E., & Kao, O. (2020). MicroRCA: Root cause localization of performance issues in microservices. In NOMS 2020—2020 IEEE/IFIP Network Operations and Management Symposium (pp. 1–9). IEEE. https://doi.org/10.1109/NOMS47738.2020.9110353

Xin, Q. (2025). Explaining OpenStack failure-injection log anomalies with retrieved normal prototypes. Emerging Information Science and Technology, 6(2), 125–146. https://doi.org/10.18196/eist.v6i2.31232

Xin, Q. (2026). Log anomaly detection with conformal alert control and evidence-grounded incident ticket generation. Aviation Electronics, Information Technology, Telecommunications, Electricals, and Controls (AVITEC), 8(2), 247–264. https://doi.org/10.28989/avitec.v8i2.3974

Xu, J., Zhang, Q., Zhong, Z., He, S., Zhang, C., Lin, Q., Pei, D., He, P., Zhang, D., & Zhang, Q. (2025). OpenRCA: Can large language models locate the root cause of software failures? In The Thirteenth International Conference on Learning Representations. https://openreview.net/forum?id=M4qNIzQYpd

Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct: Synergizing reasoning and acting in language models. In The Eleventh International Conference on Learning Representations. https://openreview.net/forum?id=WE_vluYUL-X

Yu, G., Chen, P., Chen, H., Guan, Z., Huang, Z., Jing, L., Weng, T., Sun, X., & Li, X. (2021). MicroRank: End-to-end latency issue localization with extended spectrum analysis in microservice environments. In Proceedings of the Web Conference 2021 (pp. 3087–3098). Association for Computing Machinery. https://doi.org/10.1145/3442381.3449905

Yu, G., Chen, P., Li, Y., Chen, H., Li, X., & Zheng, Z. (2023). Nezha: Interpretable fine-grained root causes analysis for microservices on multi-modal observability data. In Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (pp. 553–565). Association for Computing Machinery. https://doi.org/10.1145/3611643.3616249

Zhang, B., Sun, X., Liu, G., & Zhou, B. (2026). LLM-style DevOps copilot for cloud-native troubleshooting: Retrieval-augmented runbook generation and command-safety evaluation. Journal of Technology Informatics and Engineering, 5(2), 104–118. https://doi.org/10.51903/jtie.v5i2.534

Zhang, R., Wen, Z., Wang, C., Tang, C., Xu, P., & Jiang, Y. (2025). Quality analysis and evaluation prediction of RAG retrieval based on machine learning algorithms [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2511.19481

Zhang, S., Jin, P., Lin, Z., Sun, Y., Zhang, B., Xia, S., Li, Z., Zhong, Z., Ma, M., Jin, W., Zhang, D., Zhu, Z., & Pei, D. (2023). Robust failure diagnosis of microservice system through multimodal data. IEEE Transactions on Services Computing, 16(6), 3851–3864. https://doi.org/10.1109/TSC.2023.3290018

Zhang, S., Xia, S., Fan, W., Shi, B., Xiong, X., Zhong, Z., Ma, M., Sun, Y., & Pei, D. (2025). Failure diagnosis in microservice systems: A comprehensive survey and analysis. ACM Transactions on Software Engineering and Methodology. https://doi.org/10.1145/3715005

Zhang, W., Guo, H., Yang, J., Tian, Z., Zhang, Y., Chaoran, Y., Li, Z., Li, T., Shi, X., Zheng, L., & Zhang, B. (2024). mABC: Multi-agent blockchain-inspired collaboration for root cause analysis in micro-services architecture. In Findings of the Association for Computational Linguistics: EMNLP 2024 (pp. 4017–4033). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-emnlp.232

Zhang, X., Wang, Q., Li, M., Yuan, Y., Xiao, M., Zhuang, F., & Yu, D. (2025). TAMO: Fine-grained root cause analysis via tool-assisted LLM agent with multi-modality observation data in cloud-native systems [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2504.20462

Zhong, Z. S., Li, C., & Rao, H. (2026). Trajectory reliability prediction for generalist AI agents: Tool-use failure analysis and success forecasting on ZClawBench. Journal of Technology Informatics and Engineering, 5(1), 341–360. https://doi.org/10.51903/jtie.v5i1.539

Downloads

Published

2026-07-19