Risk-Calibrated Multi-Resource Scheduling for Disaggregated AI Inference with LLM-Generated Evidence Memos
DOI:
https://doi.org/10.61424/6zd6va79Keywords:
Disaggregated AI inference; multi-resource scheduling; GPU cluster trace; conformal prediction; histogram gradient boosting; integer programming; service-level objective; resource fragmentation; evidence-grounded language models; AIOps.Abstract
Disaggregated inference separates compute-intensive and accelerator-intensive stages, but it also turns capacity control into a coupled, multi-resource decision under time-varying demand. This study evaluates a risk-calibrated scheduling pipeline on the 2025 Alibaba trace of disaggregated deep-learning recommendation model services. The trace contains 23,871 instances from 156 long-running services over 31 days, including 16,485 compute-node instances and 7,386 heterogeneous-node instances. CPU, GPU, RDMA, memory, disk, deployment-density, and lifecycle fields were aggregated into 15-minute decision epochs. A chronological 60/20/20 design produced 1,689 training windows, 595 calibration windows, and 596 test observations; 595 held-out scheduling decisions were evaluated after lag initialization. The proposed RC-HGB-ILP policy combines a persistence-guarded histogram gradient-boosting forecast, a strictly online one-sided conformal demand envelope, and an integer node-mix planner. It was compared with reactive best-fit, a training-peak inventory, an exact one-step oracle, an uncalibrated HGB planner, and a density-constraint ablation. At 90% nominal coverage, mean empirical envelope coverage reached 94.03%. RC-HGB-ILP reduced the held-out scheduling-SLO violation rate from 18.98% for HGB-ILP to 5.08%, while increasing normalized operating cost by 979.55 cost units, or 0.098%. Relative to reactive best-fit, the violation rate fell from 9.32% to 5.08%. GPU utilization remained 90.25%, and the fragmentation index decreased to 0.31751. Paired epoch tests confirmed lower violation risk and fragmentation, with a small cost increase. An evidence-constrained memo protocol converted saved metrics into seven operational summaries; all 34 numeric claims matched their evidence records, all claims carried evidence identifiers, and all comparison directions were correct. The results show that calibrated upper-demand envelopes can absorb forecast drift at modest capacity cost while preserving high accelerator use, and that bounded evidence memos can expose the resulting trade-offs without numerical drift.
References
Alibaba Group. (2025). Alibaba cluster trace program: Traces for GPU-disaggregated deep learning recommendation models [Data set]. GitHub. https://github.com/alibaba/clusterdata/tree/master/cluster-trace-gpu-v2025
Angelopoulos, A. N., & Bates, S. (2023). Conformal prediction: A gentle introduction. Foundations and Trends in Machine Learning, 16(4), 494–591. https://doi.org/10.1561/2200000101
Bai, J., & Wu, Q. (2026). Privacy-safe marketing mix modeling and budget optimization under identifier loss: A controlled simulation study. International Journal of Electronic Communication Systems, 6(1). https://doi.org/10.24042/ijecs.v6i1.30533
Bai, J., Chen, S., Zheng, D., & Kuo, M.-J. (2026). Interpretable attack-chain stage detection from AWS CloudTrail event sequences via linear models and HMM smoothing. Information, Electrical and Electronics Engineering, 6(1), 28–43. https://doi.org/10.33474/infotron.v6i1.24923
Bai, J., Wang, H., Wu, Q., & Zhang, B. (2026). Privacy-robust incrementality estimation in cookieless settings via uplift modeling: Reproducible evidence from the Hillstrom E-Mail Experiment. Journal of Technology Informatics and Engineering, 5(1), 17–38. https://doi.org/10.51903/jtie.v5i1.468
Barber, R. F., Candès, E. J., Ramdas, A., & Tibshirani, R. J. (2023). Conformal prediction beyond exchangeability. The Annals of Statistics, 51(2), 816–845. https://doi.org/10.1214/23-AOS2276
Chang, X., Lu, Y., & Zhong, Z. S. (2026). Review-grounded explainable recommendation with faithfulness evaluation on Amazon Reviews. Journal of Electrical Engineering and Computer Science, 11(1), 9–22. https://doi.org/10.54732/jeecs.v11i1.2
Chekuri, C., & Khanna, S. (2004). On multidimensional packing problems. SIAM Journal on Computing, 33(4), 837–851. https://doi.org/10.1137/S0097539799356265
Chen, S., He, S., & Sun, E. (2024). Risk-bounded GPU resource oversubscription via conformal demand envelopes in production AI clusters. Journal of Advanced Computing Systems, 4(5), 119–134. https://doi.org/10.69987/JACS.2024.40509
Chen, Y., & Xu, H. (2026). Trust-calibrated multilingual RAG for humanitarian information platforms: Empirical evaluation on OMoS-QA for migration information access. International Journal of Graphic Design, 4(1), 141–164. https://doi.org/10.51903/ijgd.v4i1.3552
Chen, Y., Shetty, M., Somashekar, G., Ma, M., Simmhan, Y., Mace, J., Bansal, C., Wang, R., & Rajmohan, S. (2025). AIOpsLab: A holistic framework to evaluate AI agents for enabling autonomous clouds. Proceedings of Machine Learning and Systems, 7. https://proceedings.mlsys.org/paper_files/paper/2025/hash/d1f9e4a9f109b6e8b75ed362736f22ec-Abstract-Conference.html
Chen, Y., Zhang, Y., & Sherman, M. (2024). Going concern and bankruptcy prediction under extreme class imbalance: Cost-sensitive learning, resampling, and focal loss with explainable financial-ratio portraits. Journal of Advanced Computing Systems, 4(4), 80–96. https://doi.org/10.69987/JACS.2024.40407
Chen, Y., Zhang, Y., Chau, D., & Sherman, M. (2023). Credit card default risk tiering with probability calibration and uncertainty-driven rejection: A reproducible study on the UCI Credit Card Clients Dataset. Journal of Advanced Computing Systems, 3(4), 31–47. https://doi.org/10.69987/JACS.2023.30403
Chen, Y., Zhou, S., & Lin, E. (2025). Accounting-aware evidence retrieval for institutional due diligence of tokenized trade receivable RWA. Journal of Technology Informatics and Engineering, 4(3), 649–663. https://doi.org/10.51903/jtie.v4i3.542
Crankshaw, D., Sela, G.-E., Mo, X., Zumar, C., Stoica, I., Gonzalez, J. E., & Tumanov, A. (2020). InferLine: Latency-aware provisioning and scaling for prediction serving pipelines. In Proceedings of the 11th ACM Symposium on Cloud Computing (pp. 477–491). Association for Computing Machinery. https://doi.org/10.1145/3419111.3421285
Delimitrou, C., Bambos, N., & Kozyrakis, C. (2013). QoS-aware admission control in heterogeneous datacenters. In Proceedings of the 10th International Conference on Autonomic Computing (pp. 291–296). USENIX Association. https://www.usenix.org/conference/icac13/technical-sessions/presentation/delimitrou
Es, S., James, J., Espinosa-Anke, L., & Schockaert, S. (2024). RAGAs: Automated evaluation of retrieval augmented generation. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations (pp. 150–158). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.eacl-demo.16
Gao, L., Dai, Z., Pasupat, P., Chen, A., Chaganty, A. T., Fan, Y., Zhao, V., Lao, N., Lee, H., Juan, D.-C., & Guu, K. (2023). RARR: Researching and revising what language models say, using language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 16477–16508). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.acl-long.910
Gao, T., Yen, H., Yu, J., & Chen, D. (2023). Enabling large language models to generate text with citations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (pp. 6465–6488). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.emnlp-main.398
Ghodsi, A., Zaharia, M., Hindman, B., Konwinski, A., Shenker, S., & Stoica, I. (2011). Dominant resource fairness: Fair allocation of multiple resource types. In Proceedings of the 8th USENIX Symposium on Networked Systems Design and Implementation. USENIX Association. https://www.usenix.org/conference/nsdi11/dominant-resource-fairness-fair-allocation-multiple-resource-types
Gibbs, I., & Candès, E. (2021). Adaptive conformal inference under distribution shift. Advances in Neural Information Processing Systems, 34, 1660–1672. https://proceedings.neurips.cc/paper/2021/hash/0d441de75945e5acbc865406fc9a2559-Abstract.html
Grandl, R., Ananthanarayanan, G., Kandula, S., Rao, S., & Akella, A. (2014). Multi-resource packing for cluster schedulers. In Proceedings of the 2014 ACM Conference on SIGCOMM (pp. 455–466). Association for Computing Machinery. https://doi.org/10.1145/2619239.2626334
Gu, J., Chowdhury, M., Shin, K. G., Zhu, Y., Jeon, M., Qian, J., Liu, H., & Guo, C. (2019). Tiresias: A GPU cluster manager for distributed deep learning. In Proceedings of the 16th USENIX Symposium on Networked Systems Design and Implementation (pp. 485–500). USENIX Association. https://www.usenix.org/conference/nsdi19/presentation/gu
Gujarati, A., Karimi, R., Alzayat, S., Hao, W., Kaufmann, A., Vigfusson, Y., & Mace, J. (2020). Serving DNNs like Clockwork: Performance predictability from the bottom up. In Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation (pp. 443–462). USENIX Association. https://www.usenix.org/conference/osdi20/presentation/gujarati
Harlap, A., Tumanov, A., Chung, A., Ganger, G. R., & Gibbons, P. B. (2017). Proteus: Agile ML elasticity through tiered reliability in dynamic resource markets. In Proceedings of the 12th European Conference on Computer Systems (pp. 589–604). Association for Computing Machinery. https://doi.org/10.1145/3064176.3064182
He, S., Chang, X., & Sun, E. (2024). Cross-cloud transfer learning for AI training capacity forecasting under workload and topology distribution shift. Journal of Advanced Computing Systems, 4(1), 100–120. https://doi.org/10.69987/JACS.2024.40108
He, S., Li, C., & Rao, H. (2025). Few-shot cold-start workload forecasting for new AI inference tenants with time-series foundation models. Journal of Technology Informatics and Engineering, 4(1), 306–324. https://doi.org/10.51903/jtie.v4i1.546
He, S., Nie, J., & Li, C. (2026). Power-aware inventory planning for AI infrastructure using job-level forecasting and LLM workload explanations. Journal of Technology Informatics and Engineering, 5(1), 341–359. https://doi.org/10.51903/jtie.v5i1.548
He, S., Tu, H., & Liu, I. (2023). Safe PD capacity forecasting with time-series foundation models and calibrated uncertainty for heterogeneous GPU clusters. Journal of Advanced Computing Systems, 3(4), 48–66. https://doi.org/10.69987/JACS.2023.30404
Huang, Y., Yang, Z., Xing, J., Dai, Y., Qiu, Y., Wu, D., Lai, F., & Chen, A. (2024). A disaggregation approach to embedding recommendation systems. arXiv. https://doi.org/10.48550/arXiv.2410.12794
Jin, J. (2025a). Calibrated resume-job matching for trustworthy LLM-assisted recruiter screening: Pairwise matching, probability calibration, and selective refusal on two public recruitment datasets. Journal of Technology Informatics and Engineering, 4(3), 625–648. https://doi.org/10.51903/jtie.v4i3.529
Jin, J. (2025b). Evidence-chain reliable RAG: Hallucination detection, source attribution, and deterministic provenance explanations. Journal of Technology Informatics and Engineering, 4(2), 520–533. https://doi.org/10.51903/jtie.v4i2.535
Jin, J. (2025c). LLM-style evidence cards for scientific search interfaces: A UI/UX design framework for retrieval transparency, ranking trust, and visual evidence hierarchy. International Journal of Graphic Design, 3(2), 397–414. https://doi.org/10.51903/ijgd.v3i2.3698
Jin, J., Huang, T., & Lu, S. (2024a). A model-risk-friendly probability of default workflow: Calibration, distribution-free uncertainty quantification, and SHAP explanations on the UCI Credit Card Default Dataset. Journal of Advanced Computing Systems, 4(6), 74–85. https://doi.org/10.69987/JACS.2024.40606
Jin, J., Huang, T., & Lu, S. (2024b). Cost-sensitive learning, simulated PU learning, and one-class autoencoding for extreme-imbalance credit card fraud detection. Journal of Advanced Computing Systems, 4(6), 64–73. https://doi.org/10.69987/JACS.2024.40605
Johnson, D. S., Demers, A., Ullman, J. D., Garey, M. R., & Graham, R. L. (1974). Worst-case performance bounds for simple one-dimensional packing algorithms. SIAM Journal on Computing, 3(4), 299–325. https://doi.org/10.1137/0203025
Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., & Liu, T.-Y. (2017). LightGBM: A highly efficient gradient boosting decision tree. Advances in Neural Information Processing Systems, 30, 3146–3154. https://proceedings.neurips.cc/paper/6907-lightgbm-a-highly-efficient-gradient-boosting-decision-tree
Kuo, M.-J., Zheng, D., & Hires, J. (2025). Federated topic-preference learning for knowledge-grounded chat with differential privacy. Journal of Technology Informatics and Engineering, 4(2), 385–401. https://doi.org/10.51903/jtie.v4i2.502
Li, C., Bai, J., & Wang, S. (2024). Evidence-chain reliable RAG: Word-level hallucination detection, source attribution, and provenance explanation for LLM applications. Journal of Advanced Computing Systems, 4(2), 76–92. https://doi.org/10.69987/JACS.2024.40207
Li, C., Liu, G., & Zhao, Z. (2026). Cost-aware LLM-style routing for AIOps log analysis: Log parsing, anomaly detection, fault diagnosis, and incident summarization on LogEval task files. Journal of Technology Informatics and Engineering, 5(2), 91–103. https://doi.org/10.51903/jtie.v5i2.538
Li, J., & Zhou, A. (2026). Multi-regulation RAG for AI product counsel: A legal governance framework for cross-border digital commerces. Rule of Law Studies Journal, 2(2), 105–123. https://doi.org/10.64780/rolsj.v2i2.225
Li, Y. (2024). Findable then explainable: Retrieval–summary integration for code intelligence on a lightweight CodeSearchNet subset. Journal of Advanced Computing Systems, 4(7), 65–82. https://doi.org/10.69987/JACS.2024.40706
Li, Y., & Lu, S. (2025). Language-guided feature selection for DDoS and intrusion detection on CICIDS2017. Journal of Technology Informatics and Engineering, 4(1), 284–305. https://doi.org/10.51903/jtie.v4i1.531
Li, Y., Lu, S., & Zhao, L. (2025). LLM-as-design-critic: Aligning AI-generated UI feedback with human graphic design judgment. International Journal of Graphic Design, 3(1), 196–215. https://doi.org/10.51903/ijgd.v3i1.3661
Li, Z., Zhang, K., & Wong, A. (2026). Numerical-reasoning guardrails for a quant research assistant: A compact reproducible benchmark using SEC and FRED data. Journal of Technology Informatics and Engineering, 5(2), 75–90. https://doi.org/10.51903/jtie.v5i2.541
Li, Z., Zhou, S., & Zhou, Z. (2025). Financial risk dashboard design for institutional RWA investors: Visual hierarchy, chart comprehension, and explainability in FinChart-Bench. International Journal of Graphic Design, 3(1), 196–210. https://doi.org/10.51903/ijgd.v3i1.3715
Liu, G., He, S., & Liu, I. (2023). LLM-augmented multi-source root cause attribution for CPU and network faults in microservices. Journal of Advanced Computing Systems, 3(6), 39–57. https://doi.org/10.69987/JACS.2023.30604
Liu, G., He, S., & Wong, H. (2025). LLM-compatible visual brief cards for AI infrastructure capacity dashboards: A UI/UX framework for turning forecast risk into graphic design decisions. International Journal of Graphic Design, 3(1), 196–213. https://doi.org/10.51903/ijgd.v3i1.3723
Liu, G., Li, C., & Zhang, E. (2024). OpsLLM for cloud incident triage: Bilingual RAG-based root cause analysis and alert summarization for AI infrastructure operations. Journal of Advanced Computing Systems, 4(4), 97–111. https://doi.org/10.69987/JACS.2024.40408
Liu, X., Zhao, Y., Liu, S., Li, X., Zhu, Y., Liu, X., & Jin, X. (2024). MuxFlow: Efficient GPU sharing in production-level clusters with more than 10000 GPUs. Science China Information Sciences, 67, Article 222101. https://doi.org/10.1007/s11432-024-4227-2
Lu, S., & Zhou, D. (2024). TinyLLM-assisted intrusion detection for real-time IoT networks. Journal of Advanced Computing Systems, 4(8), 72–87. https://doi.org/10.69987/JACS.2024.40809
Lu, Y., Zhou, H., & Zhang, Y. (2025). A constrained, data-driven budgeting framework integrating macro demand forecasting and marketing response modeling. Journal of Technology Informatics and Engineering, 4(3), 493–520. https://doi.org/10.51903/jtie.v4i3.466
Mao, H., Alizadeh, M., Menache, I., & Kandula, S. (2016). Resource management with deep reinforcement learning. In Proceedings of the 15th ACM Workshop on Hot Topics in Networks (pp. 50–56). Association for Computing Machinery. https://doi.org/10.1145/3005745.3005750
Mao, H., Schwarzkopf, M., Venkatakrishnan, S. B., Meng, Z., & Alizadeh, M. (2019). Learning scheduling algorithms for data processing clusters. In Proceedings of the ACM Special Interest Group on Data Communication (pp. 270–288). Association for Computing Machinery. https://doi.org/10.1145/3341302.3342080
Meng, S., Chen, J., & Zheng, I. (2026). LLM-inspired offline reranking for financial search: Query rewriting, hybrid retrieval, and listwise relevance ranking on FiQA. Journal of Technology Informatics and Engineering, 5(1), 361–378. https://doi.org/10.51903/jtie.v5i1.537
Min, S., Krishna, K., Lyu, X., Lewis, M., Yih, W.-t., Koh, P. W., Iyyer, M., Zettlemoyer, L., & Hajishirzi, H. (2023). FActScore: Fine-grained atomic evaluation of factual precision in long-form text generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (pp. 12076–12100). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.emnlp-main.741
Mu, J., Lu, Y., & Hwang, E. (2026). Structured visual brief interfaces for advertising design: A UI/UX framework for turning creative intentions into designer-editable graphic design cards. International Journal of Graphic Design, 4(1), 192–208. https://doi.org/10.51903/ijgd.v4i1.3702
Mu, J., Lu, Y., & Smith, M. (2023). LLM-assisted incrementality (uplift) modeling for email advertising: From feature interactions to interpretable audience–creative–channel policies. Journal of Advanced Computing Systems, 3(1), 31–48. https://doi.org/10.69987/JACS.2023.30103
Mu, J., Ye, T., & Patel, P. (2025). Offline counterfactual evaluation for advertising and recommendation slot policies: A reproducible study on the Open Bandit Dataset (Small). Journal of Technology Informatics and Engineering, 4(3), 521–543. https://doi.org/10.51903/jtie.v4i3.500
Narayanan, D., Santhanam, K., Kazhamiaka, F., Phanishayee, A., & Zaharia, M. (2020). Heterogeneity-aware cluster scheduling policies for deep learning workloads. In Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation (pp. 481–498). USENIX Association. https://www.usenix.org/conference/osdi20/presentation/narayanan-deepak
Nie, J., & Zheng, D. (2024). Noisy-neighbor-aware VM degradation risk modeling with unsupervised residual fusion. Journal of Advanced Computing Systems, 4(4), 112–123. https://doi.org/10.69987/JACS.2024.40409
Nie, J., Liu, G., Li, C., & Zou, T. (2026). Evidence-constrained incident visualization cards for distributed cloud logs: A UI/UX framework for turning Hadoop, OpenStack, and ZooKeeper logs into actionable SRE design interfaces. International Journal of Graphic Design, 4(1), 179–185. https://doi.org/10.51903/ijgd.v4i1.3703
Patel, P., Choukse, E., Zhang, C., Shah, A., Goiri, Í., Maleki, S., & Bianchini, R. (2024). Splitwise: Efficient generative LLM inference using phase splitting. In 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture. IEEE. https://doi.org/10.1109/ISCA59077.2024.00019
Qiu, H., Mao, W., Patke, A., Cui, S., Jha, S., Wang, C., Franke, H., Kalbarczyk, Z., Başar, T., & Iyer, R. K. (2024). Power-aware deep learning model serving with μ-Serve. In 2024 USENIX Annual Technical Conference (pp. 75–93). USENIX Association. https://www.usenix.org/conference/atc24/presentation/qiu
Romano, Y., Patterson, E., & Candès, E. J. (2019). Conformalized quantile regression. Advances in Neural Information Processing Systems, 32, 3543–3553. https://proceedings.neurips.cc/paper/2019/hash/5103c3584b063c431bd1268e9b5e76fb-Abstract.html
Su, W., Chen, S., & Qian, E. (2026). Narrative-aware scientific claim verification agent with evidence ranking for ClimateCheck. Journal of Technology Informatics and Engineering, 5(1), 327–340. https://doi.org/10.51903/jtie.v5i1.549
Su, W., Chen, S., & Zhao, C. (2025). Budgeted multi-hop retrieval agent for compositional question answering: A retrieval-policy evaluation on the official MultiHop-RAG benchmark. Journal of Technology Informatics and Engineering, 4(3), 649–662. https://doi.org/10.51903/jtie.v4i3.543
Su, W., Rao, H., & Ma, E. (2026). Privacy and data-integrity risk cards for LLM agents: A UI/UX design framework for secure human oversight under prompt-injection attacks. International Journal of Graphic Design, 4(1), 186–191. https://doi.org/10.51903/ijgd.v4i1.3699
Sun, X., Lu, Y., & Chen, J. (2023). Controllable long-term user memory for multi-session dialogue: Confidence-gated writing, time-aware retrieval-augmented generation, and update/forgetting. Journal of Advanced Computing Systems, 3(8), 9–24. https://doi.org/10.69987/JACS.2023.30802
Sun, X., Zhong, Z. S., & Wu, Q. (2026). Retrieval-grounded HDFS log anomaly detection and deterministic failure narrative generation. Journal of Computational Systems and Applications, 3(1), 15–30. https://doi.org/10.64229/j6d7fr94
Tu, H., Zhao, S., & Zhou, A. (2025). Visual brief cards for advertising design: A structured UI/UX framework for turning creative intentions into graphic design decisions. International Journal of Graphic Design, 3(1), 210–226. https://doi.org/10.51903/ijgd.v3i1.3714
Wang, B., He, Y., Shui, Z., Xin, Q., & Lei, H. (2024). Predictive optimization of DDoS attack mitigation in distributed systems using machine learning. In Proceedings of the 6th International Conference on Computing and Data Science (CDS 2024) (pp. 89–94).
Wang, C., Wen, Z., Zhang, R., Xu, P., & Jiang, Y. (2025). GPU memory requirement prediction for deep learning task based on bidirectional gated recurrent unit optimization transformer. In 2025 5th International Conference on Artificial Intelligence, Virtual Reality and Visualization (AIVRV). IEEE. https://doi.org/10.1109/AIVRV67401.2025.11350369
Weng, Q., Yang, L., Yu, Y., Wang, W., Tang, X., Yang, G., & Zhang, L. (2023). Beware of fragmentation: Scheduling GPU-sharing workloads with fragmentation gradient descent. In 2023 USENIX Annual Technical Conference (pp. 995–1008). USENIX Association. https://www.usenix.org/conference/atc23/presentation/weng
Wu, C.-J., Raghavendra, R., Gupta, U., Acun, B., Ardalani, N., Maeng, K., Chang, G., Aga, F., Huang, J., Bai, C., Gschwind, M., Gupta, A., Ott, M., Melnikov, A., Candido, S., Brooks, D., Chauhan, G., Lee, B., Lee, H.-H. S., . . . Hazelwood, K. (2022). Sustainable AI: Environmental implications, challenges and opportunities. Proceedings of Machine Learning and Systems, 4, 795–813. https://proceedings.mlsys.org/paper_files/paper/2022/hash/462211f67c7d858f663355eff93b745e-Abstract.html
Wu, Q., Meng, S., & Zhao, J. (2025). Text-grounded LLM-assisted design rationale interfaces: Turning advertising layout metadata into explainable UI/UX decision cards. International Journal of Graphic Design, 3(1), 216–240. https://doi.org/10.51903/ijgd.v3i1.3713
Xiao, W., Ren, S., Li, Y., Zhang, Y., Hou, P., Li, Z., Feng, Y., Lin, W., & Jia, Y. (2020). AntMan: Dynamic scaling on GPU clusters for deep learning. In Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation (pp. 533–548). USENIX Association. https://www.usenix.org/conference/osdi20/presentation/xiao
Xin, Q. (2025a). Explaining OpenStack failure-injection log anomalies with retrieved normal prototypes. Emerging Information Science and Technology, 6(2), 125–146. https://doi.org/10.18196/eist.v6i2.31232
Xin, Q. (2025b). Hybrid cloud architecture for efficient and cost-effective large language model deployment. Journal of Information Systems and Informatics, 7(3), 2182–2195. https://doi.org/10.51519/journalisi.v7i3.1170
Xin, Q. (2025c). Uncertainty-aware late fusion for 3D perception (confidence calibration + fusion rule learning). Journal of Technology Informatics and Engineering, 4(1), 215–238. https://doi.org/10.51903/jtie.v4i1.485
Xin, Q. (2026a). Behavior retrieval plus response generation for interpretable conversational personalized recommendation. IJEEPSE, 9(2), 120–136. https://doi.org/10.31258/ijeepse.9.2.120-136
Xin, Q. (2026b). Explainable and fair credit risk scoring with counterfactual explanations: A reproducible evaluation on the German Credit Dataset (HELOC-motivated). Journal of Information Technology, 14(2). https://doi.org/10.32664/j-intech.v14i02.2228
Xin, Q. (2026c). Host-based intrusion detection with system call sequences: Window localization and forensic narratives. AVITEC, 8(2). https://doi.org/10.28989/avitec.v8i2.3973
Xin, Q. (2026d). Log anomaly detection with conformal alert control and evidence-grounded incident ticket generation. AVITEC, 8(2), 247. https://doi.org/10.28989/avitec.v8i2.3974
Xin, Q. (2026e). Probabilistic bike-sharing demand forecasting under changing weather and seasonal regimes with transformer-based models. Transport Findings. https://doi.org/10.32866/001c.157499
Xin, Q. (2026f). Self-supervised customer representation learning for segmentation and next-purchase prediction on UCI Online Retail. Journal of Information Technology, 14(1). https://doi.org/10.32664/j-intech.v14i01.2229
Xin, Q. (2026g). Self-supervised log anomaly detection with LogBERT-style transformers: Full empirical evaluation on a reproducible SynHDFS benchmark. Journal of Electrical Engineering and Computer Science, 11(1), 23–35. https://doi.org/10.54732/jeecs.v11i1.3
Xin, Q., Xu, Z., Guo, L., Zhao, F., & Wu, B. (2024). IoT traffic classification and anomaly detection method based on deep autoencoders. In Proceedings of the 6th International Conference on Computing and Data Science (CDS 2024).
Xu, H., Chen, Y., & Med, A. (2025). Automatic detection and explanation of dark patterns from interface microcopy: Empirical comparison of BERT-style encoders, RoBERTa-style encoders, and LLM-style decoders on the ec-darkpattern dataset. Journal of Technology Informatics and Engineering, 4(3), 590–612. https://doi.org/10.51903/jtie.v4i3.491
Xu, K., Zhou, H., Zheng, H., Zhu, M., & Xin, Q. (2024). Intelligent classification and personalized recommendation of e-commerce products based on machine learning. In Proceedings of the 6th International Conference on Computing and Data Science (ICCDS 2024).
Yang, L., Wang, Y., Yu, Y., Weng, Q., Dong, J., Liu, K., Zhang, C., Zi, Y., Li, H., Zhang, Z., Wang, N., Dong, Y., Zheng, M., Xi, L., Lu, X., Ye, L., Yang, G., Fu, B., Lan, T., . . . Wang, W. (2025). GPU-disaggregated serving for deep learning recommendation models at scale. In Proceedings of the 22nd USENIX Symposium on Networked Systems Design and Implementation (pp. 847–863). USENIX Association. https://www.usenix.org/conference/nsdi25/presentation/yang
Ye, T., Mu, J., & Hunter, J. (2026). Off-policy evaluation and conservative policy selection for slot-level dynamic bidding and ranking on the Open Bandit Dataset (Small). Journal of Technology Informatics and Engineering, 5(1), 178–199. https://doi.org/10.51903/jtie.v5i1.503
You, J., Chung, J.-W., & Chowdhury, M. (2023). Zeus: Understanding and optimizing GPU energy consumption of DNN training. In Proceedings of the 20th USENIX Symposium on Networked Systems Design and Implementation (pp. 119–139). USENIX Association. https://www.usenix.org/conference/nsdi23/presentation/you
Yu, P., & Chowdhury, M. (2020). Salus: Fine-grained GPU sharing primitives for deep learning applications. Proceedings of Machine Learning and Systems, 2, 98–111. https://proceedings.mlsys.org/paper_files/paper/2020/hash/d9cd83bc91b8c36a0c7c0fcca59228f2-Abstract.html
Zhang, B., Rao, H., & Zhao, D. (2024). Evidence-grounded RAG for cloud-native DevOps: Hallucination-resistant AIOps question answering over private operations documents. Journal of Advanced Computing Systems, 4(3), 109–125. https://doi.org/10.69987/JACS.2024.40308
Zhang, B., Ren, Y., & Zou, J. (2025). LLM-style explainable e-commerce recommendation cards: A UI/UX design framework for trust-calibrated product recommendation. International Journal of Graphic Design, 3(2), 381–396. https://doi.org/10.51903/ijgd.v3i2.3697
Zhang, K., Chen, Y., & Qian, A. (2025). Evidence-grounded accounting disclosure review cards: A visual communication framework for LLM-style explanations over SEC financial statements and notes. International Journal of Graphic Design, 3(2), 395. https://doi.org/10.51903/ijgd.v3i2.3710
Zhang, R., Wen, Z., Wang, C., Tang, C., Xu, P., & Jiang, Y. (2025). Quality analysis and evaluation prediction of RAG retrieval based on machine learning algorithms. arXiv. https://doi.org/10.48550/arXiv.2511.19481
Zhao, S., Bai, J., & Roberson, D. (2025). Multi-horizon GPU demand forecasting with workload semantics and operational risk curves: An empirical study on Alibaba Clusterdata GPU Trace. Journal of Technology Informatics and Engineering, 4(3), 544–571. https://doi.org/10.51903/jtie.v4i3.498
Zhao, S., Ren, Y., & Chang, X. (2026). Profit-aware spot GPU admission control with cost-sensitive loss and evidence-grounded policy memos for AI workload supply-demand matching. Journal of Technology Informatics and Engineering, 5(2), 45–59. https://doi.org/10.51903/jtie.v5i2.545
Zheng, D., & Li, C. (2024). Behavior-level jailbreak resistance via multi-stage refusal + utility preservation. Journal of Advanced Computing Systems, 4(1), 83–99. https://doi.org/10.69987/JACS.2024.40107
Zheng, D., Li, C., & Davidson, H. (2023). Continual red-teaming for in-the-wild jailbreaks via online guardrail updates and guardrail distillation. Journal of Advanced Computing Systems, 3(2), 35–49. https://doi.org/10.69987/JACS.2023.30203
Zheng, D., Zhang, B., & Geibel, J. (2024). VerifySafe: Toxicity-safe agent responses under adversarial prompts with evidence-based self-verification. Journal of Advanced Computing Systems, 4(1), 67–82. https://doi.org/10.69987/JACS.2024.40106
Zhong, Y., Liu, S., Chen, J., Hu, J., Zhu, Y., Liu, X., Jin, X., & Zhang, H. (2024). DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving. In Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation (pp. 193–210). USENIX Association. https://www.usenix.org/conference/osdi24/presentation/zhong-yinmin
Zhong, Z. S., & Ling, S. (2024). Uncertainty quantification of spectral estimator and MLE for orthogonal group synchronization. arXiv. https://arxiv.org/abs/2408.05944
Zhong, Z. S., Chen, J., Zhong, E., & Sun, X. (2025). Evidence-calibrated RAG for unanswerable question answering: Retrieval coverage, abstention calibration, and hallucination-proxy analysis on SQuAD 2.0. Journal of Technology Informatics and Engineering, 4(2), 502–520. https://doi.org/10.51903/jtie.v4i2.536
Zhong, Z. S., Li, C., & Rao, H. (2026). Trajectory reliability prediction for generalist AI agents: Tool-use failure analysis and success forecasting on ZClawBench. Journal of Technology Informatics and Engineering, 5(1), 341–360. https://doi.org/10.51903/jtie.v5i1.539
Zhong, Z. S., Pan, X., & Lei, Q. (2025). Bridging domains with approximately shared features. In Proceedings of the 28th International Conference on Artificial Intelligence and Statistics (Vol. 258, pp. 559–567). PMLR. https://proceedings.mlr.press/v258/zhong25a.html
Zhou, H., & Zhang, K. (2025). News-based uncertainty and macro-market fusion for VIX direction forecasting: Evidence from 2015–2024 FRED panel. Journal of Technology Informatics and Engineering, 4(2), 487–501. https://doi.org/10.51903/jtie.v4i2.540
Zhou, S., Chen, Y., & Lee, K. (2026). Accounting-aware evidence-constrained agents for disclosure, settlement, and secondary-market risk monitoring in tokenized. Journal of Technology Informatics and Engineering, 5(2), 60–74. https://doi.org/10.51903/jtie.v5i2.544
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Justin Yang (Author)

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.