Trust-Calibrated Enterprise Service Copilots: Cost-Aware Agentic RAG for Customer-Support Knowledge Management
DOI:
https://doi.org/10.61424/y3my3g85Keywords:
Enterprise knowledge management, retrieval-augmented generation, customer support, selective prediction, reciprocal-rank fusion, latent semantic analysis, cost-aware routingAbstract
Enterprise service copilots must retrieve relevant knowledge, provide traceable answers, control unsupported responses, and operate within practical latency and context budgets. This study evaluated these requirements on WixQA using the complete 6,221-article knowledge base, 6,221 Synthetic questions, 200 ExpertWritten questions, and 200 Simulated questions. The pipeline compared Okapi BM25, a 256-dimensional Dense-LSA retriever, reciprocal-rank fusion, a supervised logistic reranker, a top-1 article-match confidence gate, budget-aware stopping, and deterministic extractive evidence cards. Learned components used grouped five-fold out-of-fold evaluation on Synthetic, followed by frozen external evaluation. On the pooled external sets, Recall@5 was 0.4704 for BM25, 0.3704 for Dense-LSA, 0.4946 for hybrid retrieval, and 0.4838 after reranking. The adaptive policy answered 56.5% of ExpertWritten and 48.5% of Simulated requests; answered-case token F1 was 0.3051 and 0.2927, respectively. Median retrieved context fell to 71 and 63 analysis tokens. The Synthetic gate attained 0.0999 selective top-1 error at 0.8883 coverage, but frozen external error rose to 0.5664 and 0.6907. All tested financial scenarios produced negative ROI, although referral yielded the least negative base-case ratio. Hybrid retrieval improved external evidence recall, but synthetic confidence and benchmark efficiency did not support autonomous service. Representative calibration, explicit referral, and operational measurement remain necessary.
References
Asai, A., Wu, Z., Wang, Y., Sil, A., & Hajishirzi, H. (2024). Self-RAG: Learning to retrieve, generate, and critique through self-reflection. The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=hSyW5go0v8
Bai, J., & Wu, Q. (2026). Privacy-safe marketing mix modeling and budget optimization under identifier loss: A controlled simulation study. International Journal of Electronics and Communication Systems, 6(1). https://doi.org/10.24042/ijecs.v6i1.30533
Bai, J., Chen, S., Zheng, D., & Kuo, M.-J. (2026). Interpretable attack-chain stage detection from AWS CloudTrail event sequences via linear models and HMM smoothing. Information, Electrical and Electronics Engineering, 6(1), 28-43. https://doi.org/10.33474/infotron.v6i1.24923
Bai, J., Wang, H., Wu, Q., & Zhang, B. (2026). Privacy-robust incrementality estimation in cookieless settings via uplift modeling: Reproducible evidence from the Hillstrom e-mail experiment. Journal of Technology Informatics and Engineering, 5(1), 17-38. https://doi.org/10.51903/jtie.v5i1.468
Brynjolfsson, E., Li, D., & Raymond, L. R. (2025). Generative AI at work. The Quarterly Journal of Economics, 140(2), 889–942. https://doi.org/10.1093/qje/qjae044
Chang, X., Lu, Y., & Zhong, Z. S. (2026). Review-grounded explainable recommendation with faithfulness evaluation on Amazon Reviews. JEECS (Journal of Electrical Engineering and Computer Sciences), 11(1), 9-22. https://doi.org/10.54732/jeecs.v11i1.2
Chen, L., Zaharia, M., & Zou, J. (2024). FrugalGPT: How to use large language models while reducing cost and improving performance. Transactions on Machine Learning Research. https://openreview.net/forum?id=cSimKw5p6R
Chen, S., He, S., & Sun, E. (2024). Risk-bounded GPU resource oversubscription via conformal demand envelopes in production AI clusters. Journal of Advanced Computing Systems, 4(5), 119-134. https://doi.org/10.69987/JACS.2024.40509
Chen, Y., & Xu, H. (2026). Trust-calibrated multilingual RAG for humanitarian information platforms: Empirical evaluation on OMoS-QA for migration information access. International Journal of Graphic Design, 4(1), 141-164. https://doi.org/10.51903/ijgd.v4i1.3552
Chen, Y., Zhang, Y., & Sherman, M. (2024). Going concern and bankruptcy prediction under extreme class imbalance: Cost-sensitive learning, resampling, and focal loss with explainable financial-ratio portraits. Journal of Advanced Computing Systems, 4(4), 80-96. https://doi.org/10.69987/JACS.2024.40407
Chen, Y., Zhang, Y., Chau, D., & Sherman, M. (2023). Credit card default risk tiering with probability calibration and uncertainty-driven rejection: A reproducible study on the UCI credit card clients dataset. Journal of Advanced Computing Systems, 3(4), 31-47. https://doi.org/10.69987/JACS.2023.30403
Cohen, D., Burg, L., Pykhnivskyi, S., Gur, H., Kovynov, S., Atzmon, O., & Barkan, G. (2025). WixQA: A multi-dataset benchmark for enterprise retrieval-augmented generation. arXiv. https://doi.org/10.48550/arXiv.2505.08643
Cormack, G. V., Clarke, C. L. A., & Buettcher, S. (2009). Reciprocal rank fusion outperforms Condorcet and individual rank learning methods. In Proceedings of the 32nd International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 758–759). ACM. https://doi.org/10.1145/1571941.1572114
Deerwester, S., Dumais, S. T., Furnas, G. W., Landauer, T. K., & Harshman, R. (1990). Indexing by latent semantic analysis. Journal of the American Society for Information Science, 41(6), 391–407. https://doi.org/10.1002/(SICI)1097-4571(199009)41:6%3C391::AID-ASI1%3E3.0.CO;2-9
DeLone, W. H., & McLean, E. R. (2003). The DeLone and McLean model of information systems success: A ten-year update. Journal of Management Information Systems, 19(4), 9–30. https://doi.org/10.1080/07421222.2003.11045748
Ding, D., Mallick, A., Wang, C., Sim, R., Mukherjee, S., Rühle, V., Lakshmanan, L. V. S., & Awadallah, A. H. (2024). Hybrid LLM: Cost-efficient and quality-aware query routing. The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=02f3mUtqnM
Es, S., James, J., Espinosa Anke, L., & Schockaert, S. (2024). RAGAS: Automated evaluation of retrieval augmented generation. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations (pp. 150–158). ACL. https://doi.org/10.18653/v1/2024.eacl-demo.16
Feng, S., Patel, S. S., Wan, H., & Joshi, S. (2021). MultiDoc2Dial: Modeling dialogues grounded in multiple documents. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (pp. 6162–6176). ACL. https://doi.org/10.18653/v1/2021.emnlp-main.498
Feng, S., Wan, H., Gunasekara, C., Patel, S. S., Joshi, S., & Lastras, L. (2020). Doc2Dial: A goal-oriented document-grounded dialogue dataset. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (pp. 8118–8128). ACL. https://doi.org/10.18653/v1/2020.emnlp-main.652
Gao, T., Yen, H., Yu, J., & Chen, D. (2023). Enabling large language models to generate text with citations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (pp. 6465–6488). ACL. https://doi.org/10.18653/v1/2023.emnlp-main.398
Geifman, Y., & El-Yaniv, R. (2017). Selective classification for deep neural networks. Advances in Neural Information Processing Systems, 30. https://proceedings.neurips.cc/paper/2017/hash/4a8423d5e91fda00bb7e46540e2b0cf1-Abstract.html
Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning (pp. 1321–1330). PMLR. https://proceedings.mlr.press/v70/guo17a.html
He, J., Chen, P., Wu, C., Liang, S., Li, Y., Tan, G., Wen, X., & Zhang, C. (2026). An end-to-end framework for building large language models for software operations. arXiv. https://doi.org/10.48550/arXiv.2605.02906
He, S., Li, C., & Rao, H. (2025). Few-shot cold-start workload forecasting for new AI inference tenants with time-series foundation models. Journal of Technology Informatics and Engineering, 4(1), 306-324. https://doi.org/10.51903/jtie.v4i1.546
He, S., Nie, J., & Li, C. (2026). Power-aware inventory planning for AI infrastructure using job-level forecasting and LLM workload explanations. Journal of Technology Informatics and Engineering, 5(1), 341-359. https://doi.org/10.51903/jtie.v5i1.548
He, S., Tu, H., & Liu, I. (2023). Safe PD capacity forecasting with time-series foundation models and calibrated uncertainty for heterogeneous GPU clusters. Journal of Advanced Computing Systems, 3(4), 48-66. https://doi.org/10.69987/JACS.2023.30404
Jeong, S., Baek, J., Cho, S., Hwang, S. J., & Park, J. (2024). Adaptive-RAG: Learning to adapt retrieval-augmented large language models through question complexity. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 7036–7050). ACL. https://doi.org/10.18653/v1/2024.naacl-long.389
Jin, J. (2025a). Evidence-chain reliable RAG: Hallucination detection, source attribution, and deterministic provenance explanations. Journal of Technology Informatics and Engineering, 4(2), 520-533. https://doi.org/10.51903/jtie.v4i2.535
Jin, J. (2025b). LLM-style evidence cards for scientific search interfaces: A UI/UX design framework for retrieval transparency, ranking trust, and visual evidence hierarchy. International Journal of Graphic Design, 3(2), 397-414. https://doi.org/10.51903/ijgd.v3i2.3698
Jin, J., Huang, T., & Lu, S. (2024). A model-risk-friendly probability of default workflow: Calibration, distribution-free uncertainty quantification, and SHAP explanations on the UCI credit card default dataset. Journal of Advanced Computing Systems, 4(6), 74-85. https://doi.org/10.69987/JACS.2024.40606
Khan, A. F., Khan, A. A., Mohamed, A., Ali, H., Moolinti, S., Haroon, S., Tahir, U., Fazzini, M., Butt, A. R., & Anwar, A. (2025). LADs: Leveraging large language models for AI-driven DevOps. arXiv. https://doi.org/10.48550/arXiv.2502.20825
Kuo, M.-J., Zheng, D., & Hires, J. (2025). Federated topic-preference learning for knowledge-grounded chat with differential privacy. Journal of Technology Informatics and Engineering, 4(2), 385-401. https://doi.org/10.51903/jtie.v4i2.502
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9474. https://proceedings.neurips.cc/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html
Li, C., Bai, J., & Wang, S. (2024). Evidence-chain reliable RAG: Word-level hallucination detection, source attribution, and provenance explanation for LLM applications. Journal of Advanced Computing Systems, 4(2), 76-92. https://doi.org/10.69987/JACS.2024.40207
Li, C., Liu, G., & Zhao, Z. (2026). Cost-aware LLM-style routing for AIOps log analysis: Log parsing, anomaly detection, fault diagnosis, and incident summarization on LogEval task files. Journal of Technology Informatics and Engineering, 5(2), 91-103. https://doi.org/10.51903/jtie.v5i2.538
Li, J., & Zhou, A. (2026). Multi-regulation RAG for AI product counsel: A legal governance framework for cross-border digital commerce. Rule of Law Studies Journal, 2(2), 105-123. https://doi.org/10.64780/rolsj.v2i2.225
Li, Y. (2024). Findable then explainable: Retrieval-summary integration for code intelligence on a lightweight CodeSearchNet subset. Journal of Advanced Computing Systems, 4(7), 65-82. https://doi.org/10.69987/JACS.2024.40706
Li, Y., & Lu, S. (2025). Language-guided feature selection for DDoS and intrusion detection on CICIDS2017. Journal of Technology Informatics and Engineering, 4(1), 284-305. https://doi.org/10.51903/jtie.v4i1.531
Li, Y., Lu, S., & Zhao, L. (2025). LLM-as-design-critic: Aligning AI-generated UI feedback with human graphic design judgment. International Journal of Graphic Design, 3(1), 196-215. https://doi.org/10.51903/ijgd.v3i1.3661
Li, Z., Zhang, K., & Wong, A. (2026). Numerical-reasoning guardrails for a quant research assistant: A compact reproducible benchmark using SEC and FRED data. Journal of Technology Informatics and Engineering, 5(2), 75-90. https://doi.org/10.51903/jtie.v5i2.541
Li, Z., Zhou, S., & Zhou, Z. (2025). Financial risk dashboard design for institutional RWA investors: Visual hierarchy, chart comprehension, and explainability in FinChart-Bench. International Journal of Graphic Design, 3(1), 196-210. https://doi.org/10.51903/ijgd.v3i1.3715
Liu, G., He, S., & Liu, I. (2023). LLM-augmented multi-source root cause attribution for CPU and network faults in microservices. Journal of Advanced Computing Systems, 3(6), 39-57. https://doi.org/10.69987/JACS.2023.30604
Liu, G., He, S., & Wong, H. (2025). LLM-compatible visual brief cards for AI infrastructure capacity dashboards: A UI/UX framework for turning forecast risk into graphic design decisions. International Journal of Graphic Design, 3(1), 196-213. https://doi.org/10.51903/ijgd.v3i1.3723
Lu, S., & Zhou, D. (2024). TinyLLM-assisted intrusion detection for real-time IoT networks. Journal of Advanced Computing Systems, 4(8), 72-87. https://doi.org/10.69987/JACS.2024.40809
Lu, Y., Zhou, H., & Zhang, Y. (2025). A constrained, data-driven budgeting framework integrating macro demand forecasting and marketing response modeling. Journal of Technology Informatics and Engineering, 4(3). https://doi.org/10.51903/jtie.v4i3.466
Meng, S., Chen, J., & Zheng, I. (2026). LLM-inspired offline reranking for financial search: Query rewriting, hybrid retrieval, and listwise relevance ranking on FiQA. Journal of Technology Informatics and Engineering, 5(1), 361-378. https://doi.org/10.51903/jtie.v5i1.537
Mu, J., Lu, Y., & Hwang, E. (2026). Structured visual brief interfaces for advertising design: A UI/UX framework for turning creative intentions into designer-editable graphic design cards. International Journal of Graphic Design, 4(1), 192-208. https://doi.org/10.51903/ijgd.v4i1.3702
Mu, J., Lu, Y., & Smith, M. (2023). LLM-assisted incrementality (uplift) modeling for email advertising: From feature interactions to interpretable audience-creative-channel policies. Journal of Advanced Computing Systems, 3(1), 31-48. https://doi.org/10.69987/JACS.2023.30103
Mu, J., Ye, T., & Patel, P. (2025). Offline counterfactual evaluation for advertising and recommendation slot policies: A reproducible study on the Open Bandit Dataset (Small). Journal of Technology Informatics and Engineering, 4(3), 521-543. https://doi.org/10.51903/jtie.v4i3.500
Nie, J., & Zheng, D. (2024). Noisy-neighbor-aware VM degradation risk modeling with unsupervised residual fusion. Journal of Advanced Computing Systems, 4(4), 112-123. https://doi.org/10.69987/JACS.2024.40409
Nie, J., Liu, G., Li, C., & Zou, T. (2026). Evidence-constrained incident visualization cards for distributed cloud logs: A UI/UX framework for turning Hadoop, OpenStack, and ZooKeeper logs into actionable SRE design interfaces. International Journal of Graphic Design, 4(1), 179-185. https://doi.org/10.51903/ijgd.v4i1.3703
Nogueira, R., & Cho, K. (2019). Passage re-ranking with BERT. arXiv. https://doi.org/10.48550/arXiv.1901.04085
Ong, I., Almahairi, A., Wu, V., Chiang, W.-L., Wu, T., Gonzalez, J. E., Kadous, M. W., & Stoica, I. (2025). RouteLLM: Learning to route LLMs from preference data. The Thirteenth International Conference on Learning Representations. https://openreview.net/forum?id=8sSqNntaMr
Rajpurkar, P., Zhang, J., Lopyrev, K., & Liang, P. (2016). SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing (pp. 2383–2392). ACL. https://doi.org/10.18653/v1/D16-1264
Robertson, S., & Zaragoza, H. (2009). The probabilistic relevance framework: BM25 and beyond. Foundations and Trends in Information Retrieval, 3(4), 333–389. https://doi.org/10.1561/1500000019
Ru, D., Qiu, L., Hu, X., Zhang, T., Shi, P., Chang, S., Jiayang, C., Wang, C., Sun, S., Li, H., Zhang, Z., Wang, B., Jiang, J., He, T., Wang, Z., Liu, P., Zhang, Y., & Zhang, Z. (2024). RAGChecker: A fine-grained framework for diagnosing retrieval-augmented generation. Advances in Neural Information Processing Systems, 37, 21999–22027. https://doi.org/10.52202/079017-0692
Saad-Falcon, J., Khattab, O., Potts, C., & Zaharia, M. (2024). ARES: An automated evaluation framework for retrieval-augmented generation systems. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 338–354). ACL. https://doi.org/10.18653/v1/2024.naacl-long.20
Su, W., Chen, S., & Qian, E. (2026). Narrative-aware scientific claim verification agent with evidence ranking for ClimateCheck. Journal of Technology Informatics and Engineering, 5(1), 327-340. https://doi.org/10.51903/jtie.v5i1.549
Su, W., Chen, S., & Zhao, C. (2025). Budgeted multi-hop retrieval agent for compositional question answering: A retrieval-policy evaluation on the official MultiHop-RAG benchmark. Journal of Technology Informatics and Engineering, 4(3), 649-662. https://doi.org/10.51903/jtie.v4i3.543
Su, W., Rao, H., & Ma, E. (2026). Privacy and data-integrity risk cards for LLM agents: A UI/UX design framework for secure human oversight under prompt-injection attacks. International Journal of Graphic Design, 4(1), 186-191. https://doi.org/10.51903/ijgd.v4i1.3699
Sun, X., Lu, Y., & Chen, J. (2023). Controllable long-term user memory for multi-session dialogue: Confidence-gated writing, time-aware retrieval-augmented generation, and update/forgetting. Journal of Advanced Computing Systems, 3(8), 9-24. https://doi.org/10.69987/JACS.2023.30802
Sun, X., Zhong, Z. S., & Wu, Q. (2026). Retrieval-grounded HDFS log anomaly detection and deterministic failure narrative generation. Journal of Computer Systems and Applications, 3(1), 15-30. https://doi.org/10.64229/j6d7fr94
Tang, Y., & Yang, Y. (2024). MultiHop-RAG: Benchmarking retrieval-augmented generation for multi-hop queries. arXiv. https://doi.org/10.48550/arXiv.2401.15391
Trivedi, H., Balasubramanian, N., Khot, T., & Sabharwal, A. (2023). Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (pp. 10014–10037). ACL. https://doi.org/10.18653/v1/2023.acl-long.557
Tu, H., Zhao, S., & Zhou, A. (2025). Visual brief cards for advertising design: A structured UI/UX framework for turning creative intentions into graphic design decisions. International Journal of Graphic Design, 3(1), 210-226. https://doi.org/10.51903/ijgd.v3i1.3714
Wang, B., He, Y., Shui, Z., Xin, Q., & Lei, H. (2024). Predictive optimization of DDoS attack mitigation in distributed systems using machine learning. In Proceedings of the 6th International Conference on Computing and Data Science (CDS 2024) (pp. 89-94).
Wang, C., Wen, Z., Zhang, R., Xu, P., & Jiang, Y. (2025). GPU memory requirement prediction for deep learning tasks based on bidirectional gated recurrent unit optimization transformer. In 2025 5th International Conference on Artificial Intelligence, Virtual Reality and Visualization (AIVRV). IEEE. https://doi.org/10.1109/AIVRV67401.2025.11350369
Wang, L., Yang, N., Huang, X., Jiao, B., Yang, L., Jiang, D., Majumder, R., & Wei, F. (2022). Text embeddings by weakly-supervised contrastive pre-training. arXiv. https://doi.org/10.48550/arXiv.2212.03533
Wix.com. (2025). WixQA [Data set]. Hugging Face. https://huggingface.co/datasets/Wix/WixQA
Wu, Q., Meng, S., & Zhao, J. (2025). Text-grounded LLM-assisted design rationale interfaces: Turning advertising layout metadata into explainable UI/UX decision cards. International Journal of Graphic Design, 3(1), 216-240. https://doi.org/10.51903/ijgd.v3i1.3713
Xin, Q. (2025a). Explaining OpenStack failure-injection log anomalies with retrieved normal prototypes. Emerging Information Science and Technology, 6(2), 125-146. https://doi.org/10.18196/eist.v6i2.31232
Xin, Q. (2025b). Hybrid cloud architecture for efficient and cost-effective large language model deployment. Journal of Information Systems and Informatics, 7(3), 2182-2195. https://doi.org/10.51519/journalisi.v7i3.1170
Xin, Q. (2025c). Uncertainty-aware late fusion for 3D perception (confidence calibration + fusion rule learning). Journal of Technology Informatics and Engineering, 4(1), 215-238. https://doi.org/10.51903/jtie.v4i1.485
Xin, Q. (2026a). Auditable automated essay scoring and formative feedback: A rubric-grounded pipeline for secondary and higher education. Journal of Applied Artificial Intelligence in Education, 2(1), 1-19. https://doi.org/10.66053/jaaie.v2i1.348
Xin, Q. (2026b). Early-warning analytics with LLM intervention rationales for student retention decisions: Classroom interaction modeling with xAPI-Edu-Data and dropout/success prediction. Interdisciplinary Journal of Pedagogical Research and Media Technology, 2(1). https://doi.org/10.64268/inspire.v2i1.117
Xin, Q. (2026c). Explainable and fair credit risk scoring with counterfactual explanations: A reproducible evaluation on the German Credit dataset (HELOC-motivated). Journal of Information and Technology, 14(2), 215-231. https://doi.org/10.32664/j-intech.v14i02.2228
Xin, Q. (2026d). Host-based intrusion detection with system call sequences: Window localization and forensic narratives. Aviation Electronics, Information Technology, Telecommunications, Electricals, and Controls, 8(2). https://doi.org/10.28989/avitec.v8i2.3973
Xin, Q. (2026e). Log anomaly detection with conformal alert control and evidence-grounded incident ticket generation. Aviation Electronics, Information Technology, Telecommunications, Electricals, and Controls, 8(2), 247-264. https://doi.org/10.28989/avitec.v8i2.3974
Xin, Q. (2026f). Self-supervised customer representation learning for segmentation and next-purchase prediction on UCI Online Retail. Journal of Information and Technology, 14(1), 20-37. https://doi.org/10.32664/j-intech.v14i01.2229
Xin, Q. (2026g). Self-supervised log anomaly detection with LogBERT-style transformers: Full empirical evaluation on a reproducible SynHDFS benchmark. JEECS (Journal of Electrical Engineering and Computer Sciences), 11(1), 23-35. https://doi.org/10.54732/jeecs.v11i1.3
Xin, Q., Xu, Z., Guo, L., Zhao, F., & Wu, B. (2024). IoT traffic classification and anomaly detection method based on deep autoencoders. In Proceedings of the 6th International Conference on Computing and Data Science (CDS 2024).
Xu, H., Chen, Y., & Med, A. (2025). Automatic detection and explanation of dark patterns from interface microcopy: Empirical comparison of BERT-style encoders, RoBERTa-style encoders, and LLM-style decoders on the ec-darkpattern dataset. Journal of Technology Informatics and Engineering, 4(3), 590-612. https://doi.org/10.51903/jtie.v4i3.491
Xu, K., Zhou, H., Zheng, H., Zhu, M., & Xin, Q. (2024). Intelligent classification and personalized recommendation of e-commerce products based on machine learning. In Proceedings of the 6th International Conference on Computing and Data Science (ICCDS 2024).
Xu, X., Weytjens, H., Zhang, D., Lu, Q., Weber, I., & Zhu, L. (2025). RAGOps: An enterprise framework for LLM lifecycle management. arXiv. https://doi.org/10.48550/arXiv.2506.03401
Yao, S., Shinn, N., Razavi, P., & Narasimhan, K. (2024). τ-bench: A benchmark for tool-agent-user interaction in real-world domains. arXiv. https://doi.org/10.48550/arXiv.2406.12045
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct: Synergizing reasoning and acting in language models. The Eleventh International Conference on Learning Representations. https://openreview.net/forum?id=WE_vluYUL-X
Ye, T., Mu, J., & Hunter, J. (2026). Off-policy evaluation and conservative policy selection for slot-level dynamic bidding and ranking on the Open Bandit Dataset (Small). Journal of Technology Informatics and Engineering, 5(1), 178-199. https://doi.org/10.51903/jtie.v5i1.503
Zhang, B., Rao, H., & Zhao, D. (2024). Evidence-grounded RAG for cloud-native DevOps: Hallucination-resistant AIOps question answering over private operations documents. Journal of Advanced Computing Systems, 4(3), 109-125. https://doi.org/10.69987/JACS.2024.40308
Zhang, B., Ren, Y., & Zou, J. (2025). LLM-style explainable e-commerce recommendation cards: A UI/UX design framework for trust-calibrated product recommendation. International Journal of Graphic Design, 3(2), 381-396. https://doi.org/10.51903/ijgd.v3i2.3697
Zhang, B., Sun, X., Liu, G., & Zhou, B. (2026). LLM-style DevOps copilot for cloud-native troubleshooting: Retrieval-augmented runbook generation and command-safety evaluation. Journal of Technology Informatics and Engineering, 5(2), 104-118. https://doi.org/10.51903/jtie.v5i2.534
Zhang, K., Chen, Y., & Qian, A. (2025). Evidence-grounded accounting disclosure review cards: A visual communication framework for LLM-style explanations over SEC financial statements and notes. International Journal of Graphic Design, 3(2), 395. https://doi.org/10.51903/ijgd.v3i2.3710
Zhang, R., Wen, Z., Wang, C., Tang, C., Xu, P., & Jiang, Y. (2025). Quality analysis and evaluation prediction of RAG retrieval based on machine learning algorithms. arXiv. https://doi.org/10.48550/arXiv.2511.19481
Zhang, Y., & Zhang, H. (2025a). A therapist-facing session copilot for live counseling support: Reasoning-guided retrieval and ranking from multi-turn counseling dialogues. Journal of Technology Informatics and Engineering, 4(2), 464-486. https://doi.org/10.51903/jtie.v4i2.547
Zhang, Y., & Zhang, H. (2025b). Visualizing the right counseling support: Evidence-linked recommendation cards for explainable mental health intake interfaces. International Journal of Graphic Design, 3(1), 214-229. https://doi.org/10.51903/ijgd.v3i1.3722
Zhang, Y., & Zhou, Z. (2026). Strategy-aware therapist imitation for emotional support dialogues: A reproducible ESConv study for LLM response control. Advances in Educational Technology and Psychology, 10(2), 92-97. https://doi.org/10.23977/aetp.2026.100213
Zhang, Y., Liao, Q. V., & Bellamy, R. K. E. (2020). Effect of confidence and explanation on accuracy and trust calibration in AI-assisted decision making. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency (pp. 295–305). ACM. https://doi.org/10.1145/3351095.3372852
Zhao, S., Bai, J., & Roberson, D. (2025). Multi-horizon GPU demand forecasting with workload semantics and operational risk curves: An empirical study on Alibaba Clusterdata GPU trace. Journal of Technology Informatics and Engineering, 4(3), 544-571. https://doi.org/10.51903/jtie.v4i3.498
Zhao, S., Ren, Y., & Chang, X. (2026). Profit-aware spot GPU admission control with cost-sensitive loss and evidence-grounded policy memos for AI workload supply-demand matching. Journal of Technology Informatics and Engineering, 5(2), 45-59. https://doi.org/10.51903/jtie.v5i2.545
Zheng, D., & Li, C. (2024). Behavior-level jailbreak resistance via multi-stage refusal and utility preservation. Journal of Advanced Computing Systems, 4(1), 83-99. https://doi.org/10.69987/JACS.2024.40107
Zheng, D., Li, C., & Davidson, H. (2023). Continual red-teaming for in-the-wild jailbreaks via online guardrail updates and guardrail distillation. Journal of Advanced Computing Systems, 3(2), 35-49. https://doi.org/10.69987/JACS.2023.30203
Zheng, D., Zhang, B., & Geibel, J. (2024). VerifySafe: Toxicity-safe agent responses under adversarial prompts with evidence-based self-verification. Journal of Advanced Computing Systems, 4(1), 67-82. https://doi.org/10.69987/JACS.2024.40106
Zhong, Z. S., & Ling, S. (2024). Improved theoretical guarantee for rank aggregation via spectral method. Information and Inference: A Journal of the IMA, 13(3), iaae020. https://doi.org/10.1093/imaiai/iaae020
Zhong, Z. S., Chen, J., Zhong, E., & Sun, X. (2025). Evidence-calibrated RAG for unanswerable question answering: Retrieval coverage, abstention calibration, and hallucination-proxy analysis on SQuAD 2.0. Journal of Technology Informatics and Engineering, 4(2), 502-520. https://doi.org/10.51903/jtie.v4i2.536
Zhong, Z. S., Li, C., & Rao, H. (2026). Trajectory reliability prediction for generalist AI agents: Tool-use failure analysis and success forecasting on ZClawBench. Journal of Technology Informatics and Engineering, 5(1), 341-360. https://doi.org/10.51903/jtie.v5i1.539
Zhong, Z. S., Pan, X., & Lei, Q. (2025). Bridging domains with approximately shared features. In Proceedings of the 28th International Conference on Artificial Intelligence and Statistics (Vol. 258, pp. 559-567). PMLR. https://proceedings.mlr.press/v258/zhong25a.html
Zhou, B., Jin, J., & Zhao, D. (2025). Calibrated resume-job matching for trustworthy LLM-assisted recruiter screening: Pairwise matching, probability calibration, and selective refusal on two public recruitment datasets. Journal of Technology Informatics and Engineering, 4(3), 625-648. https://doi.org/10.51903/jtie.v4i3.529
Zhou, B., Wang, H., & Chang, X. (2025). Risk-calibrated patient-facing AI safety cards: A UI/UX design framework for rubric-based medical risk communication. International Journal of Graphic Design, 3(2), 365-380. https://doi.org/10.51903/ijgd.v3i2.3696
Zhou, S., Chen, Y., & Lee, K. (2026). Accounting-aware evidence-constrained agents for disclosure, settlement, and secondary-market risk monitoring in tokenized RWA infrastructure. Journal of Technology Informatics and Engineering, 5(2), 60-74. https://doi.org/10.51903/jtie.v5i2.544
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Kevin Wu (Author)

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.