Continual Red-Teaming and Guardrail Distillation for Tool-Using LLM Agents: Prompt-Injection Resistance with Utility Preservation

Authors

  • Wesley Gao Author

DOI:

https://doi.org/10.61424/hewtat90

Keywords:

Tool-using agents; indirect prompt injection; AgentDojo; action gating; guardrail distillation; continual red teaming; security–utility trade-off.

Abstract

Tool-using language-model agents can convert indirect prompt injection into consequential actions, making guardrail quality a joint security, utility, and efficiency problem. This study evaluates a ReAct-style control, native tool filtering, deterministic self-verification, a calibrated action gate, and a distilled guardrail on AgentDojo v0.1.22. The evaluation covers 97 benign tasks, 629 canonical attack cases, ten attack formulations, and 32,456 recorded actions, of which 20,953 were assigned labels under the stated rules and 1,767 were unsafe. Archived end-to-end runs and fail-closed state replay were analyzed separately. The ReAct control reached 47.69% attack success, 50.08% attacked utility, and 69.07% benign utility. Native tool filtering reduced attack success to 6.84% while raising attacked and benign utility to 56.28% and 72.16%. The learned action gate produced a 0.00–0.95% attack-success bound, but attacked utility fell to 18.92–22.58% and false refusal rose to 46.27–47.76%. Distillation reduced serialized size from 146.40 to 3.52 KiB and probability-scoring time from 358.12 to 1.31 microseconds per action, yet lowered AUROC from 0.9971 to 0.9009 and produced a 71.64–73.13% false-refusal bound. Continual rehearsal also failed to retain protection: final seen-family unsafe recall was 21.71% with a 70.54-point peak-to-current seen-recall loss. The results show that compactness and low attack success do not establish a useful guardrail unless authorization fidelity and retention are measured concurrently.

References

Andriushchenko, M., Souly, A., Dziemian, M., Duenas, D., Lin, M., Wang, J., Hendrycks, D., Zou, A., Kolter, Z., Fredrikson, M., Gal, Y., & Davies, X. (2025). AgentHarm: A benchmark for measuring harmfulness of LLM agents. In The Thirteenth International Conference on Learning Representations. https://openreview.net/forum?id=AC5n7xHuR1

Bai, J., Chen, S., Zheng, D., & Kuo, M.-J. (2026). Interpretable attack-chain stage detection from AWS CloudTrail event sequences via linear models and HMM smoothing. Informatics, Electrical and Electronics Engineering (Infotron), 6(1), 28⁠–⁠43. https://doi.org/10.33474/infotron.v6i1.24923

Chang, X., Lu, Y., & Zhong, Z. S. (2026). Review-grounded explainable recommendation with faithfulness evaluation on Amazon Reviews. JEECS (Journal of Electrical Engineering and Computer Sciences), 11(1), 9⁠–⁠22. https://doi.org/10.54732/jeecs.v11i1.2

Chen, S., He, S., & Sun, E. (2024). Risk-bounded GPU resource oversubscription via conformal demand envelopes in production AI clusters. Journal of Advanced Computing Systems, 4(5), 119⁠–⁠134. https://doi.org/10.69987/JACS.2024.40509

Chen, S., Piet, J., Sitawarin, C., & Wagner, D. (2025). StruQ: Defending against prompt injection with structured queries. In 34th USENIX Security Symposium (USENIX Security 25) (pp. 2383⁠–⁠2400). USENIX Association. https://www.usenix.org/conference/usenixsecurity25/presentation/chen-sizhe

Chen, S., Zharmagambetov, A., Mahloujifar, S., Chaudhuri, K., Wagner, D., & Guo, C. (2025). SecAlign: Defending against prompt injection with preference optimization. In Proceedings of the 2025 ACM SIGSAC Conference on Computer and Communications Security (pp. 2833⁠–⁠2847). Association for Computing Machinery. https://doi.org/10.1145/3719027.3744836

Chen, Y., & Xu, H. (2026). Trust-calibrated multilingual RAG for humanitarian information platforms: Empirical evaluation on OMoS-QA for migration information access. International Journal of Graphic Design, 4(1), 141⁠–⁠164. https://doi.org/10.51903/ijgd.v4i1.3552

Chen, Y., Zhang, Y., & Sherman, M. (2024). Going concern and bankruptcy prediction under extreme class imbalance: Cost-sensitive learning, resampling, and focal loss with explainable financial-ratio portraits. Journal of Advanced Computing Systems, 4(4), 80⁠–⁠96. https://doi.org/10.69987/JACS.2024.40407

Chen, Y., Zhang, Y., Chau, D., & Sherman, M. (2023). Credit card default risk tiering with probability calibration and uncertainty-driven rejection: A reproducible study on the UCI Credit Card Clients dataset. Journal of Advanced Computing Systems, 3(4), 31⁠–⁠47. https://doi.org/10.69987/JACS.2023.30403

Chen, Y., Zhou, S., & Lin, E. (2025). Accounting-aware evidence retrieval for institutional due diligence of tokenized trade receivable RWA. Journal of Technology Informatics and Engineering, 4(3). https://doi.org/10.51903/jtie.v4i3.542

Debenedetti, E., Shumailov, I., Fan, T., Hayes, J., Carlini, N., Fabian, D., Kern, C., Shi, C., Terzis, A., & Tramèr, F. (2025). Defeating prompt injections by design. arXiv. https://doi.org/10.48550/arXiv.2503.18813

Debenedetti, E., Zhang, J., Balunović, M., Beurer-Kellner, L., Fischer, M., & Tramèr, F. (2024). AgentDojo: A dynamic environment to evaluate prompt injection attacks and defenses for LLM agents. In The Thirty-Eighth Conference on Neural Information Processing Systems Datasets and Benchmarks Track. https://openreview.net/forum?id=m1YYAQjO3w

Di Palo, F., Singhi, P., & Fadlallah, B. H. (2024). Performance-guided LLM knowledge distillation for efficient text classification at scale. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 3675⁠–⁠3687). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.emnlp-main.215

Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., & Fritz, M. (2023). Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (pp. 79⁠–⁠90). Association for Computing Machinery. https://doi.org/10.1145/3605764.3623985

Han, S., Rao, K., Ettinger, A., Jiang, L., Lin, B. Y., Lambert, N., Choi, Y., & Dziri, N. (2024). WildGuard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of LLMs. In Advances in Neural Information Processing Systems (Vol. 37, pp. 8093⁠–⁠8131). https://doi.org/10.52202/079017-0261

He, S., Chang, X., & Sun, E. (2024). Cross-cloud transfer learning for AI training capacity forecasting under workload and topology distribution shift. Journal of Advanced Computing Systems, 4(1), 100⁠–⁠120. https://doi.org/10.69987/JACS.2024.40108

He, S., Li, C., & Rao, H. (2025). Few-shot cold-start workload forecasting for new AI inference tenants with time-series foundation models. Journal of Technology Informatics and Engineering, 4(1), 306⁠–⁠324. https://doi.org/10.51903/jtie.v4i1.546

He, S., Nie, J., & Li, C. (2026). Power-aware inventory planning for AI infrastructure using job-level forecasting and LLM workload explanations. Journal of Technology Informatics and Engineering, 5(1). https://doi.org/10.51903/jtie.v5i1.548

He, S., Tu, H., & Liu, I. (2023). Safe PD capacity forecasting with time-series foundation models and calibrated uncertainty for heterogeneous GPU clusters. Journal of Advanced Computing Systems, 3(4), 48⁠–⁠66. https://doi.org/10.69987/JACS.2023.30404

Hines, K., Lopez, G., Hall, M., Zarfati, F., Zunger, Y., & Kiciman, E. (2024). Defending against indirect prompt injection attacks with spotlighting. arXiv. https://doi.org/10.48550/arXiv.2403.14720

Inan, H., Upasani, K., Chi, J., Rungta, R., Iyer, K., Mao, Y., Tontchev, M., Hu, Q., Fuller, B., Testuggine, D., & Khabsa, M. (2023). Llama Guard: LLM-based input-output safeguard for human-AI conversations. arXiv. https://doi.org/10.48550/arXiv.2312.06674

Jia, F., Wu, T., Qin, X., & Squicciarini, A. (2025). The Task Shield: Enforcing task alignment to defend against indirect prompt injection in LLM agents. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 29680⁠–⁠29697). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.acl-long.1435

Jin, J. (2025a). Calibrated resume-job matching for trustworthy LLM-assisted recruiter screening: Pairwise matching, probability calibration, and selective refusal on two public recruitment datasets. Journal of Technology Informatics and Engineering, 4(3), 625⁠–⁠648. https://doi.org/10.51903/jtie.v4i3.529

Jin, J. (2025b). Evidence-chain reliable RAG: Hallucination detection, source attribution, and deterministic provenance explanations. Journal of Technology Informatics and Engineering, 4(2). https://doi.org/10.51903/jtie.v4i2.535

Jin, J. (2025c). LLM-style evidence cards for scientific search interfaces: A UI/UX design framework for retrieval transparency, ranking trust, and visual evidence hierarchy. International Journal of Graphic Design, 3(2), 397⁠–⁠414. https://doi.org/10.51903/ijgd.v3i2.3698

Jin, J., Huang, T., & Lu, S. (2024a). Cost-sensitive learning, simulated PU learning, and one-class autoencoding for extreme-imbalance credit card fraud detection. Journal of Advanced Computing Systems, 4(6), 64⁠–⁠73. https://doi.org/10.69987/JACS.2024.40605

Jin, J., Huang, T., & Lu, S. (2024b). A model-risk-friendly probability of default workflow: Calibration, distribution-free uncertainty quantification, and SHAP explanations on the UCI Credit Card Default dataset. Journal of Advanced Computing Systems, 4(6), 74⁠–⁠85. https://doi.org/10.69987/JACS.2024.40606

Kuo, M.-J., Zheng, D., & Hires, J. (2025). Federated topic-preference learning for knowledge-grounded chat with differential privacy. Journal of Technology Informatics and Engineering, 4(2), 385⁠–⁠401. https://doi.org/10.51903/jtie.v4i2.502

Li, C., Bai, J., & Wang, S. (2024). Evidence-chain reliable RAG: Word-level hallucination detection, source attribution, and provenance explanation for LLM applications. Journal of Advanced Computing Systems, 4(2), 76⁠–⁠92. https://doi.org/10.69987/JACS.2024.40207

Li, C., Liu, G., & Zhao, Z. (2026). Cost-aware LLM-style routing for AIOps log analysis: Log parsing, anomaly detection, fault diagnosis, and incident summarization on LogEval task files. Journal of Technology Informatics and Engineering, 5(2), 91⁠–⁠103. https://doi.org/10.51903/jtie.v5i2.538

Li, C., Zhou, B., & Gao, K. (2025). Risk-calibrated patient-facing AI safety cards: A UI/UX benchmark for explainable medical AI response interfaces. International Journal of Graphic Design, 3(2). https://doi.org/10.51903/ijgd.v3i2.3709

Li, H., Liu, X., Zhang, N., & Xiao, C. (2025). PIGuard: Prompt injection guardrail via mitigating overdefense for free. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 30420⁠–⁠30437). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.acl-long.1468

Li, J., & Zhou, A. (2026). Multi-regulation RAG for AI product counsel: A legal governance framework for cross-border digital commerces. Rule of Law Studies Journal, 2(2), 105⁠–⁠123. https://doi.org/10.64780/rolsj.v2i2.225

Li, Y. (2024). Findable then explainable: Retrieval–summary integration for code intelligence on a lightweight CodeSearchNet subset. Journal of Advanced Computing Systems, 4(7), 65⁠–⁠82. https://doi.org/10.69987/JACS.2024.40706

Li, Y., & Lu, S. (2025). Language-guided feature selection for DDoS and intrusion detection on CICIDS2017. Journal of Technology Informatics and Engineering, 4(1), 284⁠–⁠305. https://doi.org/10.51903/jtie.v4i1.531

Li, Z., Zhang, K., & Wong, A. (2026). Numerical-reasoning guardrails for a quant research assistant: A compact reproducible benchmark using SEC and FRED data. Journal of Technology Informatics and Engineering, 5(2), 75⁠–⁠90. https://doi.org/10.51903/jtie.v5i2.541

Li, Z., Zhou, S., & Zhou, Z. (2025). Financial risk dashboard design for institutional RWA investors: Visual hierarchy, chart comprehension, and explainability in FinChart-Bench. International Journal of Graphic Design, 3(1). https://doi.org/10.51903/ijgd.v3i1.3715

Liu, G., He, S., & Liu, I. (2023). LLM-augmented multi-source root cause attribution for CPU and network faults in microservices. Journal of Advanced Computing Systems, 3(6), 39⁠–⁠57. https://doi.org/10.69987/JACS.2023.30604

Liu, G., He, S., & Wong, H. (2025). LLM-compatible visual brief cards for AI infrastructure capacity dashboards: A UI/UX framework for turning forecast risk into graphic design decisions. International Journal of Graphic Design, 3(1). https://doi.org/10.51903/ijgd.v3i1.3723

Liu, G., Li, C., & Zhang, E. (2024). OpsLLM for cloud incident triage: Bilingual RAG-based root cause analysis and alert summarization for AI infrastructure operations. Journal of Advanced Computing Systems, 4(4), 97⁠–⁠111. https://doi.org/10.69987/JACS.2024.40408

Liu, Y., Deng, G., Li, Y., Wang, K., Wang, Z., Wang, X., Zhang, T., Liu, Y., Wang, H., Zheng, Y., & Liu, Y. (2023). Prompt injection attack against LLM-integrated applications. arXiv. https://doi.org/10.48550/arXiv.2306.05499

Liu, Y., Jia, Y., Geng, R., Jia, J., & Gong, N. Z. (2024). Formalizing and benchmarking prompt injection attacks and defenses. In 33rd USENIX Security Symposium (USENIX Security 24) (pp. 1831⁠–⁠1847). USENIX Association. https://www.usenix.org/conference/usenixsecurity24/presentation/liu-yupei

Lu, S., & Zhou, D. (2024). TinyLLM-assisted intrusion detection for real-time IoT networks. Journal of Advanced Computing Systems, 4(8), 72⁠–⁠87. https://doi.org/10.69987/JACS.2024.40809

Mazeika, M., Phan, L., Yin, X., Zou, A., Wang, Z., Mu, N., Sakhaee, E., Li, N., Basart, S., Li, B., Forsyth, D., & Hendrycks, D. (2024). HarmBench: A standardized evaluation framework for automated red teaming and robust refusal. Proceedings of Machine Learning Research, 235, 35181⁠–⁠35224. https://proceedings.mlr.press/v235/mazeika24a.html

Meng, S., Chen, J., & Zheng, I. (2026). LLM-inspired offline reranking for financial search: Query rewriting, hybrid retrieval, and listwise relevance ranking on FiQA. Journal of Technology Informatics and Engineering, 5(1), 361⁠–⁠378. https://doi.org/10.51903/jtie.v5i1.537

Mi, G., Ye, T., & Wood, D. (2025). A lightweight medical foundation model for cross-modal multi-task pretraining and parameter-efficient few-shot transfer on MedMNIST. Journal of Technology Informatics and Engineering, 4(3), 572⁠–⁠589. https://doi.org/10.51903/jtie.v4i3.492

Mu, J., Ye, T., & Patel, P. (2025). Offline counterfactual evaluation for advertising and recommendation slot policies: A reproducible study on the Open Bandit Dataset (Small). Journal of Technology Informatics and Engineering, 4(3), 521⁠–⁠543. https://doi.org/10.51903/jtie.v4i3.500

Nie, J., & Zheng, D. (2024). Noisy-neighbor-aware VM degradation risk modeling with unsupervised residual fusion. Journal of Advanced Computing Systems, 4(4), 112⁠–⁠123. https://doi.org/10.69987/JACS.2024.40409

Nie, J., Liu, G., Li, C., & Zou, T. (2026). Evidence-constrained incident visualization cards for distributed cloud logs: A UI/UX framework for turning Hadoop, OpenStack, and ZooKeeper logs into actionable SRE design interfaces. International Journal of Graphic Design, 4(1), 179⁠–⁠185. https://doi.org/10.51903/ijgd.v4i1.3703

Perez, E., Huang, S., Song, F., Cai, T., Ring, R., Aslanides, J., Glaese, A., McAleese, N., & Irving, G. (2022). Red teaming language models with language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (pp. 3419⁠–⁠3448). Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.emnlp-main.225

Perez, F., & Ribeiro, I. (2022). Ignore previous prompt: Attack techniques for language models. arXiv. https://doi.org/10.48550/arXiv.2211.09527

Piet, J., Alrashed, M., Sitawarin, C., Chen, S., Wei, Z., Sun, E., Alomair, B., & Wagner, D. A. (2024). Jatmo: Prompt injection defense by task-specific finetuning. In J. García-Alfaro, R. Kozik, M. Choraś, S. K. Katsikas, & G. Rios (Eds.), Computer security—ESORICS 2024 (Lecture Notes in Computer Science, Vol. 14982, pp. 105⁠–⁠124). Springer. https://doi.org/10.1007/978-3-031-70879-4_6

Ruan, Y., Dong, H., Wang, A., Pitis, S., Zhou, Y., Ba, J., Dubois, Y., Maddison, C. J., & Hashimoto, T. B. (2024). Identifying the risks of LM agents with an LM-emulated sandbox. In The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=GEcwtMk1uA

Su, W., Chen, S., & Qian, E. (2026). Narrative-aware scientific claim verification agent with evidence ranking for ClimateCheck. Journal of Technology Informatics and Engineering, 5(1), 327⁠–⁠340. https://doi.org/10.51903/jtie.v5i1.549

Su, W., Chen, S., & Zhao, C. (2025). Budgeted multi-hop retrieval agent for compositional question answering: A retrieval-policy evaluation on the official MultiHop-RAG benchmark. Journal of Technology Informatics and Engineering, 4(3). https://doi.org/10.51903/jtie.v4i3.543

Su, W., Rao, H., & Ma, E. (2026). Privacy and data-integrity risk cards for LLM agents: A UI/UX design framework for secure human oversight under prompt-injection attacks. International Journal of Graphic Design, 4(1), 186⁠–⁠191. https://doi.org/10.51903/ijgd.v4i1.3699

Sun, X., Lu, Y., & Chen, J. (2023). Controllable long-term user memory for multi-session dialogue: Confidence-gated writing, time-aware retrieval-augmented generation, and update/forgetting. Journal of Advanced Computing Systems, 3(8), 9⁠–⁠24. https://doi.org/10.69987/JACS.2023.30802

Sun, X., Zhong, Z. S., & Wu, Q. (2026). Retrieval-grounded HDFS log anomaly detection and deterministic failure narrative generation. Journal of Computational Systems and Applications, 3(1), 15⁠–⁠30. https://doi.org/10.64229/j6d7fr94

Wallace, E., Xiao, K., Leike, R., Weng, L., Heidecke, J., & Beutel, A. (2024). The instruction hierarchy: Training LLMs to prioritize privileged instructions. arXiv. https://doi.org/10.48550/arXiv.2404.13208

Wang, B., He, Y., Shui, Z., Xin, Q., & Lei, H. (2024). Predictive optimization of DDoS attack mitigation in distributed systems using machine learning. Applied and Computational Engineering, 64(1), 89⁠–⁠94. https://doi.org/10.54254/2755-2721/64/20241350

Wang, C., Wen, Z., Zhang, R., Xu, P., & Jiang, Y. (2025). GPU memory requirement prediction for deep learning task based on bidirectional gated recurrent unit optimization Transformer. In 2025 5th International Conference on Artificial Intelligence, Virtual Reality and Visualization (AIVRV). IEEE. https://doi.org/10.1109/AIVRV67401.2025.11350369

Wu, Q., Mi, G., & Wood, D. (2025). Calibration-light subject-independent motor imagery BCI via self-supervised pretraining and Conformer. Journal of Technology Informatics and Engineering, 4(1), 239⁠–⁠262. https://doi.org/10.51903/jtie.v5i1.493

Xin, Q. (2025a). Explaining OpenStack failure-injection log anomalies with retrieved normal prototypes. International Journal of Emerging Information Science and Technology, 6(2), 125⁠–⁠146. https://doi.org/10.18196/eist.v6i2.31232

Xin, Q. (2025b). Hybrid cloud architecture for efficient and cost-effective large language model deployment. Journal of Information Systems and Informatics, 7(3), 2182⁠–⁠2195. https://doi.org/10.51519/journalisi.v7i3.1170

Xin, Q. (2025c). Uncertainty-aware late fusion for 3D perception (confidence calibration + fusion rule learning). Journal of Technology Informatics and Engineering, 4(1), 215⁠–⁠238. https://doi.org/10.51903/jtie.v4i1.485

Xin, Q. (2026a). Auditable automated essay scoring and formative feedback: A rubric-grounded pipeline for secondary and higher education. Journal of Applied Artificial Intelligence in Education, 2(1), 1⁠–⁠19. https://doi.org/10.66053/jaaie.v2i1.348

Xin, Q. (2026b). Explainable and fair credit risk scoring with counterfactual explanations: A reproducible evaluation on the German Credit dataset (HELOC-motivated). Journal of Information and Technology, 14(2), 215⁠–⁠231. https://doi.org/10.32664/j-intech.v14i02.2228

Xin, Q. (2026c). Host-based intrusion detection with system call sequences: Window localization and forensic narratives. Aviation Electronics, Information Technology, Telecommunications, Electricals, and Controls (AVITEC), 8(2), 325⁠–⁠334. https://doi.org/10.28989/avitec.v8i2.3973

Xin, Q. (2026d). LiDAR–camera object-level fusion for multi-target tracking using JPDA and EKF: A reproducible empirical study on a PandaSet-parameterised five-sequence dataset. Journal of Technology Informatics and Engineering, 5(1), 54⁠–⁠76. https://doi.org/10.51903/jtie.v5i1.486

Xin, Q. (2026e). Log anomaly detection with conformal alert control and evidence-grounded incident ticket generation. Aviation Electronics, Information Technology, Telecommunications, Electricals, and Controls (AVITEC), 8(2), 247⁠–⁠264. https://doi.org/10.28989/avitec.v8i2.3974

Xin, Q. (2026f). Probabilistic bike-sharing demand forecasting under changing weather and seasonal regimes with transformer-based models. Transport Findings. https://doi.org/10.32866/001c.157499

Xin, Q. (2026g). Self-supervised log anomaly detection with LogBERT-style transformers: Full empirical evaluation on a reproducible SynHDFS benchmark. JEECS (Journal of Electrical Engineering and Computer Sciences), 11(1), 23⁠–⁠35. https://doi.org/10.54732/jeecs.v11i1.3

Xin, Q., Xu, Z., Guo, L., Zhao, F., & Wu, B. (2024). IoT traffic classification and anomaly detection method based on deep autoencoders. Applied and Computational Engineering, 69(1), 64⁠–⁠70. https://doi.org/10.54254/2755-2721/69/20241511

Xu, H., Chen, Y., & Med, A. (2025). Automatic detection and explanation of dark patterns from interface microcopy: Empirical comparison of BERT-style encoders, RoBERTa-style encoders, and LLM-style decoders on the ec-darkpattern dataset. Journal of Technology Informatics and Engineering, 4(3), 590⁠–⁠612. https://doi.org/10.51903/jtie.v4i3.491

Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct: Synergizing reasoning and acting in language models. In The Eleventh International Conference on Learning Representations. https://openreview.net/forum?id=WE_vluYUL-X

Ye, T., Chang, X., & Zhong, E. (2025). Uncertainty-aware breast ultrasound explanation cards: A visual communication framework for image-based AI diagnostic support using BreastMNIST_224. International Journal of Graphic Design, 3(2). https://doi.org/10.51903/ijgd.v3i2.3701

Ye, T., Mu, J., & Hunter, J. (2026). Off-policy evaluation and conservative policy selection for slot-level dynamic bidding and ranking on the Open Bandit Dataset (Small). Journal of Technology Informatics and Engineering, 5(1), 178⁠–⁠199. https://doi.org/10.51903/jtie.v5i1.503

Yi, J., Xie, Y., Zhu, B., Kiciman, E., Sun, G., Xie, X., & Wu, F. (2025). Benchmarking and defending against indirect prompt injection attacks on large language models. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 1 (pp. 1809⁠–⁠1820). Association for Computing Machinery. https://doi.org/10.1145/3690624.3709179

Zhan, Q., Fang, R., Panchal, H. S., & Kang, D. (2025). Adaptive attacks break defenses against indirect prompt injection attacks on LLM agents. In Findings of the Association for Computational Linguistics: NAACL 2025 (pp. 7116⁠–⁠7132). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.findings-naacl.395

Zhan, Q., Liang, Z., Ying, Z., & Kang, D. (2024). InjecAgent: Benchmarking indirect prompt injections in tool-integrated large language model agents. In Findings of the Association for Computational Linguistics: ACL 2024 (pp. 10471⁠–⁠10506). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-acl.624

Zhang, B., Rao, H., & Zhao, D. (2024). Evidence-grounded RAG for cloud-native DevOps: Hallucination-resistant AIOps question answering over private operations documents. Journal of Advanced Computing Systems, 4(3), 109⁠–⁠125. https://doi.org/10.69987/JACS.2024.40308

Zhang, B., Sun, X., Liu, G., & Zhou, B. (2026). LLM-style DevOps copilot for cloud-native troubleshooting: Retrieval-augmented runbook generation and command-safety evaluation. Journal of Technology Informatics and Engineering, 5(2), 104⁠–⁠118. https://doi.org/10.51903/jtie.v5i2.534

Zhang, L., Ma, R., & Greg, P. (2025). Digital-twin dispatching for urban mobility via spatio-temporal transformers and offline reinforcement learning. Journal of Technology Informatics and Engineering, 4(2), 337⁠–⁠363. https://doi.org/10.51903/jtie.v4i2.501

Zhang, R., Wen, Z., Wang, C., Tang, C., Xu, P., & Jiang, Y. (2025). Quality analysis and evaluation prediction of RAG retrieval based on machine learning algorithms. arXiv. https://doi.org/10.48550/arXiv.2511.19481

Zhang, Y., & Zhang, H. (2025a). A therapist-facing session copilot for live counseling support: Reasoning-guided retrieval and ranking from multi-turn counseling dialogues. Journal of Technology Informatics and Engineering, 4(2), 464⁠–⁠486. https://doi.org/10.51903/jtie.v4i2.547

Zhang, Y., & Zhang, H. (2025b). Visualizing the right counseling support: Evidence-linked recommendation cards for explainable mental health intake interfaces. International Journal of Graphic Design, 3(1). https://doi.org/10.51903/ijgd.v3i1.3722

Zhang, Y., & Zhou, Z. (2026). Strategy-aware therapist imitation for emotional support dialogues: A reproducible ESConv study for LLM response control. Advances in Educational Technology and Psychology, 10(2), 92⁠–⁠97. https://doi.org/10.23977/aetp.2026.100213

Zhao, S., Bai, J., & Roberson, D. (2025). Multi-horizon GPU demand forecasting with workload semantics and operational risk curves: An empirical study on Alibaba Clusterdata GPU Trace. Journal of Technology Informatics and Engineering, 4(3), 544⁠–⁠571. https://doi.org/10.51903/jtie.v4i3.498

Zhao, S., Ren, Y., & Chang, X. (2026). Profit-aware spot GPU admission control with cost-sensitive loss and evidence-grounded policy memos for AI workload supply-demand matching. Journal of Technology Informatics and Engineering, 5(2), 45⁠–⁠59. https://doi.org/10.51903/jtie.v5i2.545

Zheng, D., & Li, C. (2024). Behavior-level jailbreak resistance via multi-stage refusal + utility preservation. Journal of Advanced Computing Systems, 4(1), 83⁠–⁠99. https://doi.org/10.69987/JACS.2024.40107

Zheng, D., Li, C., & Davidson, H. (2023). Continual red-teaming for in-the-wild jailbreaks via online guardrail updates and guardrail distillation. Journal of Advanced Computing Systems, 3(2), 35⁠–⁠49. https://doi.org/10.69987/JACS.2023.30203

Zheng, D., Zhang, B., & Geibel, J. (2024). VerifySafe: Toxicity-safe agent responses under adversarial prompts with evidence-based self-verification. Journal of Advanced Computing Systems, 4(1), 67⁠–⁠82. https://doi.org/10.69987/JACS.2024.40106

Zhong, Z. S., Chen, J., Zhong, E., & Sun, X. (2025). Evidence-calibrated RAG for unanswerable question answering: Retrieval coverage, abstention calibration, and hallucination-proxy analysis on SQuAD 2.0. Journal of Technology Informatics and Engineering, 4(2), 502⁠–⁠520. https://doi.org/10.51903/jtie.v4i2.536

Zhong, Z. S., Li, C., & Rao, H. (2026). Trajectory reliability prediction for generalist AI agents: Tool-use failure analysis and success forecasting on ZClawBench. Journal of Technology Informatics and Engineering, 5(1). https://doi.org/10.51903/jtie.v5i1.539

Zhong, Z. S., Pan, X., & Lei, Q. (2025). Bridging domains with approximately shared features. Proceedings of Machine Learning Research, 258, 559⁠–⁠567. https://proceedings.mlr.press/v258/zhong25a.html

Zhong, Z. S., Wu, Q., & Mi, G. (2025). Uncertainty-aware medical image explanation cards: LLM-generated visual explanations for AI-assisted radiology interfaces. International Journal of Graphic Design, 3(2), 415⁠–⁠436. https://doi.org/10.51903/ijgd.v3i2.3616

Zhou, B., Li, C., & Liu, L. (2025). Risk-calibrated patient-facing AI safety cards: A UI/UX design framework for rubric-based medical risk communication. International Journal of Graphic Design, 3(2). https://doi.org/10.51903/ijgd.v3i2.3696

Zhou, B., Wang, H., & Chang, X. (2025). Distilling VMAF into an edge-deployable quality predictor: A pilot shot-level proxy with LLM-ready quality tokens. Journal of Technology Informatics and Engineering, 4(2), 447⁠–⁠463. https://doi.org/10.51903/jtie.v4i2.522

Zhou, H., & Zhang, K. (2025). News-based uncertainty and macro-market fusion for VIX direction forecasting: Evidence from 2015⁠–⁠2024 FRED panel. Journal of Technology Informatics and Engineering, 4(2), 487⁠–⁠501. https://doi.org/10.51903/jtie.v4i2.540

Zhu, K., Yang, X., Wang, J., Guo, W., & Wang, W. Y. (2025). MELON: Provable defense against indirect prompt injection attacks in AI agents. Proceedings of Machine Learning Research, 267, 80310⁠–⁠80329. https://proceedings.mlr.press/v267/zhu25z.html

Downloads

Published

2026-07-31