Controllable without Stereotyping: Cross-Context Stability, Utility Preservation, and Uncertainty in Big-Five-Conditioned LLM Dialogue
DOI:
https://doi.org/10.61424/gwmd8w43Keywords:
Big Five personality, large language models, persona control, response selection, calibration, selective abstention, personality caricatureAbstract
Personality-controllable dialogue systems must express intended traits without losing contextual fidelity or collapsing into formulaic cues. We evaluated 100,000 BIG5-CHAT responses representing 10,000 scenarios crossed with five Big Five dimensions and high and low levels. After listwise removal of 64 incomplete scenarios, 9,936 scenarios remained. Scenario-grouped training, validation, and test partitions contained 6,955, 1,490, and 1,491 scenarios. Three locally fitted response-selection strategies were compared: a fixed cue-based persona-prompt proxy, a word-TF-IDF weighted-nearest-neighbor persona retriever, and a character-ngram logistic lightweight adapter. Evaluation combined response-level discrimination, paired high-versus-low ordering, calibration, relation-wise stability, selective abstention, content-fidelity proxies, cue and name masking, length controls, and cross-trait entanglement. On the held-out test set, paired accuracy was 0.7797, 0.9896, and 0.9935 for prompt, retrieval, and adapter, respectively; corresponding AUROC values were 0.7993, 0.9927, and 0.9958. The adapter achieved the lowest Brier score (0.0234), while retrieval had the lowest expected calibration error (0.0066). Worst-relation paired accuracy remained 0.7602, 0.9824, and 0.9889. At 90% coverage, selection risk fell from 0.1892 for the prompt proxy to 0.0003 for retrieval and 0.0000 for the adapter. However, length alone ordered 67.3% of pairs, and cross-trait entanglement was higher for retrieval (0.2712) and adaptation (0.2897) than for prompting (0.0933). The findings demonstrate strong, stable corpus-level persona discrimination but also expose a control-entanglement trade-off; they do not establish demographic stereotype safety or generative-model confidence.
References
Bai, J., Wang, H., Wu, Q., & Zhang, B. (2026). Privacy-robust incrementality estimation in cookieless settings via uplift modeling: Reproducible evidence from the Hillstrom e-mail experiment. Journal of Technology Informatics and Engineering, 5(1), 17–38. https://doi.org/10.51903/jtie.v5i1.468
Bai, J., & Wu, Q. (2026). Privacy-safe marketing mix modeling and budget optimization under identifier loss: A controlled simulation study. International Journal of Electronic Communication Systems, 6(1). https://doi.org/10.24042/ijecs.v6i1.30533
Chang, X., Lu, Y., & Zhong, Z. S. (2026). Review-grounded explainable recommendation with faithfulness evaluation on Amazon reviews. Journal of Electrical Engineering and Computer Science, 11(1), 9–22. https://doi.org/10.54732/jeecs.v11i1.2
Chen, S., He, S., & Sun, E. (2024). Risk-bounded GPU resource oversubscription via conformal demand envelopes in production AI clusters. Journal of Advanced Computing Systems, 4(5), 119–134. https://doi.org/10.69987/jacs.2024.40509
Chen, Y., & Li, M. (2025). From hand-drawn sketches to interactive web prototypes: A reproducible vision-language approach with structural and visual consistency evaluation. Journal of Technology Informatics and Engineering, 4(2), 364–384. https://doi.org/10.51903/jtie.v4i2.490
Chen, Y., & Xu, H. (2026). Trust-calibrated multilingual RAG for humanitarian information platforms: Empirical evaluation on OMoS-QA for migration information access. International Journal of Graphic Design, 4(1), 141–164. https://doi.org/10.51903/ijgd.v4i1.3552
Chen, Y., Zhang, Y., Chau, D., & Sherman, M. (2023). Credit card default risk tiering with probability calibration and uncertainty-driven rejection: A reproducible study on the UCI credit card clients dataset. Journal of Advanced Computing Systems, 3(4), 31–47. https://doi.org/10.69987/jacs.2023.30403
Chen, Y., Zhou, S., & Lin, E. (2025). Accounting-aware evidence retrieval for institutional due diligence of tokenized trade receivable RWA. Journal of Technology Informatics and Engineering, 4(3), 649–663. https://doi.org/10.51903/jtie.v4i3.542
Cheng, M., Piccardi, T., & Yang, D. (2023). CoMPosT: Characterizing and evaluating caricature in LLM simulations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (pp. 10853–10875). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.emnlp-main.669
Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). QLoRA: Efficient finetuning of quantized LLMs. In Advances in Neural Information Processing Systems (Vol. 36, pp. 10088–10115). Curran Associates, Inc. https://proceedings.neurips.cc/paper_files/paper/2023/hash/1feb87871436031bdc0f2beaa62a049b-Abstract-Conference.html
Dinan, E., Fan, A., Williams, A., Urbanek, J., Kiela, D., & Weston, J. (2020). Queens are powerful too: Mitigating gender bias in dialogue generation. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 8173–8188). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.emnlp-main.656
Geifman, Y., & El-Yaniv, R. (2017). Selective classification for deep neural networks. In Advances in Neural Information Processing Systems (Vol. 30, pp. 4885–4894). Curran Associates, Inc. https://papers.nips.cc/paper/7073-selective-classification-for-deep-neural-networks
Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning (Vol. 70, pp. 1321–1330). PMLR. https://proceedings.mlr.press/v70/guo17a.html
Han, J.-E., Koh, J.-S., Seo, H.-T., Chang, D.-S., & Sohn, K.-A. (2024). PSYDIAL: Personality-based synthetic dialogue generation using large language models. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) (pp. 13321–13331). ELRA and ICCL. https://aclanthology.org/2024.lrec-main.1166/
He, S., Chang, X., & Sun, E. (2024). Cross-cloud transfer learning for AI training capacity forecasting under workload and topology distribution shift. Journal of Advanced Computing Systems, 4(1), 100–120. https://doi.org/10.69987/jacs.2024.40108
He, S., Li, C., & Rao, H. (2025). Few-shot cold-start workload forecasting for new AI inference tenants with time-series foundation models. Journal of Technology Informatics and Engineering, 4(1), 306–324. https://doi.org/10.51903/jtie.v4i1.546
He, S., Tu, H., & Liu, I. (2023). Safe PD capacity forecasting with time-series foundation models and calibrated uncertainty for heterogeneous GPU clusters. Journal of Advanced Computing Systems, 3(4), 48–66. https://doi.org/10.69987/jacs.2023.30404
Holm, S. (1979). A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics, 6(2), 65–70. https://www.jstor.org/stable/4615733
Houlsby, N., Giurgiu, A., Jastrzębski, S., Morrone, B., de Laroussilhe, Q., Gesmundo, A., Attariyan, M., & Gelly, S. (2019). Parameter-efficient transfer learning for NLP. In Proceedings of the 36th International Conference on Machine Learning (Vol. 97, pp. 2790–2799). PMLR. https://proceedings.mlr.press/v97/houlsby19a.html
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2022). LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations. https://openreview.net/forum?id=nZeVKeeFYf9
Jin, J. (2025a). Calibrated resume-job matching for trustworthy LLM-assisted recruiter screening: Pairwise matching, probability calibration, and selective refusal on two public recruitment datasets. Journal of Technology Informatics and Engineering, 4(3), 625–648. https://doi.org/10.51903/jtie.v4i3.529
Jin, J. (2025b). Evidence-chain reliable RAG: Hallucination detection, source attribution, and deterministic provenance explanations. Journal of Technology Informatics and Engineering, 4(2), 520–533. https://doi.org/10.51903/jtie.v4i2.535
Jin, J. (2025c). LLM-style evidence cards for scientific search interfaces: A UI/UX design framework for retrieval transparency, ranking trust, and visual evidence hierarchy. International Journal of Graphic Design, 3(2), 397–414. https://doi.org/10.51903/ijgd.v3i2.3698
Jin, J., Huang, T., & Lu, S. (2024). A model-risk-friendly probability of default workflow: Calibration, distribution-free uncertainty quantification, and SHAP explanations on the UCI credit card default dataset. Journal of Advanced Computing Systems, 4(6), 74–85. https://doi.org/10.69987/jacs.2024.40606
Johnson, J. A. (2014). Measuring thirty facets of the Five Factor Model with a 120-item public domain inventory: Development of the IPIP-NEO-120. Journal of Research in Personality, 51, 78–89. https://doi.org/10.1016/j.jrp.2014.05.003
Kim, H., Hessel, J., Jiang, L., West, P., Lu, X., Yu, Y., Zhou, P., Bras, R. L., Alikhani, M., Kim, G., Sap, M., & Choi, Y. (2023). SODA: Million-scale dialogue distillation with social commonsense contextualization. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (pp. 12930–12949). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.emnlp-main.799
Kuo, M.-J., Zheng, D., & Hires, J. (2025). Federated topic-preference learning for knowledge-grounded chat with differential privacy. Journal of Technology Informatics and Engineering, 4(2), 385–401. https://doi.org/10.51903/jtie.v4i2.502
Lee, S., Lim, S., Han, S., Oh, G., Chae, H., Chung, J., Kim, M., Kwak, B.-W., Lee, Y., Lee, D., Yeo, J., & Yu, Y. (2025). Do LLMs have distinct and consistent personality? TRAIT: Personality testset designed for LLMs with psychometrics. In Findings of the Association for Computational Linguistics: NAACL 2025 (pp. 8412–8452). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.findings-naacl.469
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-T., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems (Vol. 33, pp. 9459–9474). Curran Associates, Inc. https://proceedings.neurips.cc/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html
Li, C., Bai, J., & Wang, S. (2024). Evidence-chain reliable RAG: Word-level hallucination detection, source attribution, and provenance explanation for LLM applications. Journal of Advanced Computing Systems, 4(2), 76–92. https://doi.org/10.69987/jacs.2024.40207
Li, C., Liu, G., & Zhao, Z. (2026). Cost-aware LLM-style routing for AIOps log analysis: Log parsing, anomaly detection, fault diagnosis, and incident summarization on LogEval task files. Journal of Technology Informatics and Engineering, 5(2), 91–103. https://doi.org/10.51903/jtie.v5i2.538
Li, C., Zhou, B., & Gao, K. (2025). Risk-calibrated patient-facing AI safety cards: A UI/UX benchmark for explainable medical AI response interfaces. International Journal of Graphic Design, 3(2), 381–394. https://doi.org/10.51903/ijgd.v3i2.3709
Li, J., & Zhou, A. (2026). Multi-regulation RAG for AI product counsel: A legal governance framework for cross-border digital commerces. Rule of Law Studies Journal, 2(2), 105–123. https://doi.org/10.64780/rolsj.v2i2.225
Li, W. (2024). BIG5-CHAT [Data set]. Hugging Face. https://huggingface.co/datasets/wenkai-li/big5_chat
Li, W., Liu, J., Liu, A., Zhou, X., Diab, M. T., & Sap, M. (2025). BIG5-CHAT: Shaping LLM personalities through training on human-grounded data. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 20434–20471). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.acl-long.999
Li, X. L., & Liang, P. (2021). Prefix-tuning: Optimizing continuous prompts for generation. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) (pp. 4582–4597). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.acl-long.353
Li, Y. (2024). Findable then explainable: Retrieval–summary integration for code intelligence on a lightweight CodeSearchNet subset. Journal of Advanced Computing Systems, 4(7), 65–82. https://doi.org/10.69987/jacs.2024.40706
Li, Y., Lu, S., & Zhao, L. (2025). LLM-as-design-critic: Aligning AI-generated UI feedback with human graphic design judgment. International Journal of Graphic Design, 3(1), 196–215. https://doi.org/10.51903/ijgd.v3i1.3661
Li, Z., Zhang, K., & Wong, A. (2026). Numerical-reasoning guardrails for a quant research assistant: A compact reproducible benchmark using SEC and FRED data. Journal of Technology Informatics and Engineering, 5(2), 75–90. https://doi.org/10.51903/jtie.v5i2.541
Li, Z., Zhou, S., & Zhou, Z. (2025). Financial risk dashboard design for institutional RWA investors: Visual hierarchy, chart comprehension, and explainability in FinChart-Bench. International Journal of Graphic Design, 3(1), 196–210. https://doi.org/10.51903/ijgd.v3i1.3715
Liu, A., Sap, M., Lu, X., Swayamdipta, S., Bhagavatula, C., Smith, N. A., & Choi, Y. (2021). DExperts: Decoding-time controlled text generation with experts and anti-experts. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) (pp. 6691–6706). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.acl-long.522
Liu, G., He, S., & Liu, I. (2023). LLM-augmented multi-source root cause attribution for CPU and network faults in microservices. Journal of Advanced Computing Systems, 3(6), 39–57. https://doi.org/10.69987/jacs.2023.30604
Liu, G., He, S., & Wong, H. (2025). LLM-compatible visual brief cards for AI infrastructure capacity dashboards: A UI/UX framework for turning forecast risk into graphic design decisions. International Journal of Graphic Design, 3(1), 196–213. https://doi.org/10.51903/ijgd.v3i1.3723
Liu, G., Li, C., & Zhang, E. (2024). OpsLLM for cloud incident triage: Bilingual RAG-based root cause analysis and alert summarization for AI infrastructure operations. Journal of Advanced Computing Systems, 4(4), 97–111. https://doi.org/10.69987/jacs.2024.40408
Lu, S., & Zou, T. (2026). Uncertainty-aware medical vision–language classification on a lightweight MedMNIST-compatible biomedical patch benchmark. Journal of Technology Informatics and Engineering, 5(2), 1–19. https://doi.org/10.51903/jtie.v5i2.530
Madotto, A., Lin, Z., Wu, C.-S., & Fung, P. (2019). Personalizing dialogue agents via meta-learning. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (pp. 5454–5459). Association for Computational Linguistics. https://doi.org/10.18653/v1/P19-1542
Mairesse, F., Walker, M. A., Mehl, M. R., & Moore, R. K. (2007). Using linguistic cues for the automatic recognition of personality in conversation and text. Journal of Artificial Intelligence Research, 30, 457–500. https://doi.org/10.1613/jair.2349
Mazaré, P.-E., Humeau, S., Raison, M., & Bordes, A. (2018). Training millions of personalized dialogue agents. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (pp. 2775–2779). Association for Computational Linguistics. https://doi.org/10.18653/v1/D18-1298
McCrae, R. R., & John, O. P. (1992). An introduction to the Five-Factor Model and its applications. Journal of Personality, 60(2), 175–215. https://doi.org/10.1111/j.1467-6494.1992.tb00970.x
Mehri, S., & Eskenazi, M. (2020). USR: An unsupervised and reference free evaluation metric for dialog generation. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 681–707). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.64
Meng, S., Chen, J., & Zheng, I. (2026). LLM-inspired offline reranking for financial search: Query rewriting, hybrid retrieval, and listwise relevance ranking on FiQA. Journal of Technology Informatics and Engineering, 5(1), 361–378. https://doi.org/10.51903/jtie.v5i1.537
Mi, G., Ye, T., & Wood, D. (2025). A lightweight medical foundation model for cross-modal multi-task pretraining and parameter-efficient few-shot transfer on MedMNIST. Journal of Technology Informatics and Engineering, 4(3), 572–589. https://doi.org/10.51903/jtie.v4i3.492
Mu, J., Lu, Y., & Smith, M. (2023). LLM-assisted incrementality (uplift) modeling for email advertising: From feature interactions to interpretable audience–creative–channel policies. Journal of Advanced Computing Systems, 3(1), 31–48. https://doi.org/10.69987/jacs.2023.30103
Mu, J., Ye, T., & Patel, P. (2025). Offline counterfactual evaluation for advertising and recommendation slot policies: A reproducible study on the Open Bandit Dataset (small). Journal of Technology Informatics and Engineering, 4(3), 521–543. https://doi.org/10.51903/jtie.v4i3.500
Nadeem, M., Bethke, A., & Reddy, S. (2021). StereoSet: Measuring stereotypical bias in pretrained language models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) (pp. 5356–5371). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.acl-long.416
Nangia, N., Vania, C., Bhalerao, R., & Bowman, S. R. (2020). CrowS-Pairs: A challenge dataset for measuring social biases in masked language models. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) (pp. 1953–1967). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.emnlp-main.154
Nie, J., Liu, G., Li, C., & Zou, T. (2026). Evidence-constrained incident visualization cards for distributed cloud logs: A UI/UX framework for turning Hadoop, OpenStack, and ZooKeeper logs into actionable SRE design interfaces. International Journal of Graphic Design, 4(1), 179–185. https://doi.org/10.51903/ijgd.v4i1.3703
Pennebaker, J. W., & King, L. A. (1999). Linguistic styles: Language use as an individual difference. Journal of Personality and Social Psychology, 77(6), 1296–1312. https://doi.org/10.1037/0022-3514.77.6.1296
Pillutla, K., Swayamdipta, S., Zellers, R., Thickstun, J., Welleck, S., Choi, Y., & Harchaoui, Z. (2021). MAUVE: Measuring the gap between neural text and human text using divergence frontiers. In Advances in Neural Information Processing Systems (Vol. 34, pp. 4816–4828). Curran Associates, Inc. https://proceedings.neurips.cc/paper/2021/hash/260c2432a0eecc28ce03c10dadc078a4-Abstract.html
Sap, M., Gabriel, S., Qin, L., Jurafsky, D., Smith, N. A., & Choi, Y. (2020). Social bias frames: Reasoning about social and power implications of language. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 5477–5490). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.486
Serapio-García, G., Safdari, M., Crepy, C., Sun, L., Fitz, S., Romero, P., Abdulhai, M., Faust, A., & Matarić, M. J. (2025). A psychometric framework for evaluating and shaping personality traits in large language models. Nature Machine Intelligence, 7(12), 1954–1968. https://doi.org/10.1038/s42256-025-01115-6
Shu, B., Zhang, L., Choi, M., Dunagan, L., Logeswaran, L., Lee, M., Card, D., & Jurgens, D. (2024). You don't need a personality test to know these models are unreliable: Assessing the reliability of large language models on psychometric instruments. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) (pp. 5263–5281). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.naacl-long.295
Song, H., Wang, Y., Zhang, K., Zhang, W.-N., & Liu, T. (2021). BoB: BERT over BERT for training persona-based dialogue models from limited personalized data. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers) (pp. 167–177). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.acl-long.14
Soto, C. J., & John, O. P. (2017). The next Big Five Inventory (BFI-2): Developing and assessing a hierarchical model with 15 facets to enhance bandwidth, fidelity, and predictive power. Journal of Personality and Social Psychology, 113(1), 117–143. https://doi.org/10.1037/pspp0000096
Su, W., Chen, S., & Qian, E. (2026). Narrative-aware scientific claim verification agent with evidence ranking for ClimateCheck. Journal of Technology Informatics and Engineering, 5(1), 327–340. https://doi.org/10.51903/jtie.v5i1.549
Su, W., Chen, S., & Zhao, C. (2025). Budgeted multi-hop retrieval agent for compositional question answering: A retrieval-policy evaluation on the official MultiHop-RAG benchmark. Journal of Technology Informatics and Engineering, 4(3), 649–662. https://doi.org/10.51903/jtie.v4i3.543
Su, W., Rao, H., & Ma, E. (2026). Privacy and data-integrity risk cards for LLM agents: A UI/UX design framework for secure human oversight under prompt-injection attacks. International Journal of Graphic Design, 4(1), 186–191. https://doi.org/10.51903/ijgd.v4i1.3699
Sun, X., Lu, Y., & Chen, J. (2023). Controllable long-term user memory for multi-session dialogue: Confidence-gated writing, time-aware retrieval-augmented generation, and update/forgetting. Journal of Advanced Computing Systems, 3(8), 9–24. https://doi.org/10.69987/jacs.2023.30802
Sun, X., Zhong, Z. S., & Wu, Q. (2026). Retrieval-grounded HDFS log anomaly detection and deterministic failure narrative generation. Journal of Computer Systems and Applications, 3(1), 15–30. https://doi.org/10.64229/j6d7fr94
Tian, K., Mitchell, E., Zhou, A., Sharma, A., Rafailov, R., Yao, H., Finn, C., & Manning, C. D. (2023). Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (pp. 5433–5442). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.emnlp-main.330
Wu, Q., Meng, S., & Zhao, J. (2025). Text-grounded LLM-assisted design rationale interfaces: Turning advertising layout metadata into explainable UI/UX decision cards. International Journal of Graphic Design, 3(1), 216–240. https://doi.org/10.51903/ijgd.v3i1.3713
Wu, Q., Mi, G., & Wood, D. (2025). Calibration-light subject-independent motor imagery BCI via self-supervised pretraining and conformer. Journal of Technology Informatics and Engineering, 4(1), 239–262. https://doi.org/10.51903/jtie.v5i1.493
Xin, Q. (2025a). Explaining OpenStack failure-injection log anomalies with retrieved normal prototypes. Emerging Information Science and Technology, 6(2). https://doi.org/10.18196/eist.v6i2.31232
Xin, Q. (2025b). Uncertainty-aware late fusion for 3D perception (confidence calibration + fusion rule learning). Journal of Technology Informatics and Engineering, 4(1), 215–238. https://doi.org/10.51903/jtie.v4i1.485
Xin, Q. (2026a). Auditable automated essay scoring and formative feedback: A rubric-grounded pipeline for secondary and higher education. Journal of Artificial Intelligence and Education, 2(1). https://doi.org/10.66053/jaaie.v2i1.348
Xin, Q. (2026b). Behavior retrieval plus response generation for interpretable conversational personalized recommendation. International Journal of Electrical, Energy and Power System Engineering, 9(2), 120–136. https://doi.org/10.31258/ijeepse.9.2.120-136
Xin, Q. (2026c). Early-warning analytics with LLM intervention rationales for student retention decisions: Classroom interaction modeling with xAPI-edu-data and dropout/success prediction. Interdisciplinary Journal of Pedagogical Research and Media Technology, 2(1). https://doi.org/10.64268/inspire.v2i1.117
Xin, Q. (2026d). Explainable and fair credit risk scoring with counterfactual explanations: A reproducible evaluation on the German Credit dataset (HELOC-motivated). Journal of Information and Technology, 14(2). https://doi.org/10.32664/j-intech.v14i02.2228
Xin, Q. (2026e). Log anomaly detection with conformal alert control and evidence-grounded incident ticket generation. AVITEC, 8(2), 247. https://doi.org/10.28989/avitec.v8i2.3974
Xin, Q. (2026f). Self-supervised customer representation learning for segmentation and next-purchase prediction on UCI online retail. Journal of Information and Technology, 14(1). https://doi.org/10.32664/j-intech.v14i01.2229
Xin, Q. (2026g). Self-supervised log anomaly detection with LogBERT-style transformers: Full empirical evaluation on a reproducible SynHDFS benchmark. Journal of Electrical Engineering and Computer Science, 11(1), 23–35. https://doi.org/10.54732/jeecs.v11i1.3
Xiong, M., Hu, Z., Lu, X., Li, Y., Fu, J., He, J., & Hooi, B. (2024). Can LLMs express their uncertainty? An empirical evaluation of confidence elicitation in LLMs. In The Twelfth International Conference on Learning Representations. https://openreview.net/forum?id=gjeQKFxFpZ
Xu, H., Chen, Y., & Med, A. (2025). Automatic detection and explanation of dark patterns from interface microcopy: Empirical comparison of BERT-style encoders, RoBERTa-style encoders, and LLM-style decoders on the ec-darkpattern dataset. Journal of Technology Informatics and Engineering, 4(3), 590–612. https://doi.org/10.51903/jtie.v4i3.491
Ye, T., Chang, X., & Zhong, E. (2025). Uncertainty-aware breast ultrasound explanation cards: A visual communication framework for image-based AI diagnostic support using BreastMNIST_224. International Journal of Graphic Design, 3(2), 365–380. https://doi.org/10.51903/ijgd.v3i2.3701
Ye, T., Mu, J., & Hunter, J. (2026). Off-policy evaluation and conservative policy selection for slot-level dynamic bidding and ranking on the Open Bandit Dataset (small). Journal of Technology Informatics and Engineering, 5(1), 178–199. https://doi.org/10.51903/jtie.v5i1.503
Zhang, B., Rao, H., & Zhao, D. (2024). Evidence-grounded RAG for cloud-native DevOps: Hallucination-resistant AIOps question answering over private operations documents. Journal of Advanced Computing Systems, 4(3), 109–125. https://doi.org/10.69987/jacs.2024.40308
Zhang, B., Ren, Y., & Zou, J. (2025). LLM-style explainable e-commerce recommendation cards: A UI/UX design framework for trust-calibrated product recommendation. International Journal of Graphic Design, 3(2), 381–396. https://doi.org/10.51903/ijgd.v3i2.3697
Zhang, B., Sun, X., Liu, G., & Zhou, B. (2026). LLM-style DevOps copilot for cloud-native troubleshooting: Retrieval-augmented runbook generation and command-safety evaluation. Journal of Technology Informatics and Engineering, 5(2), 104–118. https://doi.org/10.51903/jtie.v5i2.534
Zhang, J. (2025). From general human activity recognition to volleyball-oriented wearable transfer learning: Cross-dataset evidence from UCI HAR and WISDM for domain adaptation and edge deployment. Journal of Technology Informatics and Engineering, 4(1), 263–283. https://doi.org/10.51903/jtie.v4i1.524
Zhang, K., Chen, Y., & Qian, A. (2025). Evidence-grounded accounting disclosure review cards: A visual communication framework for LLM-style explanations over SEC financial statements and notes. International Journal of Graphic Design, 3(2), 395. https://doi.org/10.51903/ijgd.v3i2.3710
Zhang, R., Wen, Z., Wang, C., Tang, C., Xu, P., & Jiang, Y. (2025). Quality analysis and evaluation prediction of RAG retrieval based on machine learning algorithms [Preprint]. arXiv. https://doi.org/10.48550/arxiv.2511.19481
Zhang, S., Dinan, E., Urbanek, J., Szlam, A., Kiela, D., & Weston, J. (2018). Personalizing dialogue agents: I have a dog, do you have pets too? In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 2204–2213). Association for Computational Linguistics. https://doi.org/10.18653/v1/P18-1205
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., & Artzi, Y. (2020). BERTScore: Evaluating text generation with BERT. In International Conference on Learning Representations. https://openreview.net/forum?id=SkeHuCVFDr
Zhang, Y., & Zhang, H. (2025a). A therapist-facing session copilot for live counseling support: Reasoning-guided retrieval and ranking from multi-turn counseling dialogues. Journal of Technology Informatics and Engineering, 4(2), 464–486. https://doi.org/10.51903/jtie.v4i2.547
Zhang, Y., & Zhang, H. (2025b). Visualizing the right counseling support: Evidence-linked recommendation cards for explainable mental health intake interfaces. International Journal of Graphic Design, 3(1), 214–229. https://doi.org/10.51903/ijgd.v3i1.3722
Zhang, Y., & Zhou, Z. (2026). Strategy-aware therapist imitation for emotional support dialogues: A reproducible ESConv study for LLM response control. Advances in Educational Technology and Psychology, 10(2), 92–97. https://doi.org/10.23977/aetp.2026.100213
Zhao, S., Bai, J., & Roberson, D. (2025). Multi-horizon GPU demand forecasting with workload semantics and operational risk curves: An empirical study on Alibaba clusterdata GPU trace. Journal of Technology Informatics and Engineering, 4(3), 544–571. https://doi.org/10.51903/jtie.v4i3.498
Zheng, D., & Li, C. (2024). Behavior-level jailbreak resistance via multi-stage refusal + utility preservation. Journal of Advanced Computing Systems, 4(1), 83–99. https://doi.org/10.69987/jacs.2024.40107
Zheng, D., Li, C., & Davidson, H. (2023). Continual red-teaming for in-the-wild jailbreaks via online guardrail updates and guardrail distillation. Journal of Advanced Computing Systems, 3(2), 35–49. https://doi.org/10.69987/jacs.2023.30203
Zheng, D., Zhang, B., & Geibel, J. (2024). VerifySafe: Toxicity-safe agent responses under adversarial prompts with evidence-based self-verification. Journal of Advanced Computing Systems, 4(1), 67–82. https://doi.org/10.69987/jacs.2024.40106
Zhong, Z. S., Chen, J., Zhong, E., & Sun, X. (2025). Evidence-calibrated RAG for unanswerable question answering: Retrieval coverage, abstention calibration, and hallucination-proxy analysis on SQuAD 2.0. Journal of Technology Informatics and Engineering, 4(2), 502–520. https://doi.org/10.51903/jtie.v4i2.536
Zhong, Z. S., Li, C., & Rao, H. (2026). Trajectory reliability prediction for generalist AI agents: Tool-use failure analysis and success forecasting on ZClawBench. Journal of Technology Informatics and Engineering, 5(1), 341–360. https://doi.org/10.51903/jtie.v5i1.539
Zhong, Z. S., & Ling, S. (2024a). Improved theoretical guarantee for rank aggregation via spectral method. Information and Inference: A Journal of the IMA, 13(3), iaae020. https://doi.org/10.1093/imaiai/iaae020
Zhong, Z. S., & Ling, S. (2024b). Uncertainty quantification of spectral estimator and MLE for orthogonal group synchronization [Preprint]. arXiv. https://arxiv.org/abs/2408.05944
Zhong, Z. S., Pan, X., & Lei, Q. (2025). Bridging domains with approximately shared features. In Proceedings of the 28th International Conference on Artificial Intelligence and Statistics (Vol. 258, pp. 559–567). PMLR.
Zhong, Z. S., Wu, Q., & Mi, G. (2025). Uncertainty-aware medical image explanation cards: LLM-generated visual explanations for AI-assisted radiology interfaces. International Journal of Graphic Design, 3(2), 415–436. https://doi.org/10.51903/ijgd.v3i2.3616
Zhou, B., Li, C., & Liu, L. (2025). Risk-calibrated patient-facing AI safety cards: A UI/UX design framework for rubric-based medical risk communication. International Journal of Graphic Design, 3(2), 365–380. https://doi.org/10.51903/ijgd.v3i2.3696
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Toby Ma (Author)

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.