
Linggang KONG, Xiaofeng ZHONG, Jie CHEN, Haoran FU, Yongjie WANG. Multi-perspective consistency checking for large language model hallucination detection: a black-box zero-resource approach[J]. Frontiers of Information Technology & Electronic Engineering,in press.https://doi.org/10.1631/FITEE.2500180 @article{title="Multi-perspective consistency checking for large language model hallucination detection: a black-box zero-resource approach", %0 Journal Article TY - JOUR
多视角一致性校验的大语言模型幻觉检测:一种黑盒零资源方法1国防科技大学电子对抗学院,中国合肥市,230037 2安徽省网络空间安全态势感知与评估重点实验室,中国合肥市,230037 摘要:大语言模型(LLM)凭借其卓越的自然语言处理与生成能力,已被广泛应用于各个领域。然而,LLM时不时会生成与事实相悖的内容,即所谓幻觉,这为其在现实场景中的应用带来严峻挑战。为提升LLM的可靠性,在LLM生成过程中检测幻觉现象至关重要。常用于检测幻觉的方法包括获取外部知识或检查模型内部状态,但这需要对LLM进行白盒访问或依赖可靠的专家知识资源,对终端用户而言存在较高门槛。为解决这些挑战,我们提出一种基于多视角一致性校验的黑盒零资源检测方法,用于识别LLM的幻觉现象。该方法通过融合查询与响应的多视角一致性分数,有效缓解了LLM过度自信问题。与依赖单一视角的检测方法相比,我们的方法在多个数据集和不同LLM上均展现出更优的幻觉检测性能。值得注意的是,在一个LLM幻觉率为94.7%的实验场景中,相较单视角一致性方法,我们的方法将平均准确率(B-ACC)提升2.3个百分点,并实现0.832的曲线下面积(AUC),全程无需依赖任何外部资源。 关键词组: Darkslateblue:Affiliate; Royal Blue:Author; Turquoise:Article
Reference[1]Cheng FP, Zouhar V, Arora S, et al., 2024. RELIC: investigating large language model responses using self-consistency. Proc CHI Conf on Human Factors in Computing Systems, Article 647. [2]Cheng XX, Li JY, Zhao WX, et al., 2024. Small agent can also rock! Empowering small language models as hallucination detector. Proc Conf on Empirical Methods in Natural Language Processing, p.14600-14615. [3]Chern IC, Chern S, Chen SQ, et al., 2023. FacTool: factuality detection in generative AI—a tool augmented framework for multi-task and multi-domain scenarios. https://arxiv.org/abs/2307.13528 [4]Devlin J, Chang MW, Lee K, et al., 2019. BERT: pre-training of deep bidirectional transformers for language understanding. Proc Conf of the North American Chapter of the Association for Computational Linguistics, p.4171-4186. [5]Du XF, Xiao CW, Li YX, 2024. HaloScope: harnessing unlabeled LLM generations for hallucination detection. Proc 38th Int Conf on Neural Information Processing Systems, Article 3270. [6]Efron B, Tibshirani RJ, 1993. An Introduction to the Bootstrap. Chapman & Hall, New York, USA. [7]Farquhar S, Kossen J, Kuhn L, et al., 2024. Detecting hallucinations in large language models using semantic entropy. Nature, 630(8017):625-630. [8]Fu JL, Ng SK, Jiang ZB, et al., 2024. GPTScore: evaluate as you desire. Proc Conf of the North American Chapter of the Association for Computational Linguistics, p.6556-6576. [9]Guan XY, Liu YJ, Lin HY, et al., 2024. Mitigating large language model hallucinations via autonomous knowledge graph-based retrofitting. Proc 38th Annual AAAI Conf on Artificial Intelligence, p.18126-18134. [10]Honovich O, Aharoni R, Herzig J, et al., 2022. TRUE: re-evaluating factual consistency evaluation. Proc Conf of the North American Chapter of the Association for Computational Linguistics, p.3905-3920. [11]Hu XM, Zhang YM, Peng R, et al., 2024. Embedding and gradient say wrong: a white-box method for hallucination detection. Proc Conf on Empirical Methods in Natural Language Processing, p.1950-1959. [12]Huang XM, Li S, Yu MX, et al., 2024. Uncertainty in language models: assessment through rank-calibration. Proc Conf on Empirical Methods in Natural Language Processing, p.284-312. [13]Kadavath S, Conerly T, Askell A, et al., 2022. Language models (mostly) know what they know. https://arxiv.org/abs/2207.05221 [14]Li JY, Cheng XX, Zhao WX, et al., 2023. HaluEval: a large-scale hallucination evaluation benchmark for large language models. Proc Conf on Empirical Methods in Natural Language Processing, p.6449-6464. [15]Li JY, Chen J, Ren RY, et al., 2024. The dawn after the dark: an empirical study on factuality hallucination in large language models. Proc 62nd Annual Meeting of the Association for Computational Linguistics, p.10879-10899. [16]Li TJ, Li Z, Zhang Y, 2024. Improving faithfulness of large language models in summarization via sliding generation and self-consistency. Proc Joint Int Conf on Computational Linguistics, p.8804-8817. [17]Liang X, Song SC, Zheng ZF, et al., 2024. Internal consistency and self-feedback in large language models: a survey. https://arxiv.org/abs/2407.14507 [18]Lin ZC, Guan SY, Zhang WD, et al., 2024. Towards trustworthy LLMs: a review on debiasing and dehallucinating in large language models. Artif Intell Rev, 57(9):243. [19]Liu HF, Huang HG, Gu XM, et al., 2025. On calibration of LLM-based guard models for reliable content moderation. Proc 13th Int Conf on Learning Representations. [20]Luo YW, Yang Y, 2024. Large language model and domain-specific model collaboration for smart education. Front Inform Technol Electron Eng, 25(3):333-341. [21]Manakul P, Liusie A, Gales M, 2023. SelfCheckGPT: zero-resource black-box hallucination detection for generative large language models. Proc Conf on Empirical Methods in Natural Language Processing, p.9004-9017. [22]Mündler N, He JX, Jenko S, et al., 2024. Self-contradictory hallucinations of large language models: evaluation, detection and mitigation. Proc 12th Int Conf on Learning Representations. [23]Reimers N, Gurevych I, 2019. Sentence-BERT: sentence embeddings using Siamese BERT-networks. Proc Conf on Empirical Methods in Natural Language Processing and the 9th Int Joint Conf on Natural Language Processing, p.3982-3992. [24]Sadat M, Zhou ZY, Lange L, et al., 2023. DelucionQA: detecting hallucinations in domain-specific question answering. Proc Findings of the Association for Computational Linguistics, p.822-835. [25]Tian K, Mitchell E, Zhou A, et al., 2023. Just ask for calibration: strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. Proc Conf on Empirical Methods in Natural Language Processing, p.5433-5442. [26]Verspoor K, 2024. ‘Fighting fire with fire’—using LLMs to combat LLM hallucinations. Nature, 630(8017):569-570. [27]Wan FQ, Huang XT, Cui LY, et al., 2024. Knowledge verification to nip hallucination in the bud. Proc Conf on Empirical Methods in Natural Language Processing, p.2616-2633. [28]Zhang JX, Li ZH, Das K, et al., 2023. SAC3: reliable hallucination detection in black-box language models via semantic-aware cross-check consistency. Proc Findings of the Association for Computational Linguistics, p.15445-15458. [29]Zhang SL, Yu T, Feng Y, 2024. TruthX: alleviating hallucinations by editing large language models in truthful space. Proc 62nd Annual Meeting of the Association for Computational Linguistics, p.8908-8949. [30]Zhang TY, Kishore V, Wu F, et al., 2020. BERTScore: evaluating text generation with BERT. Proc 8th Int Conf on Learning Representations. [31]Zhang XY, Peng BL, Tian Y, et al., 2024. Self-alignment for factuality: mitigating hallucinations in LLMs via self-evaluation. Proc 62nd Annual Meeting of the Association for Computational Linguistics, p.1946-1965. [32]Zhang Y, Li YF, Cui LY, et al., 2023. Siren’s song in the AI ocean: a survey on hallucination in large language models. https://arxiv.org/abs/2309.01219 CLC number: TP18 On-line Access: 2026-01-08 Received: 2025-03-21 Revision Accepted: 2025-10-08 Crosschecked: 2026-01-08 Cited: 0 Clicked: 1436 Journal of Zhejiang University-SCIENCE, 38 Zheda Road, Hangzhou
310027, China
Tel: +86-571-87952783; E-mail: cjzhang@zju.edu.cn Copyright © 2000 - 2026 Journal of Zhejiang University-SCIENCE | ||||||||||||||
Open peer comments: Debate/Discuss/Question/Opinion
<1>