ENGINEERING Information Technology & Electronic Engineering  2026 Vol.27 No.7 P.1-13

http://doi.org/10.1631/ENG.ITEE.2026.0083


CCU-Bench: a systematic benchmark for cross-cultural understanding in large language models


Author(s):  Wei ZHANG, Ting YIN, Jia CHEN, Wenjie NING, Xiaoyan GU, Kai LIU, Hengnian QI, Wei CHEN

Affiliation(s):  1. School of Computer and Computing Science, Hangzhou City University, Hangzhou 310015, China more

Corresponding email(s):   chenvis@zju.edu.cn

Key Words:  Benchmark, Large language models (LLMs), Cross-cultural understanding, LLM evaluation and benchmarks


Wei ZHANG, Ting YIN, Jia CHEN, Wenjie NING, Xiaoyan GU, Kai LIU, Hengnian QI, Wei CHEN. CCU-Bench: a systematic benchmark for cross-cultural understanding in large language models[J]. Journal of Zhejiang University Science C, 2026, 27(7): 1-13.

@article{title="CCU-Bench: a systematic benchmark for cross-cultural understanding in large language models",
author="Wei ZHANG, Ting YIN, Jia CHEN, Wenjie NING, Xiaoyan GU, Kai LIU, Hengnian QI, Wei CHEN",
journal="Journal of Zhejiang University Science C",
volume="27",
number="7",
pages="1-13",
year="2026",
publisher="Zhejiang University Press & Springer",
doi="10.1631/ENG.ITEE.2026.0083"
}

%0 Journal Article
%T CCU-Bench: a systematic benchmark for cross-cultural understanding in large language models
%A Wei ZHANG
%A Ting YIN
%A Jia CHEN
%A Wenjie NING
%A Xiaoyan GU
%A Kai LIU
%A Hengnian QI
%A Wei CHEN
%J Frontiers of Information Technology & Electronic Engineering
%V 27
%N 7
%P 1-13
%@ 1869-1951
%D 2026
%I Zhejiang University Press & Springer
%DOI 10.1631/ENG.ITEE.2026.0083

TY - JOUR
T1 - CCU-Bench: a systematic benchmark for cross-cultural understanding in large language models
A1 - Wei ZHANG
A1 - Ting YIN
A1 - Jia CHEN
A1 - Wenjie NING
A1 - Xiaoyan GU
A1 - Kai LIU
A1 - Hengnian QI
A1 - Wei CHEN
J0 - Frontiers of Information Technology & Electronic Engineering
VL - 27
IS - 7
SP - 1
EP - 13
%@ 1869-1951
Y1 - 2026
PB - Zhejiang University Press & Springer
ER -
DOI - 10.1631/ENG.ITEE.2026.0083


Abstract: 
Large language models (LLMs) have achieved strong performance on many language tasks, but they still struggle with culturally grounded symbolic reasoning. Existing benchmarks have not systematically evaluated the progressive capability chain required for cross-cultural understanding, which involves intra-cultural symbolic understanding, cross-cultural symbolic alignment, and cross-cultural conflict identification. To address this gap, we propose CCU-Bench, a systematic benchmark for evaluating cross-cultural understanding in LLMs. Grounded in Hofstede’s cultural dimensions theory, CCU-Bench focuses on three culturally distinct contexts, China, Japan, and Mexico, and is organized around three dimensions, image, connotation, and emotion. The benchmark is constructed through three stages, data collection and filtering, question–answer (QA) generation, and quality control, resulting in a high-quality evaluation set of 3029 QA pairs in five formats. Experiments on 19 mainstream LLMs show that current models perform unsatisfactorily on this task, achieving an average score of 57.8%. Closed-source models consistently outperform open-source models, while cross-cultural symbolic alignment remains the most challenging sub-task. Further error analysis reveals that the dominant failures stem from deficiencies in cultural knowledge, biases in intent mapping, and weak higher-order reasoning about cross-cultural conflicts, rather than simple instruction-following issues. These findings highlight persistent limitations of current LLMs in culturally grounded reasoning and demonstrate that CCU-Bench provides a standardized benchmark for culturally aware artificial intelligence research.

CCU-Bench:大语言模型跨文化理解的系统性基准

张玮1,尹婷1,陈佳1,宁文洁1,古晓燕2,刘凯2,祁亨年3,陈为2
1浙大城市学院计算机与计算科学学院,中国杭州市,310015
2浙江大学计算机辅助设计与图形系统全国重点实验室,中国杭州市,310058
3湖州师范大学信息工程学院,中国湖州市,313000
摘要:大型语言模型在许多语言任务上表现优异,但在基于文化的符号推理方面仍面临挑战。现有基准测试尚未系统评估跨文化理解所需的渐进式能力链条,该链条涉及文化内符号理解、跨文化符号对齐以及跨文化冲突识别。为填补这一空白,提出一个用于评估大语言模型跨文化理解能力的系统性基准—CCU-Bench。基于霍夫斯泰德文化维度理论,CCU-Bench聚焦中国、日本和墨西哥这3种具有差异性的文化语境,围绕图像、内涵和情感3个维度展开。该基准通过数据收集与筛选、问答生成和质量控制3个阶段构建,最终形成包含3029个问答对、涵盖5种题型的高质量评估集。对19款主流大型语言模型的实验表明,现有模型在跨文化理解任务上表现不尽如人意,平均得分仅为57.8%。闭源模型表现整体优于开源模型,而跨文化符号对齐仍是最具挑战性的子任务。进一步的错误分析揭示,主要的失败源于文化知识匮乏、意图映射偏差以及对跨文化冲突的高阶推理能力不足,而非简单的指令遵循问题。这些发现凸显了当前大语言模型在文化根基推理方面存在局限性,证明CCU-Bench可为文化感知人工智能研究提供一个标准化基准测试平台。

关键词:基准;大语言模型;跨文化理解;大语言模型评测与基准

Darkslateblue:Affiliate; Royal Blue:Author; Turquoise:Article

Reference

[1]Adilazuarda MF, Mukherjee S, Lavania P, et al., 2024. Towards measuring and modeling “Culture” in LLMs: a survey. Proc Conf on Empirical Methods in Natural Language Processing, p.15763-15784.

[2]An CX, Gong SS, Zhong M, et al., 2024. L-Eval: instituting standardized evaluation for long context language models. Proc 62nd Annual Meeting of the Association for Computational Linguistics, p.14388-14411.

[3]Anil R, Borgeaud S, Alayrac J, et al., 2023. Gemini: a family of highly capable multimodal models.

[4]Anthropic, 2025. Introducing Claude Sonnet 4.5. https://www.anthropic.com/news/claude-sonnet-4-5 [Accessed on Sept. 29, 2025].

[5]Arora A, Kaffee LA, Augenstein I, 2023. Probing pre-trained language models for cross-cultural differences in values. Proc 1st Workshop on Cross-Cultural Considerations in NLP, p.114-130.

[6]Austin J, Odena A, Nye M, et al., 2021. Program synthesis with large language models.

[7]Bi XJ, Li S, Xing JY, 2026. A weakly supervised preference alignment framework for robust ancient Chinese translation. Patt Recogn, 179:113562.

[8]Cao Y, Zhou L, Lee S, et al., 2023. Assessing cross-cultural alignment between ChatGPT and human societies: an empirical study. Proc 1st Workshop on Cross-Cultural Considerations in NLP, p.53-67.

[9]Chen M, Tworek J, Jun H, et al., 2021. Evaluating large language models trained on code.

[10]Chiu YY, Jiang LW, Lin BY, et al., 2024. CulturalBench: a robust, diverse, and challenging cultural benchmark by human-AI CulturalTeaming.

[11]Dhar S, Shamir L, 2021. Evaluation of the benchmark datasets for testing the efficacy of deep convolutional neural networks. Vis Inform, 5(3):92-101.

[12]Durmus E, Nguyen K, Liao TI, et al., 2023. Towards measuring the representation of subjective global opinions in language models.

[13]Eberhard W, 1988. Dictionary of Chinese Symbols: Hidden Symbols in Chinese Life and Thought. Routledge, London, UK.

[14]Feng SB, Park CY, Liu YH, et al., 2023. From pretraining data to language models to downstream tasks: tracking the trails of political biases leading to unfair NLP models. Proc 61st Annual Meeting of the Association for Computational Linguistics, p.11737-11762.

[15]Guo MQ, Hwa R, Kovashka A, 2023. Decoding symbolism in language models. Proc 61st Annual Meeting of the Association for Computational Linguistics, p.3311-3324.

[16]He C, Huang Y, Lan H, et al., 2026. AI-assisted assessment of higher education quality: a visual analytical approach. Vis Inform, 10:100306.

[17]Hofstede GH, 2001. Culture’s Consequences: Comparing Values, Behaviors, Institutions, and Organizations Across Nations. Sage Publications, Thousand Oaks, USA.

[18]Hogan A, Blomqvist E, Cochez M, et al., 2022. Knowledge graphs. ACM Comput Surv, 54(4):71.

[19]Hu Y, Song Z, Yu J, et al., 2025. TimeJudge: empowering video-LLMs as zero-shot judges for temporal consistency in video captions. Front Inform Technol Electron Eng, 26(11):2204-2214.

[20]Ikegami Y, 1991. The Empire of Signs: Semiotic Essays on Japanese Culture. John Benjamins Publishing, Amsterdam, the Netherlands.

[21]Johnson RL, Pistilli G, Menédez-González N, et al., 2022. The ghost in the machine has an American accent: value conflict in GPT-3.

[22]Kabir M, Ahmed T, Rahman MM, et al., 2026. XCR-Bench: a multi-task benchmark for evaluating cultural reasoning in LLMs.

[23]Kharchenko J, Roosta T, Chadha A, et al., 2025. How well do LLMs represent values across cultures? Empirical analysis of LLM responses based on Hofstede cultural dimensions.

[24]Kong L, Zhong X, Chen J, et al., 2025. Multi-perspective consistency checking for large language model hallucination detection: a black-box zero-resource approach. Front Inform Technol Electron Eng, 26(11):2298-2309.

[25]Kusnick J, Andersen NS, Liem J, et al., 2026. A survey on visualization-based storytelling in digital humanities and cultural heritage. Vis Inform, 10(2):100312.

[26]Li C, Chen MZ, Wang JD, et al., 2024. CultureLLM: incorporating cultural differences into large language models. Proc 38th Int Conf on Neural Information Processing Systems, Article 2693.

[27]Liu RB, Wei J, Liu FY, et al., 2024. Best practices and lessons learned on synthetic data.

[28]Lu JG, Song LL, Zhang LD, 2025. Cultural tendencies in generative AI. Nat Hum Behav, p.2360-2369.

[29]Masoud RI, Liu ZQ, Ferianc M, et al., 2025. Cultural alignment in large language models: an explanatory analysis based on Hofstede’s cultural dimensions. Proc 31st Int Conf on Computational Linguistics, p.8474-8503.

[30]Mei C, 2023. National Museum of Anthropology, México. Guangxi Normal University Press, Guilin, China (in Chinese).

[31]Myung J, Lee N, Zhou Y, et al., 2024. BLEND: a benchmark for LLMs on everyday knowledge in diverse cultures and languages. Proc 38th Int Conf on Neural Information Processing Systems, Article 2483.

[32]Nguyen TP, Razniewski S, Varde A, et al., 2023. Extracting cultural commonsense knowledge at scale. Proc ACM Web Conf, p.1907-1917.

[33]OpenAI, 2025. OpenAI GPT-5 system card.

[34]Peng YZ, Zhang GR, Zhang MS, et al., 2026. LMM-R1: empowering 3B LMMs with strong reasoning abilities with two-stage rule-based RL. Front Comput Sci, 20:1-28.

[35]Prabhakaran V, Qadri R, Hutchinson B, 2022. Cultural incongruencies in artificial intelligence.

[36]Qwen Team, 2025. Qwen3 technical report.

[37]Rahman H, Salam H, 2026. CCD-Bench: probing cultural conflict in large language model decision-making. Proc 40th AAAI Conf on Artificial Intelligence, p.39125-39133.

[38]Romero D, Lyu CY, Wibowo HA, et al., 2024. CVQA: culturally-diverse multilingual visual question answering benchmark. 38th Conf on Neural Information Processing Systems, p.11479-11505.

[39]Rystrøm J, Kirk HR, Hale SA, 2025. Multilingual != multicultural: evaluating gaps between multilingual capabilities and cultural alignment in LLMs. Proc 1st Interdisciplinary Workshop on Observations of Misunderstood, Misguided and Malicious Use of Language Models, p.74-85.

[40]Santurkar S, Durmus E, Ladhak F, et al., 2023. Whose opinions do language models reflect? Proc 40th Int Conf on Machine Learning, p.29971-30004.

[41]Shen SQ, Logeswaran L, Lee M, et al., 2024. Understanding the capabilities and limitations of large language models for cultural commonsense. Proc Conf on the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, p.5668-5680.

[42]Shi WY, Li R, Zhang YT, et al., 2024. CultureBank: an online community-driven knowledge base towards culturally aware language technologies. Findings of the Association for Computational Linguistics, p.4996-5025.

[43]Tam ZR, Wu CK, Chiu YY, et al., 2025. Language matters: how do multilingual input and reasoning paths affect large reasoning models?

[44]Tang Y, Qiao L, Yin L, et al., 2025. Training large-scale language models with limited GPU memory: a survey. Front Inform Technol Electron Eng, 26(3):309-331.

[45]Team GLM, 2024. ChatGLM: a family of large language models from GLM-130B to GLM-4 all tools.

[46]Wang W, Yang Y, Pan Y, 2025. Visual knowledge in the big model era: retrospect and prospect. Front Inform Technol Electron Eng, 26(1):1-19.

[47]Wang YZ, Kordi Y, Mishra S, et al., 2023. Self-instruct: aligning language models with self-generated instructions. Proc 61st Annual Meeting of the Association for Computational Linguistics, p.13484-13508.

[48]Wu ZY, Chen XK, Pan ZZ, et al., 2024. DeepSeek-VL2: mixture-of-experts vision-language models for advanced multimodal understanding.

[49]Xu GH, Liu JY, Yan M, et al., 2023. CValues: measuring the values of Chinese large language models from safety to responsibility.

[50]Xu SY, Dong WL, Guo ZS, et al., 2024. Exploring multilingual concepts of human values in large language models: is value alignment consistent, transferable and controllable across languages? Findings of the Association for Computational Linguistics, p.1771-1793.

[51]Yeh RS, 1988. On Hofstede’s treatment of Chinese and Japanese values. Asia Pac J Manag, 6(1):149-160.

[52]Zhang W, Zhang J, Wong K, et al., 2024a. Computational approaches for traditional Chinese painting: from the “six principles of painting” perspective. J Comput Sci Technol, 39(2):269-285.

[53]Zhang W, Kam-Kwai W, Chen Y, et al., 2024b. ScrollTimes: tracing the provenance of paintings as a window into history. IEEE Trans Vis Comput Graph, 30(6):2981-2994.

[54]Zhang W, Kam-Kwai W, Xu BY, et al., 2025. CultiVerse: towards cross-cultural understanding for paintings with large language model. Proc 33rd ACM Int Conf on Multimedia, p.6710-6719.

[55]Zhang W, Gu X, Liu H, et al., 2026. DAVA: decoding art with visual analytics through feature modeling and multi-agent collaboration. IEEE Trans Vis Comput Graph, 32(3):2773-2786.

[56]Zhou J, Ke P, Qiu XP, et al., 2024. ChatGPT: potential, prospects, and limitations. Front Inform Technol Electron Eng, 25(1):6-11.

[57]Zhu JC, Zhu MH, Rui RT, et al., 2027. Evolutionary perspectives on the evaluation of LLM-based AI agents: a comprehensive survey. Front Comput Sci, 21:2101341.

Open peer comments: Debate/Discuss/Question/Opinion

<1>

Please provide your name, email address and a comment





Full Text:   <1>

Summary:  <1>

CLC number: TP391.1

On-line Access: 2026-08-12

Received: 2026-03-27

Revision Accepted: 2026-06-23

Crosschecked: 2026-08-12

Cited: 0

Clicked: 1

Citations:  Bibtex RefMan EndNote GB/T7714

 ORCID:

Wei ZHANG

0000-0002-8321-4607

Ting YIN

0009-0009-6706-6622

Jia CHEN

0009-0002-6905-6005

Wenjie NING

0009-0008-4871-4586

Xiaoyan GU

0009-0009-5379-985X

Kai LIU

0009-0003-5786-6489

Hengnian QI

0000-0002-1927-7160

Wei CHEN

0000-0002-8365-4741

Journal of Zhejiang University-SCIENCE, 38 Zheda Road, Hangzhou 310027, China
Tel: +86-571-87952783; E-mail: cjzhang@zju.edu.cn
Copyright © 2000 - 2026 Journal of Zhejiang University-SCIENCE