
Wei ZHANG, Ting YIN, Jia CHEN, Wenjie NING, Xiaoyan GU, Kai LIU, Hengnian QI, Wei CHEN. CCU-Bench: a systematic benchmark for cross-cultural understanding in large language models[J]. Journal of Zhejiang University Science C, 2026, 27(7): 1-13.
@article{title="CCU-Bench: a systematic benchmark for cross-cultural understanding in large language models",
author="Wei ZHANG, Ting YIN, Jia CHEN, Wenjie NING, Xiaoyan GU, Kai LIU, Hengnian QI, Wei CHEN",
journal="Journal of Zhejiang University Science C",
volume="27",
number="7",
pages="1-13",
year="2026",
publisher="Zhejiang University Press & Springer",
doi="10.1631/ENG.ITEE.2026.0083"
}
%0 Journal Article
%T CCU-Bench: a systematic benchmark for cross-cultural understanding in large language models
%A Wei ZHANG
%A Ting YIN
%A Jia CHEN
%A Wenjie NING
%A Xiaoyan GU
%A Kai LIU
%A Hengnian QI
%A Wei CHEN
%J Frontiers of Information Technology & Electronic Engineering
%V 27
%N 7
%P 1-13
%@ 1869-1951
%D 2026
%I Zhejiang University Press & Springer
%DOI 10.1631/ENG.ITEE.2026.0083
TY - JOUR
T1 - CCU-Bench: a systematic benchmark for cross-cultural understanding in large language models
A1 - Wei ZHANG
A1 - Ting YIN
A1 - Jia CHEN
A1 - Wenjie NING
A1 - Xiaoyan GU
A1 - Kai LIU
A1 - Hengnian QI
A1 - Wei CHEN
J0 - Frontiers of Information Technology & Electronic Engineering
VL - 27
IS - 7
SP - 1
EP - 13
%@ 1869-1951
Y1 - 2026
PB - Zhejiang University Press & Springer
ER -
DOI - 10.1631/ENG.ITEE.2026.0083
Abstract: Large language models (LLMs) have achieved strong performance on many language tasks, but they still struggle with culturally grounded symbolic reasoning. Existing benchmarks have not systematically evaluated the progressive capability chain required for cross-cultural understanding, which involves intra-cultural symbolic understanding, cross-cultural symbolic alignment, and cross-cultural conflict identification. To address this gap, we propose CCU-Bench, a systematic benchmark for evaluating cross-cultural understanding in LLMs. Grounded in Hofstede’s cultural dimensions theory, CCU-Bench focuses on three culturally distinct contexts, China, Japan, and Mexico, and is organized around three dimensions, image, connotation, and emotion. The benchmark is constructed through three stages, data collection and filtering, question–answer (QA) generation, and quality control, resulting in a high-quality evaluation set of 3029 QA pairs in five formats. Experiments on 19 mainstream LLMs show that current models perform unsatisfactorily on this task, achieving an average score of 57.8%. Closed-source models consistently outperform open-source models, while cross-cultural symbolic alignment remains the most challenging sub-task. Further error analysis reveals that the dominant failures stem from deficiencies in cultural knowledge, biases in intent mapping, and weak higher-order reasoning about cross-cultural conflicts, rather than simple instruction-following issues. These findings highlight persistent limitations of current LLMs in culturally grounded reasoning and demonstrate that CCU-Bench provides a standardized benchmark for culturally aware artificial intelligence research.
[1]Adilazuarda MF, Mukherjee S, Lavania P, et al., 2024. Towards measuring and modeling “Culture” in LLMs: a survey. Proc Conf on Empirical Methods in Natural Language Processing, p.15763-15784.
[2]An CX, Gong SS, Zhong M, et al., 2024. L-Eval: instituting standardized evaluation for long context language models. Proc 62nd Annual Meeting of the Association for Computational Linguistics, p.14388-14411.
[3]Anil R, Borgeaud S, Alayrac J, et al., 2023. Gemini: a family of highly capable multimodal models.
[4]Anthropic, 2025. Introducing Claude Sonnet 4.5. https://www.anthropic.com/news/claude-sonnet-4-5 [Accessed on Sept. 29, 2025].
[5]Arora A, Kaffee LA, Augenstein I, 2023. Probing pre-trained language models for cross-cultural differences in values. Proc 1st Workshop on Cross-Cultural Considerations in NLP, p.114-130.
[6]Austin J, Odena A, Nye M, et al., 2021. Program synthesis with large language models.
[7]Bi XJ, Li S, Xing JY, 2026. A weakly supervised preference alignment framework for robust ancient Chinese translation. Patt Recogn, 179:113562.
[8]Cao Y, Zhou L, Lee S, et al., 2023. Assessing cross-cultural alignment between ChatGPT and human societies: an empirical study. Proc 1st Workshop on Cross-Cultural Considerations in NLP, p.53-67.
[9]Chen M, Tworek J, Jun H, et al., 2021. Evaluating large language models trained on code.
[10]Chiu YY, Jiang LW, Lin BY, et al., 2024. CulturalBench: a robust, diverse, and challenging cultural benchmark by human-AI CulturalTeaming.
[11]Dhar S, Shamir L, 2021. Evaluation of the benchmark datasets for testing the efficacy of deep convolutional neural networks. Vis Inform, 5(3):92-101.
[12]Durmus E, Nguyen K, Liao TI, et al., 2023. Towards measuring the representation of subjective global opinions in language models.
[13]Eberhard W, 1988. Dictionary of Chinese Symbols: Hidden Symbols in Chinese Life and Thought. Routledge, London, UK.
[14]Feng SB, Park CY, Liu YH, et al., 2023. From pretraining data to language models to downstream tasks: tracking the trails of political biases leading to unfair NLP models. Proc 61st Annual Meeting of the Association for Computational Linguistics, p.11737-11762.
[15]Guo MQ, Hwa R, Kovashka A, 2023. Decoding symbolism in language models. Proc 61st Annual Meeting of the Association for Computational Linguistics, p.3311-3324.
[16]He C, Huang Y, Lan H, et al., 2026. AI-assisted assessment of higher education quality: a visual analytical approach. Vis Inform, 10:100306.
[17]Hofstede GH, 2001. Culture’s Consequences: Comparing Values, Behaviors, Institutions, and Organizations Across Nations. Sage Publications, Thousand Oaks, USA.
[18]Hogan A, Blomqvist E, Cochez M, et al., 2022. Knowledge graphs. ACM Comput Surv, 54(4):71.
[19]Hu Y, Song Z, Yu J, et al., 2025. TimeJudge: empowering video-LLMs as zero-shot judges for temporal consistency in video captions. Front Inform Technol Electron Eng, 26(11):2204-2214.
[20]Ikegami Y, 1991. The Empire of Signs: Semiotic Essays on Japanese Culture. John Benjamins Publishing, Amsterdam, the Netherlands.
[21]Johnson RL, Pistilli G, Menédez-González N, et al., 2022. The ghost in the machine has an American accent: value conflict in GPT-3.
[22]Kabir M, Ahmed T, Rahman MM, et al., 2026. XCR-Bench: a multi-task benchmark for evaluating cultural reasoning in LLMs.
[23]Kharchenko J, Roosta T, Chadha A, et al., 2025. How well do LLMs represent values across cultures? Empirical analysis of LLM responses based on Hofstede cultural dimensions.
[24]Kong L, Zhong X, Chen J, et al., 2025. Multi-perspective consistency checking for large language model hallucination detection: a black-box zero-resource approach. Front Inform Technol Electron Eng, 26(11):2298-2309.
[25]Kusnick J, Andersen NS, Liem J, et al., 2026. A survey on visualization-based storytelling in digital humanities and cultural heritage. Vis Inform, 10(2):100312.
[26]Li C, Chen MZ, Wang JD, et al., 2024. CultureLLM: incorporating cultural differences into large language models. Proc 38th Int Conf on Neural Information Processing Systems, Article 2693.
[27]Liu RB, Wei J, Liu FY, et al., 2024. Best practices and lessons learned on synthetic data.
[28]Lu JG, Song LL, Zhang LD, 2025. Cultural tendencies in generative AI. Nat Hum Behav, p.2360-2369.
[29]Masoud RI, Liu ZQ, Ferianc M, et al., 2025. Cultural alignment in large language models: an explanatory analysis based on Hofstede’s cultural dimensions. Proc 31st Int Conf on Computational Linguistics, p.8474-8503.
[30]Mei C, 2023. National Museum of Anthropology, México. Guangxi Normal University Press, Guilin, China (in Chinese).
[31]Myung J, Lee N, Zhou Y, et al., 2024. BLEND: a benchmark for LLMs on everyday knowledge in diverse cultures and languages. Proc 38th Int Conf on Neural Information Processing Systems, Article 2483.
[32]Nguyen TP, Razniewski S, Varde A, et al., 2023. Extracting cultural commonsense knowledge at scale. Proc ACM Web Conf, p.1907-1917.
[33]OpenAI, 2025. OpenAI GPT-5 system card.
[34]Peng YZ, Zhang GR, Zhang MS, et al., 2026. LMM-R1: empowering 3B LMMs with strong reasoning abilities with two-stage rule-based RL. Front Comput Sci, 20:1-28.
[35]Prabhakaran V, Qadri R, Hutchinson B, 2022. Cultural incongruencies in artificial intelligence.
[36]Qwen Team, 2025. Qwen3 technical report.
[37]Rahman H, Salam H, 2026. CCD-Bench: probing cultural conflict in large language model decision-making. Proc 40th AAAI Conf on Artificial Intelligence, p.39125-39133.
[38]Romero D, Lyu CY, Wibowo HA, et al., 2024. CVQA: culturally-diverse multilingual visual question answering benchmark. 38th Conf on Neural Information Processing Systems, p.11479-11505.
[39]Rystrøm J, Kirk HR, Hale SA, 2025. Multilingual != multicultural: evaluating gaps between multilingual capabilities and cultural alignment in LLMs. Proc 1st Interdisciplinary Workshop on Observations of Misunderstood, Misguided and Malicious Use of Language Models, p.74-85.
[40]Santurkar S, Durmus E, Ladhak F, et al., 2023. Whose opinions do language models reflect? Proc 40th Int Conf on Machine Learning, p.29971-30004.
[41]Shen SQ, Logeswaran L, Lee M, et al., 2024. Understanding the capabilities and limitations of large language models for cultural commonsense. Proc Conf on the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, p.5668-5680.
[42]Shi WY, Li R, Zhang YT, et al., 2024. CultureBank: an online community-driven knowledge base towards culturally aware language technologies. Findings of the Association for Computational Linguistics, p.4996-5025.
[43]Tam ZR, Wu CK, Chiu YY, et al., 2025. Language matters: how do multilingual input and reasoning paths affect large reasoning models?
[44]Tang Y, Qiao L, Yin L, et al., 2025. Training large-scale language models with limited GPU memory: a survey. Front Inform Technol Electron Eng, 26(3):309-331.
[45]Team GLM, 2024. ChatGLM: a family of large language models from GLM-130B to GLM-4 all tools.
[46]Wang W, Yang Y, Pan Y, 2025. Visual knowledge in the big model era: retrospect and prospect. Front Inform Technol Electron Eng, 26(1):1-19.
[47]Wang YZ, Kordi Y, Mishra S, et al., 2023. Self-instruct: aligning language models with self-generated instructions. Proc 61st Annual Meeting of the Association for Computational Linguistics, p.13484-13508.
[48]Wu ZY, Chen XK, Pan ZZ, et al., 2024. DeepSeek-VL2: mixture-of-experts vision-language models for advanced multimodal understanding.
[49]Xu GH, Liu JY, Yan M, et al., 2023. CValues: measuring the values of Chinese large language models from safety to responsibility.
[50]Xu SY, Dong WL, Guo ZS, et al., 2024. Exploring multilingual concepts of human values in large language models: is value alignment consistent, transferable and controllable across languages? Findings of the Association for Computational Linguistics, p.1771-1793.
[51]Yeh RS, 1988. On Hofstede’s treatment of Chinese and Japanese values. Asia Pac J Manag, 6(1):149-160.
[52]Zhang W, Zhang J, Wong K, et al., 2024a. Computational approaches for traditional Chinese painting: from the “six principles of painting” perspective. J Comput Sci Technol, 39(2):269-285.
[53]Zhang W, Kam-Kwai W, Chen Y, et al., 2024b. ScrollTimes: tracing the provenance of paintings as a window into history. IEEE Trans Vis Comput Graph, 30(6):2981-2994.
[54]Zhang W, Kam-Kwai W, Xu BY, et al., 2025. CultiVerse: towards cross-cultural understanding for paintings with large language model. Proc 33rd ACM Int Conf on Multimedia, p.6710-6719.
[55]Zhang W, Gu X, Liu H, et al., 2026. DAVA: decoding art with visual analytics through feature modeling and multi-agent collaboration. IEEE Trans Vis Comput Graph, 32(3):2773-2786.
[56]Zhou J, Ke P, Qiu XP, et al., 2024. ChatGPT: potential, prospects, and limitations. Front Inform Technol Electron Eng, 25(1):6-11.
[57]Zhu JC, Zhu MH, Rui RT, et al., 2027. Evolutionary perspectives on the evaluation of LLM-based AI agents: a comprehensive survey. Front Comput Sci, 21:2101341.
CLC number: TP391.1
On-line Access: 2026-08-12
Received: 2026-03-27
Revision Accepted: 2026-06-23
Crosschecked: 2026-08-12
Cited: 0
Clicked: 1
Open peer comments: Debate/Discuss/Question/Opinion
<1>