Benchmarking Multi-National Value Alignment for Large Language Models

Shi, Weijie; Ju, Chengyi; Liu, Chengzhong; Ji, Jiaming; Zhang, Jipeng; Zhang, Ruiyuan; Zhu, Jia; Xu, Jiajie; Yang, Yaodong; Han, Sirui; Guo, Yike

Computer Science > Computation and Language

arXiv:2504.12911 (cs)

[Submitted on 17 Apr 2025 (v1), last revised 19 Apr 2025 (this version, v2)]

Title:Benchmarking Multi-National Value Alignment for Large Language Models

Authors:Weijie Shi, Chengyi Ju, Chengzhong Liu, Jiaming Ji, Jipeng Zhang, Ruiyuan Zhang, Jia Zhu, Jiajie Xu, Yaodong Yang, Sirui Han, Yike Guo

View PDF HTML (experimental)

Abstract:Do Large Language Models (LLMs) hold positions that conflict with your country's values? Occasionally they do! However, existing works primarily focus on ethical reviews, failing to capture the diversity of national values, which encompass broader policy, legal, and moral considerations. Furthermore, current benchmarks that rely on spectrum tests using manually designed questionnaires are not easily scalable.
To address these limitations, we introduce NaVAB, a comprehensive benchmark to evaluate the alignment of LLMs with the values of five major nations: China, the United States, the United Kingdom, France, and Germany. NaVAB implements a national value extraction pipeline to efficiently construct value assessment datasets. Specifically, we propose a modeling procedure with instruction tagging to process raw data sources, a screening process to filter value-related topics and a generation process with a Conflict Reduction mechanism to filter non-conflicting this http URL conduct extensive experiments on various LLMs across countries, and the results provide insights into assisting in the identification of misaligned scenarios. Moreover, we demonstrate that NaVAB can be combined with alignment techniques to effectively reduce value concerns by aligning LLMs' values with the target country.

Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2504.12911 [cs.CL]
	(or arXiv:2504.12911v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2504.12911

Submission history

From: Weijie Shi [view email]
[v1] Thu, 17 Apr 2025 13:01:38 UTC (9,975 KB)
[v2] Sat, 19 Apr 2025 04:07:54 UTC (9,975 KB)

Computer Science > Computation and Language

Title:Benchmarking Multi-National Value Alignment for Large Language Models

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Benchmarking Multi-National Value Alignment for Large Language Models

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators