Assessing Social Alignment: Do Personality-Prompted Large Language Models Behave Like Humans?

Zakazov, Ivan; Boronski, Mikolaj; Drudi, Lorenzo; West, Robert

Computer Science > Computers and Society

arXiv:2412.16772 (cs)

[Submitted on 21 Dec 2024]

Title:Assessing Social Alignment: Do Personality-Prompted Large Language Models Behave Like Humans?

Authors:Ivan Zakazov, Mikolaj Boronski, Lorenzo Drudi, Robert West

View PDF HTML (experimental)

Abstract:The ongoing revolution in language modelling has led to various novel applications, some of which rely on the emerging "social abilities" of large language models (LLMs). Already, many turn to the new "cyber friends" for advice during pivotal moments of their lives and trust them with their deepest secrets, implying that accurate shaping of LLMs' "personalities" is paramount. Leveraging the vast diversity of data on which LLMs are pretrained, state-of-the-art approaches prompt them to adopt a particular personality. We ask (i) if personality-prompted models behave (i.e. "make" decisions when presented with a social situation) in line with the ascribed personality, and (ii) if their behavior can be finely controlled. We use classic psychological experiments - the Milgram Experiment and the Ultimatum Game - as social interaction testbeds and apply personality prompting to GPT-3.5/4/4o-mini/4o. Our experiments reveal failure modes of the prompt-based modulation of the models' "behavior", thus challenging the feasibility of personality prompting with today's LLMs.

Comments:	Accepted to NeurIPS 2024 Workshop on Behavioral Machine Learning
Subjects:	Computers and Society (cs.CY); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as:	arXiv:2412.16772 [cs.CY]
	(or arXiv:2412.16772v1 [cs.CY] for this version)
	https://doi.org/10.48550/arXiv.2412.16772

Submission history

From: Ivan Zakazov [view email]
[v1] Sat, 21 Dec 2024 20:58:19 UTC (345 KB)

Computer Science > Computers and Society

Title:Assessing Social Alignment: Do Personality-Prompted Large Language Models Behave Like Humans?

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computers and Society

Title:Assessing Social Alignment: Do Personality-Prompted Large Language Models Behave Like Humans?

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators