C3AI: Crafting and Evaluating Constitutions for Constitutional AI

Kyrychenko, Yara; Zhou, Ke; Bogucka, Edyta; Quercia, Daniele

doi:10.1145/3696410.3714705

Computer Science > Artificial Intelligence

arXiv:2502.15861 (cs)

[Submitted on 21 Feb 2025]

Title:C3AI: Crafting and Evaluating Constitutions for Constitutional AI

Authors:Yara Kyrychenko, Ke Zhou, Edyta Bogucka, Daniele Quercia

View PDF HTML (experimental)

Abstract:Constitutional AI (CAI) guides LLM behavior using constitutions, but identifying which principles are most effective for model alignment remains an open challenge. We introduce the C3AI framework (\textit{Crafting Constitutions for CAI models}), which serves two key functions: (1) selecting and structuring principles to form effective constitutions before fine-tuning; and (2) evaluating whether fine-tuned CAI models follow these principles in practice. By analyzing principles from AI and psychology, we found that positively framed, behavior-based principles align more closely with human preferences than negatively framed or trait-based principles. In a safety alignment use case, we applied a graph-based principle selection method to refine an existing CAI constitution, improving safety measures while maintaining strong general reasoning capabilities. Interestingly, fine-tuned CAI models performed well on negatively framed principles but struggled with positively framed ones, in contrast to our human alignment results. This highlights a potential gap between principle design and model adherence. Overall, C3AI provides a structured and scalable approach to both crafting and evaluating CAI constitutions.

Comments:	This has been accepted for the Web Conference 2025
Subjects:	Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as:	arXiv:2502.15861 [cs.AI]
	(or arXiv:2502.15861v1 [cs.AI] for this version)
	https://doi.org/10.48550/arXiv.2502.15861
Related DOI:	https://doi.org/10.1145/3696410.3714705

Submission history

From: Ke Zhou [view email]
[v1] Fri, 21 Feb 2025 10:26:42 UTC (1,243 KB)

Computer Science > Artificial Intelligence

Title:C3AI: Crafting and Evaluating Constitutions for Constitutional AI

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Artificial Intelligence

Title:C3AI: Crafting and Evaluating Constitutions for Constitutional AI

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators