Large Language Models Sometimes Generate Purely Negatively-Reinforced Text

Roger, Fabien

Computer Science > Machine Learning

arXiv:2306.07567 (cs)

[Submitted on 13 Jun 2023 (v1), last revised 16 Jun 2023 (this version, v2)]

Title:Large Language Models Sometimes Generate Purely Negatively-Reinforced Text

Authors:Fabien Roger

View PDF

Abstract:When using adversarial training, it is common practice to train against the most egregious failures. However, this might imply using examples with sensitive information (such as leaked passwords or security vulnerabilities) as training data. One might assume that language models trained with gradient descent never generate text snippets which were only present in examples associated with the lowest possible reward. In this paper, we show that this assumption is wrong: in some situations, large language models do learn from such negatively-reinforced examples. We present a specific training setup that enables Pythia-160M to guess passwords 13% more often than it would by guessing randomly, despite only showing it these passwords on examples where the model is incentivized to not output these passwords. Our code is available at this http URL

Comments:	6 pages, 5 figures, LaTeX; added a related work section
Subjects:	Machine Learning (cs.LG); Computation and Language (cs.CL)
Cite as:	arXiv:2306.07567 [cs.LG]
	(or arXiv:2306.07567v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2306.07567

Submission history

From: Fabien Roger [view email]
[v1] Tue, 13 Jun 2023 06:40:37 UTC (928 KB)
[v2] Fri, 16 Jun 2023 16:23:21 UTC (930 KB)

Computer Science > Machine Learning

Title:Large Language Models Sometimes Generate Purely Negatively-Reinforced Text

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Large Language Models Sometimes Generate Purely Negatively-Reinforced Text

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators