Sparse Autoencoders for Hypothesis Generation

Movva, Rajiv; Peng, Kenny; Garg, Nikhil; Kleinberg, Jon; Pierson, Emma

Computer Science > Computation and Language

arXiv:2502.04382 (cs)

[Submitted on 5 Feb 2025]

Title:Sparse Autoencoders for Hypothesis Generation

Authors:Rajiv Movva, Kenny Peng, Nikhil Garg, Jon Kleinberg, Emma Pierson

View PDF HTML (experimental)

Abstract:We describe HypotheSAEs, a general method to hypothesize interpretable relationships between text data (e.g., headlines) and a target variable (e.g., clicks). HypotheSAEs has three steps: (1) train a sparse autoencoder on text embeddings to produce interpretable features describing the data distribution, (2) select features that predict the target variable, and (3) generate a natural language interpretation of each feature (e.g., "mentions being surprised or shocked") using an LLM. Each interpretation serves as a hypothesis about what predicts the target variable. Compared to baselines, our method better identifies reference hypotheses on synthetic datasets (at least +0.06 in F1) and produces more predictive hypotheses on real datasets (~twice as many significant findings), despite requiring 1-2 orders of magnitude less compute than recent LLM-based methods. HypotheSAEs also produces novel discoveries on two well-studied tasks: explaining partisan differences in Congressional speeches and identifying drivers of engagement with online headlines.

Comments:	First two authors contributed equally; working paper
Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Computers and Society (cs.CY)
Cite as:	arXiv:2502.04382 [cs.CL]
	(or arXiv:2502.04382v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2502.04382

Submission history

From: Rajiv Movva [view email]
[v1] Wed, 5 Feb 2025 18:58:02 UTC (6,695 KB)

Computer Science > Computation and Language

Title:Sparse Autoencoders for Hypothesis Generation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Sparse Autoencoders for Hypothesis Generation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators