A Poisson convolution model for characterizing topical content with word frequency and exclusivity

Airoldi, Edoardo M; Bischof, Jonathan M

Computer Science > Machine Learning

arXiv:1206.4631 (cs)

[Submitted on 18 Jun 2012 (v1), last revised 28 Jul 2014 (this version, v3)]

Title:A Poisson convolution model for characterizing topical content with word frequency and exclusivity

Authors:Edoardo M Airoldi, Jonathan M Bischof

View PDF

Abstract:An ongoing challenge in the analysis of document collections is how to summarize content in terms of a set of inferred themes that can be interpreted substantively in terms of topics. The current practice of parametrizing the themes in terms of most frequent words limits interpretability by ignoring the differential use of words across topics. We argue that words that are both common and exclusive to a theme are more effective at characterizing topical content. We consider a setting where professional editors have annotated documents to a collection of topic categories, organized into a tree, in which leaf-nodes correspond to the most specific topics. Each document is annotated to multiple categories, at different levels of the tree. We introduce a hierarchical Poisson convolution model to analyze annotated documents in this setting. The model leverages the structure among categories defined by professional editors to infer a clear semantic description for each topic in terms of words that are both frequent and exclusive. We carry out a large randomized experiment on Amazon Turk to demonstrate that topic summaries based on the FREX score are more interpretable than currently established frequency based summaries, and that the proposed model produces more efficient estimates of exclusivity than with currently models. We also develop a parallelized Hamiltonian Monte Carlo sampler that allows the inference to scale to millions of documents.

Comments:	Originally appeared in ICML2012
Subjects:	Machine Learning (cs.LG); Computation and Language (cs.CL); Information Retrieval (cs.IR); Methodology (stat.ME); Machine Learning (stat.ML)
Cite as:	arXiv:1206.4631 [cs.LG]
	(or arXiv:1206.4631v3 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1206.4631

Submission history

From: Edoardo Airoldi [view email] [via ICML2012 proxy]
[v1] Mon, 18 Jun 2012 15:11:38 UTC (1,961 KB)
[v2] Wed, 12 Dec 2012 17:32:26 UTC (4,101 KB)
[v3] Mon, 28 Jul 2014 03:02:39 UTC (5,310 KB)

Computer Science > Machine Learning

Title:A Poisson convolution model for characterizing topical content with word frequency and exclusivity

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:A Poisson convolution model for characterizing topical content with word frequency and exclusivity

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators