Learning Poisson Binomial Distributions

Daskalakis, Constantinos; Diakonikolas, Ilias; Servedio, Rocco A.

Computer Science > Data Structures and Algorithms

arXiv:1107.2702 (cs)

[Submitted on 13 Jul 2011 (v1), last revised 17 Feb 2015 (this version, v4)]

Title:Learning Poisson Binomial Distributions

Authors:Constantinos Daskalakis, Ilias Diakonikolas, Rocco A. Servedio

View PDF

Abstract:We consider a basic problem in unsupervised learning: learning an unknown \emph{Poisson Binomial Distribution}. A Poisson Binomial Distribution (PBD) over $\{0,1,\dots,n\}$ is the distribution of a sum of $n$ independent Bernoulli random variables which may have arbitrary, potentially non-equal, expectations. These distributions were first studied by S. Poisson in 1837 \cite{Poisson:37} and are a natural $n$-parameter generalization of the familiar Binomial Distribution. Surprisingly, prior to our work this basic learning problem was poorly understood, and known results for it were far from optimal.
We essentially settle the complexity of the learning problem for this basic class of distributions. As our first main result we give a highly efficient algorithm which learns to $\eps$-accuracy (with respect to the total variation distance) using $\tilde{O}(1/\eps^3)$ samples \emph{independent of $n$}. The running time of the algorithm is \emph{quasilinear} in the size of its input data, i.e., $\tilde{O}(\log(n)/\eps^3)$ bit-operations. (Observe that each draw from the distribution is a $\log(n)$-bit string.) Our second main result is a {\em proper} learning algorithm that learns to $\eps$-accuracy using $\tilde{O}(1/\eps^2)$ samples, and runs in time $(1/\eps)^{\poly (\log (1/\eps))} \cdot \log n$. This is nearly optimal, since any algorithm {for this problem} must use $\Omega(1/\eps^2)$ samples. We also give positive and negative results for some extensions of this learning problem to weighted sums of independent Bernoulli random variables.

Comments:	Revised full version. Improved sample complexity bound of O~(1/eps^2)
Subjects:	Data Structures and Algorithms (cs.DS); Machine Learning (cs.LG); Statistics Theory (math.ST)
Cite as:	arXiv:1107.2702 [cs.DS]
	(or arXiv:1107.2702v4 [cs.DS] for this version)
	https://doi.org/10.48550/arXiv.1107.2702

Submission history

From: Ilias Diakonikolas [view email]
[v1] Wed, 13 Jul 2011 23:30:39 UTC (43 KB)
[v2] Fri, 15 Jul 2011 06:03:55 UTC (43 KB)
[v3] Sat, 7 Dec 2013 17:37:52 UTC (51 KB)
[v4] Tue, 17 Feb 2015 01:45:53 UTC (53 KB)

Computer Science > Data Structures and Algorithms

Title:Learning Poisson Binomial Distributions

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Data Structures and Algorithms

Title:Learning Poisson Binomial Distributions

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators