A Hierarchical Dirichlet Process Mixture Model for Haplotype Reconstruction from Multi-Population Data

Sohn, Kyung-Ah; Xing, Eric P.

Statistics > Machine Learning

arXiv:0812.4648v1 (stat)

[Submitted on 26 Dec 2008 (this version), latest version 20 Aug 2009 (v2)]

Title:A Hierarchical Dirichlet Process Mixture Model for Haplotype Reconstruction from Multi-Population Data

Authors:Kyung-Ah Sohn, Eric P. Xing

View PDF

Abstract: Uncovering the haplotypes of single nucleotide polymorphisms is essential for many biological and medical applications. While it is uncommon for the genotype data to be pooled from multiple ethnically distinct populations, few existing programs have explicitly leverage the individual ethnic information for haplotype inference. In this paper, we present a new haplotype inference program, Haploi, which makes use of such information and is readily applicable to genotype sequences with thousands of SNPs from heterogeneous populations, with competent and sometimes superior speed and accuracy comparing to the state-of-the-art programs. Underlying Haploi is a new haplotype distribution model based on a nonparametric Bayesian formalism known as the hierarchical Dirichlet process, which represents a tractable surrogate to the coalescent process. The proposed model is exchangeable, unbounded, and capable of coupling demographic information of different populations. It offers a well-founded statistical framework for posterior inference of individual haplotypes, the size and configuration of haplotype ancestor pools, and other parameters of interest given genotype data.

Comments:	33 pages. 7 figures. To appear in the Annals of Applied Statistics
Subjects:	Machine Learning (stat.ML); Genomics (q-bio.GN); Quantitative Methods (q-bio.QM); Applications (stat.AP); Methodology (stat.ME)
Cite as:	arXiv:0812.4648 [stat.ML]
	(or arXiv:0812.4648v1 [stat.ML] for this version)
	https://doi.org/10.48550/arXiv.0812.4648

Submission history

From: Kyung-Ah Sohn [view email]
[v1] Fri, 26 Dec 2008 06:40:01 UTC (207 KB)
[v2] Thu, 20 Aug 2009 08:17:58 UTC (555 KB)

Statistics > Machine Learning

Title:A Hierarchical Dirichlet Process Mixture Model for Haplotype Reconstruction from Multi-Population Data

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Statistics > Machine Learning

Title:A Hierarchical Dirichlet Process Mixture Model for Haplotype Reconstruction from Multi-Population Data

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators