Conflation of short identity-by-descent segments bias their inferred length distribution

Chiang, Charleston W. K.; Ralph, Peter; Novembre, John

doi:10.1534/g3.116.027581

Quantitative Biology > Populations and Evolution

arXiv:1410.5313 (q-bio)

[Submitted on 20 Oct 2014 (v1), last revised 18 Aug 2015 (this version, v2)]

Title:Conflation of short identity-by-descent segments bias their inferred length distribution

Authors:Charleston W.K. Chiang, Peter Ralph, John Novembre

View PDF

Abstract:Identity-by-descent (IBD) is a fundamental concept in genetics with many applications. In a common definition, two haplotypes are said to contain an IBD segment if they share a segment that is inherited from a recent shared common ancestor without intervening recombination. Long IBD segments (> 1cM) can be efficiently detected by a number of algorithms using high-density SNP array data from a population sample. However, these approaches detect IBD based on contiguous segments of identity-by-state, and such segments may exist due to the conflation of smaller, nearby IBD segments. We quantified this effect using coalescent simulations, finding that nearly 40% of inferred segments 1-2cM long are results of conflations of two or more shorter segments, under demographic scenarios typical for modern humans. This biases the inferred IBD segment length distribution, and so can affect downstream inferences. We observed this conflation effect universally across different IBD detection programs and human demographic histories, and found inference of segments longer than 2cM to be much more reliable (less than 5% conflation rate). As an example of how this can negatively affect downstream analyses, we present and analyze a novel estimator of the de novo mutation rate using IBD segments, and demonstrate that the biased length distribution of the IBD segments due to conflation can lead to inflated estimates if the conflation is not modeled. Understanding the conflation effect in detail will make its correction in future methods more tractable.

Subjects:	Populations and Evolution (q-bio.PE)
Cite as:	arXiv:1410.5313 [q-bio.PE]
	(or arXiv:1410.5313v2 [q-bio.PE] for this version)
	https://doi.org/10.48550/arXiv.1410.5313
Journal reference:	G3 May 1, 2016 vol. 6 no. 5 1287-1296
Related DOI:	https://doi.org/10.1534/g3.116.027581

Submission history

From: Charleston Chiang [view email]
[v1] Mon, 20 Oct 2014 15:25:19 UTC (835 KB)
[v2] Tue, 18 Aug 2015 01:58:51 UTC (608 KB)

Quantitative Biology > Populations and Evolution

Title:Conflation of short identity-by-descent segments bias their inferred length distribution

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Quantitative Biology > Populations and Evolution

Title:Conflation of short identity-by-descent segments bias their inferred length distribution

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators