Convolutional Embedded Networks for Population Scale Clustering and Bio-ancestry Inferencing

Karim, Md. Rezaul; Cochez, Michael; Zappa, Achille; Sahay, Ratnesh; Beyan, Oya; Schuhmann, Dietrich-Rebholz; Decker, Stefan

Computer Science > Machine Learning

arXiv:1805.12218 (cs)

[Submitted on 30 May 2018 (v1), last revised 19 Apr 2020 (this version, v2)]

Title:Convolutional Embedded Networks for Population Scale Clustering and Bio-ancestry Inferencing

Authors:Md. Rezaul Karim, Michael Cochez, Achille Zappa, Ratnesh Sahay, Oya Beyan, Dietrich-Rebholz Schuhmann, Stefan Decker

View PDF

Abstract:The study of genetic variants can help find correlating population groups to identify cohorts that are predisposed to common diseases and explain differences in disease susceptibility and how patients react to drugs. Machine learning algorithms are increasingly being applied to identify interacting GVs to understand their complex phenotypic traits. Since the performance of a learning algorithm not only depends on the size and nature of the data but also on the quality of underlying representation, deep neural networks can learn non-linear mappings that allow transforming GVs data into more clustering and classification friendly representations than manual feature selection. In this paper, we proposed convolutional embedded networks in which we combine two DNN architectures called convolutional embedded clustering and convolutional autoencoder classifier for clustering individuals and predicting geographic ethnicity based on GVs, respectively. We employed CAE-based representation learning on 95 million GVs from the 1000 genomes and Simons genome diversity projects. Quantitative and qualitative analyses with a focus on accuracy and scalability show that our approach outperforms state-of-the-art approaches such as VariantSpark and ADMIXTURE. In particular, CEC can cluster targeted population groups in 22 hours with an adjusted rand index of 0.915, the normalized mutual information of 0.92, and the clustering accuracy of 89%. Contrarily, the CAE classifier can predict the geographic ethnicity of unknown samples with an F1 and Mathews correlation coefficient(MCC) score of 0.9004 and 0.8245, respectively. To provide interpretations of the predictions, we identify significant biomarkers using gradient boosted trees(GBT) and SHAP. Overall, our approach is transparent and faster than the baseline methods, and scalable for 5% to 100% of the full human genome.

Comments:	This article is under review in IEEE/ACM Transactions on Computational Biology and Bioinformatics. It is based on a workshop paper discussed at the Extended Semantic Web Conference (ESWC'2017) workshop on Semantic Web Solutions for Large-scale Biomedical Data Analytics (SeWeBMeDA), Slovenia, May, 28-29, 2017
Subjects:	Machine Learning (cs.LG); Quantitative Methods (q-bio.QM); Machine Learning (stat.ML)
Cite as:	arXiv:1805.12218 [cs.LG]
	(or arXiv:1805.12218v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1805.12218

Submission history

From: Md. Rezaul Karim [view email]
[v1] Wed, 30 May 2018 20:30:13 UTC (7,523 KB)
[v2] Sun, 19 Apr 2020 19:18:51 UTC (2,030 KB)

Computer Science > Machine Learning

Title:Convolutional Embedded Networks for Population Scale Clustering and Bio-ancestry Inferencing

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Convolutional Embedded Networks for Population Scale Clustering and Bio-ancestry Inferencing

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators