Targeted Learning with Daily EHR Data

Sofrygin, Oleg; Zhu, Zheng; Schmittdiel, Julie A; Adams, Alyce S.; Grant, Richard W.; van der Laan, Mark J.; Neugebauer, Romain

Statistics > Applications

arXiv:1705.09874v1 (stat)

[Submitted on 27 May 2017 (this version), latest version 14 Dec 2018 (v2)]

Title:Targeted Learning with Daily EHR Data

Authors:Oleg Sofrygin, Zheng Zhu, Julie A Schmittdiel, Alyce S. Adams, Richard W. Grant, Mark J. van der Laan, Romain Neugebauer

View PDF

Abstract:Electronic health records (EHR) data provide a cost and time-effective opportunity to conduct cohort studies of the effects of multiple time-point interventions in the diverse patient population found in real-world clinical settings. Because the computational cost of analyzing EHR data at daily (or more granular) scale can be quite high, a pragmatic approach has been to partition the follow-up into coarser intervals of pre-specified length. Current guidelines suggest employing a 'small' interval, but the feasibility and practical impact of this recommendation has not been evaluated and no formal methodology to inform this choice has been developed. We start filling these gaps by leveraging large-scale EHR data from a diabetes study to develop and illustrate a fast and scalable targeted learning approach that allows to follow the current recommendation and study its practical impact on inference. More specifically, we map daily EHR data into four analytic datasets using 90, 30, 15 and 5-day intervals. We apply a semi-parametric and doubly robust estimation approach, the longitudinal TMLE, to estimate the causal effects of four dynamic treatment rules with each dataset, and compare the resulting inferences. To overcome the computational challenges presented by the size of these data, we propose a novel TMLE implementation, the 'long-format TMLE', and rely on the latest advances in scalable data-adaptive machine-learning software, xgboost and h2o, for estimation of the TMLE nuisance parameters.

Subjects:	Applications (stat.AP); Computation (stat.CO); Machine Learning (stat.ML)
Cite as:	arXiv:1705.09874 [stat.AP]
	(or arXiv:1705.09874v1 [stat.AP] for this version)
	https://doi.org/10.48550/arXiv.1705.09874

Submission history

From: Oleg Sofrygin [view email]
[v1] Sat, 27 May 2017 22:43:08 UTC (51 KB)
[v2] Fri, 14 Dec 2018 21:53:24 UTC (86 KB)

Statistics > Applications

Title:Targeted Learning with Daily EHR Data

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Statistics > Applications

Title:Targeted Learning with Daily EHR Data

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators