| Title: | Domain-aware matrix completion for phenotype imputation using electronic health record data with applications in genomic research |
| Journal: | The Annals of Applied Statistics |
| Published: | 1 Jun 2026 |
| DOI: | https://doi.org/10.1214/26-aoas2165 |
| Title: | Domain-aware matrix completion for phenotype imputation using electronic health record data with applications in genomic research |
| Journal: | The Annals of Applied Statistics |
| Published: | 1 Jun 2026 |
| DOI: | https://doi.org/10.1214/26-aoas2165 |
WARNING: the interactive features of this website use CSS3, which your browser does not support. To use the full features of this website, please update your browser.
This document contains (1) technical results on the strong convexity and Lipschitz continuous gradient of the covImpute objective function; (2) a discussion of the numerical complexity of the fast gradient method; (3) an overview of existing phenotype prediction methods; (4) a sensitivity analysis of genetic covariance perturbations in UKBB analyses; (5) further simulations evaluating FPR under heterogeneous genetic structure and covariance misspecification; (6) additional figures and tables from simulations and real-data analyses on UKBB. We provide Python code for the proposed method and for replicating the simulation results. The code for covImpute is also available on Github: https://github.com/Slangevar/covImpute. Large-scale biobanks and electronic health records (EHR) offer great opportunities for next-generation genetic studies. However, missing phenotype data is a pervasive feature of EHR, leading to low power of such studies. One promising solution is prediction-powered inference, where statistical or machine learning models are employed to impute phenotypes prior to performing genetic analyses. Although many such methods exist, they tend to be generic and do not incorporate domain-aware knowledge to optimize their performance for downstream genetic analyses. We propose a novel matrix completion method, covImpute, which, unlike generic matrix completion methods such as softImpute, incorporates external information in the form of a genetic covariance matrix among phenotypic features and imputes missing entries with latent genetic components. We compare covImpute with existing methods, including a domain-aware liability threshold model LTPI, and generic softImpute and deep learning autoencoder models in simulations under different missingness mechanisms with respect to power in downstream genetic analyses. In applications to several diseases in UK Biobank, we show that genetically informed methods, such as covImpute and LTPI, can perform substantially better in terms of power of genetic association studies relative to generic imputation models currently in use. Moreover, compared to LTPI, covImpute's flexible framework for incorporating external covariance information provides a more general approach with applicability beyond genetics.</p>
Enabling scientific discoveries that improve human health