Open Access
2021 Estimating high-dimensional covariance and precision matrices under general missing dependence
Seongoh Park, Xinlei Wang, Johan Lim
Author Affiliations +
Electron. J. Statist. 15(2): 4868-4915 (2021). DOI: 10.1214/21-EJS1892

Abstract

A sample covariance matrix S of completely observed data is the key statistic in a large variety of multivariate statistical procedures, such as structured covariance/precision matrix estimation, principal component analysis, and testing of equality of mean vectors. However, when the data are partially observed, the sample covariance matrix from the available data is biased and does not provide valid multivariate procedures. To correct the bias, a simple adjustment method called inverse probability weighting (IPW) has been used in previous research, yielding the IPW estimator. The estimator can play the role of S in the missing data context, thus replacing S in off-the-shelf multivariate procedures such as the graphical lasso algorithm. However, theoretical properties (e.g. concentration) of the IPW estimator have been only established in earlier work under very simple missing structures; every variable of each sample is independently subject to missingness with equal probability. We investigate the deviation of the IPW estimator when observations are partially observed under general missing dependency. We prove the optimal convergence rate Op( logpn) of the IPW estimator based on the element-wise maximum norm, even when two unrealistic assumptions (known mean and/or missing probabilities) frequently assumed to be known in the past work are relaxed. The optimal rate is especially crucial in estimating a precision matrix, because of the “meta-theorem” [26] that claims the rate of the IPW estimator governs that of the resulting precision matrix estimator. In the simulation study, we discuss one of practically important issues, non-positive semi-definiteness of the IPW estimator, and compare the estimator with imputation methods.

Funding Statement

Seongoh Park is supported by the Sungshin Women’s University Research Grant of H20210143, Johan Lim is supported by the National Research Foundation of Korea (NRF) grant funded by the Korea government (MSIT) (NRF-2017R1A2B2012264).

Acknowledgments

We appreciate the two anonymous reviewers and the Associate Editor whose comments helped us improve this manuscript significantly.

Citation

Download Citation

Seongoh Park. Xinlei Wang. Johan Lim. "Estimating high-dimensional covariance and precision matrices under general missing dependence." Electron. J. Statist. 15 (2) 4868 - 4915, 2021. https://doi.org/10.1214/21-EJS1892

Information

Received: 1 April 2021; Published: 2021
First available in Project Euclid: 19 October 2021

Digital Object Identifier: 10.1214/21-EJS1892

Subjects:
Primary: 62H12
Secondary: 60E15

Keywords: convergence rate , Covariance matrix , dependent missing structure , element-wise maximum norm , inverse probability weighting

Vol.15 • No. 2 • 2021
Back to Top