Multilevel Stochastic Optimization for Imputation in Massive Medical Data Records

被引:1
|
作者
Li, Wenrui [1 ]
Wang, Xiaoyu [1 ]
Sun, Yuetian [1 ]
Milanovic, Snezana [1 ,2 ]
Kon, Mark [1 ]
Castrillon-Candas, Julio Enrique [1 ]
机构
[1] Boston Univ, Dept Math & Stat, Boston, MA 02215 USA
[2] Sunov Pharmaceut, Marlborough, MA 01752 USA
基金
美国国家科学基金会;
关键词
Covariance matrices; Optimization; Stochastic processes; Deep learning; Iterative methods; Costs; Big Data; Best linear unbiased predictor; computational applied mathematics; machine learning; massive datasets; numerical stability; APPROXIMATION; EQUATIONS; PDES;
D O I
10.1109/TBDATA.2023.3328433
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
It has long been a recognized problem that many datasets contain significant levels of missing numerical data. A potentially critical predicate for application of machine learning methods to datasets involves addressing this problem. However, this is a challenging task. In this article, we apply a recently developed multi-level stochastic optimization approach to the problem of imputation in massive medical records. The approach is based on computational applied mathematics techniques and is highly accurate. In particular, for the Best Linear Unbiased Predictor (BLUP) this multi-level formulation is exact, and is significantly faster and more numerically stable. This permits practical application of Kriging methods to data imputation problems for massive datasets. We test this approach on data from the National Inpatient Sample (NIS) data records, Healthcare Cost and Utilization Project (HCUP), Agency for Healthcare Research and Quality. Numerical results show that the multi-level method significantly outperforms current approaches and is numerically robust. It has superior accuracy as compared with methods recommended in the recent report from HCUP. Benchmark tests show up to 75% reductions in error. Furthermore, the results are also superior to recent state of the art methods such as discriminative deep learning.
引用
收藏
页码:122 / 131
页数:10
相关论文
共 50 条
  • [21] Salvaging Data Records with Missing Data: Data Imputation using the Multivariate t Distribution
    Hooke, Melissa
    Mrozinski, Joseph
    DiNicola, Michael
    2021 IEEE AEROSPACE CONFERENCE (AEROCONF 2021), 2021,
  • [22] OPTIMIZATION OF MULTILEVEL ORGANIZATION OF DATA
    SHVETSOV, VI
    ENGINEERING CYBERNETICS, 1978, 16 (04): : 18 - 22
  • [23] Embracing Massive Medical Data
    Chou, Yu-Cheng
    Zhou, Zongwei
    Yuille, Alan
    MEDICAL IMAGE COMPUTING AND COMPUTER ASSISTED INTERVENTION - MICCAI 2024, PT I, 2024, 15001 : 24 - 35
  • [24] Multiple Imputation of Multilevel Missing Data-Rigor Versus Simplicity
    Drechsler, Joerg
    JOURNAL OF EDUCATIONAL AND BEHAVIORAL STATISTICS, 2015, 40 (01) : 69 - 95
  • [25] Multiple imputation of incomplete multilevel data using Heckman selection models
    Munoz, Johanna
    Efthimiou, Orestis
    Audigier, Vincent
    de Jong, Valentijn M. T.
    Debray, Thomas P. A.
    STATISTICS IN MEDICINE, 2024, 43 (03) : 514 - 533
  • [26] Multiple Imputation of Multilevel Missing Data: An Introduction to the R Package pan
    Grund, Simon
    Luedtke, Oliver
    Robitzsch, Alexander
    SAGE OPEN, 2016, 6 (04):
  • [27] Multiple Imputation of Missing Data in Multilevel Designs: A Comparison of Different Strategies
    Luedtke, Oliver
    Robitzsch, Alexander
    Grund, Simon
    PSYCHOLOGICAL METHODS, 2017, 22 (01) : 141 - 165
  • [28] Multiple imputation by chained equations for systematically and sporadically missing multilevel data
    Resche-Rigon, Matthieu
    White, Ian R.
    STATISTICAL METHODS IN MEDICAL RESEARCH, 2018, 27 (06) : 1634 - 1649
  • [29] An Imputation Approach to Electronic Medical Records Based on Time Series and Feature Association
    Yin, Y. F.
    Yuan, Z. W.
    Yang, J. X.
    Bao, X. J.
    12TH ASIAN-PACIFIC CONFERENCE ON MEDICAL AND BIOLOGICAL ENGINEERING, VOL 2, APCMBE 2023, 2024, 104 : 259 - 276
  • [30] Applying Stochastic Process Model to Imputation of Censored Longitudinal Data
    Zhbannikov, Ilya
    Arbeev, Konstantin
    Yashin, Anatoliy
    ACM-BCB'18: PROCEEDINGS OF THE 2018 ACM INTERNATIONAL CONFERENCE ON BIOINFORMATICS, COMPUTATIONAL BIOLOGY, AND HEALTH INFORMATICS, 2018, : 457 - 464