Multilevel Stochastic Optimization for Imputation in Massive Medical Data Records

被引:1
|
作者
Li, Wenrui [1 ]
Wang, Xiaoyu [1 ]
Sun, Yuetian [1 ]
Milanovic, Snezana [1 ,2 ]
Kon, Mark [1 ]
Castrillon-Candas, Julio Enrique [1 ]
机构
[1] Boston Univ, Dept Math & Stat, Boston, MA 02215 USA
[2] Sunov Pharmaceut, Marlborough, MA 01752 USA
基金
美国国家科学基金会;
关键词
Covariance matrices; Optimization; Stochastic processes; Deep learning; Iterative methods; Costs; Big Data; Best linear unbiased predictor; computational applied mathematics; machine learning; massive datasets; numerical stability; APPROXIMATION; EQUATIONS; PDES;
D O I
10.1109/TBDATA.2023.3328433
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
It has long been a recognized problem that many datasets contain significant levels of missing numerical data. A potentially critical predicate for application of machine learning methods to datasets involves addressing this problem. However, this is a challenging task. In this article, we apply a recently developed multi-level stochastic optimization approach to the problem of imputation in massive medical records. The approach is based on computational applied mathematics techniques and is highly accurate. In particular, for the Best Linear Unbiased Predictor (BLUP) this multi-level formulation is exact, and is significantly faster and more numerically stable. This permits practical application of Kriging methods to data imputation problems for massive datasets. We test this approach on data from the National Inpatient Sample (NIS) data records, Healthcare Cost and Utilization Project (HCUP), Agency for Healthcare Research and Quality. Numerical results show that the multi-level method significantly outperforms current approaches and is numerically robust. It has superior accuracy as compared with methods recommended in the recent report from HCUP. Benchmark tests show up to 75% reductions in error. Furthermore, the results are also superior to recent state of the art methods such as discriminative deep learning.
引用
收藏
页码:122 / 131
页数:10
相关论文
共 50 条
  • [1] MULTIPLE IMPUTATION FOR CATEGORICAL VARIABLES IN MULTILEVEL DATA
    Kottage, Helani Dilshara
    BULLETIN OF THE AUSTRALIAN MATHEMATICAL SOCIETY, 2022, 106 (02) : 349 - 350
  • [2] An efficient approach for imputation and classification of medical data values using class-based clustering of medical records
    Yelipe, UshaRani
    Porika, Sammulal
    Golla, Madhu
    COMPUTERS & ELECTRICAL ENGINEERING, 2018, 66 : 487 - 504
  • [3] Imputation of Mixed Data With Multilevel Singular Value Decomposition
    Husson, Francois
    Josse, Julie
    Narasimhan, Balasubramanian
    Robin, Genevieve
    JOURNAL OF COMPUTATIONAL AND GRAPHICAL STATISTICS, 2019, 28 (03) : 552 - 566
  • [4] An Imputation Measure For Data Imputation and Disease Classification of Medical Datasets
    Aljawarneh, Shadi
    Radhakrishna, Vangipuram
    Kumar, Gunupudi Rajesh
    INTERNATIONAL CONFERENCE ON KEY ENABLING TECHNOLOGIES (KEYTECH 2019), 2019, 2146
  • [5] Multiple Imputation for Multilevel Data with Continuous and Binary Variables
    Audigier, Vincent
    White, Ian R.
    Jolani, Shahab
    Debray, Thomas P. A.
    Quartagno, Matteo
    Carpenter, James
    van Buuren, Stef
    Resche-Rigon, Matthieu
    STATISTICAL SCIENCE, 2018, 33 (02) : 160 - 183
  • [6] Multiple imputation of binary multilevel missing not at random data
    Hammon, Angelina
    Zinn, Sabine
    JOURNAL OF THE ROYAL STATISTICAL SOCIETY SERIES C-APPLIED STATISTICS, 2020, 69 (03) : 547 - 564
  • [7] Optimization of Electronic Medical Records for Data Mining Using a Common Data Model
    Kwong, Manlik
    Gardner, Heather L.
    Dieterle, Neil
    Rentko, Virginia
    TOPICS IN COMPANION ANIMAL MEDICINE, 2019, 37
  • [8] Missing Value Imputation in Medical Records for Remote Health Care
    Das, Sayan
    Sil, Jaya
    DATA SCIENCE AND BIG DATA ANALYTICS, 2019, 16 : 321 - 331
  • [9] Electronic medical records imputation by temporal Generative Adversarial Network
    Yin, Yunfei
    Yuan, Zheng
    Tanvir, Islam Md
    Bao, Xianjian
    BIODATA MINING, 2024, 17 (01):
  • [10] DATA IMPUTATION: AN OPTIMIZATION APPROACH.
    Cooley, Philip C.
    International Journal on Policy and Information, 1987, 11 (01): : 39 - 45