BOLT-SSI: A STATISTICAL APPROACH TO SCREENING INTERACTION EFFECTS FOR ULTRA-HIGH DIMENSIONAL DATA

被引:2
|
作者
Zhou, Min [1 ]
Dai, Mingwei [2 ,3 ]
Yao, Yuan [4 ]
Liu, Jin [5 ]
Yang, Can [6 ]
Peng, Heng [7 ]
机构
[1] Beijing Normal Univ, Hong Kong Baptist Univ United Int Coll, Zhuhai 519088, Guangdong, Peoples R China
[2] Southwestern Univ Finance & Econ, Ctr Stat Res, Chengdu 610074, Sichuan, Peoples R China
[3] Southwestern Univ Finance & Econ, Sch Stat, Chengdu 610074, Sichuan, Peoples R China
[4] Victoria Univ Wellington, Sch Math & Stat, Wellington 6012, New Zealand
[5] Duke NUS Grad Med Sch, Singapore 169857, Singapore
[6] Hong Kong Univ Sci & Technol, Kowloon, Clear Water Bay, Hong Kong, Peoples R China
[7] Hong Kong Baptist Univ, Kowloon Tong, Kowloon, Hong Kong, Peoples R China
关键词
interaction detection; trade-off between statistical efficiency and computational; complexity; NONCONCAVE PENALIZED LIKELIHOOD; VARIABLE SELECTION; REGRESSION; MODELS;
D O I
10.5705/ss.202020.0498
中图分类号
O21 [概率论与数理统计]; C8 [统计学];
学科分类号
020208 ; 070103 ; 0714 ;
摘要
Detecting the interaction effects among the predictors on the response variable is a crucial step in numerous applications. We first propose a simple method for sure screening interactions (SSI). Although its computation complexity is O(p2n), the SSI method works well for problems of moderate dimensionality (e.g., p = 103 similar to 104), without the heredity assumption. For ultrahigh-dimensional problems (e.g., p = 106), motivated by a discretization associated Boolean representation and operations and a contingency table for discrete variables, we propose a fast algorithm, called "BOLT-SSI." The statistical theory is established for SSI and BOLT-SSI, guaranteeing their sure screening property. We evaluate the performance of SSI and BOLT-SSI using comprehensive simulations and real case studies. Our numerical results demonstrate that SSI and BOLT-SSI often outperform their competitors in terms of computational efficiency and statistical accuracy. The proposed method can be applied to fully detect interactions with more than 300,000 predictors. Based on our findings, we believe there is a need to rethink the relationship between statistical accuracy and computational efficiency. We have shown that the computational performance of a statistical method can often be greatly improved by exploring the advantages of computational architecture with a tolerable loss of statistical accuracy.
引用
收藏
页码:2327 / 2358
页数:32
相关论文
共 50 条
  • [41] Bayesian Multiresolution Variable Selection for Ultra-High Dimensional Neuroimaging Data
    Zhao, Yize
    Kang, Jian
    Long, Qi
    IEEE-ACM TRANSACTIONS ON COMPUTATIONAL BIOLOGY AND BIOINFORMATICS, 2018, 15 (02) : 537 - 550
  • [42] An ultra-high throughput screening approach for an adenine transferase using fluorescence polarization
    Li, ZY
    Mehdi, S
    Patel, I
    Kawooya, J
    Judkins, M
    Zhang, WH
    Diener, K
    Lozada, A
    Dunnington, D
    JOURNAL OF BIOMOLECULAR SCREENING, 2000, 5 (01) : 31 - 37
  • [43] Combined performance of screening and variable selection methods in ultra-high dimensional data in predicting time-to-event outcomes
    Lira Pi
    Susan Halabi
    Diagnostic and Prognostic Research, 2 (1)
  • [44] A new joint screening method for right-censored time-to-event data with ultra-high dimensional covariates
    Liu, Yi
    Chen, Xiaolin
    Li, Gang
    STATISTICAL METHODS IN MEDICAL RESEARCH, 2020, 29 (06) : 1499 - 1513
  • [45] Quantile-adaptive variable screening in ultra-high dimensional varying coefficient models
    Zhang, Junying
    Zhang, Riquan
    Lu, Zhiping
    JOURNAL OF APPLIED STATISTICS, 2016, 43 (04) : 643 - 654
  • [46] Model Based Screening Embedded Bayesian Variable Selection for Ultra-high Dimensional Settings
    Li, Dongjin
    Dutta, Somak
    Roy, Vivekananda
    JOURNAL OF COMPUTATIONAL AND GRAPHICAL STATISTICS, 2023, 32 (01) : 61 - 73
  • [47] A sure independence screening procedure for ultra-high dimensional partially linear additive models
    Kazemi, M.
    Shahsavani, D.
    Arashi, M.
    JOURNAL OF APPLIED STATISTICS, 2019, 46 (08) : 1385 - 1403
  • [48] Block-diagonal precision matrix regularization for ultra-high dimensional data
    Yang, Yihe
    Dai, Hongsheng
    Pan, Jianxin
    COMPUTATIONAL STATISTICS & DATA ANALYSIS, 2023, 179
  • [49] Sparsity identification in ultra-high dimensional quantile regression models with longitudinal data
    Gao, Xianli
    Liu, Qiang
    COMMUNICATIONS IN STATISTICS-THEORY AND METHODS, 2020, 49 (19) : 4712 - 4736
  • [50] Screening of the most relevant parameters for method development in ultra-high,performance hydrophilic interaction chromatography
    Periat, Aurelie
    Debrus, Benjamin
    Rudaz, Serge
    Guillarme, Davy
    JOURNAL OF CHROMATOGRAPHY A, 2013, 1282 : 72 - 83