The Effect of Different Dimensionality Reduction Techniques on Machine Learning Overfitting Problem

被引:0
|
作者
Salam, Mustafa Abdul [1 ]
Azar, Ahmad Taher [2 ,3 ]
Elgendy, Mustafa Samy [4 ]
Fouad, Khaled Mohamed [5 ]
机构
[1] Benha Univ, Fac Comp & Artificial Intelligence, Artificial Intelligence Dept, Banha, Egypt
[2] Benha Univ, Fac Comp & Artificial Intelligence, Banha, Egypt
[3] Prince Sultan Univ, Coll Comp & Informat Sci, Riyadh, Saudi Arabia
[4] Benha Univ, Sci Comp Dept, Fac Comp & Artificial Intelligence, Banha, Egypt
[5] Benha Univ, Informat Syst Dept, Fac Comp & Artificial Intelligence, Banha, Egypt
关键词
Dimensionality reduction; feature subset selection; rough set; overfitting; underfitting; machine learning; PRINCIPAL COMPONENT ANALYSIS; ALGORITHM;
D O I
暂无
中图分类号
TP301 [理论、方法];
学科分类号
081202 ;
摘要
In most conditions, it is a problematic mission for a machine-learning model with a data record, which has various attributes, to be trained. There is always a proportional relationship between the increase of model features and the arrival to the overfitting of the susceptible model. That observation occurred since not all the characteristics are always important. For example, some features could only cause the data to be noisier. Dimensionality reduction techniques are used to overcome this matter. This paper presents a detailed comparative study of nine dimensionality reduction methods. These methods are missing-values ratio, low variance filter, highcorrelation filter, random forest, principal component analysis, linear discriminant analysis, backward feature elimination, forward feature construction, and rough set theory. The effects of used methods on both training and testing performance were compared with two different datasets and applied to three different models. These models are, Artificial Neural Network (ANN), Support Vector Machine (SVM) and Random Forest classifier (RFC). The results proved that the RFC model was able to achieve the dimensionality reduction via limiting the overfitting crisis. The introduced RFC model showed a general progress in both accuracy and efficiency against compared approaches. The results revealed that dimensionality reduction could minimize the overfitting process while holding the performance so near to or better than the original one.
引用
收藏
页码:641 / 655
页数:15
相关论文
共 50 条
  • [21] Ship Engine Model Selection by Applying Machine Learning Classification Techniques Using Imputation and Dimensionality Reduction
    Skarlatos, Kyriakos
    Papageorgiou, Grigorios
    Biris, Panagiotis
    Skamnia, Ekaterini
    Economou, Polychronis
    Bersimis, Sotirios
    JOURNAL OF MARINE SCIENCE AND ENGINEERING, 2024, 12 (01)
  • [22] Machine Learning Models and Dimensionality Reduction for Prediction of Polymer Properties
    Mysona, Joshua A.
    Nealey, Paul F.
    de Pablo, Juan J.
    MACROMOLECULES, 2024, 57 (05) : 1988 - 1997
  • [23] Dimensionality Reduction for Machine Learning Based IoT Botnet Detection
    Bahsi, Hayretdin
    Nomm, Sven
    La Torre, Fabio Benedetto
    2018 15TH INTERNATIONAL CONFERENCE ON CONTROL, AUTOMATION, ROBOTICS AND VISION (ICARCV), 2018, : 1857 - 1862
  • [24] Dimensionality Reduction with Extreme Learning Machine Based on Manifold Preserving
    Li, Canyao
    Lv, Jujian
    Zhao, Huimin
    Chen, Rongjun
    Zhan, Jin
    Li, Kaihan
    ADVANCES IN BRAIN INSPIRED COGNITIVE SYSTEMS, 2020, 11691 : 128 - 138
  • [25] Dimensionality Reduction for Sentiment Classification using Machine Learning Classifiers
    Islam, Mazharul
    Anjum, Aftab
    Ahsan, Tanveer
    Wang, Lin
    2019 IEEE SYMPOSIUM SERIES ON COMPUTATIONAL INTELLIGENCE (IEEE SSCI 2019), 2019, : 3097 - 3103
  • [26] A Privacy Preserving Scheme with Dimensionality Reduction for Distributed Machine Learning
    Chen, Zhaoheng
    Omote, Kazumasa
    2021 16TH ASIA JOINT CONFERENCE ON INFORMATION SECURITY (ASIAJCIS 2021), 2021, : 45 - 50
  • [27] Materials informatics of woven fabric composites: Effect of different dimensionality reduction and learning methods
    Olfatbakhsh, T.
    Andrews, J. L.
    Milani, A. S.
    MATERIALS TODAY COMMUNICATIONS, 2022, 32
  • [28] Seeing is Learning in High Dimensions: The Synergy Between Dimensionality Reduction and Machine Learning
    Telea A.
    Machado A.
    Wang Y.
    SN Computer Science, 5 (3)
  • [29] NETWORK INTRUSION DETECTION SYSTEMS USING SUPERVISED MACHINE LEARNING CLASSIFICATION AND DIMENSIONALITY REDUCTION TECHNIQUES: A SYSTEMATIC REVIEW
    Ashi, Zein
    Aburashed, Laila
    Al-Qudah, Mahmoud
    Qusef, Abdallah
    JORDANIAN JOURNAL OF COMPUTERS AND INFORMATION TECHNOLOGY, 2021, 7 (04): : 373 - 390
  • [30] Multivariate and Dimensionality-Reduction-Based Machine Learning Techniques for Tumor Classification of RNA-Seq Data
    Al-khassaweneh, Mahmood
    Bronakowski, Mark
    Al-Sharoa, Esraa
    APPLIED SCIENCES-BASEL, 2023, 13 (23):