Feature Learning With a Divergence-Encouraging Autoencoder for Imbalanced Data Classification

被引:5
|
作者
Luo, Ruisen [1 ]
Feng, Qian [1 ]
Wang, Chen [1 ,2 ]
Yang, Xiaomei [1 ]
Tu, Haiyan [1 ]
Yu, Qin [1 ]
Fei, Shaomin [3 ,4 ]
Gong, Xiaofeng [1 ]
机构
[1] Sichuan Univ, Coll Elect Engn & Informat Technol, Chengdu 610064, Sichuan, Peoples R China
[2] UCL, Dept Comp Sci, London WC1E 6BT, England
[3] Chengdu Univ Informat Technol, Expt Ctr Elect, Chengdu 610059, Sichuan, Peoples R China
[4] DaGongBoChuang Corp, Chengdu 610005, Sichuan, Peoples R China
来源
IEEE ACCESS | 2018年 / 6卷
关键词
Imbalanced data classification; autoencoder; divergence loss; convergence analysis; alternating training paradigm;
D O I
10.1109/ACCESS.2018.2879221
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Imbalanced data exists commonly in machine learning classification applications. Popular classification algorithms are based on the assumption that data in different classes are roughly equally distributed; however, extremely skewed data, with instances from one class taking up most of the dataset, is not exceptional in practice. Thus, performance of algorithm often degrades significantly when encountering skewed data. Mitigating the problem caused by imbalanced data has been an open challenge for years, and previous researches mostly have proposed solutions from the perspectives of data re-sampling and algorithm improvement. In this paper, focusing on two-class imbalanced data, we have proposed a novel divergence-encouraging autoencoder (DEA) to explicitly learn features from both of the two classes and have designed an imbalanced data classification algorithm based on the proposed autoencoder. By encouraging maximization of divergence loss between different classes in the bottleneck layer, the proposed DEA can learn features for both majority and minority classes simultaneously. The training procedure of the proposed autoencoder is to alternately optimize reconstruction and divergence losses. After obtaining the features, we directly compute the cosine distances between the training and testing features and compare the median of distances between classes to perform classification. Experimental results illustrate that our algorithm outperform ordinary and loss-sensitive CNN models both in terms of performance evaluation metrics and convergence properties. To the best of our knowledge, this is the first paper proposed to solve the imbalanced data classification problem from the perspective of explicitly learning representations of different classes simultaneously. In addition, designing of the proposed DEA is also an innovative work, which could improve the performance of imbalanced data classification without data re-sampling and benefit future researches in the field.
引用
收藏
页码:70197 / 70211
页数:15
相关论文
共 50 条
  • [31] Imbalanced fault diagnosis of rotating machinery using autoencoder-based SuperGraph feature learning
    Liu, Jie
    Zhou, Kaibo
    Yang, Chaoying
    Lu, Guoliang
    FRONTIERS OF MECHANICAL ENGINEERING, 2021, 16 (04) : 829 - 839
  • [32] An Improved Extreme Learning Machine for Imbalanced Data Classification
    Zhang, Xiaopeng
    Qin, Liangxi
    IEEE ACCESS, 2022, 10 : 8634 - 8642
  • [33] Meta-learning for imbalanced data and classification ensemble in binary classification
    Lin, Sung-Chiang
    Chang, Yuan-chin I.
    Yang, Wei-Ning
    NEUROCOMPUTING, 2009, 73 (1-3) : 484 - 494
  • [34] A novel oversampling and feature selection hybrid algorithm for imbalanced data classification
    Feng, Fang
    Li, Kuan-Ching
    Yang, Erfu
    Zhou, Qingguo
    Han, Lihong
    Hussain, Amir
    Cai, Mingjiang
    MULTIMEDIA TOOLS AND APPLICATIONS, 2023, 82 (03) : 3231 - 3267
  • [35] Weighted ReliefF with threshold constraints of feature selection for imbalanced data classification
    Song, Yan
    Si, Weiyun
    Dai, Feifan
    Yang, Guisong
    CONCURRENCY AND COMPUTATION-PRACTICE & EXPERIENCE, 2020, 32 (14):
  • [36] A novel oversampling and feature selection hybrid algorithm for imbalanced data classification
    Fang Feng
    Kuan-Ching Li
    Erfu Yang
    Qingguo Zhou
    Lihong Han
    Amir Hussain
    Mingjiang Cai
    Multimedia Tools and Applications, 2023, 82 : 3231 - 3267
  • [37] Adaptive Learning in Imbalanced Data Streams With Unpredictable Feature Evolution
    Tu, Jiahang
    Tang, Xijia
    Gu, Shilin
    Dai, Yucong
    Fan, Ruidong
    Hou, Chenping
    IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, 2025, 37 (04) : 1527 - 1541
  • [38] Classification model for imbalanced traffic data based on secondary feature extraction
    Shen, Jian
    Xia, Jingbo
    Shan, Yong
    Wei, Zekun
    IET COMMUNICATIONS, 2017, 11 (11) : 1725 - 1731
  • [39] Iterative ensemble feature selection for multiclass classification of imbalanced microarray data
    Yang, Junshan
    Zhou, Jiarui
    Zhu, Zexuan
    Ma, Xiaoliang
    Ji, Zhen
    JOURNAL OF BIOLOGICAL RESEARCH-THESSALONIKI, 2016, 23
  • [40] ContrastNet: Unsupervised feature learning by autoencoder and prototypical contrastive learning for hyperspectral imagery classification
    Cao, Zeyu
    Li, Xiaorun
    Feng, Yueming
    Chen, Shuhan
    Xia, Chaoqun
    Zhao, Liaoying
    NEUROCOMPUTING, 2021, 460 : 71 - 83