Feature Learning With a Divergence-Encouraging Autoencoder for Imbalanced Data Classification

被引：5

作者：

Luo, Ruisen ^{[1
]}

Feng, Qian ^{[1
]}

Wang, Chen ^{[1
,2
]}

Yang, Xiaomei ^{[1
]}

Tu, Haiyan ^{[1
]}

Yu, Qin ^{[1
]}

Fei, Shaomin ^{[3
,4
]}

Gong, Xiaofeng ^{[1
]}

机构：

[1] Sichuan Univ, Coll Elect Engn & Informat Technol, Chengdu 610064, Sichuan, Peoples R China

[2] UCL, Dept Comp Sci, London WC1E 6BT, England

[3] Chengdu Univ Informat Technol, Expt Ctr Elect, Chengdu 610059, Sichuan, Peoples R China

[4] DaGongBoChuang Corp, Chengdu 610005, Sichuan, Peoples R China

来源：

IEEE ACCESS | 2018年 / 6卷

关键词：

Imbalanced data classification; autoencoder; divergence loss; convergence analysis; alternating training paradigm;

D O I：

10.1109/ACCESS.2018.2879221

中图分类号：

TP [自动化技术、计算机技术];

学科分类号：

0812 ;

摘要：

Imbalanced data exists commonly in machine learning classification applications. Popular classification algorithms are based on the assumption that data in different classes are roughly equally distributed; however, extremely skewed data, with instances from one class taking up most of the dataset, is not exceptional in practice. Thus, performance of algorithm often degrades significantly when encountering skewed data. Mitigating the problem caused by imbalanced data has been an open challenge for years, and previous researches mostly have proposed solutions from the perspectives of data re-sampling and algorithm improvement. In this paper, focusing on two-class imbalanced data, we have proposed a novel divergence-encouraging autoencoder (DEA) to explicitly learn features from both of the two classes and have designed an imbalanced data classification algorithm based on the proposed autoencoder. By encouraging maximization of divergence loss between different classes in the bottleneck layer, the proposed DEA can learn features for both majority and minority classes simultaneously. The training procedure of the proposed autoencoder is to alternately optimize reconstruction and divergence losses. After obtaining the features, we directly compute the cosine distances between the training and testing features and compare the median of distances between classes to perform classification. Experimental results illustrate that our algorithm outperform ordinary and loss-sensitive CNN models both in terms of performance evaluation metrics and convergence properties. To the best of our knowledge, this is the first paper proposed to solve the imbalanced data classification problem from the perspective of explicitly learning representations of different classes simultaneously. In addition, designing of the proposed DEA is also an innovative work, which could improve the performance of imbalanced data classification without data re-sampling and benefit future researches in the field.

引用

页码：70197 / 70211

页数：15

共 50 条

[21] An Improved Ensemble Learning for Imbalanced Data Classification
Yuan, Zhengwu
Zhao, Pu
PROCEEDINGS OF 2019 IEEE 8TH JOINT INTERNATIONAL INFORMATION TECHNOLOGY AND ARTIFICIAL INTELLIGENCE CONFERENCE (ITAIC 2019), 2019, : 408 - 411
[22] Unsupervised Feature Learning for Heart Sounds Classification Using Autoencoder
Hu, Wei
Lv, Jiancheng
Liu, Dongbo
Chen, Yao
2ND INTERNATIONAL CONFERENCE ON MACHINE VISION AND INFORMATION TECHNOLOGY (CMVIT 2018), 2018, 1004
[23] A feature selection method to handle imbalanced data in text classification
Chang, Fengxiang
Guo, Jun
Xu, Weiran
Yao, Kejun
Journal of Digital Information Management, 2015, 13 (03): : 169 - 175
[24] FIAO: Feature Information Aggregation Oversampling for imbalanced data classification
Wang, Fei
Zheng, Ming
Hu, Xiaowen
Li, Hongchao
Wang, Taochun
Chen, Fulong
APPLIED SOFT COMPUTING, 2024, 161
[25] ADA-INCVAE: Improved data generation using variational autoencoder for imbalanced classification
Huang, Kai
Wang, Xiaoguo
APPLIED INTELLIGENCE, 2022, 52 (03) : 2838 - 2853
[26] Imbalanced fault diagnosis of rotating machinery using autoencoder-based SuperGraph feature learning
Jie LIU
Kaibo ZHOU
Chaoying YANG
Guoliang LU
Frontiers of Mechanical Engineering, 2021, (04) : 829 - 839
[27] Classification of Imbalanced Data Using SMOTE and AutoEncoder Based Deep Convolutional Neural Network
Alex, Suja A.
Nayahi, J. Jesu Vedha
INTERNATIONAL JOURNAL OF UNCERTAINTY FUZZINESS AND KNOWLEDGE-BASED SYSTEMS, 2023, 31 (03) : 437 - 469
[28] ADA-INCVAE: Improved data generation using variational autoencoder for imbalanced classification
Kai Huang
Xiaoguo Wang
Applied Intelligence, 2022, 52 : 2838 - 2853
[29] Imbalanced Data Classification Method Based on Ensemble Learning
Xiang, Yu
Xie, Yongping
COMMUNICATIONS, SIGNAL PROCESSING, AND SYSTEMS, CSPS 2018, VOL III: SYSTEMS, 2020, 517 : 18 - 24
[30] Imbalanced fault diagnosis of rotating machinery using autoencoder-based SuperGraph feature learning
Jie Liu
Kaibo Zhou
Chaoying Yang
Guoliang Lu
Frontiers of Mechanical Engineering, 2021, 16 : 829 - 839

← 1 2 3 4 5 →