Three-Stage Sampling Algorithm for Highly Imbalanced Multi-Classification Time Series Datasets

被引:1
|
作者
Wang, Haoming [1 ]
机构
[1] Guangdong Univ Educ, Sch Math, Dept Stat & Financial Math, 351,Xingang Middle Rd, Guangzhou 510303, Peoples R China
来源
SYMMETRY-BASEL | 2023年 / 15卷 / 10期
关键词
imbalanced data; data preprocessing; sampling; Tomek links; DTW; SMOTE;
D O I
10.3390/sym15101849
中图分类号
O [数理科学和化学]; P [天文学、地球科学]; Q [生物科学]; N [自然科学总论];
学科分类号
07 ; 0710 ; 09 ;
摘要
To alleviate the data imbalance problem caused by subjective and objective factors, scholars have developed different data-preprocessing algorithms, among which undersampling algorithms are widely used because of their fast and efficient performance. However, when the number of samples of some categories in a multi-classification dataset is too small to be processed via sampling or the number of minority class samples is only one or two, the traditional undersampling algorithms will be less effective. In this study, we select nine multi-classification time series datasets with extremely few samples as research objects, fully consider the characteristics of time series data, and use a three-stage algorithm to alleviate the data imbalance problem. In stage one, random oversampling with disturbance items is used to increase the number of sample points; in stage two, on the basis of the latter operation, SMOTE (synthetic minority oversampling technique) oversampling is employed; in stage three, the dynamic time-warping distance is used to calculate the distance between sample points, identify the sample points of Tomek links at the boundary, and clean up the boundary noise. This study proposes a new sampling algorithm. In the nine multi-classification time series datasets with extremely few samples, the new sampling algorithm is compared with four classic undersampling algorithms, namely, ENN (edited nearest neighbours), NCR (neighborhood cleaning rule), OSS (one-side selection), and RENN (repeated edited nearest neighbors), based on the macro accuracy, recall rate, and F1-score evaluation indicators. The results are as follows: of the nine datasets selected, for the dataset with the most categories and the fewest minority class samples, FiftyWords, the accuracy of the new sampling algorithm was 0.7156, far beyond that of ENN, RENN, OSS, and NCR; its recall rate was also better than that of the four undersampling algorithms used for comparison, corresponding to 0.7261; and its F1-score was 200.71%, 188.74%, 155.29%, and 85.61% better than that of ENN, RENN, OSS, and NCR, respectively. For the other eight datasets, this new sampling algorithm also showed good indicator scores. The new algorithm proposed in this study can effectively alleviate the data imbalance problem of multi-classification time series datasets with many categories and few minority class samples and, at the same time, clean up the boundary noise data between classes.
引用
收藏
页数:14
相关论文
共 50 条
  • [31] A New Maintenance Optimization Model Based on Three-Stage Time Delay for Series Intelligent System with Intermediate Buffer
    Lv, Xiaolei
    Liu, Qinming
    Li, Zhinan
    Dong, Yifan
    Xia, Tangbin
    Chen, Xiang
    SHOCK AND VIBRATION, 2021, 2021
  • [32] Time Series Classification Based on Multi-codebook Important Time Subsequence Approximation Algorithm
    Tao, Zhiwei
    Zhang, Li
    Wang, Bangjun
    Li, Fanzhang
    NEURAL INFORMATION PROCESSING, ICONIP 2016, PT IV, 2016, 9950 : 582 - 589
  • [33] Performance evaluation and experiment of a configuration algorithm for three-stage multi-granularity optical cross-connects
    Qi, Yongmin
    Guo, Wei
    Zhang, Yi
    Zuo, Siye
    Jin, Yaohui
    Hu, Weisheng
    IEICE TRANSACTIONS ON COMMUNICATIONS, 2006, E89B (06) : 1747 - 1754
  • [34] Three-stage multi-innovation parameter estimation for an exponential autoregressive time-series model with moving average noise by using the data filtering technique
    Xu, Huan
    Ding, Feng
    Yang, Erfu
    INTERNATIONAL JOURNAL OF ROBUST AND NONLINEAR CONTROL, 2021, 31 (01) : 166 - 184
  • [35] Data-driven three-stage polymer intrinsic viscosity prediction model with long sequence time series data
    Zhang, Peng
    Bi, Jinmao
    Wang, Ming
    Zhang, Jie
    Zhao, Chuncai
    Cui, Li
    JOURNAL OF ENGINEERED FIBERS AND FABRICS, 2024, 19
  • [36] LITE-FORT: Lightweight three-stage energy theft detection based on time series forecasting of consumption patterns
    Aoufi, Souhila
    Derhab, Abdelouahid
    Guerroumi, Mohamed
    Guemmouma, Hanane
    Lazali, Halla
    ELECTRIC POWER SYSTEMS RESEARCH, 2023, 225
  • [37] A Three-stage CE-IS Monte Carlo Algorithm for Highly Reliable Composite System Reliability Evaluation Based on Screening Method
    Yan, Chao
    Luca, Giambattista Lucarelli
    Bie, Zhaohong
    Ding, Tao
    Li, Gengfeng
    2016 INTERNATIONAL CONFERENCE ON PROBABILISTIC METHODS APPLIED TO POWER SYSTEMS (PMAPS), 2016,
  • [38] Three-Stage Transfer Learning with AlexNet50 for MRI Image Multi-Class Classification with Optimal Learning Rate
    Athisayamani, Suganya
    Singh, A. Robert
    Joshi, Gyanendra Prasad
    Cho, Woong
    CMES-COMPUTER MODELING IN ENGINEERING & SCIENCES, 2025, 142 (01): : 155 - 183
  • [39] Improving long-term electricity time series forecasting in smart grid with a three-stage channel-temporal approach
    Sun, Zhao
    Song, Dongjin
    Peng, Qinke
    Li, Haozhou
    Li, Pulin
    JOURNAL OF CLEANER PRODUCTION, 2024, 468
  • [40] A three-stage adaptive memetic algorithm for multi-objective optimization of flexible assembly job-shop scheduling problem
    Zhang, Chenlu
    Feng, Jiamei
    Zhang, Mingchuan
    Yang, Lei
    Zhang, Lei
    Wang, Lin
    Zhu, Junlong
    Wu, Qingtao
    ENGINEERING APPLICATIONS OF ARTIFICIAL INTELLIGENCE, 2025, 144