Three-Stage Sampling Algorithm for Highly Imbalanced Multi-Classification Time Series Datasets

被引:1
|
作者
Wang, Haoming [1 ]
机构
[1] Guangdong Univ Educ, Sch Math, Dept Stat & Financial Math, 351,Xingang Middle Rd, Guangzhou 510303, Peoples R China
来源
SYMMETRY-BASEL | 2023年 / 15卷 / 10期
关键词
imbalanced data; data preprocessing; sampling; Tomek links; DTW; SMOTE;
D O I
10.3390/sym15101849
中图分类号
O [数理科学和化学]; P [天文学、地球科学]; Q [生物科学]; N [自然科学总论];
学科分类号
07 ; 0710 ; 09 ;
摘要
To alleviate the data imbalance problem caused by subjective and objective factors, scholars have developed different data-preprocessing algorithms, among which undersampling algorithms are widely used because of their fast and efficient performance. However, when the number of samples of some categories in a multi-classification dataset is too small to be processed via sampling or the number of minority class samples is only one or two, the traditional undersampling algorithms will be less effective. In this study, we select nine multi-classification time series datasets with extremely few samples as research objects, fully consider the characteristics of time series data, and use a three-stage algorithm to alleviate the data imbalance problem. In stage one, random oversampling with disturbance items is used to increase the number of sample points; in stage two, on the basis of the latter operation, SMOTE (synthetic minority oversampling technique) oversampling is employed; in stage three, the dynamic time-warping distance is used to calculate the distance between sample points, identify the sample points of Tomek links at the boundary, and clean up the boundary noise. This study proposes a new sampling algorithm. In the nine multi-classification time series datasets with extremely few samples, the new sampling algorithm is compared with four classic undersampling algorithms, namely, ENN (edited nearest neighbours), NCR (neighborhood cleaning rule), OSS (one-side selection), and RENN (repeated edited nearest neighbors), based on the macro accuracy, recall rate, and F1-score evaluation indicators. The results are as follows: of the nine datasets selected, for the dataset with the most categories and the fewest minority class samples, FiftyWords, the accuracy of the new sampling algorithm was 0.7156, far beyond that of ENN, RENN, OSS, and NCR; its recall rate was also better than that of the four undersampling algorithms used for comparison, corresponding to 0.7261; and its F1-score was 200.71%, 188.74%, 155.29%, and 85.61% better than that of ENN, RENN, OSS, and NCR, respectively. For the other eight datasets, this new sampling algorithm also showed good indicator scores. The new algorithm proposed in this study can effectively alleviate the data imbalance problem of multi-classification time series datasets with many categories and few minority class samples and, at the same time, clean up the boundary noise data between classes.
引用
收藏
页数:14
相关论文
共 50 条
  • [41] THS-IDPC: A three-stage hierarchical sampling method based on improved density peaks clustering algorithm for encrypted malicious traffic detection
    Chen, Liangchen
    Gao, Shu
    Liu, Baoxu
    Lu, Zhigang
    Jiang, Zhengwei
    JOURNAL OF SUPERCOMPUTING, 2020, 76 (09): : 7489 - 7518
  • [42] THS-IDPC: A three-stage hierarchical sampling method based on improved density peaks clustering algorithm for encrypted malicious traffic detection
    Liangchen Chen
    Shu Gao
    Baoxu Liu
    Zhigang Lu
    Zhengwei Jiang
    The Journal of Supercomputing, 2020, 76 : 7489 - 7518
  • [43] Three-stage Algorithm of Estimation and Fault diagnose for Closed-loop Gas Turbine Engine Systems with Unknown Time Delay
    Ren, Jia
    Ma, Hongjun
    Yang, Guanghong
    2013 25TH CHINESE CONTROL AND DECISION CONFERENCE (CCDC), 2013, : 5134 - 5139
  • [44] Automatic deflection measurement for outdoor steel structure based on digital image correlation and three-stage multi-scale clustering algorithm
    Sun, Haobo
    Huang, Yongqi
    AUTOMATION IN CONSTRUCTION, 2024, 163
  • [45] Multi-subband and Multi-subepoch Time Series Feature Learning for EEG-based Sleep Stage Classification
    An, Panfeng
    Yuan, Zhiyong
    Zhao, Jianhui
    Jiang, Xue
    Wang, Zengmao
    Du, Bo
    2021 INTERNATIONAL JOINT CONFERENCE ON BIOMETRICS (IJCB 2021), 2021,
  • [46] Adaptive swarm cluster-based dynamic multi-objective synthetic minority oversampling technique algorithm for tackling binary imbalanced datasets in biomedical data classification
    Li, Jinyan
    Fong, Simon
    Sung, Yunsick
    Cho, Kyungeun
    Wong, Raymond
    Wong, Kelvin K. L.
    BIODATA MINING, 2016, 9 : 1 - 15
  • [47] Adaptive swarm cluster-based dynamic multi-objective synthetic minority oversampling technique algorithm for tackling binary imbalanced datasets in biomedical data classification
    Jinyan Li
    Simon Fong
    Yunsick Sung
    Kyungeun Cho
    Raymond Wong
    Kelvin K. L. Wong
    BioData Mining, 9
  • [48] Multi-stage Algorithm Based on Neural Network Committee for Prediction and Search for Precursors in Multi-dimensional Time Series
    Dolenko, Sergey
    Guzhva, Alexander
    Persiantsev, Igor
    Shugai, Julia
    ARTIFICIAL NEURAL NETWORKS - ICANN 2009, PT II, 2009, 5769 : 295 - 304
  • [49] A Batch Scheduling Model for a Three-stage Flow Shop with Job and Batch Processors Considering a Sampling Inspection to Minimize Expected Total Actual Flow Time
    Suryadhini, Pratya Poeri
    Sukoyo, Sukoyo
    Suprayogi, Suprayogi
    Halim, Abdul Hakim
    JOURNAL OF INDUSTRIAL ENGINEERING AND MANAGEMENT-JIEM, 2021, 14 (03): : 520 - 537
  • [50] A novel real-time multi-step forecasting system with a three-stage data preprocessing strategy for containerized freight market
    Yin, Kedong
    Guo, Hongbo
    Yang, Wendong
    EXPERT SYSTEMS WITH APPLICATIONS, 2024, 246