Co-Training Semi-Supervised Deep Learning for Sentiment Classification of MOOC Forum Posts

被引:17
|
作者
Chen, Jing [1 ]
Feng, Jun [1 ,2 ]
Sun, Xia [1 ]
Liu, Yang [1 ]
机构
[1] Northwest Univ, Sch Informat Sci & Technol, Xian 710127, Shaanxi, Peoples R China
[2] Northwest Univ, State Prov Joint Engn & Res Ctr Adv Networking &, Sch Informat Sci & Technol, Xian 710127, Shaanxi, Peoples R China
来源
SYMMETRY-BASEL | 2020年 / 12卷 / 01期
基金
中国国家自然科学基金;
关键词
co-training; semi-supervised learning; sentiment classification; asymmetric data; MOOC;
D O I
10.3390/sym12010008
中图分类号
O [数理科学和化学]; P [天文学、地球科学]; Q [生物科学]; N [自然科学总论];
学科分类号
07 ; 0710 ; 09 ;
摘要
Sentiment classification of forum posts of massive open online courses is essential for educators to make interventions and for instructors to improve learning performance. Lacking monitoring on learners' sentiments may lead to high dropout rates of courses. Recently, deep learning has emerged as an outstanding machine learning technique for sentiment classification, which extracts complex features automatically with rich representation capabilities. However, deep neural networks always rely on a large amount of labeled data for supervised training. Constructing large-scale labeled training datasets for sentiment classification is very laborious and time consuming. To address this problem, this paper proposes a co-training, semi-supervised deep learning model for sentiment classification, leveraging limited labeled data and massive unlabeled data simultaneously to achieve performance comparable to those methods trained on massive labeled data. To satisfy the condition of two views of co-training, we encoded texts into vectors from views of word embedding and character-based embedding independently, considering words' external and internal information. To promote the classification performance with limited data, we propose a double-check strategy sample selection method to select samples with high confidence to augment the training set iteratively. In addition, we propose a mixed loss function both considering the labeled data with asymmetric and unlabeled data. Our proposed method achieved a 89.73% average accuracy and an 93.55% average F1-score, about 2.77% and 3.2% higher than baseline methods. Experimental results demonstrate the effectiveness of the proposed model trained on limited labeled data, which performs much better than those trained on massive labeled data.
引用
收藏
页数:24
相关论文
共 50 条
  • [31] Co-Training Semi-Supervised Active Learning Algorithm based on Noise Filter
    Chen Ya-bi
    Zhan Yong-zhao
    PROCEEDINGS OF THE 2009 WRI GLOBAL CONGRESS ON INTELLIGENT SYSTEMS, VOL III, 2009, : 524 - 528
  • [32] Three-Way Co-Training with Pseudo Labels for Semi-Supervised Learning
    Wang, Liuxin
    Gao, Can
    Zhou, Jie
    Wen, Jiajun
    MATHEMATICS, 2023, 11 (15)
  • [33] Multi-Label Learning with Co-Training Based on Semi-Supervised Regression
    Xu, Meixiang
    Sun, Fuming
    Jiang, Xiaojun
    2014 INTERNATIONAL CONFERENCE ON SECURITY, PATTERN ANALYSIS, AND CYBERNETICS (SPAC), 2014, : 175 - 180
  • [34] Temporal-Frequency Co-training for Time Series Semi-supervised Learning
    Liu, Zhen
    Ma, Qianli
    Ma, Peitian
    Wang, Linghao
    THIRTY-SEVENTH AAAI CONFERENCE ON ARTIFICIAL INTELLIGENCE, VOL 37 NO 7, 2023, : 8923 - 8931
  • [35] Semi-Supervised Co-Training Model Using Convolution and Transformer for Hyperspectral Image Classification
    Zhao, Feng
    Song, Xiqun
    Zhang, Junjie
    Liu, Hanqiang
    IEEE GEOSCIENCE AND REMOTE SENSING LETTERS, 2024, 21
  • [36] HIGH ACCURATE INTERNET TRAFFIC CLASSIFICATION BASED ON CO-TRAINING SEMI-SUPERVISED CLUSTERING
    Li, Xiang
    Qi, Feng
    Yu, Li Kun
    Qiu, Xue Song
    PROCEEDINGS OF THE 2010 INTERNATIONAL CONFERENCE ON ADVANCED INTELLIGENCE AND AWARENESS INTERNET, AIAI2010, 2010, : 193 - 197
  • [37] Uncertainty-aware deep co-training for semi-supervised medical image segmentation
    Zheng, Xu
    Fu, Chong
    Xie, Haoyu
    Chen, Jialei
    Wang, Xingwei
    Sham, Chiu-Wing
    COMPUTERS IN BIOLOGY AND MEDICINE, 2022, 149
  • [38] Co-Training based Semi-Supervised Web Spam Detection
    Wang, Wei
    Lee, Xiao-Dong
    Hu, An-Lei
    Geng, Guang-Gang
    2013 10TH INTERNATIONAL CONFERENCE ON FUZZY SYSTEMS AND KNOWLEDGE DISCOVERY (FSKD), 2013, : 789 - 793
  • [39] A Semi-supervised Learning Approach for Microblog Sentiment Classification
    Yu, Zhiwei
    Wong, Raymond K.
    Chi, Chi-Hung
    Chen, Fang
    2015 IEEE INTERNATIONAL CONFERENCE ON SMART CITY/SOCIALCOM/SUSTAINCOM (SMARTCITY), 2015, : 339 - 344
  • [40] A Mutually Attentive Co-Training Framework for Semi-Supervised Recognition
    Min, Shaobo
    Chen, Xuejin
    Xie, Hongtao
    Zha, Zheng-Jun
    Zhang, Yongdong
    IEEE TRANSACTIONS ON MULTIMEDIA, 2021, 23 : 899 - 910