Web Spam Detection Based on Improved Tri-training

被引:0
|
作者
Li, Hailong [1 ]
机构
[1] Beihang Univ, Sch Comp Sci & Engn, Beijing 100191, Peoples R China
关键词
web spam; search engine; web spam detection; tri-training; co-training; feature view;
D O I
暂无
中图分类号
TP301 [理论、方法];
学科分类号
081202 ;
摘要
Web spamming is the deliberate manipulation of search engine indexes to make a page get high ranking than which it deserved considering its true value. Since the evolution of web spam, a new based on machine learning algorithm web spam detection method which has self-learning ability has emerged. Web spam detection is viewed as a binary classification learning problem. Because labeled training examples are fairly expensive to obtain which need the participation of experts in this field and labor costs, how to fully utilize a large number of unlabeled web page examples on the web is a challenge faced by web spam detection. In this paper, we present a web spam detection algorithm according to improve tri-training. It uses a small amount of labeled examples and a large number of unlabeled examples to train classifiers, which can reduce the cost of labeled examples and improve the learning performance. Both web page content features and link features are used in this paper.
引用
收藏
页码:61 / 65
页数:5
相关论文
共 50 条
  • [31] Offline data-driven evolutionary optimization based on tri-training
    Huang, Pengfei
    Wang, Handing
    Jin, Yaochu
    SWARM AND EVOLUTIONARY COMPUTATION, 2021, 60
  • [32] Tri-training for Dependency Parsing Domain Adaptation
    Jiang, Shu
    Li, Zuchao
    Zhao, Hai
    Lu, Bao-Liang
    Wang, Rui
    ACM TRANSACTIONS ON ASIAN AND LOW-RESOURCE LANGUAGE INFORMATION PROCESSING, 2022, 21 (03)
  • [33] Asymmetric Tri-training for Unsupervised Domain Adaptation
    Saito, Kuniaki
    Ushiku, Yoshitaka
    Harada, Tatsuya
    INTERNATIONAL CONFERENCE ON MACHINE LEARNING, VOL 70, 2017, 70
  • [34] Semi-supervised Software Defect Prediction Model Based on Tri-training
    Meng, Fanqi
    Cheng, Wenying
    Wang, Jingdong
    KSII Transactions on Internet and Information Systems, 2021, 15 (11) : 4028 - 4042
  • [35] 基于Tri-training的半监督SVM
    李昆仑
    张伟
    代运娜
    计算机工程与应用, 2009, 45 (22) : 103 - 106
  • [36] 基于特征变换的Tri-Training算法
    赵文亮
    郭华平
    范明
    计算机工程, 2014, 40 (05) : 183 - 187+191
  • [37] Tri-Training for authorship attribution with limited training data: a comprehensive study
    Qian, Tieyun
    Liu, Bing
    Chen, Li
    Peng, Zhiyong
    Zhong, Ming
    He, Guoliang
    Li, Xuhui
    Xu, Gang
    NEUROCOMPUTING, 2016, 171 : 798 - 806
  • [38] Research on Chinese Medical Entity Recognition Based on Multi-Neural Network Fusion and Improved Tri-Training Algorithm
    Qi, Renlong
    Lv, Pengtao
    Zhang, Qinghui
    Wu, Meng
    APPLIED SCIENCES-BASEL, 2022, 12 (17):
  • [39] Three-way decision-based tri-training with entropy minimization
    Pan, Linchao
    Gao, Can
    Zhou, Jie
    INFORMATION SCIENCES, 2022, 610 : 33 - 51
  • [40] Entity Extraction of Adverse Drug Reaction on Social Media Based on Tri-training
    He, Zhongbo
    Yan, Xin
    Xu, Guangyi
    Zhang, Jinpeng
    Deng, Zhongying
    Computer Engineering and Applications, 60 (03): : 177 - 186