Web Spam Detection Based on Improved Tri-training

被引:0
|
作者
Li, Hailong [1 ]
机构
[1] Beihang Univ, Sch Comp Sci & Engn, Beijing 100191, Peoples R China
关键词
web spam; search engine; web spam detection; tri-training; co-training; feature view;
D O I
暂无
中图分类号
TP301 [理论、方法];
学科分类号
081202 ;
摘要
Web spamming is the deliberate manipulation of search engine indexes to make a page get high ranking than which it deserved considering its true value. Since the evolution of web spam, a new based on machine learning algorithm web spam detection method which has self-learning ability has emerged. Web spam detection is viewed as a binary classification learning problem. Because labeled training examples are fairly expensive to obtain which need the participation of experts in this field and labor costs, how to fully utilize a large number of unlabeled web page examples on the web is a challenge faced by web spam detection. In this paper, we present a web spam detection algorithm according to improve tri-training. It uses a small amount of labeled examples and a large number of unlabeled examples to train classifiers, which can reduce the cost of labeled examples and improve the learning performance. Both web page content features and link features are used in this paper.
引用
收藏
页码:61 / 65
页数:5
相关论文
共 50 条
  • [21] Revisiting Tri-training of Dependency Parsers
    Wagner, Joachim
    Foster, Jennifer
    2021 CONFERENCE ON EMPIRICAL METHODS IN NATURAL LANGUAGE PROCESSING (EMNLP 2021), 2021, : 9457 - 9473
  • [22] Tri-training and MapReduce-based massive data learning
    Guo, Mao-Zu
    Deng, Chao
    Liu, Yang
    Li, Ping
    INTERNATIONAL JOURNAL OF GENERAL SYSTEMS, 2011, 40 (04) : 355 - 380
  • [23] Tri-training based learning from positive and unlabeled data
    Zhang, Bangzuo
    Zuo, Wanli
    2008 INTERNATIONAL SYMPOSIUM ON INFORMATION PROCESSING AND 2008 INTERNATIONAL PACIFIC WORKSHOP ON WEB MINING AND WEB-BASED APPLICATION, 2008, : 640 - 644
  • [24] A Novel Semi-supervised SVM Based on Tri-training
    Li, KunLun
    Zhang, Wei
    Ma, Xiaotao
    Cao, Zheng
    Zhang, Chao
    2008 INTERNATIONAL SYMPOSIUM ON INTELLIGENT INFORMATION TECHNOLOGY APPLICATION, VOL III, PROCEEDINGS, 2008, : 47 - +
  • [25] Improved tri-training method for identifying user abnormal behavior based on adaptive golden jackal algorithm
    Wang, Kun
    Gao, Jinggeng
    Kang, Xiaohua
    Li, Huan
    AIP ADVANCES, 2023, 13 (03)
  • [26] Semi-Supervised PolSAR Image Classification Based on Improved Tri-Training With a Minimum Spanning Tree
    Wang, Shuang
    Guo, Yanhe
    Hua, Wenqiang
    Liu, Xinan
    Song, Guoxin
    Hou, Biao
    Jiao, Licheng
    IEEE TRANSACTIONS ON GEOSCIENCE AND REMOTE SENSING, 2020, 58 (12): : 8583 - 8597
  • [27] Tri-Training for Authorship Attribution with Limited Training Data
    Qian, Tieyun
    Liu, Bing
    Chen, Li
    Peng, Zhiyong
    PROCEEDINGS OF THE 52ND ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS, VOL 2, 2014, : 345 - 351
  • [28] Co-Training based Semi-Supervised Web Spam Detection
    Wang, Wei
    Lee, Xiao-Dong
    Hu, An-Lei
    Geng, Guang-Gang
    2013 10TH INTERNATIONAL CONFERENCE ON FUZZY SYSTEMS AND KNOWLEDGE DISCOVERY (FSKD), 2013, : 789 - 793
  • [29] Trust Prediction Based on Extreme Learning Machine and Asymmetric Tri-Training
    Wang, Yan
    Tong, Xiangrong
    IEEE ACCESS, 2021, 9 : 64358 - 64367
  • [30] Tri-training algorithm based on cross entropy and K-nearest neighbors for network intrusion detection
    Zhao, Jia
    Li, Song
    Wu, Runxiu
    Zhang, Yiying
    Zhang, Bo
    Han, Longzhe
    KSII TRANSACTIONS ON INTERNET AND INFORMATION SYSTEMS, 2022, 16 (12): : 3889 - 3903