Web Spam Detection Based on Improved Tri-training

被引:0
|
作者
Li, Hailong [1 ]
机构
[1] Beihang Univ, Sch Comp Sci & Engn, Beijing 100191, Peoples R China
关键词
web spam; search engine; web spam detection; tri-training; co-training; feature view;
D O I
暂无
中图分类号
TP301 [理论、方法];
学科分类号
081202 ;
摘要
Web spamming is the deliberate manipulation of search engine indexes to make a page get high ranking than which it deserved considering its true value. Since the evolution of web spam, a new based on machine learning algorithm web spam detection method which has self-learning ability has emerged. Web spam detection is viewed as a binary classification learning problem. Because labeled training examples are fairly expensive to obtain which need the participation of experts in this field and labor costs, how to fully utilize a large number of unlabeled web page examples on the web is a challenge faced by web spam detection. In this paper, we present a web spam detection algorithm according to improve tri-training. It uses a small amount of labeled examples and a large number of unlabeled examples to train classifiers, which can reduce the cost of labeled examples and improve the learning performance. Both web page content features and link features are used in this paper.
引用
收藏
页码:61 / 65
页数:5
相关论文
共 50 条
  • [41] Tri-training and data editing based semi-supervised clustering algorithm
    Deng, Chao
    Guo, Mao Zu
    MICAI 2006: ADVANCES IN ARTIFICIAL INTELLIGENCE, PROCEEDINGS, 2006, 4293 : 641 - +
  • [42] Biomedical Named Entity Recognition with Tri-training learning
    Cai, YueHong
    Cheng, XianYi
    PROCEEDINGS OF THE 2009 2ND INTERNATIONAL CONFERENCE ON BIOMEDICAL ENGINEERING AND INFORMATICS, VOLS 1-4, 2009, : 2178 - +
  • [43] HMM-BASED TRI-TRAINING ALGORITHM IN HUMAN ACTIVITY RECOGNITION WITH SMARTPHONE
    Xie, Bin
    Wu, Qing
    2012 IEEE 2nd International Conference on Cloud Computing and Intelligent Systems (CCIS) Vols 1-3, 2012, : 109 - 113
  • [44] Semi-supervised Software Defect Prediction Model Based on Tri-training
    Meng, Fanqi
    Cheng, Wenying
    Wang, Jingdong
    KSII TRANSACTIONS ON INTERNET AND INFORMATION SYSTEMS, 2021, 15 (11): : 4028 - 4042
  • [45] Tri-training and data editing based semi-supervised clustering algorithm
    Deng, Chao
    Guo, Mao-Zu
    Ruan Jian Xue Bao/Journal of Software, 2008, 19 (03): : 663 - 673
  • [46] A Reliable Application of MPC for Securing the Tri-Training Algorithm
    Kurniawan, Hendra
    Mambo, Masahiro
    IEEE ACCESS, 2023, 11 : 34718 - 34735
  • [47] Network Traffic Classification Using Tri-training Based on Statistical Flow Characteristics
    Zhao, Shuyuan
    Zhang, Yongzheng
    Chang, Peng
    2017 16TH IEEE INTERNATIONAL CONFERENCE ON TRUST, SECURITY AND PRIVACY IN COMPUTING AND COMMUNICATIONS / 11TH IEEE INTERNATIONAL CONFERENCE ON BIG DATA SCIENCE AND ENGINEERING / 14TH IEEE INTERNATIONAL CONFERENCE ON EMBEDDED SOFTWARE AND SYSTEMS, 2017, : 323 - 330
  • [48] Semisupervised classification of hyperspectral images based on tri-training algorithm with enhanced diversity
    Cui, Ying
    Song, Guojiao
    Wang, Xueting
    Lu, Zhongjun
    Wang, Liguo
    JOURNAL OF APPLIED REMOTE SENSING, 2017, 11
  • [49] 基于Tri-training的主动学习算法
    张雁
    吴保国
    吕丹桔
    林英
    计算机工程, 2014, 40 (06) : 215 - 218+229
  • [50] Multi-Source Tri-Training Transfer Learning
    Cheng, Yuhu
    Wang, Xuesong
    Cao, Ge
    IEICE TRANSACTIONS ON INFORMATION AND SYSTEMS, 2014, E97D (06): : 1668 - 1672