A Focused Crawler for Dark Web Forums

被引:58
|
作者
Fu, Tianjun [1 ]
Abbasi, Ahmed [2 ]
Chen, Hsinchun [1 ]
机构
[1] Univ Arizona, Dept Management Informat Syst, Artificial Intelligence Lab, Tucson, AZ 85721 USA
[2] Univ Wisconsin, Sheldon B Lubar Sch Business, Milwaukee, WI 53201 USA
来源
JOURNAL OF THE AMERICAN SOCIETY FOR INFORMATION SCIENCE AND TECHNOLOGY | 2010年 / 61卷 / 06期
关键词
SPIDER; LINK;
D O I
10.1002/asi.21323
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
The unprecedented growth of the Internet has given rise to the Dark Web, the problematic facet of the Web associated with cybercrime, hate, and extremism. Despite the need for tools to collect and analyze Dark Web forums, the covert nature of this part of the Internet makes traditional Web crawling techniques insufficient for capturing such content. In this study, we propose a novel crawling system designed to collect Dark Web forum content. The system uses a human-assisted accessibility approach to gain access to Dark Web forums. Several URL ordering features and techniques enable efficient extraction of forum postings. The system also includes an incremental crawler coupled with a recall-improvement mechanism intended to facilitate enhanced retrieval and updating of collected content. Experiments conducted to evaluate the effectiveness of the human-assisted accessibility approach and the recall-improvement-based, incremental-update procedure yielded favorable results. The human-assisted approach significantly improved access to Dark Web forums while the incremental crawler with recall improvement also outperformed standard periodic- and incremental-update approaches. Using the system, we were able to collect over 100 Dark Web forums from three regions. A case study encompassing link and content analysis of collected forums was used to illustrate the value and importance of gathering and analyzing content from such online communities.
引用
收藏
页码:1213 / 1231
页数:19
相关论文
共 50 条
  • [1] Keyword Focused Web Crawler
    Agre, Gunjan H.
    Mahajan, Nikita V.
    2015 2ND INTERNATIONAL CONFERENCE ON ELECTRONICS AND COMMUNICATION SYSTEMS (ICECS), 2015, : 1089 - 1092
  • [2] Smart Focused Web Crawler for Hidden Web
    Kaur, Sawroop
    Geetha, G.
    INFORMATION AND COMMUNICATION TECHNOLOGY FOR COMPETITIVE STRATEGIES, 2019, 40 : 419 - 427
  • [3] A Framework of a Hybrid Focused Web Crawler
    Sun, Yixue
    Jin, Peiquan
    Yue, Lihua
    2008 SECOND INTERNATIONAL CONFERENCE ON FUTURE GENERATION COMMUNICATION AND NETWORKING SYMPOSIA, VOLS 1-5, PROCEEDINGS, 2008, : 146 - 149
  • [4] An algorithm OFC for the focused web crawler
    Zhu, Qiang
    PROCEEDINGS OF 2007 INTERNATIONAL CONFERENCE ON MACHINE LEARNING AND CYBERNETICS, VOLS 1-7, 2007, : 4059 - 4063
  • [5] Focused Web Crawler for Indonesian Recipes
    Alfarisy, Gusti Ahmad Fanshuri
    Bachtiar, Fitra A.
    2017 INTERNATIONAL CONFERENCE ON SUSTAINABLE INFORMATION ENGINEERING AND TECHNOLOGY (SIET), 2017, : 196 - 202
  • [6] Intelligent Crawler for Web Forums based on Improved Regular Expressions
    Pavkovic, Milos
    Protic, Jelica
    2013 21ST TELECOMMUNICATIONS FORUM (TELFOR), 2013, : 817 - 820
  • [7] Keyword query based focused Web crawler
    Kumar, Manish
    Bindal, Ankit
    Gautam, Robin
    Bhatia, Rajesh
    6TH INTERNATIONAL CONFERENCE ON SMART COMPUTING AND COMMUNICATIONS, 2018, 125 : 584 - 590
  • [8] Dark Web Forums Portal: Searching and Analyzing Jihadist Forums
    Zhang, Yulei
    Zeng, Shuo
    Fan, Li
    Dang, Yan
    Larson, Catherine A.
    Chen, Hsinchun
    ISI: 2009 IEEE INTERNATIONAL CONFERENCE ON INTELLIGENCE AND SECURITY INFORMATICS, 2009, : 71 - 76
  • [9] LEARNING-based Focused WEB Crawler
    Kumar, Naresh
    Aggarwal, Dhruv
    IETE JOURNAL OF RESEARCH, 2023, 69 (04) : 2037 - 2045
  • [10] Weakly supervised learning for an effective focused web crawler
    Dhanith, P. R. Joe
    Saeed, Khalid
    Rohith, G.
    Raja, S. P.
    ENGINEERING APPLICATIONS OF ARTIFICIAL INTELLIGENCE, 2024, 132