Ducky : A Data Extraction System for Various Structured Web Documents

被引:2
|
作者
Kanaoka, Kei [1 ]
Fujii, Yotaro [1 ]
Toyama, Motomichi [1 ]
机构
[1] Keio Univ, Dept Comp Sci, Yokohama, Kanagawa, Japan
关键词
Data Extraction; Web scraping; Web Wrapper; CSS selector;
D O I
10.1145/2628194.2628244
中图分类号
TP301 [理论、方法];
学科分类号
081202 ;
摘要
The World Wide Web has become a primary source of information. Therefore, extracting data from Web sources has become a key technology. In this paper, we introduce a semi-automatic system Ducky : including a Web Wrapper which extracts data from Web sources and translates them into structured data. In Ducky, by defining a configuration file consisting in several parameters (URL of the Web page, CSS selectors which locates the data to retrieve and so on.), users do not need to write Web scraping programs at all. The definition is simple, yet can extract data flexibly from various structured Web pages. Additionally, Ducky provides a Web API and various output data formats: XML, JSON, CSV. Finally, experimentations confirmed that Ducky can accurately extract data from 22 different structured Web sources.
引用
收藏
页码:342 / 347
页数:6
相关论文
共 50 条
  • [21] Automatic Opinion Extraction from Web Documents
    Shandilya, Shishir K.
    Jain, Suresh
    2009 INTERNATIONAL CONFERENCE ON COMPUTER AND AUTOMATION ENGINEERING, PROCEEDINGS, 2009, : 351 - 355
  • [22] WebDB: a system for querying semi-structured data on the Web
    Li, WS
    Shim, J
    Candan, KS
    JOURNAL OF VISUAL LANGUAGES AND COMPUTING, 2002, 13 (01): : 3 - 33
  • [23] Unsupervised Relation Extraction from Web Documents
    Eichler, Kathrin
    Hemsen, Holmer
    Neumann, Guenter
    SIXTH INTERNATIONAL CONFERENCE ON LANGUAGE RESOURCES AND EVALUATION, LREC 2008, 2008, : 1674 - 1679
  • [24] A Model for Geographic Knowledge Extraction on Web Documents
    Campelo, Claudio E. C.
    Baptista, Claudio de Souza
    ADVANCES IN CONCEPTUAL MODELING - CHALLENGES PERSPECTIVES, 2009, 5833 : 317 - 326
  • [25] An Efficient Mechanism for Deep Web Data Extraction Based on Tree-Structured Web Pattern Matching
    Ahamed, B. Bazeer
    Yuvaraj, D.
    Shitharth, S.
    Mirza, Olfat M.
    Alsobhi, Aisha
    Yafoz, Ayman
    WIRELESS COMMUNICATIONS & MOBILE COMPUTING, 2022, 2022
  • [26] Automatic Content Extraction on Semi-Structured Documents
    dos Santos, Jose Eduardo Bastos
    11TH INTERNATIONAL CONFERENCE ON DOCUMENT ANALYSIS AND RECOGNITION (ICDAR 2011), 2011, : 1235 - 1239
  • [27] Information extraction from the structured part of office documents
    Hao, XL
    Wang, JTL
    Ng, PA
    INFORMATION SCIENCES, 1996, 91 (3-4) : 245 - 274
  • [28] DocTr: Document Transformer for Structured Information Extraction in Documents
    Liao, Haofu
    RoyChowdhury, Aruni
    Li, Weijian
    Bansal, Ankan
    Zhang, Yuting
    Tu, Zhuowen
    Satzoda, Ravi Kumar
    Manmatha, R.
    Mahadevan, Vijay
    2023 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2023), 2023, : 19527 - 19537
  • [29] A Loosely Coupled Interactive Web Data Extraction System
    Su, Jui-Yuan
    Chen, Lung-Pin
    Wu, I-Chen
    JOURNAL OF INTERNET TECHNOLOGY, 2010, 11 (02): : 237 - 249
  • [30] A SYSTEM FOR INTERACTIVE VIEWING OF STRUCTURED DOCUMENTS
    WITTEN, IH
    BRAMWELL, B
    COMMUNICATIONS OF THE ACM, 1985, 28 (03) : 280 - 288