Ducky : A Data Extraction System for Various Structured Web Documents

被引：2

作者：

Kanaoka, Kei ^{[1
]}

Fujii, Yotaro ^{[1
]}

Toyama, Motomichi ^{[1
]}

机构：

[1] Keio Univ, Dept Comp Sci, Yokohama, Kanagawa, Japan

来源：

PROCEEDINGS OF THE 18TH INTERNATIONAL DATABASE ENGINEERING AND APPLICATIONS SYMPOSIUM (IDEAS14) | 2014年

关键词：

Data Extraction; Web scraping; Web Wrapper; CSS selector;

D O I：

10.1145/2628194.2628244

中图分类号：

TP301 [理论、方法];

学科分类号：

081202 ;

摘要：

The World Wide Web has become a primary source of information. Therefore, extracting data from Web sources has become a key technology. In this paper, we introduce a semi-automatic system Ducky : including a Web Wrapper which extracts data from Web sources and translates them into structured data. In Ducky, by defining a configuration file consisting in several parameters (URL of the Web page, CSS selectors which locates the data to retrieve and so on.), users do not need to write Web scraping programs at all. The definition is simple, yet can extract data flexibly from various structured Web pages. Additionally, Ducky provides a Web API and various output data formats: XML, JSON, CSV. Finally, experimentations confirmed that Ducky can accurately extract data from 22 different structured Web sources.

引用

页码：342 / 347

页数：6

共 50 条

[21] Automatic Opinion Extraction from Web Documents
Shandilya, Shishir K.
Jain, Suresh
2009 INTERNATIONAL CONFERENCE ON COMPUTER AND AUTOMATION ENGINEERING, PROCEEDINGS, 2009, : 351 - 355
[22] WebDB: a system for querying semi-structured data on the Web
Li, WS
Shim, J
Candan, KS
JOURNAL OF VISUAL LANGUAGES AND COMPUTING, 2002, 13 (01): : 3 - 33
[23] Unsupervised Relation Extraction from Web Documents
Eichler, Kathrin
Hemsen, Holmer
Neumann, Guenter
SIXTH INTERNATIONAL CONFERENCE ON LANGUAGE RESOURCES AND EVALUATION, LREC 2008, 2008, : 1674 - 1679
[24] A Model for Geographic Knowledge Extraction on Web Documents
Campelo, Claudio E. C.
Baptista, Claudio de Souza
ADVANCES IN CONCEPTUAL MODELING - CHALLENGES PERSPECTIVES, 2009, 5833 : 317 - 326
[25] An Efficient Mechanism for Deep Web Data Extraction Based on Tree-Structured Web Pattern Matching
Ahamed, B. Bazeer
Yuvaraj, D.
Shitharth, S.
Mirza, Olfat M.
Alsobhi, Aisha
Yafoz, Ayman
WIRELESS COMMUNICATIONS & MOBILE COMPUTING, 2022, 2022
[26] Automatic Content Extraction on Semi-Structured Documents
dos Santos, Jose Eduardo Bastos
11TH INTERNATIONAL CONFERENCE ON DOCUMENT ANALYSIS AND RECOGNITION (ICDAR 2011), 2011, : 1235 - 1239
[27] Information extraction from the structured part of office documents
Hao, XL
Wang, JTL
Ng, PA
INFORMATION SCIENCES, 1996, 91 (3-4) : 245 - 274
[28] DocTr: Document Transformer for Structured Information Extraction in Documents
Liao, Haofu
RoyChowdhury, Aruni
Li, Weijian
Bansal, Ankan
Zhang, Yuting
Tu, Zhuowen
Satzoda, Ravi Kumar
Manmatha, R.
Mahadevan, Vijay
2023 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2023), 2023, : 19527 - 19537
[29] A Loosely Coupled Interactive Web Data Extraction System
Su, Jui-Yuan
Chen, Lung-Pin
Wu, I-Chen
JOURNAL OF INTERNET TECHNOLOGY, 2010, 11 (02): : 237 - 249
[30] A SYSTEM FOR INTERACTIVE VIEWING OF STRUCTURED DOCUMENTS
WITTEN, IH
BRAMWELL, B
COMMUNICATIONS OF THE ACM, 1985, 28 (03) : 280 - 288

← 1 2 3 4 5 →