Exploring DBpedia and Wikipedia for Portuguese Semantic Relationship Extraction

被引:0
|
作者
Batista, David S. [1 ,2 ]
Forte, David [1 ,2 ]
Silva, Rui [1 ,2 ]
Martins, Bruno [1 ,2 ]
Silva, Mario J. [1 ,2 ]
机构
[1] Inst Super Tecn, Lisbon, Portugal
[2] INESC ID, Lisbon, Portugal
来源
LINGUAMATICA | 2013年 / 5卷 / 01期
关键词
Relation Extraction; Information Extraction;
D O I
暂无
中图分类号
H0 [语言学];
学科分类号
030303 ; 0501 ; 050102 ;
摘要
The identification of semantic relationships, as expressed between named entities in text, is an important step for extracting knowledge from large document collections, such as the Web. Previous works have addressed this task for the English language through supervised learning techniques for automatic classification. The current state of the art involves the use of learning methods based on string kernels (Kim et al., 2010; Zhao e Grishman, 2005). However, such approaches require manually annotated training data for each type of semantic relationship, and have scalability problems when tens or hundreds of different types of relationships have to be extracted. This article discusses an approach for distantly supervised relation extraction over texts written in the Portuguese language, which uses an efficient technique for measuring similarity between relation instances, based on minwise hashing (Broder, 1997) and on locality sensitive hashing (Rajaraman e Ullman, 2011). In the proposed method, the training examples are automatically collected from Wikipedia, corresponding to sentences that express semantic relationships between pairs of entities extracted from DBPedia. These examples are represented as sets of character quadgrams and other representative elements. The sets are indexed in a data structure that implements the idea of locality-sensitive hashing. To check which semantic relationship is expressed between a given pair of entities referenced in a sentence, the most similar training examples are searched, based on an approximation to the Jaccard coefficient, obtained through min-hashing. The relation class is assigned with basis on the weighted votes of the most similar examples. Tests with a dataset from Wikipedia validate the suitability of the proposed method, showing, for instance, that the method is able to extract 10 different types of semantic relations, 8 of them corresponding to asymmetric relations, with an average score of 55.6%, measured in terms of F-1.
引用
收藏
页码:41 / 57
页数:17
相关论文
共 50 条
  • [41] Semantic Tag Cloud Generation via DBpedia
    Mirizzi, Roberto
    Ragone, Azzurra
    Di Noia, Tommaso
    Di Sciascio, Eugenio
    E-COMMERCE AND WEB TECHNOLOGIES, 2010, 61 : 36 - 48
  • [42] Semantic Annotation for Web Services Based on DBpedia
    Zhang, Zhen
    Chen, Shizhan
    Feng, Zhiyong
    2013 IEEE SEVENTH INTERNATIONAL SYMPOSIUM ON SERVICE-ORIENTED SYSTEM ENGINEERING (SOSE 2013), 2013, : 280 - 285
  • [43] Semantic Linking of Learning Object Repositories to DBpedia
    Lama, Manuel
    Vidal, Juan C.
    Otero-Garcia, Estefania
    Bugarin, Alberto
    Barro, Senen
    EDUCATIONAL TECHNOLOGY & SOCIETY, 2012, 15 (04): : 47 - 61
  • [44] CQL Grammars for Lexical and Semantic Information Extraction for Portuguese and Italian
    Barbero, Chiara
    COMPUTATIONAL PROCESSING OF THE PORTUGUESE LANGUAGE, PROPOR 2022, 2022, 13208 : 376 - 386
  • [45] Automatic Annotation of the Catalan Wikipedia: Exploring the Semantic Space via multiple NERC systems
    Atserias, Jordi
    Domingo, Judith
    Rodriguez-Penagos, Carlos
    Sunol, Teresa
    PROCESAMIENTO DEL LENGUAJE NATURAL, 2010, (45): : 169 - 173
  • [46] DBpedia - A large-scale, multilingual knowledge base extracted from Wikipedia
    Lehmann, Jens
    Isele, Robert
    Jakob, Max
    Jentzsch, Anja
    Kontokostas, Dimitris
    Mendes, Pablo N.
    Hellmann, Sebastian
    Morsey, Mohamed
    van Kleef, Patrick
    Auer, Soeren
    Bizer, Christian
    SEMANTIC WEB, 2015, 6 (02) : 167 - 195
  • [47] MMKG: An approach to generate metallic materials knowledge graph based on DBpedia and Wikipedia
    Zhang, Xiaoming
    Liu, Xin
    Li, Xin
    Pan, Dongyu
    COMPUTER PHYSICS COMMUNICATIONS, 2017, 211 : 98 - 112
  • [48] Qsense Learning Semantic Web Concepts by Querying DBpedia
    Panu, Andrei
    Buraga, Sabin C.
    Alboaie, Lenuta
    PROCEEDINGS OF THE 10TH INTERNATIONAL CONFERENCE ON E-BUSINESS (ICE-B 2013), 2013, : 351 - 356
  • [49] Semantic convergence of Wikipedia articles
    Thomas, Christopher
    Sheth, Amit P.
    PROCEEDINGS OF THE IEEE/WIC/ACM INTERNATIONAL CONFERENCE ON WEB INTELLIGENCE: WI 2007, 2007, : 600 - 606
  • [50] A semantic approach for question answering using DBpedia and WordNet
    Sengloiluean, Kittiphong
    Arch-int, Ngamnij
    Arch-int, Somjit
    Thongkrau, Theerayut
    PROCEEDINGS OF 2017 14TH INTERNATIONAL JOINT CONFERENCE ON COMPUTER SCIENCE AND SOFTWARE ENGINEERING (JCSSE), 2017,