Exploring DBpedia and Wikipedia for Portuguese Semantic Relationship Extraction

被引:0
|
作者
Batista, David S. [1 ,2 ]
Forte, David [1 ,2 ]
Silva, Rui [1 ,2 ]
Martins, Bruno [1 ,2 ]
Silva, Mario J. [1 ,2 ]
机构
[1] Inst Super Tecn, Lisbon, Portugal
[2] INESC ID, Lisbon, Portugal
来源
LINGUAMATICA | 2013年 / 5卷 / 01期
关键词
Relation Extraction; Information Extraction;
D O I
暂无
中图分类号
H0 [语言学];
学科分类号
030303 ; 0501 ; 050102 ;
摘要
The identification of semantic relationships, as expressed between named entities in text, is an important step for extracting knowledge from large document collections, such as the Web. Previous works have addressed this task for the English language through supervised learning techniques for automatic classification. The current state of the art involves the use of learning methods based on string kernels (Kim et al., 2010; Zhao e Grishman, 2005). However, such approaches require manually annotated training data for each type of semantic relationship, and have scalability problems when tens or hundreds of different types of relationships have to be extracted. This article discusses an approach for distantly supervised relation extraction over texts written in the Portuguese language, which uses an efficient technique for measuring similarity between relation instances, based on minwise hashing (Broder, 1997) and on locality sensitive hashing (Rajaraman e Ullman, 2011). In the proposed method, the training examples are automatically collected from Wikipedia, corresponding to sentences that express semantic relationships between pairs of entities extracted from DBPedia. These examples are represented as sets of character quadgrams and other representative elements. The sets are indexed in a data structure that implements the idea of locality-sensitive hashing. To check which semantic relationship is expressed between a given pair of entities referenced in a sentence, the most similar training examples are searched, based on an approximation to the Jaccard coefficient, obtained through min-hashing. The relation class is assigned with basis on the weighted votes of the most similar examples. Tests with a dataset from Wikipedia validate the suitability of the proposed method, showing, for instance, that the method is able to extract 10 different types of semantic relations, 8 of them corresponding to asymmetric relations, with an average score of 55.6%, measured in terms of F-1.
引用
收藏
页码:41 / 57
页数:17
相关论文
共 50 条
  • [31] Semantic relatedness in DBpedia: A comparative and experimental assessment
    Formica, Anna
    Taglino, Francesco
    INFORMATION SCIENCES, 2023, 621 : 474 - 505
  • [32] WikiQA - A Question Answering System on Wikipedia using Freebase, DBpedia and Infobox
    Abbas, Faheem
    Malik, Muhammad Kamran
    Rashid, Muhammad Umair
    Zafar, Rizwan
    2016 SIXTH INTERNATIONAL CONFERENCE ON INNOVATIVE COMPUTING TECHNOLOGY (INTECH), 2016, : 185 - 193
  • [33] Building up ontologies from the English wikipedia and comparing with YAGO and DBpedia
    Kawakami T.
    Morita T.
    Yamaguchi T.
    Transactions of the Japanese Society for Artificial Intelligence, 2020, 35 (04) : 1 - 14
  • [34] DBpedia FlexiFusion the Best of Wikipedia > Wikidata > Your Data
    Frey, Johannes
    Hofer, Marvin
    Obraczka, Daniel
    Lehmann, Jens
    Hellmann, Sebastian
    SEMANTIC WEB - ISWC 2019, PT II, 2019, 11779 : 96 - 112
  • [35] WC3:Wikipedia Category Consistency Checker based on DBPedia
    Yoshioka, Masaharu
    Loban, Rhett
    2015 11TH INTERNATIONAL CONFERENCE ON SIGNAL-IMAGE TECHNOLOGY & INTERNET-BASED SYSTEMS (SITIS), 2015, : 712 - 718
  • [36] Semantic Wonder Cloud: Exploratory Search in DBpedia
    Mirizzi, Roberto
    Ragone, Azzurra
    Di Noia, Tommaso
    Di Sciascio, Eugenio
    CURRENT TRENDS IN WEB ENGINEERING, 2010, 6385s : 138 - 149
  • [37] Automatic extraction of semantic relationships for WordNet by means of pattern learning from Wikipedia
    Ruiz-Casado, M
    Alfonseca, E
    Castells, P
    NATURAL LANGUAGE PROCESSING AND INFORMATION SYSTEMS, PROCEEDINGS, 2005, 3513 : 67 - 79
  • [38] SEMANTIC RELATEDNESS IN DBPEDIA: A COMPARATIVE AND EXPERIMENTAL ASSESSMENT
    Formica, Anna
    Taglino, Francesco
    arXiv, 2023,
  • [39] Semantic Relation Extraction by Conditional Random Fields from Turkish Wikipedia Pages
    Girgin, Canan
    Diri, Banu
    2014 22ND SIGNAL PROCESSING AND COMMUNICATIONS APPLICATIONS CONFERENCE (SIU), 2014, : 136 - 139
  • [40] DBpedia Based SAWSDL for Semantic Web Services
    Yadav, Hemant N.
    Patel, Ravi V.
    2015 2ND INTERNATIONAL CONFERENCE ON COMPUTING FOR SUSTAINABLE GLOBAL DEVELOPMENT (INDIACOM), 2015, : 35 - 39