Predicting structured metadata from unstructured metadata

被引:5
|
作者
Posch, Lisa [1 ,2 ]
Panahiazar, Maryam [3 ]
Dumontier, Michel [3 ]
Gevaert, Olivier [3 ]
机构
[1] GESIS Leibniz Inst Social Sci, Cologne, Germany
[2] Univ Koblenz Landau, Inst Web Sci & Technol, Koblenz, Germany
[3] Stanford Univ, Dept Med, Stanford Ctr Biomed Informat Res, Stanford, CA 94305 USA
关键词
D O I
10.1093/database/baw080
中图分类号
Q [生物科学];
学科分类号
07 ; 0710 ; 09 ;
摘要
Enormous amounts of biomedical data have been and are being produced by investigators all over the world. However, one crucial and limiting factor in data reuse is accurate, structured and complete description of the data or data about the data-defined as metadata. We propose a framework to predict structured metadata terms from unstructured metadata for improving quality and quantity of metadata, using the Gene Expression Omnibus (GEO) microarray database. Our framework consists of classifiers trained using term frequency-inverse document frequency (TF-IDF) features and a second approach based on topics modeled using a Latent Dirichlet Allocation model (LDA) to reduce the dimensionality of the unstructured data. Our results on the GEO database show that structured metadata terms can be the most accurately predicted using the TF-IDF approach followed by LDA both outperforming the majority vote baseline. While some accuracy is lost by the dimensionality reduction of LDA, the difference is small for elements with few possible values, and there is a large improvement over the majority classifier baseline. Overall this is a promising approach for metadata prediction that is likely to be applicable to other datasets and has implications for researchers interested in biomedical metadata curation and metadata prediction.
引用
收藏
页数:9
相关论文
共 50 条
  • [41] Metadata
    Hider, Philip
    AUSTRALIAN ACADEMIC & RESEARCH LIBRARIES, 2008, 39 (04) : 296 - 297
  • [42] Metadata
    Bullimore, Alan
    JOURNAL OF PEDAGOGIC DEVELOPMENT, 2016, 6 (03):
  • [43] METADATA
    Lange, Holley R.
    TECHNICAL SERVICES QUARTERLY, 2010, 27 (01) : 139 - 141
  • [44] METADATA
    Molanphy, Emily
    JOURNAL OF WEB LIBRARIANSHIP, 2009, 3 (03) : 297 - 298
  • [45] Metadata
    Ibrus, Indrek
    INTERNATIONAL JOURNAL OF COMMUNICATION, 2017, 11 : 2227 - 2230
  • [46] Metadata
    Mardis, Marcia
    INTERNATIONAL JOURNAL OF COMMUNICATION, 2018, 12 : 1153 - 1156
  • [47] Metadata
    Miraglia, Elizabeth
    LIBRARY RESOURCES & TECHNICAL SERVICES, 2017, 61 (01): : 64 - 65
  • [48] Metadata
    Hahn, Jim
    LIBRARY JOURNAL, 2016, 141 (02) : 91 - 91
  • [49] Metadata
    Gilchrist, Alan
    JOURNAL OF DOCUMENTATION, 2009, 65 (04) : 708 - 711
  • [50] Metadata
    Vellucci, SL
    ANNUAL REVIEW OF INFORMATION SCIENCE AND TECHNOLOGY, 1998, 33 : 187 - 222