Scalable Probabilistic Similarity Ranking in Uncertain Databases

被引:22
|
作者
Bernecker, Thomas [1 ]
Kriegel, Hans-Peter [1 ]
Mamoulis, Nikos [2 ]
Renz, Matthias [1 ]
Zuefle, Andreas [1 ]
机构
[1] Univ Munich, Inst Informat, D-80538 Munich, Germany
[2] Univ Hong Kong, Dept Comp Sci, Hong Kong, Hong Kong, Peoples R China
关键词
Uncertain databases; probabilistic ranking; similarity search; TOP-K QUERIES;
D O I
10.1109/TKDE.2010.78
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
This paper introduces a scalable approach for probabilistic top-k similarity ranking on uncertain vector data. Each uncertain object is represented by a set of vector instances that is assumed to be mutually exclusive. The objective is to rank the uncertain data according to their distance to a reference object. We propose a framework that incrementally computes for each object instance and ranking position, the probability of the object falling at that ranking position. The resulting rank probability distribution can serve as input for several state-of-the-art probabilistic ranking models. Existing approaches compute this probability distribution by applying the Poisson binomial recurrence technique of quadratic complexity. In this paper, we theoretically as well as experimentally show that our framework reduces this to a linear-time complexity while having the same memory requirements, facilitated by incremental accessing of the uncertain vector instances in increasing order of their distance to the reference object. Furthermore, we show how the output of our method can be used to apply probabilistic top-k ranking for the objects, according to different state-of-the-art definitions. We conduct an experimental evaluation on synthetic and real data, which demonstrates the efficiency of our approach.
引用
收藏
页码:1234 / 1246
页数:13
相关论文
共 50 条
  • [21] Summarizing Uncertain Transaction Databases by Probabilistic Tiles
    Liu, Chunyang
    Chen, Ling
    2016 INTERNATIONAL JOINT CONFERENCE ON NEURAL NETWORKS (IJCNN), 2016, : 4375 - 4382
  • [22] On Pruning for Top-K Ranking in Uncertain Databases
    Wang, Chonghai
    Yuan, Li Yan
    You, Jia-Huai
    Zaiane, Osmar R.
    Pei, Jian
    PROCEEDINGS OF THE VLDB ENDOWMENT, 2011, 4 (10): : 598 - 609
  • [23] Probabilistic Similarity Search for Uncertain Time Series
    Assfalg, Johannes
    Kriegel, Hans-Peter
    Kroeger, Peer
    Benz, Matthias
    SCIENTIFIC AND STATISTICAL DATABASE MANAGEMENT, PROCEEDINGS, 2009, 5566 : 435 - 443
  • [24] Scalable Graph Similarity Search in Large Graph Databases
    Kiran, P.
    Sivadasan, Naveen
    PROCEEDINGS OF THE 2015 IEEE RECENT ADVANCES IN INTELLIGENT COMPUTATIONAL SYSTEMS (RAICS), 2015, : 207 - 211
  • [25] Probabilistic Inverse Ranking Queries over Uncertain Data
    Lian, Xiang
    Chen, Lei
    DATABASE SYSTEMS FOR ADVANCED APPLICATIONS, PROCEEDINGS, 2009, 5463 : 35 - 50
  • [26] Graph similarity search on large uncertain graph databases
    Yuan, Ye
    Wang, Guoren
    Chen, Lei
    Wang, Haixun
    VLDB JOURNAL, 2015, 24 (02): : 271 - 296
  • [27] Graph similarity search on large uncertain graph databases
    Ye Yuan
    Guoren Wang
    Lei Chen
    Haixun Wang
    The VLDB Journal, 2015, 24 : 271 - 296
  • [28] Probabilistic group nearest neighbor queries in uncertain databases
    Lian, Xiang
    Chen, Lei
    IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING, 2008, 20 (06) : 809 - 824
  • [29] Interactive Mining of Probabilistic Frequent Patterns in Uncertain Databases
    Lin, Ming-Yen
    Fu, Cheng-Tai
    Hsueh, Sue-Chen
    INTERNATIONAL JOURNAL OF UNCERTAINTY FUZZINESS AND KNOWLEDGE-BASED SYSTEMS, 2022, 30 (02) : 263 - 283
  • [30] Mining Probabilistic Frequent Closed Itemsets in Uncertain Databases
    Tang, Peiyi
    Peterson, Erich A.
    PROCEEDINGS OF THE 49TH ANNUAL ASSOCIATION FOR COMPUTING MACHINERY SOUTHEAST CONFERENCE (ACMSE '11), 2011, : 86 - 91