Upper-Bound Approximations for Dynamic Pruning

被引:21
|
作者
Macdonald, Craig [1 ]
Ounis, Iadh [1 ]
Tonellotto, Nicola [2 ]
机构
[1] Univ Glasgow, Dept Comp Sci, Lilybank Gardens, Glasgow G12 8QQ, Lanark, Scotland
[2] CNR, ISTI, I-56124 Pisa, Italy
基金
欧盟第七框架计划;
关键词
Performance; Experimentation; Dynamic pruning; upper bounds;
D O I
10.1145/2037661.2037662
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
Dynamic pruning strategies for information retrieval systems can increase querying efficiency without decreasing effectiveness by using upper bounds to safely omit scoring documents that are unlikely to make the final retrieved set. Often, such upper bounds are pre-calculated at indexing time for a given weighting model. However, this precludes changing, adapting or training the weighting model without recalculating the upper bounds. Instead, upper bounds should be approximated at querying time from various statistics of each term to allow on-the-fly adaptation of the applied retrieval strategy. This article, by using uniform notation, formulates the problem of determining a term upper-bound given a weighting model and discusses the limitations of existing approximations. Moreover, we propose an upper-bound approximation using a constrained nonlinear maximization problem. We prove that our proposed upper-bound approximation does not impact the retrieval effectiveness of several modern weighting models from various different families. We also show the applicability of the approximation for the Markov Random Field proximity model. Finally, we empirically examine how the accuracy of the upper-bound approximation impacts the number of postings scored and the resulting efficiency in the context of several large Web test collections.
引用
收藏
页数:28
相关论文
共 50 条