Investigating Bloom Filters for Web Archives' Holdings

被引:2
|
作者
Klein, Martin [1 ]
Balakireva, Lyudmila [1 ]
Holub, Karolina [2 ]
Celjak, Drazenko [3 ]
Rudomino, Ingeborg [2 ]
机构
[1] Los Alamos Natl Lab, Los Alamos, NM 87545 USA
[2] Natl & Univ Lib Zagreb, Zagreb, Croatia
[3] Univ Zagreb Univ, Comp Ctr, Zagreb, Croatia
关键词
bloom filters; web archives; web archive profiling; index sharing;
D O I
10.1145/3529372.3530934
中图分类号
TP [自动化技术、计算机技术];
学科分类号
0812 ;
摘要
What web archives hold is often opaque to the public and even experts in the domain struggle to provide precise assessments. Given the increasing need for and use of crawled and archived web resources, discovery of individual records as well as sharing of entire holdings are pressing use cases. We investigate Bloom Filters (BFs) and their applicability to address these use cases. We experiment with and analyze parameters for their creation, measure their performance, outline an approach for scalability, and describe various pilot implementations that showcase their potential to meet our needs. BFs come with beneficial characteristics and hence have enjoyed popularity in various domains. We highlight their suitability for web archiving use cases and how they can contribute to very fast and accurate search services.
引用
收藏
页数:10
相关论文
共 50 条