Cost-Aware Region-Level Data Placement in Multi-Tiered Parallel I/O Systems

被引:10
|
作者
He, Shuibing [1 ,2 ]
Wang, Yang [2 ]
Li, Zheng [3 ]
Sun, Xian-He [4 ]
Xu, Chenzhong [2 ]
机构
[1] Wuhan Univ, Comp Sch, State Key Lab Software Engn, Wuhan 430072, Hubei, Peoples R China
[2] Chinese Acad Sci, Shenzhen Inst Adv Technol, Xueyuan Blvd 1068, Shenzhen 518055, Peoples R China
[3] Western Illinois Univ, Sch Comp Sci, Macomb, IL 61455 USA
[4] IIT, Dept Comp Sci, Chicago, IL 60616 USA
基金
美国国家科学基金会;
关键词
Parallel I/O system; parallel file system; data placement; solid state drive; SCHEME; CACHE;
D O I
10.1109/TPDS.2016.2636837
中图分类号
TP301 [理论、方法];
学科分类号
081202 ;
摘要
Multi-tiered Parallel I/O systems that combine traditional HDDs with emerging SSDs mitigate the cost burden of SSDs while benefiting from their superior I/O performance. While a multi-tiered parallel I/O system is promising for data-intensive applications in high-performance (HPC) domains, placing data on each tier of the system to achieve high I/O performance remains a challenge. In this paper, we propose a cost-aware region-level (CARL) data placement scheme in multi-tiered parallel I/O systems. CARL divides a large file into several small regions, and then places regions on different types of servers based on region access costs. CARL includes a static policy S-CARL and a dynamic policy D-CARL. For applications whose I/O access patterns are completely known, S-CARL calculates the region costs within the entire workload duration, and uses a static data placement scheme to selectively place regions on the proper servers. To adapt to applications whose access patterns are unknown in advance, D-CARL uses a dynamic data placement scheme which migrates data among different servers within each time window. We have implemented CARL under MPI-IO library and OrangeFS parallel file system environment. Our evaluation with representative benchmarks and an application shows that CARL is both feasible and able to improve I/O performance significantly.
引用
收藏
页码:1853 / 1865
页数:13
相关论文
共 13 条
  • [11] A Novel Cost-aware Data Placement Strategy for Edge-Cloud Collaborative Smart Systems
    Zhang, Yifei
    Xu, Jia
    Liu, Xiao
    Pan, Wuzhen
    Li, Xuejun
    2023 IEEE 16TH INTERNATIONAL CONFERENCE ON CLOUD COMPUTING, CLOUD, 2023, : 450 - 456
  • [12] A Holistic Heterogeneity-Aware Data Placement Scheme for Hybrid Parallel I/O Systems
    He, Shuibing
    Li, Zheng
    Zhou, Jiang
    Yin, Yanlong
    Xu, Xiaohua
    Chen, Yong
    Sun, Xian-He
    IEEE TRANSACTIONS ON PARALLEL AND DISTRIBUTED SYSTEMS, 2020, 31 (04) : 830 - 842
  • [13] A Cost-Effective Distribution-Aware Data Replication Scheme for Parallel I/O Systems
    He, Shuibing
    Sun, Xian-He
    IEEE TRANSACTIONS ON COMPUTERS, 2018, 67 (10) : 1374 - 1387