Faithfully Rounded Floating-point Computations

被引：8

作者：

Lange, Marko ^{[1
]}

Rump, Siegfried M. ^{[1
,2
]}

机构：

[1] Waseda Univ, Fac Sci & Engn, Shinjuku Ku, 3-4-1 Okubo, Tokyo 1698555, Japan

[2] Hamburg Univ Technol, Inst Reliable Comp, Schwarzenberg Campus 3, D-21071 Hamburg, Germany

来源：

ACM TRANSACTIONS ON MATHEMATICAL SOFTWARE | 2020年 / 46卷 / 03期

基金：

日本科学技术振兴机构;

关键词：

Double-double; inaccurate cancellation; rigorous error bounds; ACCURATE;

D O I：

10.1145/3290955

中图分类号：

TP31 [计算机软件];

学科分类号：

081202 ; 0835 ;

摘要：

We present a pair arithmetic for the four basic operations and square root. It can be regarded as a simplified, more-efficient double-double arithmetic. The central assumption on the underlying arithmetic is the first standard model for error analysis for operations on a discrete set of real numbers. Neither do we require a floating-point grid nor a rounding to nearest property. Based on that, we define a relative rounding error unit u and prove rigorous error bounds for the computed result of an arbitrary arithmetic expression depending on u, the size of the expression, and possibly a condition measure. In the second part of this note, we extend the error analysis by examining requirements to ensure faithfully rounded outputs and apply our results to IEEE 754 standard conform floating-point systems. For a class of mathematical expressions, using an IEEE 754 standard conform arithmetic with base beta, the result is proved to be faithfully rounded for up to 1/root beta u - 2 operations. Our findings cover a number of previously published algorithms to compute faithfully rounded results, among them Horner's scheme, products, sums, dot products, or Euclidean norm. Beyond that, several other problems can be analyzed, such as polynomial interpolation, orientation problems, Householder transformations, or the smallest singular value of Hilbert matrices of large size.

引用

页数：20

共 50 条

[41] Floating-point tricks
Blinn, JF
IEEE COMPUTER GRAPHICS AND APPLICATIONS, 1997, 17 (04) : 80 - 84
[42] FLOATING-POINT REPLY
WILLIAMS, A
DR DOBBS JOURNAL, 1993, 18 (13): : 10 - 10
[43] FLOATING-POINT ARITHMETICS
WADEY, WG
JOURNAL OF THE ACM, 1960, 7 (02) : 129 - 139
[44] Floating-point verification
Harrison, J
FM 2005: FORMAL METHODS, PROCEEDINGS, 2005, 3582 : 529 - 532
[45] Floating-point verification
Harrison, John
JOURNAL OF UNIVERSAL COMPUTER SCIENCE, 2007, 13 (05) : 629 - 638
[46] FLOATING-POINT COUNTERS
KALUGIN, VV
INSTRUMENTS AND EXPERIMENTAL TECHNIQUES, 1982, 25 (05) : 1131 - 1133
[47] TAPERED FLOATING POINT - NEW FLOATING-POINT REPRESENTATION
MORRIS, R
IEEE TRANSACTIONS ON COMPUTERS, 1971, C 20 (12) : 1578 - &
[48] 24-BIT SINGLE-CHIP MULTIPLIER EASES FLOATING-POINT COMPUTATIONS
不详
EDN MAGAZINE-ELECTRICAL DESIGN NEWS, 1978, 23 (11): : 156 - 156
[49] An area- and energy-efficient hybrid architecture for floating-point FFT computations
Wang, Mingyu
Li, Zhaolin
MICROPROCESSORS AND MICROSYSTEMS, 2019, 65 : 14 - 22
[50] Using Floating-Point Intervals for Non-Modular Computations in Residue Number System
Isupov, Konstantin
IEEE ACCESS, 2020, 8 : 58603 - 58619

← 1 2 3 4 5 →