Energy-Efficient Approximate Edge Inference Systems

被引：6

作者：

Ghosh, Soumendu Kumar ^{[1
]}

Raha, Arnab ^{[2
]}

Raghunathan, Vijay ^{[1
]}

机构：

[1] Purdue Univ, Elmore Family Sch Elect & Comp Engn, 610 Purdue Mall, W Lafayette, IN 47907 USA

[2] Intel Corp, 2200 Mission Coll Blvd, Santa Clara, CA 95054 USA

来源：

ACM TRANSACTIONS ON EMBEDDED COMPUTING SYSTEMS | 2023年 / 22卷 / 04期

关键词：

Approximate computing; approximate systems; deep learning; DRAM; edge AI; edge-to-cloud computing; energy efficiency; quality-aware pruning; quality-energy tradeoff; CMOS IMAGE SENSOR; NEURAL-NETWORKS; PERFORMANCE; CHALLENGES;

D O I：

10.1145/3589766

中图分类号：

TP3 [计算技术、计算机技术];

学科分类号：

0812 ;

摘要：

The rapid proliferation of the Internet of Things and the dramatic resurgence of artificial intelligence based application workloads have led to immense interest in performing inference on energy-constrained edge devices. Approximate computing (a design paradigm that trades off a small degradation in application quality for disproportionate energy savings) is a promising technique to enable energy-efficient inference at the edge. This article introduces the concept of an approximate edge inference system (AxIS) and proposes a systematic methodology to perform joint approximations between different subsystems in a deep neural network (DNN)-based edge inference system, leading to significant energy benefits compared to approximating individual subsystems in isolation. We use a smart camera system that executes various DNN-based image classification and object detection applications to illustrate how the sensor, memory, compute, and communication subsystems can all be approximated synergistically. We demonstrate our proposed methodology using two variants of a smart camera system: (a) Cam(Edge), where the DNN is executed locally on the edge device, and (b) CamCloud, where the edge device sends the captured image to a remote cloud server that executes the DNN. We have prototyped such an approximate inference system using an Intel Stratix IV GX-based Terasic TR4-230 FPGA development board. Experimental results obtained using six large DNNs and four compact DNNs running image classification applications demonstrate significant energy savings (approximate to 1.6x-4.7x for large DNNs and approximate to 1.5x-3.6x for small DNNs), for minimal (<1%) loss in application-level quality. Furthermore, results using four object detection DNNs exhibit energy savings of approximate to 1.5x-5.2x for similar quality loss. Compared to approximating a single subsystem in isolation, AxIS achieves 1.05x-3.25x gains in energy savings for image classification and 1.35x-4.2x gains for object detection on average, for minimal (<1%) application-level quality loss.

引用

页数：50

共 50 条

[21] Energy-efficient computing with approximate multipliers
Pilipovic, Ratko
Bulic, Patricio
Lotric, Uros
ELEKTROTEHNISKI VESTNIK, 2022, 89 (03): : 117 - 123
[22] Energy-Efficient Approximate MAC Unit
Adams, Elizabeth
Venkatachalam, Suganthi
Ko, Seok-Bum
2019 IEEE INTERNATIONAL SYMPOSIUM ON CIRCUITS AND SYSTEMS (ISCAS), 2019,
[23] Resource Provision for Energy-efficient Mobile Edge Computing Systems
Chang, Peiliang
Miao, Guowang
2018 IEEE GLOBAL COMMUNICATIONS CONFERENCE (GLOBECOM), 2018,
[24] Approximate Computing for Energy-efficient Error-resilient Multimedia Systems
Roy, Kaushik
PROCEEDINGS OF THE 2013 IEEE 16TH INTERNATIONAL SYMPOSIUM ON DESIGN AND DIAGNOSTICS OF ELECTRONIC CIRCUITS & SYSTEMS (DDECS), 2013, : 5 - 6
[25] EXTREM-EDGE-EXtensions To RISC-V for Energy-efficient ML inference at the EDGE of IoT
Verma, Vaibhav
Tracy II, Tommy
Stan, Mircea R.
SUSTAINABLE COMPUTING-INFORMATICS & SYSTEMS, 2022, 35
[26] An Energy-Efficient and Approximate Accelerator Design for Real-Time Canny Edge Detection
Leonardo Bandeira Soares
Julio Oliveira
Eduardo Antonio César da Costa
Sergio Bampi
Circuits, Systems, and Signal Processing, 2020, 39 : 6098 - 6120
[27] An Energy-Efficient and Approximate Accelerator Design for Real-Time Canny Edge Detection
Soares, Leonardo Bandeira
Oliveira, Julio
da Costa, Eduardo Antonio Cesar
Bampi, Sergio
CIRCUITS SYSTEMS AND SIGNAL PROCESSING, 2020, 39 (12) : 6098 - 6120
[28] eNODE: Energy-Efficient and Low-Latency Edge Inference and Training of Neural ODEs
Zhu, Junkang
Tao, Yaoyu
Zhang, Zhengya
2023 IEEE INTERNATIONAL SYMPOSIUM ON HIGH-PERFORMANCE COMPUTER ARCHITECTURE, HPCA, 2023, : 802 - 813
[29] Energy-efficient cooperative inference via adaptive deep neural network splitting at the edge
Labriji, Ibtissam
Merluzzi, Mattia
Airod, Fatima Ezzahra
Strinati, Emilio Calvanese
ICC 2023-IEEE INTERNATIONAL CONFERENCE ON COMMUNICATIONS, 2023, : 1712 - 1717
[30] A Converting Autoencoder Toward Low-latency and Energy-efficient DNN Inference at the Edge
Mahmud, Hasanul
Kang, Peng
Desai, Kevin
Lama, Palden
Prasad, Sushil K.
2024 IEEE INTERNATIONAL PARALLEL AND DISTRIBUTED PROCESSING SYMPOSIUM WORKSHOPS, IPDPSW 2024, 2024, : 592 - 599

← 1 2 3 4 5 →