Search arXiv⌕ Search

arXiv · 2610.12111

RAPID-DNN: Reliability-Aware Partitioning for Robust DNN Inference Deployment on Edge AI

Abstract

Deep neural networks (DNNs) deployed on heterogeneous edge accelerators are increasingly exposed to runtime hardware faults. Tight power and resource budgets push these platforms toward aggressive voltage scaling and limit the use of hardware fault mitigation, while process variation, thermal stress, and device aging raise fault rates further. The resulting transient bit level corruptions propagate through inference and can sharply reduce prediction accuracy. Existing DNN partitioning frameworks optimize latency and energy under fault free assumptions, so their deployment strategies often fail to sustain reliable inference in realistic operating conditions. This paper presents RAPID DNN, a reliability aware partitioning framework that adds runtime reliability as a third optimization objective alongside latency and energy. RAPID DNN first characterizes layer wise reliability through systematic fault injection in the weight and activation domains to identify vulnerability critical layers. The resulting sensitivity profile guides a three objective NSGA II optimizer that jointly minimizes inference latency, energy consumption, and expected fault induced accuracy degradation. We evaluate RAPID DNN on AlexNet, SqueezeNet, ResNet18, VGG16, and MobileNetV2 using heterogeneous Eyeriss and SIMBA accelerator profiles. Across fault models and injection probabilities, it consistently improves inference robustness over the fault agnostic baseline. Under the representative mixed random fault configuration, RAPID DNN raises average Top 1 accuracy by 9.81 percentage points and reduces average expected fault induced accuracy degradation by 36.86%, at a cost of 2.38% average latency overhead and 3.34% average energy overhead. These results show that treating reliability as a partitioning objective enables robust DNN deployment on heterogeneous edge accelerators under runtime hardware faults.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mukta Debnath, Krishnendu Guha, Debasri Saha, Amlan Chakrabarti, Susmita Sur-Kolay. 2026-10-08. RAPID-DNN: Reliability-Aware Partitioning for Robust DNN Inference Deployment on Edge AI. https://arxiv.org/abs/2610.12111

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

An RRAM-based Hardware Implementation of a Radial Basis Function Neuron for Edge Classifiers

The deployment of modern machine learning (ML) solutions on resource-constrained edge devices highlights implementation challenges. This is especially true for extreme edge applications that include safety-critical components, such as autonomous navigation tasks. This paper demonstrates an artificial neural network (ANN) design leveraging Metal-Oxide Resistive RAM (RRAM) -based Analogue Content Addressable Memory (ACAM) as an efficient hardware substrate for performing metric-based classification and online adaptation on the edge. The proposed design is based on a custom Template piXeL (TXL) cell used for building the ACAM module, where each TXL cell acts as a configurable receptive field neuron. These cells employ a Radial Basis activation function to calculate the distance of an input from the programmed receptive field. The TXL can be organised into dense arrays for calculating the distance of a high-dimensional input against all stored prototypes, effectively performing fast and energy efficient similarity search. This hardware engine enables on-the-fly learning, where the receptive field parameters can be tuned to track domain shift. Through simulation of the proposed TXL-RBF classifier we can achieve 89.1\% accuracy on the MNIST dataset while consuming 185fJ per cell per operation when operating at 10MHz.

cs.ET↗

Perspective: Content-addressable memories as a computing primitive for today's AI and beyond

Modern artificial intelligence is predominantly executed on computing architectures optimized for dense linear algebra. While this has enabled the success of contemporary neural networks and motivated compute-in-memory (CIM) architectures, a growing class of artificial intelligence (AI) workloads depends on associative retrieval, identifying stored information by content or similarity rather than by explicit memory addresses. Such operations are central to transformer attention, tree-based inference, genomic search, and other retrieval-intensive applications, yet remain inefficiently supported by conventional memory systems. In this Perspective, we argue that content-addressable memories (CAMs) provide a complementary hardware primitive for associative processing in AI. We review their ability to perform massively parallel in-memory matching, discuss how emerging memory technologies can improve density and energy efficiency, and identify hierarchical search, application-specific architectures, hardware-aware learning, and heterogeneous integration with CIM as key directions for scalable associative computing. Together, these developments suggest that associative retrieval should complement linear algebra as a foundational computing primitive for future AI systems.

cs.ET↗

Analytical Modeling of Open- and Closed-Loop Dispersive Molecular Communication Channels with Pulsatile Flow

Molecular communication (MC) is a communication paradigm in which information is conveyed through the release, propagation, and reception of molecules. Many envisioned healthcare applications of MC are expected to operate inside the human body, where the cardiovascular system (CVS) may serve as the physical propagation environment and molecular transport is governed by diffusion and blood flow. Although blood flow is inherently pulsatile, most analytical MC channel models assume steady flow. In this paper, we develop a time-variant analytical model for dispersive MC channels with pulsatile flow. We derive the straight-duct response as a Normal distribution with time-variant mean and variance, capturing the combined effects of diffusion and pulsatile flow, and extend it to closed-loop channels through a wrapped-Normal representation. The model is validated against three-dimensional (3D) particle-based simulations (PBSs) for synthetic and physiologically motivated velocity waveforms. We further derive a closed-form first-order approximation for the first-arrival peak time and introduce the nondimensional indicator $S_{\mathrm{Rx}}$ for assessing the applicability of the reduced one-dimensional (1D) model. Our results show that pulsatility has the strongest influence when molecular transport is predominantly advective and the temporal flow variations are not averaged out, whereas stronger diffusion and faster pulsations reduce its impact on the received signal. Moreover, a PBS-based parameter sweep supports $S_{\mathrm{Rx}} \ge 2$ as a practical criterion for the applicability of the proposed analytical model. Finally, relating the analytical assumptions to representative blood-vessel classes shows that model applicability depends not only on the transport conditions but also on physiological properties such as vessel rigidity, geometric uniformity, and blood rheology.

cs.ET↗