Search arXiv⌕ Search

arXiv · 2609.37399

MEDEM: Multi-Engine DL Accelerator Design Methodology

Abstract

Multi-engine deep learning (DL) accelerators are becoming increasingly prevalent as they address the heterogeneity and growing complexity of modern DL workloads. To efficiently process diverse DL workloads, these accelerators must incorporate combinations of engines with complementary capabilities to match the distinct computational characteristics of these workloads' heterogeneous kernels. However, existing multi-engine DL accelerator design approaches lack a systematic methodology, leaving fundamental questions unresolved. These include how to co-design engines for workloads with diverse computational characteristics and which engine combinations minimize aggregate execution costs (such as time or energy) across such workloads. Addressing these questions requires efficient exploration of exponentially large design spaces. To address these questions systematically, this work proposes Multi-Engine DL Accelerator Design Methodology (MEDEM). MEDEM defines generic engine abstractions, co-designs candidate instances (engines), and selects a combination of co-designed engines to minimize aggregate execution cost given diverse DL workloads and a resource budget. MEDEM encompasses a set of design strategies that efficiently navigate the exponentially large design spaces of engine co-design and combination selection, identifying highly optimized multi-engine accelerators. A comprehensive evaluation demonstrates that MEDEM identifies accelerators that outperform state-of-the-art designs, delivering geometric-mean improvements of up to 4.84x in energy-delay product (EDP) and 1.59x in throughput. The improvements are achieved using different resource budgets, demonstrating MEDEM's scalability, and using 51 single- and multi-model DL workloads, demonstrating its generalizability.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Fareed Qararyah, Mohammad Ali Maleki, Pedro Trancoso. 2026-09-29. MEDEM: Multi-Engine DL Accelerator Design Methodology. https://arxiv.org/abs/2609.37399

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs

Emerging agentic large language model (LLM) workloads are driving rapidly growing demand for memory capacity and bandwidth. Different phases of inference, such as prefill and decode, have distinct requirements. Industry is responding by combining heterogeneous accelerators into interconnected systems, as exemplified by NVIDIA's Vera Rubin platform, where each device has its own memory architecture. The range of available memory technologies is also expanding. High-density on-chip SRAM, HBM, LPDDR, GDDR, and emerging options such as high-bandwidth flash (HBF) each offer different trade-offs in capacity, bandwidth, and power. Identifying efficient memory architectures for next-generation inference accelerators remains challenging because the design space spans workload characteristics, NPU design choices, and memory system designs. To address this challenge, we present MemExplorer, a new memory system synthesizer for heterogeneous NPU systems. MemExplorer provides a unified way to model memory technologies at different levels of the hierarchy, including on-chip and off-chip memory. It automatically selects an efficient heterogeneous memory system alongside NPU design choices, such as matrix engine size, to balance throughput and power across prefill and decode devices in a multi-device system. For agentic workloads under the same power budget, MemExplorer achieves up to 2.3 times the energy efficiency of the baseline NPU and 3.23 times that of an H100 in the prefill-only setting. At equivalent performance targets in the decode setting, it delivers up to 1.93 times and 2.72 times the power efficiency of the baseline NPU and H100, respectively.

cs.AR↗

PoisonCap: Efficient Hierarchical Temporal Safety for CHERI

In this paper, we present PoisonCap: scalable temporal safety with strict use-after-free protection and initialisation safety for CHERI systems. Efficient memory safety is an increasing priority for programming languages, operating systems, and hardware designs, and CHERI is a leading hardware/software system that provides native spatial safety and a foundation for temporal memory safety. Cornucopia Reloaded, the current state-of-the-art CHERI temporal safety solution, provides use-after-reallocation safety instead of stronger use-after-free safety, and is not able to enforce initialisation safety. We show that a new 'poison' capability format can be used to enforce strict use-after-free and initialisation safety, and also to communicate memory state to the microarchitecture for efficient cache management of quarantined memory. We enable elegant delegation of memory poisoning privilege using capability bounds to allow nested allocators to enforce safety on their consumers without disturbing upstream allocators. PoisonCap can replace the Cornucopia shadow bitmap, and also automatically zeros memory on reallocation, or optionally traps on read-before-write to enforce initialisation safety. As a result, it incurs no fundamental overhead relative to a Cornucopia baseline that zeros before reallocation, strengthening CHERI temporal safety without performance overhead.

cs.AR↗

BEACON: A Versatile Accelerator for Computational Pathology Applications

While accelerators for AI have seen great commercial success, it is challenging to replicate that success for other specialized domains due to a number of factors. We make the case that barriers for new accelerators can be lowered by starting with a baseline AI accelerator, and adding minimal logic to support new operators demanded by new specialized domains. This leads to a versatile chip that can be manufactured at high volume and deployed for a range of popular applications. We refer to this as the AI+X approach. This paper explores its potential for the emerging domain of Computational Pathology, which involves analysis of large whole-slide tissue images with a multi-stage pipeline. The pipeline requires support for a number of different kernels and operators - early stages perform segmentation and feature extraction, followed by graph creation with k nearest neighbor (kNN) algorithms, and finally inference with an iterative graph convolutional network (GCN) that alternates between Aggregation and Combination. We show that these stages execute inefficiently on a range of baseline CPU, GPU, AI, and GCN accelerators. That inefficiency is addressed with a combination of software re-structuring and small modifications to a baseline systolic AI accelerator. Many of the above kernels can be mapped to a systolic accelerator by offering a flexible datapath between processing elements and register access mechanisms. We add support for feature aggregation, load balanced execution, Euclidean distance calculation, binning, and counter aggregation. This additional flexibility and logic grows the area of a baseline AI chiplet by 1.1x, but by avoiding the memory wall and offering high parallelism, the proposed accelerator BEACON yields over an order of magnitude higher throughput for Computational Pathology than baseline CPU and GPU platforms.

cs.AR↗