Search arXiv⌕ Search

arXiv · 2609.37021

Low-level optimizations in high-level HDLs: Is there a benefit?

Abstract

This paper explores the applicability of functional programming to the design of Application-specific Integrated Circuits (ASICs). We investigate the impact of designing ASICs using high-level, abstract Hardware Description Language (HDL) features versus employing low-level optimizations on the area of the synthesized circuits. The aim is to determine whether using low-level optimizations is beneficial and, if so, whether it is worth the added implementation effort. To carry out the investigation, we implement an unsigned bit-serial multiply-accumulate (MAC) unit in 16 different configurations using the functional HDL Clash. We make use of both low-level bit-manipulation techniques as well as Clash's high-level constructs for circuit design. The experimental evaluation shows that some high-level constructs of Clash have negligible influence on the resulting circuit size, suggesting that using the full power of functional programming is a viable approach to hardware design. To evaluate the impact of the used HDL itself, we also implemented versions of the MAC in Verilog. The experiments clearly show that there seems to be an inherent overhead in using Clash compared to Verilog code written by a seasoned engineer.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Oliver Keszocze, Tjark Petersen, Arved Friedemann, Matthias Bo Stuart. 2026-09-29. Low-level optimizations in high-level HDLs: Is there a benefit?. https://arxiv.org/abs/2609.37021

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs

Emerging agentic large language model (LLM) workloads are driving rapidly growing demand for memory capacity and bandwidth. Different phases of inference, such as prefill and decode, have distinct requirements. Industry is responding by combining heterogeneous accelerators into interconnected systems, as exemplified by NVIDIA's Vera Rubin platform, where each device has its own memory architecture. The range of available memory technologies is also expanding. High-density on-chip SRAM, HBM, LPDDR, GDDR, and emerging options such as high-bandwidth flash (HBF) each offer different trade-offs in capacity, bandwidth, and power. Identifying efficient memory architectures for next-generation inference accelerators remains challenging because the design space spans workload characteristics, NPU design choices, and memory system designs. To address this challenge, we present MemExplorer, a new memory system synthesizer for heterogeneous NPU systems. MemExplorer provides a unified way to model memory technologies at different levels of the hierarchy, including on-chip and off-chip memory. It automatically selects an efficient heterogeneous memory system alongside NPU design choices, such as matrix engine size, to balance throughput and power across prefill and decode devices in a multi-device system. For agentic workloads under the same power budget, MemExplorer achieves up to 2.3 times the energy efficiency of the baseline NPU and 3.23 times that of an H100 in the prefill-only setting. At equivalent performance targets in the decode setting, it delivers up to 1.93 times and 2.72 times the power efficiency of the baseline NPU and H100, respectively.

cs.AR↗

PoisonCap: Efficient Hierarchical Temporal Safety for CHERI

In this paper, we present PoisonCap: scalable temporal safety with strict use-after-free protection and initialisation safety for CHERI systems. Efficient memory safety is an increasing priority for programming languages, operating systems, and hardware designs, and CHERI is a leading hardware/software system that provides native spatial safety and a foundation for temporal memory safety. Cornucopia Reloaded, the current state-of-the-art CHERI temporal safety solution, provides use-after-reallocation safety instead of stronger use-after-free safety, and is not able to enforce initialisation safety. We show that a new 'poison' capability format can be used to enforce strict use-after-free and initialisation safety, and also to communicate memory state to the microarchitecture for efficient cache management of quarantined memory. We enable elegant delegation of memory poisoning privilege using capability bounds to allow nested allocators to enforce safety on their consumers without disturbing upstream allocators. PoisonCap can replace the Cornucopia shadow bitmap, and also automatically zeros memory on reallocation, or optionally traps on read-before-write to enforce initialisation safety. As a result, it incurs no fundamental overhead relative to a Cornucopia baseline that zeros before reallocation, strengthening CHERI temporal safety without performance overhead.

cs.AR↗

BEACON: A Versatile Accelerator for Computational Pathology Applications

While accelerators for AI have seen great commercial success, it is challenging to replicate that success for other specialized domains due to a number of factors. We make the case that barriers for new accelerators can be lowered by starting with a baseline AI accelerator, and adding minimal logic to support new operators demanded by new specialized domains. This leads to a versatile chip that can be manufactured at high volume and deployed for a range of popular applications. We refer to this as the AI+X approach. This paper explores its potential for the emerging domain of Computational Pathology, which involves analysis of large whole-slide tissue images with a multi-stage pipeline. The pipeline requires support for a number of different kernels and operators - early stages perform segmentation and feature extraction, followed by graph creation with k nearest neighbor (kNN) algorithms, and finally inference with an iterative graph convolutional network (GCN) that alternates between Aggregation and Combination. We show that these stages execute inefficiently on a range of baseline CPU, GPU, AI, and GCN accelerators. That inefficiency is addressed with a combination of software re-structuring and small modifications to a baseline systolic AI accelerator. Many of the above kernels can be mapped to a systolic accelerator by offering a flexible datapath between processing elements and register access mechanisms. We add support for feature aggregation, load balanced execution, Euclidean distance calculation, binning, and counter aggregation. This additional flexibility and logic grows the area of a baseline AI chiplet by 1.1x, but by avoiding the memory wall and offering high parallelism, the proposed accelerator BEACON yields over an order of magnitude higher throughput for Computational Pathology than baseline CPU and GPU platforms.

cs.AR↗