Search arXiv⌕ Search

arXiv · 2610.01867

ZTA-Q: an Open-source RISC-V Platform for Accurate Quantized CNN Inference

Abstract

Low-precision inference is widely adopted in edge AI to reduce computational cost and memory footprint. However, existing open-source accelerator platforms provide limited end-to-end support for CNNs following the standard TensorFlow Lite integer inference scheme. This paper presents ZTA-Q, an open-source RISC-V-based platform that enables accurate deployment of TensorFlow Lite INT8 models. In addition to extending operator support, ZTA-Q provides a configurable post-processing datapath for studying how circuit-level approximations, including reduced multiplier precision, shared shift scaling, and simplified rounding, affect model accuracy. The proposed system is implemented on a Digilent Arty A7-100T FPGA and operates at 83.3 MHz. Evaluations on representative CNN models show that with LUT, register, and DSP overheads of 26.3%, 12.6%, and 150%, respectively, ZTA-Q limits the degradation in both top-1 and top-5 accuracy to within 0.25 percentage points.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Yike Li, Ajay Kumar M, Vishnu PS, Dimitrios S. Nikolopoulos, Bo Ji, Hans Vandierendonck, Deepu John. 2026-10-01. ZTA-Q: an Open-source RISC-V Platform for Accurate Quantized CNN Inference. https://arxiv.org/abs/2610.01867

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

From Physical Devices to RTL Models: Abstraction and Validation in Hardware Engineering

This paper introduces the foundational principles underlying hardware engineering models and argues that abstraction is their defining characteristic. Because abstraction necessarily omits detail and constrains what engineers can build, models are inherently incomplete in specific respects - or, as George Box famously observed, "All models are wrong, but some are useful". At the same time, abstraction is essential for simplification, which is key to managing complexity. More abstract models also tend to simulate faster because fewer details must be considered. This paper subsequently examines a range of abstraction methods in digital design - sometimes referred to as design disciplines - including lumped models, value-discrete models, and time-discrete models. Together with constraints that define the validity of the abstraction and design guidelines, these abstraction methods establish design disciplines. This paper further relates these forms of abstraction to pre-clustered design elements such as transistors, gates, registers, and transfer functions. These pre-clustered elements define abstraction levels, such as the gate level, and are presented as a key enabler of increased design productivity.

cs.AR↗

U-Sonic: An Open-Source 8-Channel Ultrasound Transmit IP in a 130 nm RISC-V SoC

Miniaturized ultrasound (US) probes require programmable and synchronized transmit (TX) excitation across multiple elements, while existing compact platforms often rely on limited microcontroller (MCU) pulse generators or closed-source fixed-function pulser devices. We present U-Sonic, an open-source digital US TX peripheral integrated into a 32-bit RISC-V system-on-chip (SoC). The implemented SoC integrates 8 pulser cores, while the parameterized architecture supports up to 16 channels. Each core generates single- or dual-tone bursts with programmable period, duty cycle, pulse count, polarity, and idle level, together with optional inverted stop pulses for active damping. A shared memory-mapped Open Bus Interface (OBI) enables synchronous start and stop of arbitrary channel subsets and supports composite bipolar, gated, and three-level excitation schemes. Functional correctness was verified in Verilator against a Python golden model over 4379 checked cycles across directed and randomized configurations, and confirmed on a Terasic DE10-Lite field-programmable gate array (FPGA). The design was synthesized and placed-and-routed in IHP 130 nm. The post-layout area in kilo gate equivalents (kGE), scales as 1.65 kGE plus 1.66 kGE per channel. The 8-channel instance occupies 14.9 kGE, corresponding to approximately 14.3% of the 104 kGE SoC. The register-transfer level (RTL), register descriptions, verification collateral, and software support are released as open source.

cs.AR↗

Open-Source Multi-Wire SPI Readout for Wearable Ultrasound Probes

Wearable ultrasound probes must transfer increasingly large acquisition payloads while maintaining compact, low-power electronics. In TinyProbe, the current bottleneck in data transfer occurs between the acquisition FPGA and the wireless system controller. This work presents an open-source, multi-wire SPI readout interface that uses serial command and address phases followed by a build-time-selectable dual- or quad-lane payload phase that is intended to address this bottleneck by increasing the potential bandwidth over the wifi limit while retaining compatibility with the Microcontroller-centric wearable US architecture. The interface emulates a serial flash memory, enabling compatibility with a broad range of microcontroller families and their existing peripheral interfaces. On the FPGA, the data path connects the existing acquisition FIFOs to the SPI interface through clock-domain crossing, sample reshaping, and packing into 32-bit words. Dual-SPI readout is integrated into the existing IGLOO2/SiWG917 TinyProbe architecture and verified at an SCLK frequency of 5 MHz. A separate Kria K26 testbed is used to characterize the FPGA SPI interface independently of the acquisition and wireless subsystems, demonstrating error-free transfers at SCLK frequencies up to 66 MHz. These measurements identify the SiWG917 multi-lane SPI implementation as the next bandwidth-limiting component and motivate a future upgrade of the system controller. The HDL and MCU implementations are released under a permissive open-source license.

cs.AR↗