Search arXiv⌕ Search

arXiv · 2609.36111

Automated Pre-Silicon Verification of High-Speed DDR5 and LPDDR5/6 Memory Controllers: Closed-Loop Timing, Mode Register, and PHY Synchronization in UVM

Abstract

External memory interfaces (such as LPDDR5/4 and DDR5) are essential components of modern mobile, cloud, and enterprise computing systems. While memory manufacturers focus on physical DRAM die development, the vast majority of semiconductor firms design custom application ASICs that require a dedicated Memory Controller to interface with these standardized external memories. Consequently, pre-silicon design verification of the Memory Controller RTL -- acting as the Device Under Test (DUT) against a third-party DRAM Verification IP (VIP) -- is a ubiquitous and critical challenge across the global semiconductor industry. This article presents an automated, pre-silicon configuration and closed-loop initialization framework for Memory Controller verification. The proposed solution parses JEDEC timing and configuration parameters directly from the DRAM VIP's database files (such as Denali SOMA files) to configure the Memory Controller DUT's registers, while a custom Mode Register Register Abstraction Layer (MR-RAL) tracks volatile DRAM VIP states in real-time. Furthermore, a dynamic, re-compilation-free PHY initialization flow randomizes interface parameters directly in the testbench, executes on-the-fly configuration generation via system calls, and parses the resulting register write sequences at runtime. By integrating these automated methodologies into pre-silicon verification flows, engineers can eliminate setup overhead, prevent false protocol violations, and enable comprehensive randomized testing of complex PHY and Memory Controller configurations.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Manan Patel, Anirban Majumder. 2026-09-28. Automated Pre-Silicon Verification of High-Speed DDR5 and LPDDR5/6 Memory Controllers: Closed-Loop Timing, Mode Register, and PHY Synchronization in UVM. https://arxiv.org/abs/2609.36111

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

zkPHIRE: A Programmable Accelerator for ZKPs over HIgh-degRee, Expressive Gates

Zero-Knowledge Proofs (ZKPs) have emerged as a powerful tool for secure and privacy-preserving computation. ZKPs enable one party to convince another of a statement's validity without revealing anything else. This capability has profound implications in many domains, including machine learning, blockchain, image authentication, and electronic voting. Despite their potential, ZKPs have seen limited deployment because of their exceptionally high computational overhead, which manifests primarily during proof generation. To mitigate these overheads, a (growing) body of researchers has proposed hardware accelerators and GPU implementations of both kernels and complete protocols. Prior art spans a wide variety of ZKP schemes that vary significantly in computational overhead, proof size, verifier cost, protocol setup, and trust. The latest and widely used ZKP protocols are intentionally designed to balance these trade-offs. One particular challenge in modern ZKP systems is supporting complex, high-degree gates using the SumCheck protocol. We address this challenge with a novel programmable accelerator to efficiently handle arbitrary custom gates via SumCheck. Our accelerator achieves upwards of $1000\times$ geomean speedup over CPU-based SumChecks across a range of gate types. We include this unit in zkPHIRE, a programmable, full-system accelerator that accelerates the HyperPlonk protocol. zkPHIRE achieves $1486\times$ geomean speedup over CPU and $11.87\times$ geomean speedup over the state-of-the-art at iso-area. Together, these results demonstrate compelling performance while scaling to large problem sizes (upwards of $2^{30}$ constraints) and maintaining small proof sizes ($4-5$ KB).

cs.AR↗

HBFSim: Fast and Faithful Simulation of High-Bandwidth Flash Under Real GPU Execution

High-Bandwidth Flash (HBF) places high-capacity NAND beside HBM to relieve the memory-capacity bottleneck of LLM inference, yet its system-level behavior cannot be evaluated before hardware becomes available. Cycle-level GPU simulators are too slow for production-scale models. Trace replay has a further shortcoming: it cannot capture the allocation, migration, and execution changes induced by different HBM-HBF configurations. Our key insight is that HBF need not be evaluated by simulating the GPU: only the program-visible effects of HBF need to be modeled. And only a real LLM workload running on real hardware can answer the arguments about HBF. Hence the modeled service has to be injected into that running program, and the injection must not destroy the GPU concurrency that would hide the original I/O latency. We present HBFSim, an open-source HBF simulator that executes LLM workloads on a real GPU while modeling HBF timing, thermal, and other behaviors online. HBFSim rewrites the PTX of the workload's kernels and routes accesses inside a registered address range into the HBF simulator. It supports asynchronous TMA transfers and capacities beyond physical GPU memory. HBFSim leaves the model's run unaffected across ordinary-memory, TMA, and capacity-mode tests. The delay it injects matches the delay requested to within 0.152%. We also design a coupled thermal module that puts HBF, HBM, and the GPU in one advanced package, which is important for answering how severe the hot throttling problem becomes after HBF runs for a long time. Experiments with Qwen3-30B show how package heating, HBM-HBF allocation, and shared MoE demand jointly constrain the design space of future HBF accelerators. The source code of HBFSim is available at https://github.com/SlugLab/hbfsim/.

cs.AR↗

MemExplorer: Navigating the Heterogeneous Memory Design Space for Agentic Inference NPUs

Emerging agentic large language model (LLM) workloads are driving rapidly growing demand for memory capacity and bandwidth. Different phases of inference, such as prefill and decode, have distinct requirements. Industry is responding by combining heterogeneous accelerators into interconnected systems, as exemplified by NVIDIA's Vera Rubin platform, where each device has its own memory architecture. The range of available memory technologies is also expanding. High-density on-chip SRAM, HBM, LPDDR, GDDR, and emerging options such as high-bandwidth flash (HBF) each offer different trade-offs in capacity, bandwidth, and power. Identifying efficient memory architectures for next-generation inference accelerators remains challenging because the design space spans workload characteristics, NPU design choices, and memory system designs. To address this challenge, we present MemExplorer, a new memory system synthesizer for heterogeneous NPU systems. MemExplorer provides a unified way to model memory technologies at different levels of the hierarchy, including on-chip and off-chip memory. It automatically selects an efficient heterogeneous memory system alongside NPU design choices, such as matrix engine size, to balance throughput and power across prefill and decode devices in a multi-device system. For agentic workloads under the same power budget, MemExplorer achieves up to 2.3 times the energy efficiency of the baseline NPU and 3.23 times that of an H100 in the prefill-only setting. At equivalent performance targets in the decode setting, it delivers up to 1.93 times and 2.72 times the power efficiency of the baseline NPU and H100, respectively.

cs.AR↗