Search arXiv⌕ Search

arXiv · 2610.09292

An Interleaved Parallel Dependent Quantization Hardware Architecture for H.266/VVC

Abstract

While dependent quantization in H.266/VVC delivers a high compression ratio, its strong serial nature and high complexity result in poor real-time performance, making it difficult to deploy in practical scenarios. To improve the real-time performance of dependent quantization with minimal degradation to its compression performance, we propose a interleaved parallel dependent quantization hardware architecture with low BDBR loss, which achieves four-channel parallel dependent quantization by time-division multiplexing most combinational logic. This architecture adopts the proposed intra-CG context simplification scheme and the encoding scheme that skips decAbslevel during rate estimation. The proposed design incurs a BDBR loss of only 0.42% under the All Intra configuration and 0.38% under the Random Access configuration, respectively. Implemented in Verilog HDL, the dependent quantization hardware architecture achieves quantization speeds of 4K@33.2, 91.4, 331.2, and 454.1 fps at QP = 22, 27, 32, and 37, respectively, when implemented on the Xilinx XCZU19 FPGA. When implemented on ASIC using the TSMC 28nm process standard cell library, the corresponding quantization speeds reach 4K@84.0, 231.0, 837.0, and 1147.7 fps.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jun Zhang, Weizhi Bian, Hao Zhang. 2026-10-07. An Interleaved Parallel Dependent Quantization Hardware Architecture for H.266/VVC. https://arxiv.org/abs/2610.09292

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

MiX: Micro-Inverted-Scaling for End-to-End Low-Bit Vision-Language Model Acceleration

The deployment of Vision-Language Models (VLMs) on edge devices is severely bottlenecked by memory bandwidth, necessitating aggressive sub-8-bit quantization. Since edge accelerators are strictly constrained by area and power, they require end-to-end quantized models. However, the extreme dynamic range gap between multi-modal tokens causes standard block formats to suffer "microscaling collapse," where a single massive outlier hijacks the shared exponent, underflowing surrounding elements and destroying attention maps. To break this bottleneck, we propose Micro-Inverted-Scaling (MiX), a novel format that mathematically inverts the microscaling paradigm: rather than grouping multiple mantissas under one shared exponent, MiX groups private, per-element exponents under a single shared mantissa. To handle asymmetric VLM outlier topologies, we introduce an adaptive dual-format (MiX-MX) inference framework. By algebraically factoring out the shared MiX mantissa, this framework maps to a custom accelerator, replacing multipliers with efficient shifters. Evaluated end-to-end on multiple VLMs, our 4.5-bit MiX formulation exhibits equivalent or superior accuracy on multi-modal benchmarks compared to NVFP4. Simultaneously, the MiX accelerator delivers a 25% improvement in area efficiency over the NVFP4 baseline and a 2.3-4.5x speedup with 1.4-2.9x energy reduction across models compared to the state-of-the-art accelerator Focus, proving the inverted-scaling datapath is physically superior for efficient VLM deployment.

cs.AR↗

DynaTE: Accelerating Diffusion LLMs via Dynamic Token Execution

Diffusion-based LLMs (dLLMs) have recently emerged as a promising alternative to autoregressive (AR) LLMs by enabling bidirectional parallel refinement, alleviating the sequential decoding bottleneck of AR generation. However, their parallel iterative refinement mismatches AR accelerators optimized for sequential decoding and their discrete token generation differs from DiT accelerators designed for continuous denoising. Recent dLLM accelerators have explored workload-specific optimizations to reduce vocabulary processing overhead and redundant computation across denoising iterations. However, these approaches retain all tokens in parallel execution, despite varying token refinement utility and execution requirements. This paper presents DynaTE, a hardware--software co-design architecture that dynamically adapts accelerator execution to evolving token states during dLLM decoding. DynaTE first enables adaptive token execution by skipping low-utility token computation, while a dimension-reconfigurable PE array maintains high utilization under varying active-token patterns. Second, DynaTE exploits dynamic token dependencies through FLDD to refine a small number of locally dependent tokens within the current iteration, reducing the overall number of denoising iterations, while a Merge--Split--Merge dataflow hides the resulting serial overhead. Third, a streaming vocabulary engine interleaves multiple token streams from the LM head to accommodate irregular output variations caused by selective token computation and uneven vocabulary-selection demands. Evaluated on two representative dLLMs, DynaTE achieves 2.05--2.78$\times$ speedup and 2.99--3.93$\times$ higher energy efficiency over state-of-the-art dLLM accelerators, while delivering 2.55$\times$ speedup and 6.07$\times$ higher energy efficiency over Jetson AGX Orin.

cs.AR↗

Bridging the Last Mile of Circuit Design: PostEDA-Bench, a Hierarchical Benchmark for PPA Convergence and DRC Fixing

LLM-based agents are increasingly applied to the "last mile" of Electronic Design Automation (EDA): repairing residual sign-off Design Rule Check (DRC) violations and converging Power-Performance-Area (PPA) targets after tool runs. Existing EDA-LLM benchmarks, however, omit DRC fixing entirely and rely on flat hierarchies tied to a single toolchain. We introduce PostEDA-Bench, a hierarchical benchmark with 145 tasks across DRC-Essential, DRC-Reasoning, PPA-Mono, and PPA-Multi, supported by EDA toolchains with machine-checkable evaluation. Across eight commercial and open-source LLMs under multiple agent scaffolds, we find that agents handle synthetic DRC-Essential and single-objective PPA-Mono reasonably well but degrade sharply on the more practical DRC-Reasoning, where the best success rate is 36.66%, and PPA-Multi, where the best success rate is 20.00%; vision augmentation consistently enhances DRC-Bench; and trade-off reasoning, rather than knob knowledge, is the dominant PPA-Multi bottleneck.

cs.AR↗