Search arXiv⌕ Search

arXiv · 2610.03061

Divide and conquer: Scalable performance and energy in MCM GPUs

Abstract

Multi-chip-module (MCM) GPUs offer a promising path to scale compute capability beyond monolithic designs by integrating multiple chiplets on a common package. However, the impact of disaggregation on performance scalability and energy consumption remains underexplored. The design space grows rapidly across dimensions such as SMs per chiplet, chiplet count, and interconnection network. The inter-chiplet network is particularly critical, as it determines whether additional compute resources translate into performance gains. This limited understanding leaves industry and research without clear guidance on the performance and energy trade-offs of MCM GPU scaling. In this work, we investigate whether distributing compute and memory capability across multiple chiplets offers a more scalable alternative to concentrating resources. We quantify their effects on performance, energy, and efficiency and examine how inter-chiplet topology influences scalability at different system sizes. Our results demonstrate that a 16 chiplet Torus configuration with 256 SMs delivers a remarkable $2.40\times$ performance improvement over a state-of-the-art MCM architecture with the same compute capability, while simultaneously reducing energy consumption by $4.45\times$. These substantial gains provide evidence that disaggregation is a first-order architectural factor and will be critical to unlocking the performance and energy-efficiency potential of next-generation GPUs.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Mario Ibáñez Bolado, Borja Pérez Pavón, Jose Luis Bosque Orero, Julio Ramón Beivide. 2026-10-02. Divide and conquer: Scalable performance and energy in MCM GPUs. https://arxiv.org/abs/2610.03061

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A Self-Calibrating Framework for Analog Circuit Sizing Using LLM-Derived Analytical Equations

We present a design automation framework for analog circuit sizing that produces calibrated, topology-specific analytical equations from raw circuit netlists. A large language model (LLM) derives a complete Python sizing function in which each device dimension is traceable to a specific design rationale - a form of interpretable output absent from existing optimization-based and LLM-based sizing methods. A deterministic calibration loop extracts process-dependent parameters from a single DC operating point simulation, while a prediction-error feedback mechanism compensates for analytical inaccuracies. We validate the framework on circuits ranging from 6 to 30 transistors - spanning single-stage, current-mirror (simple and cascoded), folded-cascode, gain-boosted folded-cascode, two-stage Miller-compensated, nested-Miller-compensated, and complementary class-AB output topologies - across six process nodes from 32 nm to 180 nm. On matched-specification benchmarks, including the class-AB opamp case, the framework converges within a few simulations. Despite large initial prediction errors, convergence depends on the measurement-feedback architecture, not prediction accuracy. The one-shot calibration automatically captures process-dependent variations, enabling cross-node portability without modification, retraining, or per-process characterization.

cs.AR↗

Towards Enabling Distance-Based Memory Addressing

Approximate Nearest-Neighbor Search (ANNS) in high dimensional vector datasets is an application of significant prevalence across different AI applications. However, such an operation is significantly bandwidth limited at large workingset sizes owing to the curse of dimensionality. Traditional indices used to accelerate ANNS rely on search-space pruning as a preprocessing step to alleviate such bandwidth requirement, but such optimization occurs either at the cost of increased bandwidth-inefficiency and/or degradation of search quality. This paper proposes a data-parallel hardware/software mechanism for performing large-scale similarity search in-memory. We propose a novel algorithm to simplify the computation requirement for similarity search across various distance metrics through lightweight primitives to perform a fast and approximate data-parallel brute-force search on the entire vector space. We further build a memory system capable of executing the required operations to generate a distance metric per datapoints, which is then used to enable pruning as a post-processing step. We offer adequate software support for user control over the proposed system. By enabling such search-space pruning as a post-processing step, we achieve near-perfect recall across representative workloads while achieving orders of magnitude performance and energy improvement over state-of-the-art algorithmic approaches on million and billion-scale workloads.

cs.AR↗

Reinforcement Learning-Based Optimization of Workload-Aware Power Delivery Networks

Power Delivery Networks (PDNs) are critical components of modern VLSI chips, providing stable voltage levels while satisfying electromigration (EM) and IR-drop constraints. Conventional PDN design methodologies typically rely on worst-case assumptions, often resulting in over-provisioned networks and inefficient use of resources. This paper presents a reinforcement learning-based framework for the optimization of workload-aware PDNs. The proposed methodology first generates workload-aware PDNs using architectural power traces obtained from system-level simulations. These power traces are mapped to spatial power density distributions, enabling adaptive allocation of PDN resources according to local current demand. A reinforcement learning agent then performs wire-width optimization to minimize PDN area while maintaining EM and voltage integrity constraints. Electrical and reliability metrics are obtained using SPICE-based circuit analysis and EM lifetime estimation. Experimental evaluation is performed on a dataset of workload-aware PDNs generated from 4-, 8-, and 16-core multiprocessor floorplans using PARSEC and SPLASH-2 benchmark workloads. Furthermore, the proposed Deep Q-Network (DQN)-based optimizer reduces the average normalized PDN area by 47\% while satisfying all EM and IR-drop constraints. Compared to simulated annealing, the proposed approach achieves comparable optimization quality while providing approximately 26$\times$ faster optimization.

cs.AR↗