Search arXiv⌕ Search

arXiv subjects

Cristian Zambelli

Publications and source records attributed to Cristian Zambelli.

3 recordsLinked to original sources

The future of 3D NAND flash technology

3D NAND flash has become the foundational non-volatile storage platform of the AI era, underpinning workloads from model training to large-scale inference and cold data archival. With roadmaps now targeting kilolayer stacks and tens of trillions of devices per die, scaling is no longer governed primarily by process integration or lithography. Instead, it is increasingly constrained by the physics of charge storage itself: lateral charge migration, electrostatic coupling, read disturb, and transport limitations are entering a margin-limited regime. In this regime, their collective interaction, not any single mechanism, compresses operating margins with each generation. In this Perspective, we examine these converging bottlenecks and argue that sustaining NAND scaling will require application-specific co-optimization rather than a monolithic device roadmap. In this framework, conventional charge-trap flash, ferroelectric storage, alternative channel materials, and system-level integration are best viewed as complementary solutions targeted to distinct tiers of data-centric computing.

cond-mat.mes-hall↗

A Survey on Design Methodologies for Accelerating Deep Learning on Heterogeneous Architectures

Given their increasing size and complexity, the need for efficient execution of deep neural networks has become increasingly pressing in the design of heterogeneous High-Performance Computing (HPC) and edge platforms, leading to a wide variety of proposals for specialized deep learning architectures and hardware accelerators. The design of such architectures and accelerators requires a multidisciplinary approach combining expertise from several areas, from machine learning to computer architecture, low-level hardware design, and approximate computing. Several methodologies and tools have been proposed to improve the process of designing accelerators for deep learning, aimed at maximizing parallelism and minimizing data movement to achieve high performance and energy efficiency. This paper critically reviews influential tools and design methodologies for Deep Learning accelerators, offering a wide perspective in this rapidly evolving field. This work complements surveys on architectures and accelerators by covering hardware-software co-design, automated synthesis, domain-specific compilers, design space exploration, modeling, and simulation, providing insights into technical challenges and open research directions.

cs.AR↗

A Survey on Deep Learning Hardware Accelerators for Heterogeneous HPC Platforms

Recent trends in deep learning (DL) have made hardware accelerators essential for various high-performance computing (HPC) applications, including image classification, computer vision, and speech recognition. This survey summarizes and classifies the most recent developments in DL accelerators, focusing on their role in meeting the performance demands of HPC applications. We explore cutting-edge approaches to DL acceleration, covering not only GPU- and TPU-based platforms but also specialized hardware such as FPGA- and ASIC-based accelerators, Neural Processing Units, open hardware RISC-V-based accelerators, and co-processors. This survey also describes accelerators leveraging emerging memory technologies and computing paradigms, including 3D-stacked Processor-In-Memory, non-volatile memories like Resistive RAM and Phase Change Memories used for in-memory computing, as well as Neuromorphic Processing Units, and Multi-Chip Module-based accelerators. Furthermore, we provide insights into emerging quantum-based accelerators and photonics. Finally, this survey categorizes the most influential architectures and technologies from recent years, offering readers a comprehensive perspective on the rapidly evolving field of deep learning acceleration.

cs.AR↗