Search arXivSearch

arXiv · 2204.07503

Cryogenic Neuromorphic Hardware

Abstract

The revolution in artificial intelligence (AI) brings up an enormous storage and data processing requirement. Large power consumption and hardware overhead have become the main challenges for building next-generation AI hardware. To mitigate this, Neuromorphic computing has drawn immense attention due to its excellent capability for data processing with very low power consumption. While relentless research has been underway for years to minimize the power consumption in neuromorphic hardware, we are still a long way off from reaching the energy efficiency of the human brain. Furthermore, design complexity and process variation hinder the large-scale implementation of current neuromorphic platforms. Recently, the concept of implementing neuromorphic computing systems in cryogenic temperature has garnered intense interest thanks to their excellent speed and power metric. Several cryogenic devices can be engineered to work as neuromorphic primitives with ultra-low demand for power. Here we comprehensively review the cryogenic neuromorphic hardware. We classify the existing cryogenic neuromorphic hardware into several hierarchical categories and sketch a comparative analysis based on key performance metrics. Our analysis concisely describes the operation of the associated circuit topology and outlines the advantages and challenges encountered by the state-of-the-art technology platforms. Finally, we provide insights to circumvent these challenges for the future progression of research.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Md Mazharul Islam, Shamiul Alam, Md Shafayat Hossain, Kaushik Roy, Ahmedullah Aziz. 2022-08-24. Cryogenic Neuromorphic Hardware. https://doi.org/10.1063/5.0133515

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Composability rather than computation sets the cost of an analog EML hardware fabric

The operator eml(x, y) = exp(x) - ln(y) with the constant 1 generates the elementary functions, a continuous counterpart to NAND. Whether it yields a useful fabric had not been asked of hardware. We ask in network models, circuit simulation and SkyWater 130 nm layout. Four bipolar junctions evaluate the operator for 13 fJ, beating a width-matched digital datapath by 4-134x. The fabric assembled from them is not cheap: it loses to resource-matched baselines, and over the reals its grammar excludes trigonometry. Amplifiers holding those junctions' operating points take 74.5% of a cell's current, so a cell costs 3000 times what they spend. Extracted non-idealities cost 2.6x when a cell must hold a value and nothing when it need only be repeatable. Sharing them across cells recovers two of the three orders. The premise was that a universal primitive licenses a uniform machine. It survives in the primitive and fails in the machine.

cs.ET

Co-occurrence-Aware Quadratic Assignment for Local Feature Matching in Simultaneous Localization and Mapping

Local feature matching, which associates keypoints in two images as keypoint pairs, is fundamental to Visual Simultaneous Localization and Mapping (Visual SLAM). Nearest Neighbor (NN) search is commonly used for keypoint matching, but it has difficulty selecting correct keypoint pairs when multiple candidates have similar costs. To improve matching accuracy, this paper proposes a keypoint matching method that considers the pairwise co-occurrence of two keypoint pairs. The keypoint matching is formulated as a quadratic assignment problem, which is an NP-hard combinatorial optimization problem, making it difficult to solve quickly on conventional computers. Recently, Ising machines have been developed as computing devices capable of solving hard combinatorial optimization problems. Using a simulated bifurcation based Ising machine, the proposed method improved matching accuracy by approximately 8 percentage points over a conventional method on the HPatches dataset. Furthermore, we integrated the proposed method into ORB-SLAM3, a representative academic Visual SLAM system, and achieved a 3.78-fold improvement in absolute pose error (APE) and a 2.85-fold improvement in relative pose error (RPE) on the KITTI dataset scenes where multiple same shape objects are repeatedly arranged, which are challenging for accurate self-pose estimation by the original ORB-SLAM3.

cs.ET

System-Technology Co-Evaluation of A7 CFET and A10 NSFET Technologies from Cell Parasitics to Chip Reliability

Complementary FETs (CFETs) extend nanosheet FET (NSFET) scaling by vertically stacking n- and p-type gate-all-around (GAA) devices, thereby shrinking standard-cell area. The performance gain, however, cannot be assessed from device metrics alone, as CFET layouts also introduce larger cell-level parasitic resistance and capacitance (RC). In this work, we present a physics-based thermal- and aging-aware system-technology co-evaluation (STCO) flow to assess parasitic RCs in A7 CFET and A10 NSFET technology nodes. Our flow links calibrated device models, optimized standard-cell generation, automated GDS-to-TCAD conversion enabling accurate 3D parasitic RC extraction, full RTL-to-GDS implementation for an AI accelerator, multiphysics thermal analysis, and physics-based bias temperature instability (BTI) aging evaluation. Using the same device model for both technologies, we can isolate the impact of parasitic RCs and design at different levels of the design flow. The results of the AI accelerator design demonstrate that the A7 CFET reduces the chip area by 24.7% and the total wire length by 12%, improving the area efficiency TOPS/mm^2 by 74% relative to the baseline of the A10 NSFET. Under iso-frequency operation, results reveal that CFET voltage scaling reduces power by 68% and lowers power density from 148 W/cm^2 to 55 W/cm^2, which reduces the chip's temperature from 125 degrees C down to merely 62 degrees C. The resulting reduction in stress temperature suppresses 10-year BTI-induced degradation by 39%, reducing the required aging timing guardband by 53%.

cs.ET