Search arXivSearch

arXiv · 1912.07821

Valley-Coupled-Spintronic Non-Volatile Memories with Compute-In-Memory Support

Abstract

In this work, we propose valley-coupled spin-hall memories (VSH-MRAMs) based on monolayer WSe2. The key features of the proposed memories are (a) the ability to switch magnets with perpendicular magnetic anisotropy (PMA) via VSH effect and (b) an integrated gate that can modulate the charge/spin current (IC/IS) flow. The former attribute results in high energy efficiency (compared to the Giant-Spin Hall (GSH) effect-based devices with in-plane magnetic anisotropy (IMA) magnets). The latter feature leads to a compact access transistor-less memory array design. We experimentally measure the gate controllability of the current as well as the nonlocal resistance associated with VSH effect. Based on the measured data, we develop a simulation framework (using physical equations) to propose and analyze single-ended and differential VSH effect based magnetic memories (VSH-MRAM and DVSH-MRAM, respectively). At the array level, the proposed VSH/DVSH-MRAMs achieve 50%/ 11% lower write time, 59%/ 67% lower write energy and 35%/ 41% lower read energy at iso-sense margin, compared to single ended/differential (GSH/DGSH)-MRAMs. System level evaluation in the context of general purpose processor and intermittently-powered system shows up to 3.14X and 1.98X better energy efficiency for the proposed (D)VSH-MRAMs over (D)GSH-MRAMs respectively. Further, the differential sensing of the proposed DVSH-MRAM leads to natural and simultaneous in-memory computation of bit-wise AND and NOR logic functions. Using this feature, we design a computation-in-memory (CiM) architecture that performs Boolean logic and addition (ADD) with a single array access. System analysis performed by integrating our DVSH-MRAM: CiM in the Nios II processor across various application benchmarks shows up to 2.66X total energy savings, compared to DGSH-MRAM: CiM.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Sandeep Thirumala, Yi-Tse Hung, Shubham Jain, Arnab Raha, Niharika Thakuria, Vijay Raghunathan, Anand Raghunathan, Zhihong Chen, Sumeet Gupta. 2019-12-17. Valley-Coupled-Spintronic Non-Volatile Memories with Compute-In-Memory Support. https://doi.org/10.1109/tnano.2020.3012550

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Beyond HBM-on-GPU: Thermal Design Envelope for 3D Volumetric DRAM-on-GPU Integration

The scaling of GPUs for AI and HPC workloads is increasingly constrained by the capacity, bandwidth, and thermal limits of both 2.5D HBM-GPU and direct-stacked 3D HBM-on-GPU integration. This work establishes the thermal design envelope for 3D volumetric DRAM-on-GPU integration, in which vertically oriented DRAM dies and interleaved cooling cavities reshape heat flow and memory interfacing above the GPU. Using a package-level thermal model anchored to a consistent HBM-on-GPU baseline and driven by a realistic reticle-scale non-uniform GPU power map, we quantify the key parameters governing thermal feasibility. Stack height is the dominant limiter of peak temperature, while cooling-cavity conductivity shifts the feasible region, and mold insertion and stack orientation further modulate thermal behavior. A distributed memory-controller and network-on-chip tier introduces only a moderate thermal penalty. Although die-level parallelism increases bandwidth, the reduction in simulated training time saturates once execution becomes compute-bound. These results define a bounded co-design space across bandwidth, capacity, and thermal constraints for 3D volumetric DRAM-on-GPU integration.

cs.ET

Droop-Aware Foundation Model Power Flow

This paper develops a droop-aware extension of the GridFM power systems foundation model, embedding droop gains and frequency/voltage deadband parameters as per-bus node features to enable control-aware AC power-flow analysis. Existing power-flow datasets encode only static electrical features, conflating operating points from qualitatively different control regimes; this work resolves that gap by exposing droop and deadband parameters as structured node features, with deadband discontinuities handled through a smooth tanh approximation that preserves solver differentiability. A transformer-based graph neural network is pre-trained on masked reconstruction and fine-tuned on the resulting control-aware datasets. The framework is validated against PSCAD electromagnetic-transient simulations on a two-bus system (0.11% maximum steady-state error) and cross-validated against an independent PyPower droop solver on the IEEE 24-bus RTS. On the 24-bus system the surrogate attains R2 = 0.9996 for active generation and 0.0015 p.u. voltage-magnitude RMSE; scalability is confirmed on the IEEE 300-bus system (0.0036 p.u. RMSE, R2 = 0.9841 for voltage magnitude across 299,700 predictions). A three-mode control study further shows that the deadband widens the control-error distribution while leaving total droop compensation unchanged, establishing deadband width as an actionable node-level design feature.

cs.ET

GNN-based Path-aware multi-view Circuit Learning for Technology Mapping

Traditional technology mapping suffers from systemic inaccuracies in delay estimation due to its reliance on abstract, technology-agnostic delay models that fail to capture the nuanced timing behavior behavior of real post-mapping circuits. To address this fundamental limitation, we introduce GPA(graph neural network (GNN)-based Path-Aware multi-view circuit learning), a novel GNN framework that learns precise, data-driven delay predictions by synergistically fusing three complementary views of circuit structure: And-Inverter Graphs (AIGs)-based functional encoding, post-mapping technology emphasizes critical timing paths. Trained exclusively on real cell delays extracted from critical paths of industrial-grade post-mapping netlists, GPA learns to classify cut delays with unprecedented accuracy, directly informing smarter mapping decisions. Evaluated on the 19 EPFL combinational benchmarks, GPA achieves 19.9%, 2.1% and 4.1% average delay reduction over the conventional heuristics methods (techmap, MCH) and the prior state-of-the-art ML-based approach SLAP, respectively-without compromising area efficiency.

cs.ET