Search arXiv⌕ Search

arXiv · 2607.24396

The SpiNNaker2 chip: a many-core platform for flexible and scalable brain-inspired computing

Abstract

In deep learning, efficiency gets more and more important to compensate for the ongoing growth in model sizes and applications. Neuromorphic hardware has long been advocated as an upcoming alternative to deep networks, taking inspiration from the brain for achieving unprecedented energy efficiency. However, demonstrations of these gains only recently began to grow in complexity and real-world applicability. With SpiNNaker2, we present a chip that bridges the gap between deep networks and neuromorphic computing and allows for flexible exploration of computing approaches that combine both worlds. It features 152 processing elements equipped with an ARM M4F processor and dedicated accelerators, an extended SpiNNaker routing fabric for scalable event-based communication and a range of external interfaces for system integration, including Gbit Ethernet and an LPDDR4 memory interface. We demonstrate performance and efficiency of the SpiNNaker2 chip for neuromorphic and deep network workloads, as well as novel event-based computing approaches. For deep network workloads, the chip achieves up to 4.5 TOPS in high performance mode and up to 2.7 TOPS/W efficiency in high efficiency mode for INT8 workloads. The chip supports spiking neural networks with >150000 neurons and >1.8 billion synaptic events/s when simulated with a 1 ms time step. Its low baseline power of less than 250 mW allows for efficiency even under varying workload conditions, allowing to explore sparse and event-based modes of computation. All this demonstrates the chip's capabilities as a universal hardware platform for scalable brain-inspired computing and its combinations with mainstream deep network approaches.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Stefan Scholze, Johannes Partzsch, Sebastian Höppner, Florian Kelber, Andreas Dixius, Marco Stolba, Sirine Arfa, Marc Berthel, Georg Ellguth, Jim Garside, Hector A. Gonzalez, Stephan Hartmann, Thomas Kiel-Hocker, Dongwei Hu, Matthias Jobst, Khaleelulla Khan Nazeer, Tim Langer, Chen Liu, Gengting Liu, Matthias Lohrmann, Mantas Mikaitis, Felix Neumärker, Amirhossein Rostami, Stefan Schiefer, Tilo Schubert, Delong Shang, Bernhard Vogginger, Yexin Yan, Steve Furber, Christian Mayr. 2026-07-27. The SpiNNaker2 chip: a many-core platform for flexible and scalable brain-inspired computing. https://doi.org/10.1109/ojcas.2026.3714974

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Highly Scalable Selectorless Cryogenic Memory Array Using Ferroelectric Josephson Field-Effect Transistors

Scalable memory systems that satisfy the temperature, speed, and energy requirements of cryogenic environments are essential for the development of large-scale quantum computers. They may also benefit high-performance computing and space applications. However, existing cryogenic memory technologies often suffer from limited scalability, low operating speed, and/or high power consumption, restricting the scalability of target applications. Ferroelectric Josephson field-effect transistors (Fe-JoFETs), which combine ferroelectric polarization with the superconducting properties of Josephson junctions, offer a promising solution. The ferroelectric layer enables nonvolatile storage capability, while the Josephson junction supports high-speed, energy-efficient operations. In this work, we leverage Fe-JoFETs to develop a highly scalable, ultra-low-power, nonvolatile cryogenic memory array that does not need additional selector devices for random access. Moreover, the superconducting component of Fe-JoFET provides a binary decision during read, eliminating the need for sensing peripheral circuitry. We first develop a physics-based Verilog-A compact model for Fe-JoFETs and use it to verify the functionality of the proposed memory array. By eliminating both selector and sensing circuitry, the proposed memory architecture offers higher scalability than existing technologies. The ultra-low-power operation of this memory also makes it compatible with strict power budgets of cryogenic applications.

cs.ET↗

A fully parallel densely connected probabilistic Ising machine with inertia for real-time applications

Ising machines---special-purpose hardware for heuristically solving Ising optimization problems---based on probabilistic bits (p-bits) have been established as a promising alternative to heuristic optimization algorithms run on conventional computers. However, it has---until now---been thought that Ising spins that are connected in probabilistic Ising machines (PIMs) cannot be updated in parallel without ruining the machine's solving ability. This has presented a major challenge to realizing the potential for probabilistic Ising machines to act as fast solvers for densely connected Ising problems. In this paper, we show that it is possible to circumvent this conventional wisdom. We introduce a modified form of Ising spin dynamics for PIMs, adding an inertia term, and verify in algorithm simulations, field-programmable gate array (FPGA) emulation, and in FPGA experiments that the modified dynamics enables fully parallel, synchronous updates and at the same time improves the achieved success probability. Our evaluations were performed with various types of abstract (Max-Cut and Sherrington-Kirkpatrick model) and application-derived (multiple-input and multiple-output, MIMO detection) dense Ising benchmark instances. Performing fully parallel updates results in a speed advantage that grows superlinearly with the number of spins, giving rise to large time-to-solution reductions for practical problem sizes. For both MC and the SK model at a problem size of 200, our approach achieved an average speedup of ~34x, with the best single-instance speedup reaching 150x. As an example of the practical utility of our approach in an application where speed is critical, we co-design the algorithm dynamics and hardware implementation for MIMO detection, achieving improved detection accuracy relative to the standard linear detector and higher throughput than the conventional sequential-update PIM.

cs.ET↗

Towards clinical adoption of voice and speech as measures of health: the need for harmonization

Speech and voice are multidimensional signals that capture both communicative intent and underlying physiological processes, providing a unique, non-invasive window into health. Analyzing these signals has the potential to yield digital biomarkers that (i) provide scalable, objective measurement tools for research and clinical care and (ii) reflect the presence or progression of diverse conditions, including neurological, psychiatric, respiratory, and cardiovascular disorders. Realizing this promise, however, requires the field to overcome pervasive reproducibility and generalizability issues due to heterogeneous data collection, processing, and analysis practices. A major source of this heterogeneity is how underlying acoustic measures themselves are defined and computed. In this paper, we outline key considerations across the speech biomarker discovery lifecycle, from data collection through machine learning modeling to clinical interpretation, needed to achieve reliable, reproducible, and clinically translatable results. Chief among these is the need for harmonization efforts to start from common, precisely specified measure definitions. As a first step, we therefore provide definitions, physiological correlates, and computational implementations for a minimal, clinically interpretable set of core speech measures spanning respiration, phonation, articulation, and fluency. We close by discussing ongoing standardization efforts and the open challenges that remain in advancing the adoption of speech- and voice-based digital biomarkers.

cs.ET↗