Search arXivSearch

arXiv subjects

Themis Prodromakis

Publications and source records attributed to Themis Prodromakis.

At least 19 recordsLinked to original sources

A relational fabrication-to-modeling database for memristor devices

Resistive random access memory (RRAM) devices, also known as memristors, are highly dependent on fabrication, location on the wafer, and measurement protocols. However, openly available large-scale memristor datasets remain scarce, particularly those that combine experimental metadata with electrical measurements and preserve explicit links across the experimental workflow. Here, we present a comprehensive relational database of automated electrical characterization data from oxide-based memristors. The database links 6,190 memristor devices across TiN/HfOx/TiN, Pt/TiOx/AlOy/Pt, and Pt/TiOx/Pt stacks to 161,006 validated experiments and over 169 million electrical point records. It integrates fabrication, wafer mapping, electrical characterization, and modeling to describe electroforming, current-voltage non-linearity, memory windows (R_off/R_on), switching dynamics, and short-term volatility. Released as a normalized and indexed SQLite database with schema documentation, graphical user interfaces, and examples of empirical modeling, this resource supports provenance-aware querying, statistical analysis, and data-driven applications for the development of future memristor technologies.

cond-mat.mtrl-sci

Gen-TAS: A Generative AI-Aided Hardware-Software Task Allocation Framework for FPGA-GPP Heterogeneous Systems

FPGA-GPP heterogeneous systems combine software flexibility with the performance and energy efficiency of reconfigurable hardware. However, determining which application tasks should execute on the GPP or FPGA requires extensive expertise and design-space exploration, particularly when user objectives vary across latency, communication, resource utilisation, and power. This paper proposes Gen-TAS, a knowledge-grounded LLM framework for user-specific FPGA-GPP task allocation. By combining task-graph analysis with RAG, Gen-TAS grounds LLM reasoning in historical implementation knowledge and generates multiple explainable strategies tailored to the specified objectives. Human-in-the-loop selection and a deterministic backend connect LLM-generated decisions to reproducible FPGA SoC implementations. Experiments on CNN and SDR workloads across multiple LLMs demonstrate stable, requirement-driven allocation. Under latency-oriented objectives, implementations following the selected strategies achieve speedups of up to 2.45$\times$ and 92.53$\times$, respectively, relative to the corresponding all-GPP baselines while other objectives select strategies that trade some acceleration performance for FPGA-GPP communication, resource utilisation, or FPGA power.

cs.AR

LLMs in Digital EDA: A perspective on shifting roles from Generation to Orchestration

Electronic design automation (EDA) has advanced engineering productivity through successive generations of tooling that progressively automate synthesis, optimisation, and verification. Large language models (LLMs) extend this trajectory by enabling direct translation from design intent to hardware implementations. In most of the EDA literature, LLM-based solutions are typically assisting siloed design stages or tasks, however this obscured the drivers by which capability emerges and systems scale. In this Perspective, we instead define three hierarchical roles that reveal how capability accumulates: a Generator that produces design artifacts in a single pass, an Agent that refines outputs through iterative tool feedback, and an Orchestrator that coordinates decisions across EDA-stages. Across published systems, this reveals a syntax trap in which models are trained to produce plausible code rather than physically correct hardware, compounded by fragmented tools and loss of design context that obscure how decisions affect later stages. Comparisons across the three roles show that current approaches struggle to scale to industrial designs, motivating a shift towards a standardised, physics-aware orchestrator that connects tools and agents across the EDA flow for more reliable and accessible hardware design.

cs.AR

Vision-centric generative AI models: A software-hardware perspective

Vision generative artificial intelligence (AI) has emerged as one of the most rapidly advancing areas of deep learning. The explosion of multimodal models has made them widely associated with text-to-image applications running on large datacentres. However, vision generative models are equally needed in applications that operate under strict hardware constraints at the edge, including autonomous vehicles, agricultural sensors, and mobile devices. In this Perspective, we argue that progress in vision generative AI has been driven by output quality, with hardware evolving reactively to accommodate growing model demands. We quantify the parameter cost and energy efficiency of these models across a range of accelerator platforms, and map four generative model families against seven real-world application domains. Finally, we advocate a software-hardware co-design approach, where deployment constraints are considered from the start of the design process, ensuring that the "right model" runs on the "right hardware" to serve the "right application", making generative AI deployment sustainable and accessible across a much broader range of platforms.

cs.CV

Parallel Spatial Photonic Programming of Optoelectronic IGZO RRAM with a compact $\mu$LED Array

Optoelectronic resistive random-access memory elements (ORRAM) are critical emerging devices that leverage photonic technologies to bring the advantages of optical programming to traditionally electronic memristive platforms for neuromorphic computing and artificial intelligence. In this work, a free-space optic micro-LED ($\mu$LED) array is combined with a 2-terminal oxide semiconductor ORRAM (based on IGZO\textsubscript{Rich}/IGZO active layers), to realise parallel spatial programming of form-free memristive arrays. We report the optical and electrical programming of resistive states with potentiation/depression analysis of various stimuli parameters (pulse frequency, pulse width, pulse amplitude). Further, we demonstrate the simultaneous photonic-electronic programming of the ORRAM with optical SET (blue 450\,nm) and electrical RESET functionality. Persistent photocurrent is also observed and exploited as a pathway to fading memory or synaptic plasticity for temporal bit encoding. Finally, parallel optical $\mu$LED to ORRAM channels are demonstrated to achieve the simultaneous photonic programming of multiple devices and the writing of spatial patterns across a chip of memristive IGZO devices. This work highlights the light-enabled scalability of the optoelectronic platform and the feasibility of ORRAM to interface with spatially-multiplexed optical sources to bring neuromorphic technologies directly into applications that process and sense in the optical domain.

physics.optics

Closing the Loop on LLM-Generated RTL Assertions with Quality-Aware Formal Verification

Large language model (LLM) based assertion generation is making formal verification more accessible for Register Transfer Level (RTL) designs, but three practical issues remain. Generated properties can pass for the wrong reason, proof cost can vary widely from one design to another, and failing traces are hard to interpret. This paper presents a lightweight, open-source framework that addresses these issues in one loop. Our method combines mutation-guided refinement to reject weak assertions, including vacuous ones and those that fail to distinguish faulty behaviour, a solver-selection stage that chooses among candidate Satisfiability Modulo Theories (SMT) backends using RTL structure, and causal narrative synthesis to explain why a proof failed. Across diverse RTL designs, the framework improves confidence in generated assertions, reduces runtime variability over fixed-solver choices, and produces failure explanations that remain grounded in the counterexample trace. The results suggest that quality-aware closure, rather than assertion generation alone, is the missing step for practical LLM-assisted formal verification.

cs.AR

Non-frontal face recognition using GANs and memristor-based classifiers

Face recognition systems have advanced significantly through deep learning techniques, delivering high performance and robustness in complex scenarios. However, these approaches incur substantial computational overhead, limiting their in situ applicability in resource-constrained platforms such as drones, where they can address challenges including non-frontal facial imagery. Memristor-based neuromorphic systems have emerged as a compelling approach for edge AI applications, combining biologically inspired processing with efficient and scalable computation. In this work, we propose a facial recognition framework that addresses non-frontal pose variations by integrating lightweight generative adversarial network (GAN)-based pose frontalisation with memristor-based neuromorphic recognition. The experimental results on two datasets demonstrate the effectiveness of combining adversarial learning with memristive technology, achieving up to 96% identification accuracy. The proposed approach alleviates the computational bottlenecks of conventional AI and offers a scalable, efficient solution for face recognition in dynamic real-world environments.

cs.CV

BenDi: An Energy-Efficient Quasi-Stochastic Systolic Architecture for Edge Bioelectronics

Continuous long-term monitoring and diagnosis of biomedical signals, such as electrocardiograms (ECGs), can help mitigate an increasing threat to public health. Artificial Intelligence (AI) models, such as Convolutional Neural Networks (CNNs), provide accurate monitoring and classification for relevant diseases; however, they require more computational resources than conventional AI hardware can typically afford, especially for a resource-constrained environment on the edge. In this work, we present BenDi, an energy-efficient quasi-stochastic systolic architecture for bioelectronic systems on the edge. BenDi leverages multiple levels of energy and power optimization, ranging from circuits to software quantization, including low supply voltage, the \underline{Ben}t-Pyramid data format for quasi-stochastic multiplication, the \underline{Di}P systolic dataflow, and hardware-aware quantization, to handle CNNs with high accuracy on the edge within limited hardware budgets. The hardware implementation results, using a commercial 22nm technology, show that BenDi architecture, at 0.5 Voltage and 100 MHz, offers 3.35x smaller area and 5x higher energy efficiency, compared to state-of-the-art binary-based weight-stationary systolic architectures. Regarding Bioelectronic edge systems, BenDi achieves an order-of-magnitude improvement in energy efficiency and another order-of-magnitude improvement in area efficiency, compared to its counterparts. This significant improvement comes at the cost of 1\% to 3.3\% accuracy loss on the MIT-BIH and Apnea-ECG benchmarks, respectively, compared with conventional computing using the 32-bit floating-point format.

cs.AR

An RRAM-based Hardware Implementation of a Radial Basis Function Neuron for Edge Classifiers

The deployment of modern machine learning (ML) solutions on resource-constrained edge devices highlights implementation challenges. This is especially true for extreme edge applications that include safety-critical components, such as autonomous navigation tasks. This paper demonstrates an artificial neural network (ANN) design leveraging Metal-Oxide Resistive RAM (RRAM) -based Analogue Content Addressable Memory (ACAM) as an efficient hardware substrate for performing metric-based classification and online adaptation on the edge. The proposed design is based on a custom Template piXeL (TXL) cell used for building the ACAM module, where each TXL cell acts as a configurable receptive field neuron. These cells employ a Radial Basis activation function to calculate the distance of an input from the programmed receptive field. The TXL can be organised into dense arrays for calculating the distance of a high-dimensional input against all stored prototypes, effectively performing fast and energy efficient similarity search. This hardware engine enables on-the-fly learning, where the receptive field parameters can be tuned to track domain shift. Through simulation of the proposed TXL-RBF classifier we can achieve 89.1\% accuracy on the MNIST dataset while consuming 185fJ per cell per operation when operating at 100MHz.

cs.ET

Implementation and Performance Evaluation of CMOS-integrated Memristor-driven Flip-flop Circuits

In this work, we report implementation and performance evaluation of memristor-driven fundamental logic gates, including NOT, AND, NAND, OR, NOR, and XOR, and novel and optimized design of the sequential logic circuits, such as D flip-flop, T-flip-flop, JK-flip-flop, and SR-flip-flop. The design, implementation, and optimization of these logic circuits were performed in SPECTRE in Cadence Virtuoso and integrated with 90 nm CMOS technology node. Additionally, we discuss an optimized design of memristor-driven logic gates and sequential logic circuits, and draw a comparative analysis with the other reported state-of-the-art work on sequential circuits. Moreover, the utilized memristor framework was experimentally pre-validated with the experimental data of Y2O3-based memristive devices, which shows significantly low values of variability during switching in both device-to-device (D2D) and cycle-to-cycle (C2C) operation. The performance metrics were calculated in terms of area, power, and delay of these sequential circuits and were found to be reduced by more than ~24%, 60%, and 58%, respectively, as compared to the other state-of-the-art work on sequential circuits. Therefore, the implemented memristor-based design significantly improves the performance of various logic designs, which makes it more area and power-efficient and shows the potential of memristor in designing various low-power, low-cost, ultrafast, and compact circuits.

cs.AR

D-Legion: A Scalable Many-Core Architecture for Accelerating Matrix Multiplication in Quantized LLMs

The performance gains obtained by large language models (LLMs) are closely linked to their substantial computational and memory requirements. Quantized LLMs offer significant advantages with extremely quantized models, motivating the development of specialized architectures to accelerate their workloads. This paper proposes D-Legion, a novel scalable many-core architecture, designed using many adaptive-precision systolic array cores, to accelerate matrix multiplication in quantized LLMs. The proposed architecture consists of a set of Legions where each Legion has a group of adaptive-precision systolic arrays. D-Legion supports multiple computation modes, including quantized sparse and dense matrix multiplications. The block structured sparsity is exploited within a fully-sparse, or partially-sparse windows. In addition, memory accesses of partial summations (psums) are spatially reduced through parallel accumulators. Furthermore, data reuse is maximized through optimized scheduling techniques by multicasting matrix tiles across the Legions. A comprehensive design space exploration is performed in terms of Legion/core granularity to determine the optimal Legion configuration. Moreover, D-Legion is evaluated on attention workloads from two BitNet models, delivering up to 8.2$\times$ lower latency, up to 3.8$\times$ higher memory savings, and up to 3$\times$ higher psum memory savings compared to state-of-the-art work. D-Legion, with eight Legions and 64 total cores, achieves a peak throughput of 135.68 TOPS at a frequency of 1 GHz. A scaled version of D-Legion, with 32 Legions, is compared to Google TPUv4i, achieving up to 2.5$\times$ lower total latency, up to 2.3$\times$ higher total throughput, and up to 2.7$\times$ higher total memory savings.

cs.AR

A flexible language model-assisted electronic design automation framework

Large language models (LLMs) are transforming electronic design automation (EDA) by enhancing design stages such as schematic design, simulation, netlist synthesis, and place-and-route. Existing methods primarily focus these optimisations within isolated open-source EDA tools and often lack the flexibility to handle multiple domains, such as analogue, digital, and radio-frequency design. In contrast, modern systems require to interface with commercial EDA environments, adhere to tool-specific operation rules, and incorporate feedback from design outcomes while supporting diverse design flows. We propose a versatile framework that uses LLMs to generate files compatible with commercial EDA tools and optimise designs using power-performance-area reports. This is accomplished by guiding the LLMs with tool constraints and feedback from design outputs to meet tool requirements and user specifications. Case studies on operational transconductance amplifiers, microstrip patch antennas, and FPGA circuits show that the framework is effective as an EDA-aware assistant, handling diverse design challenges reliably.

eess.SY

A Rapid-prototyping CMOS-RRAM Integration Strategy

Moore's law has long served the semiconductor industry as the driving force for producing ever-advancing electronics technologies. However, given the economic implications and technological challenges associated with the present semiconductor scaling constraints, a shift from a traditional more Moore approach to a beyond Moore paradigm is desirable for sustaining the current pace of innovation beyond the established development route. Resistive random-access memories (RRAM) are one such beyond Moore technology that offers many avenues for innovation, and when integrated with mature complementary metal oxide semiconductors (CMOS), can extend CMOS capabilities in a scalable and power-efficient manner, both in terms of memory and computation. Nevertheless, as emerging and established technologies fuse, existing semiconductor-optimised manufacturing faces significant challenges, while the methodologies and complexities of integration are often not highlighted in depth, or overlooked at the expense of demonstrating the application-specific integrated-technologies. In this article, we focus on the integration, and detail a cost-effective, rapid-prototyping, and technology agnostic CMOS-RRAM integration strategy that employs hybridised wafer-level and multi-reticle processing techniques, supported by a systematic increased complexity approach. Leveraging the fact that CMOS technologies can be readily realised by taking advantage of mature front-end-of-line fabrication processes offered by semiconductor foundries, we establish an in-house RRAM development program that allows to combine fundamental material and device-level knowledge with custom-designed CMOS electronics. This approach utilises fully CMOS-compatible and transferable processes, ultimately enabling a seamless transition from research and development to volume production.

cond-mat.mtrl-sci

A Multi-Channel Auditory Signal Encoder with Adaptive Resolution Using Volatile Memristors

We demonstrate and experimentally validate an end-to-end hybrid CMOS-memristor auditory encoder that realises adaptive-threshold, asynchronous delta-modulation (ADM)-based spike encoding by exploiting the inherent volatility of HfTiOx devices. A spike-triggered programming pulse rapidly raises the ADM threshold Delta (desensitisation); the device's volatility then passively lowers Delta when activity subsides (resensitisation), emphasising onsets while restoring sensitivity without static control energy. Our prototype couples an 8-channel 130 nm encoder IC to off-chip HfTiOx devices via a switch interface and an off-chip controller that monitors spike activity and issues programming events. An on-chip current-mirror transimpedance amplifier (TIA) converts device current into symmetric thresholds, enabling both sensitive and conservative encoding regimes. Evaluated with gammatone-filtered speech, the adaptive loop-at matched spike budget-sharpens onsets and preserves fine temporal detail that a fixed-Delta baseline misses; multi-channel spike cochleagrams show the same trend. Together, these results establish a practical hybrid CMOS-memristor pathway to onset-salient, spike-efficient neuromorphic audio front-ends and motivate low-power single-chip integration.

eess.AS

DISCA: A Digital In-memory Stochastic Computing Architecture Using A Compressed Bent-Pyramid Format

Nowadays, we are witnessing an Artificial Intelligence revolution that dominates the technology landscape in various application domains, such as healthcare, robotics, automotive, security, and defense. Massive-scale AI models, which mimic the human brain's functionality, typically feature millions and even billions of parameters through data-intensive matrix multiplication tasks. While conventional Von-Neumann architectures struggle with the memory wall and the end of Moore's Law, these AI applications are migrating rapidly towards the edge, such as in robotics and unmanned aerial vehicles for surveillance, thereby adding more constraints to the hardware budget of AI architectures at the edge. Although in-memory computing has been proposed as a promising solution for the memory wall, both analog and digital in-memory computing architectures suffer from substantial degradation of the proposed benefits due to various design limitations. We propose a new digital in-memory stochastic computing architecture, DISCA, utilizing a compressed version of the quasi-stochastic Bent-Pyramid data format. DISCA inherits the same computational simplicity of analog computing, while preserving the same scalability, productivity, and reliability of digital systems. Post-layout modeling results of DISCA show an energy efficiency of 3.59TOPS/W per bit at 500 MHz using a commercial 180 nm CMOS technology. Therefore, DISCA significantly improves the energy efficiency for matrix multiplication workloads by orders of magnitude if scaled and compared to its counterpart architectures.

cs.AR

Multibit Ferroelectric Memcapacitor for Non-volatile Analogue Memory and Reconfigurable Filtering

Tuneable capacitors are vital for adaptive and reconfigurable electronics, yet existing approaches require continuous bias or mechanical actuation. Here we demonstrate a voltage-programmable ferroelectric memcapacitor based on HfZrO that achieves more than eight stable, reprogrammable capacitance states (3-bit encoding) within a non-volatile window of 24~pF. The device switches at low voltages (3~V), with each state exhibiting long retention (10^5~s) and high endurance (10^6 cycles), ensuring reliable multi-level operation. At the nanoscale, multistate charge retention was directly visualised using atomic force microscopy, confirming the robustness of individual states beyond macroscopic measurements. As a proof of concept, the capacitor was integrated into a high-pass filter, where the programmed capacitive states shift the cutoff frequency over 5~kHz, establishing circuit-level viability. This work demonstrates the feasibility of CMOS-compatible, non-volatile, analogue memory based on ferroelectric HfZrO, paving the way for adaptive RF filters, reconfigurable analogue front-ends, and neuromorphic electronics.

cond-mat.mtrl-sci

ADiP: Adaptive-Precision Systolic Array for Matrix Multiplication Acceleration

Transformers are at the core of modern AI nowadays. They rely heavily on matrix multiplication and require efficient acceleration due to their substantial memory and computational requirements. Quantization plays a vital role in reducing memory usage, and can be exploited for computations by designing reconfigurable architectures that enhance matrix multiplication by dynamically adjusting the precision. This paper proposes ADiP, a novel adaptive-precision systolic array architecture designed for efficient matrix multiplication acceleration. The proposed architecture consists of $N$ $\times$ $N$ reconfigurable processing elements (PEs), along with shared shifters and accumulators. ADiP supports multiple computation modes, including symmetric single-matrix multiplication as well as asymmetric multi-matrix multiplication with a shared input matrix, thereby improving data reuse and PE utilization. By adapting to different precisions, ADiP achieves up to 4$\times$ higher throughput and up to 4$\times$ higher memory efficiency. Analytical models are developed for ADiP architecture, including latency and throughput for different architecture configurations. A comprehensive hardware design space exploration is demonstrated using commercial 22nm technology. Furthermore, ADiP is evaluated on different Transformer-based workloads from GPT-2 medium, BERT large, and BitNet-1.58B models, delivering total latency improvement up to 53.6%, and total energy improvement up to 24.4% for attention workloads in BitNet-1.58B model. At a 64$\times$64 size with reconfigurable 4,096 PEs, ADiP achieves a peak throughput of 8.192 TOPS, 16.384 TOPS, and 32.768 TOPS for 8bit$\times$8bit, 8bit$\times$4bit, and 8bit$\times$2bit operations, respectively.

cs.AR

Views: a hardware-friendly graph database model for storing semantic information

The graph database (GDB) is an increasingly common storage model for data involving relationships between entries. Beyond its widespread usage in database industries, the advantages of GDBs indicate a strong potential in constructing symbolic artificial intelligences (AIs) and retrieval-augmented generation (RAG), where knowledge of data inter-relationships takes a critical role in implementation. However, current GDB models are not optimised for hardware acceleration, leading to bottlenecks in storage capacity and computational efficiency. In this paper, we propose a hardware-friendly GDB model, called Views. We show its data structure and organisation tailored for efficient storage and retrieval of graph data and demonstrate its functional equivalence and storage performance advantage compared to represent traditional graph representations. We further demonstrate its symbolic processing abilities in semantic reasoning and cognitive modelling with practical examples and provide a short perspective on future developments.

cs.DB