Search arXivSearch

arXiv subjects

Akash Kumar

Publications and source records attributed to Akash Kumar.

At least 19 recordsLinked to original sources

Reconfigurable field-free spin Hall nano-oscillators enabled by crystallographic anisotropy in epitaxial Co/Pt

Spin Hall nano-oscillators (SHNOs) are nanoscale microwave sources for wireless communication, neuromorphic computing and oscillator-based Ising machines, but conventional devices require a global magnetic bias. Here we replace this bias through crystallographic anisotropy in epitaxial Co/Pt. Growth of hcp Co with its c-axis in the film plane produces an anisotropy field of about 0.36 T and enables field-free auto-oscillations above 10 GHz in nanoconstriction SHNOs. The active current polarity is selected by the remanent magnetization, providing nonvolatile reconfiguration of the oscillation state. Micro-focused Brillouin light scattering confirms that the nonlinear response is confined to the nanoconstriction region. Lithographic control of the angle between the current and anisotropy axes tunes the excitation threshold and drives two spectral branches from separated modes to a dominant single branch, consistent with mutual synchronization. These results establish epitaxial crystallographic anisotropy as a route to reconfigurable field-free spintronic oscillators and oscillator networks.

cond-mat.mes-hall

Spin-Orbital Hall Nano-Oscillators using PtCr/NiFe

The orbital Hall effect provides a promising route for generating angular-momentum currents beyond conventional spin Hall physics. PtCr alloys exhibit unusually large current-induced torques, but the contribution of orbital transport and the ability of these torques to sustain coherent nonlinear magnetization dynamics remain unresolved. Here we demonstrate spin-orbital Hall nano-oscillators by exploiting a homogeneous heavy-metal/light-metal alloy in which orbital Hall currents generated by Cr are converted by Pt into spin currents, producing giant spin-orbit torques. Using PtCr/NiFe heterostructures, the effective torque efficiency increases from ~0.14 in Pt/NiFe to ~0.40 in Pt0.38Cr0.62/NiFe despite substantial Pt dilution, enabling coherent auto-oscillations with the threshold current density reduced from ~ 1.07 x 10^12 to ~ 4.4 x 10^11 A m^-2. First-principles calculations show that Cr alloying suppresses the intrinsic spin Hall conductivity while enhancing the orbital Hall conductivity, and reproduce the observed torque enhancement only when orbital transport is included. Our combined experimental and first-principles results show that alloy engineering enables giant spin-orbit torques through an intrinsic orbital-mediated contribution, enabling coherent auto-oscillations without engineered multilayers and establishing a scalable materials platform for low-power nonlinear spintronic and orbitronic devices.

cond-mat.mes-hall

Phase noise analysis and control of VO$_2$-based relaxation type oscillators

VO$_2$-based relaxation oscillators form a rapidly developing field that finds applications in neuromorphic computing, Ising machines, and numerous signal processing concepts. These oscillators operate in a deeply nonlinear relaxation regime based on rapid phase transitions between insulating and metallic states in the VO$_2$ material. This process is governed by thermal effects, which lead to additional voltage fluctuations and contribute to a considerably wide spectral linewidth in the VO$_2$-based oscillator signal. In this work, we thoroughly study the phase noise in VO$_2$-based relaxation oscillators and demonstrate that the broadening of the generation spectrum linewidth at low oscillation frequencies is caused by an increased susceptibility to thermal fluctuations during the incubation phase. We explore the types of noise affecting oscillator stability and show that synchronization with an external square-wave signal improves the phase noise more effectively than a sinusoidal-shape injection locking signal.

physics.app-ph

Enhanced hydrogen response of copper-doped TiO$_2$ synthesised by helium-assisted magnetron sputtering

Cu-doped TiO$_2$ thin films for hydrogen sensing were synthesised by reactive DC magnetron sputtering in Ar/O$_2$/He mixtures, with the He fraction used as a control parameter for film growth. By combining normal-angle deposition (NAD) and glancing-angle deposition (GLAD) with post-deposition annealing, the effects of He on microstructure formation and sensor performance were examined. X-ray diffraction and electron microscopy revealed that He promotes nanostructuring, lattice expansion in as-deposited NAD films, increased porosity after annealing, and a stronger anatase character in the final oxide layers. These structural changes, which enhance the reactive surface area, lead to improved hydrogen sensing at 300\,$^\circ$C in 1~vol.\,\% H$_2$. The response of NAD films increased from 1.4 to 6.0 simply by replacing part of the argon with helium, whereas GLAD films showed only a modest increase. The observed nanostructuring is discussed in terms of a simulation-supported growth scenario involving energetic backscattered He, a reduced hammering effect, and cooling-related suppression of adatom mobility, which together favour the formation of a more open sensing layer. Helium-assisted sputtering represents a useful physical route for tailoring oxide thin films for gas-sensing applications.

cond-mat.mtrl-sci

$x$-Prediction Flow: Efficient Continuous Decoding for Masked Diffusion Language Models

Masked diffusion language models (MDLMs) generate text by iteratively unmasking tokens, but their standard decoder reduces each step to a binary action: a position is either committed to a single token or left fully masked, discarding rich predictive information rather than carrying it forward, and forcing premature, irrevocable commitments that lead to poor performance under a limited decoding budget. In this paper, we reinterpret mask prediction as a clean-state prediction ($x$-prediction) and show that it can be used to induce a continuous flow in the input embedding space. Building on this view, we propose a continuous decoding framework for MDLMs where tokens can accumulate partial progress at each diffusion step and remain revisable. To match the uneven contextual constraints across positions in language, we replace the globally synchronous schedule in image diffusion with a confidence-based asynchronous update in which the diffusion progress is token-wise accumulated. Additionally, we introduce a lightweight policy network and formulate its training as a reinforcement learning problem. Applied to pretrained LLaDA, our decoder retains 83--97% of full-budget accuracy using under 15% of the diffusion steps, largely outperforming discrete mask-prediction decoding at matched budgets.

cs.CL

Subarray based Wideband Beamforming and Variational Sparse CSI Estimation for Low-Resolution MU THz MIMO Systems

This work conceives a unified channel estimation and beamforming framework, formulated within the principles of variational Bayesian inference. Recognizing the limitations imposed by hardware constraints, frequency-dependent propagation effects, and the structural restrictions of partially connected architectures in the Terahertz (THz) band, we formulate a dual-wideband channel model incorporating root raised cosine (RRC) pulse shape to account its band-limited nature. To further address the nonlinear distortions introduced by low-resolution ADCs, Bussgang decomposition is employed, enabling a tractable linearized inference process. Unlike conventional techniques, the proposed method accommodates both on-grid and off-grid angular domains, capturing spatial sparsity with improved resolution and robustness. The multi-user (MU) Bayesian Cram\'er-Rao lower bound is also derived to benchmark the performance of the proposed estimator. Moreover, the framework incorporates a true time delay (TTD)-based hybrid transceiver design that inherently compensates for the beam-squint effect; a frequency-dependent angular deviation that arises due to the fixedphase nature of the conventional beamformer in wideband systems, thereby ensuring accurate directional alignment across all subcarriers. Extensive simulation results validate the effectiveness of the proposed variational Bayesian inference-based estimator and the TTD-enabled beamforming architecture, highlighting their robustness and performance gains under practical wideband THz system.

eess.SP

Reducing the Randomness in Partition Oracles for Bounded Degree Minor-Free Graphs

Consider a bounded-degree graph $G$ that belongs to a minor-closed family (such as planar graphs). Such a graph has a hyperfinite decomposition, wherein, for a sufficiently small $\varepsilon > 0$, one can remove $\varepsilon dn$ edges to obtain connected components of size independent of $n$. (As usual, $n$ is the number of vertices and $d$ is the degree bound.) In a seminal result, Hassidim-Kelner-Nguyen-Onak (FOCS 2009) introduced the partition oracle, a procedure that provides local access to a hyperfinite decomposition. The partition oracle computes the component containing an input vertex $v$ with query complexity (to $G$) independent of $n$. Remarkably, this is done without any preprocessing on $G$. The coordination is done purely through a shared random seed. Despite a line of work on optimizing the query complexity of partition oracles, there were no attempts to bound the size of the random seed. All existing partition oracles use a random seed of size $\Omega(n)$, which technically implies a linear setup time. Any blackbox derandomization would likely need $\Omega(\log^2n)$ uniform random bits. A natural question is whether the random seed can also have length independent of $n$. We prove the $poly(d\varepsilon^{-1})$-query partition oracles of Kumar-Seshadhri-Stolman can be implemented with a random seed of $poly(d\varepsilon^{-1}) \cdot \log n$ length. To get a deeper understanding on the randomness complexity, we consider a more general model where the vertex labels come from the universe $[N]$, where $N \geq n$. In this setting, we prove that any partition oracle even for cycles requires $\omega_N(1)$ random bits.

cs.DS

Topologically Driven Giant Effective Spin Mixing Conductance in Antiferromagnetic FeSn/Py Heterostructures

The topological semimetal FeSn antiferromagnet, characterized by its kagome lattice, two-dimensional flat bands, and Dirac-like surface states, holds immense promise for spintronic applications. In this work, for the first time, we investigate the spin pumping behavior in epitaxial-FeSn/Py (Ni$_{80}$Fe$_{20}$) heterostructures. We report a giant effective spin mixing conductance (g$^{\uparrow \downarrow}_{\mathrm{eff}}$) of $(116\pm 7)$~nm$^{-2}$, which is nearly one order of magnitude higher than that of standard Pt/Py heterostructures. The insertion of a 3 nm Al spacer layer results in a two-fold reduction in the effective damping, confirming the interfacial origin of the large g$^{\uparrow\downarrow}_{\mathrm{eff}}$. Consistently, we observe an order-of-magnitude higher inverse spin Hall effect voltage in the FeSn/Py system compared to a reference Pt/Py film stack. We attribute the giant g$^{\uparrow\downarrow}_{\mathrm{eff}}$ to the direct interfacing of the Py layer with the topologically active [001]-kagome surface of epitaxial-FeSn. These findings establish the critical role of topologically active interfaces for advanced quantum-material-based spintronic devices.

cond-mat.mes-hall

VISTA: Video Interaction Spatio-Temporal Analysis Benchmark

Existing benchmarks for Vision-Language Models (VLMs) primarily evaluate spatio-temporal understanding on simple single-action videos, closed attribute sets and restricted entity types, failing to capture the freeform, multi-action interactions between diverse entities which characterize real-world video understanding. Furthermore, the lack of a systematic framework for analyzing model failures across complementary spatio-temporal axes hinders comprehensive evaluation. To address these gaps, we introduce VISTA, a Video Interaction Spatio-Temporal Analysis benchmark designed for open-set, multi-entity and multi-action spatio-temporal understanding in VLMs. VISTA decomposes videos into interpretable entities, their associated actions, and relational dynamics, enabling multi-axis diagnostics and unified assessment of relational, spatial, and temporal understanding. Our benchmark integrates multiple datasets into a single interaction-aware taxonomy and comprises ~12K curated video-query pairs spanning diverse scenes and complexities. We systematically evaluate 11 state-of-the-art VLMs on VISTA, and break down aggregate performance across our taxonomy to reveal shortcomings and pronounced spatio-temporal biases obscured by traditional metrics. By providing detailed, taxonomy-driven diagnostics on a challenging dataset, VISTA offers a nuanced framework to guide advances in model design, pretraining strategies, and evaluation protocols. Overall, VISTA is the first, large-scale, interaction-aware diagnostic benchmark for spatio-temporal understanding in VLMs.

cs.CV

Autonomous Systems Dependability in the era of AI: Design Challenges in Safety, Security, Reliability and Certification

The design of embedded safety-critical systems such as those used in next-generation automotive and autonomous platforms, is increasingly challenged by escalating system complexity, hardware-software heterogeneity, and the integration of intelligent, data-driven components. Ensuring dependability in such systems requires a holistic approach that spans multiple abstraction layers and encompasses both design- and run-time assurance. Traditional methods for reliability, safety, and security management often fall short in addressing the dynamic and uncertain behaviors introduced by Artificial Intelligence (AI) and Machine Learning (ML) components, especially under stringent real-time, power, and safety constraints. While AI and ML offer powerful predictive, adaptive, and self-optimizing capabilities that can enhance system dependability, their inherent non-determinism, data-dependence, and lack of formal guarantees introduce new challenges for verification, validation, and certification. This paper explores emerging methodologies, architectures, and frameworks for designing dependable autonomous and embedded systems in the era of AI. It highlight advances in reliability modeling, secure system design, and certification approaches that account for imperfect, learning-enabled components, aiming to bridge the gap between AI innovation and certifiable system-level dependability.

cs.AI

AnTi-MiCS: Analytical Framework for Bounding Time in Embedded Mixed-Criticality Systems

In Mixed-Criticality (MC) systems, although the high Worst-Case Execution Time (WCET) serves as a conservative upper bound representing the task's maximum execution time under all conditions, obtaining a low WCET is essential for representing realistic executions and improving utilization and Quality-of-Service (QoS). Nevertheless, determining appropriate low WCET(s) for lower-criticality (LO) modes poses a significant challenge. Opting for a very low value of this WCET enhances processor utilization by scheduling more tasks in LO mode. Conversely, employing a larger WCET ensures fewer mode switches, thereby enhancing QoS, albeit at the cost of processor utilization. This paper proposes an analytical approach, AnTi-MiCS, to determine the appropriate low WCET through design-time analysis of task executions. In some cases, a single low WCET may not be adequate to capture large variations in the execution time distribution, for example, in scenarios like bimodal distributions. Therefore, we further propose a scalable approach, MulTi-MiCS, to compute multiple appropriate low WCETs. This approach exploits the temporal correlation between subsequent inputs presented to the application. Experimental results, conducted on a real platform with embedded real-time benchmarks, demonstrate the efficacy of our proposed scheme, in which QoS is improved by 30.27% on average while reducing utilization waste by 35.89%, compared to existing approaches. Besides, MulTi-MiCS improves QoS by 6.41% compared to AnTi-MiCS while reducing utilization waste by 8.23%.

cs.DC

Machine Learning-Based Characterization of Solar p-Mode Frequency Shifts during Solar Cycle 25

The solar interior is probed by the properties of the Sun's acoustic oscillations (p-modes) observed on the solar surface. The frequencies of these p-modes measured in the last three decades show long term variation similar to the 11 year cyclic behaviour exhibited by 10.7 cm radio flux, sunspot numbers and other solar activity indices. It is also now established that the cyclic behavior of some of the solar proxies are connected with geomagnetic activities and have implications for space weather. Hence, in recent years efforts have been made using machine-learning methods to forecast these solar proxies with a view to improve our understanding of space weather. Developing a comparable method for forecasting p-mode frequency shifts is therefore of interest for two reasons. Firstly, it will facilitate future investigations into its potential role in tracing energy drivers from the Sun's interior to the geospace response by improving models of solar interior dynamics to coronal and heliospheric plasma conditions. In other words, it will help establish a more robust and quantitative link between the Sun's interior and its exterior. Secondly, it may provide us with an independent indicator or an early indicator of ascending and descending phase of solar activity which might be useful for space weather forecasting. In this article, we develop and apply the standard time-series analysis and machine-learning based methods to characterise p-mode frequency shifts for the remaining solar cycle 25.

astro-ph.SR

Terahertz Beamforming and Group Sparse Channel Estimation Relying on Low-Resolution ADCs in MU Hybrid MIMO systems

A unified beamforming and channel estimation framework relying on Bayesian learning is conceived. Recognizing the limitations imposed by low-resolution analog-to-digital converter (ADCs) and frequency-dependent propagation effects occurring in the Terahertz (THz) band, we formulate a dual-wideband channel model incorporating root raised cosine (RRC) pulse shaping. To address the non-linear distortions introduced by low-resolution ADCs, Bussgang decomposition is employed, leading to a tractable linearized inference process. By leveraging the shared sparsity inherent in a multi-user (MU) scenario of THz systems, we propose a Hierarchical Bayesian Group-sparse Regression (HBG-SR) based channel learning technique that exploits the group-sparse structure of THz band channels. The estimated dominant angle-of-arrival/ angle-of-departure (AoA/AoD) indices are then exploited for appropriately configuring the true-time-delay (TTD) elements in the hybrid transceiver, enabling precise beam alignment across subcarriers and the effective compensation of the beam-squint effect occurring in wideband THz systems. Extensive simulation results validate the efficiency of the proposed channel estimator and the TTD-aided beamforming architecture, highlighting their robustness and performance gains under practical wideband THz system constraints.

eess.SP

Two-Stage Hybrid Transceiver Design Relying on Low-Resolution ADCs in Partially Connected MU Terahertz (THz) MIMO Systems

A two-stage hybrid transceiver is designed by considering a partially connected architecture at the base station (BS) for a low-resolution multi-user (MU) THz massive multiple input multiple output (MIMO) system. Due to its high bandwidth coupled with a high number of antennas, the THz band suffers from the deleterious spatial-wideband and frequency-wideband effects jointly termed as the dual-wideband effect. To address this undesired phenomenon, we rigorously model the THz MIMO channel at each subarray corresponding to each user by incorporating the absorption, reflection, and free-space losses. Subsequently, a novel beamforming technique is proposed that employs only a few true time delay (TTD) lines for eliminating the beam-split effect, which is the manifestation of the spatial-wideband effect in the frequency domain. Our simulation results demonstrate a performance improvement of around 13% in terms of spectral efficiency over the existing state-of-the-art techniques.

eess.SP

MCEL: Margin-Based Cross-Entropy Loss for Error-Tolerant Quantized Neural Networks

Robustness to bit errors is a key requirement for the reliable use of neural networks (NNs) on emerging approximate computing platforms and error-prone memory technologies. A common approach to achieve bit error tolerance in NNs is injecting bit flips during training according to a predefined error model. While effective in certain scenarios, training-time bit flip injection introduces substantial computational overhead, often degrades inference accuracy at high error rates, and scales poorly for larger NN architectures. These limitations make error injection an increasingly impractical solution for ensuring robustness on future approximate computing platforms and error-prone memory technologies. In this work, we investigate the mechanisms that enable NNs to tolerate bit errors without relying on error-aware training. We establish a direct connection between bit error tolerance and classification margins at the output layer. Building on this insight, we propose a novel loss function, the Margin Cross-Entropy Loss (MCEL), which explicitly promotes logit-level margin separation while preserving the favorable optimization properties of the standard cross-entropy loss. Furthermore, MCEL introduces an interpretable margin parameter that allows robustness to be tuned in a principled manner. Extensive experimental evaluations across multiple datasets of varying complexity, diverse NN architectures, and a range of quantization schemes demonstrate that MCEL substantially improves bit error tolerance, up to 15 % in accuracy for an error rate of 1 %. Our proposed MCEL method is simple to implement, efficient, and can be integrated as a drop-in replacement for standard CEL. It provides a scalable and principled alternative to training-time bit flip injection, offering new insights into the origins of NN robustness and enabling more efficient deployment on approximate computing and memory systems.

cs.LG

BiKA: Kolmogorov-Arnold-Network-inspired Ultra Lightweight Neural Network Hardware Accelerator

Lightweight neural network accelerators are essential for edge devices with limited resources and power constraints. While quantization and binarization can efficiently reduce hardware cost, they still rely on the conventional Artificial Neural Network (ANN) computation pattern. The recently proposed Kolmogorov-Arnold Network (KAN) presents a novel network paradigm built on learnable nonlinear functions. However, it is computationally expensive for hardware deployment. Inspired by KAN, we propose BiKA, a multiply-free architecture that replaces nonlinear functions with binary, learnable thresholds, introducing an extremely lightweight computational pattern that requires only comparators and accumulators. Our FPGA prototype on Ultra96-V2 shows that BiKA reduces hardware resource usage by 27.73% and 51.54% compared with binarized and quantized neural network systolic array accelerators, while maintaining competitive accuracy. BiKA provides a promising direction for hardware-friendly neural network design on edge devices.

cs.AR

Bitwise Systolic Array Architecture for Runtime-Reconfigurable Multi-precision Quantized Multiplication on Hardware Accelerators

Neural network accelerators have been widely applied to edge devices for complex tasks like object tracking, image recognition, etc. Previous works have explored the quantization technologies in related lightweight accelerator designs to reduce hardware resource consumption. However, low precision leads to high accuracy loss in inference. Therefore, mixed-precision quantization becomes an alternative solution by applying different precision in different layers to trade off resource consumption and accuracy. Because regular designs for multiplication on hardware cannot support the precision reconfiguration for a multi-precision Quantized Neural Network (QNN) model in runtime, we propose a runtime reconfigurable multi-precision multi-channel bitwise systolic array design for QNN accelerators. We have implemented and evaluated our work on the Ultra96 FPGA platform. Results show that our work can achieve 1.3185 to 3.5671 times speedup in inferring mixed-precision models and has less critical path delay, supporting a higher clock frequency (250MHz).

cs.AR

GRAU: Generic Reconfigurable Activation Unit Design for Neural Network Hardware Accelerators

With the continuous growth of neural network scales, low-precision quantization is widely used in edge accelerators. Classic multi-threshold activation hardware requires 2^n thresholds for $n$-bit outputs, causing a rapid increase in hardware cost as precision increases. We propose a reconfigurable activation hardware, GRAU, based on piecewise linear fitting, where the segment slopes are approximated by powers of two. Our design requires only basic comparators and 1-bit right shifters, supporting mixed-precision quantization and nonlinear functions such as SiLU. Compared with multi-threshold activators, GRAU reduces LUT consumption by over 90%, achieving higher hardware efficiency, flexibility, and scalability. The best trade-off is usually achieved with 6-8 segments, while complex nonlinearities under aggressive low-cost settings may suffer larger accuracy degradation.

cs.AR