Search arXivSearch

arXiv subjects

Yogesh Kumar

Publications and source records attributed to Yogesh Kumar.

At least 19 recordsLinked to original sources

Strictly Causal Streaming Video Anomaly Detection with a Theoretically-Grounded State-Space Core

Recent work has applied Mamba style state space models (SSMs) to video anomaly detection, yet existing approaches still rely on buffering clips or windows internally, lack a theoretical account of how temporal memory relates to detection latency, and benchmark efficiency only through GPU throughput rather than the edge hardware these methods are intended to target. We introduce a strictly causal streaming anomaly detector whose fixed size state is updated in O(1) time and memory per incoming frame, with no lookahead and no clip buffering. Its temporal core is a diagonal linear state space recurrence with an input and state dependent decay gate, trained self supervised through causal next embedding prediction on a frozen visual backbone. We derive a closed form relationship between the recurrence decay spectrum and both detection delay and the shortest anomaly it can reliably capture, then validate empirically on UCSD Ped2 and CUHK Avenue. The settling delay bound predicted from the learned base decay (57 to 59 frames) sits far above the measured detection delay (1.6 and 18.4 frames), showing that the event boundary gate, not the base decay, governs responsiveness. We further report end to end latency and throughput measured directly on Apple M3 Pro hardware, 0.74 ms and 0.77 ms per frame (over 1300 FPS), rather than simulated GPU numbers. With an untuned initial configuration the method reaches 67.9 percent and 70.2 percent frame level AUC on Ped2 and Avenue, trailing prior non causal SSM baselines in accuracy. Ablations over decay rate, state size, and gating reveal that the gate contribution is dataset size dependent, hurting accuracy on the smaller Ped2 training set but helping on the larger Avenue one. Closing this accuracy gap and extending evaluation to a third, larger benchmark are immediate next steps.

cs.AI

Catching Hallucinated Citations in Video-LLM Question Answering: A Self-Verification Pipeline and Verifier Ablation Study

Video question answering systems built on vision-language models often produce timestamped claims with high confidence even when unsupported by the cited frame. This deceptive hallucination arises because timestamps imply grounding without ensuring correctness, increasing user trust but not accuracy. We introduce a pipeline that closes this loop. A retrieval-augmented language model drafts answers with per-claim timestamp citations, and each cited frame is independently re-examined before being shown to the user. We compare against a plain baseline and ablate three verification designs, evaluated on both Apple Silicon (MLX) and Google Colab (HF Transformers, CUDA). Directly asking the vision model whether a frame supports a claim fails completely (0% catch rate on 40 claims) due to sycophancy. Blind re-captioning plus a general LLM judge improves results but is unstable, oscillating between 0% and 100% flagged depending on prompt phrasing. Replacing that judge with a small natural language inference model yields a stable, interpretable verifier that catches 79% of fabricated claims on adversarial false-premise questions while leaving true claims untouched. We release the full pipeline, evaluation harness, and implementations for both Apple Silicon and Colab. Code is available at https://github.com/yogesh-iitj/grounded-video-qa.

cs.CV

Revisiting In-Medium QCD Effects on Spin-Polarized Strange Quark Stars

We investigate the properties of exotic Strange Quark Matter with spin polarization and the complex configuration of Strange Quark Stars using a phenomenological MIT Bag Model, enhanced by incorporating a QCD informed running strange quark mass dependent on the chemical potential. The effective mass using quasiparticle approach is utilized to understand the framework of SQM. The resulting Equation of State is constructed for two distinct parameter sets, yielding an energy per baryon below the iron limit for stable configurations and thus supporting the Bodmer Witten Terazawa hypothesis that SQM could be the actual ground state of exotic matter. This framework is then used to address the Tolman Oppenheimer Volkoff equations to study the effect of spin polarization on stellar properties. The model predicts that the maximum stellar mass increases with the degree of spin polarization, a result that diverges from previous constant mass models. Furthermore, the model predictions for mass, radius, and surface redshift are in excellent agreement with observational constraints for the compact object Vela X one. This agreement validates our theoretical approach and strengthens the candidacy of Vela X one as a Strange Quark Star. Overall, the model results highlight the importance of in medium QCD effects in describing dense matter.

hep-ph

XGRVFL-MV: Residual-Coupled Graph-Embedded Multi-View Random Vector Functional Link Network with FleXi Guardian Loss

Random Vector Functional Link (RVFL) networks provide an efficient randomized learning framework for classification. Existing multi-view RVFL methods utilize complementary information from multiple views. However, preserving view-specific geometric structure, limiting the influence of large prediction residuals, and modeling relationships between multiple views remain challenging. This paper proposes a Residual-Coupled Graph-Embedded Multi-View RVFL model with fleXi guardian loss (XGRVFL-MV) for multi-view classification. The proposed model constructs RVFL representation for each view, incorporates graph embedding with intrinsic and penalty graphs constructed using the Local Fisher Discriminant Analysis weighting scheme. It also uses the bounded and asymmetric FleXi Guardian (XG) loss for residual learning. A residual-coupling term is introduced to encourage consistency among view-specific prediction residuals while preserving view-specific representations. The resulting optimization problem is solved using an inversion-free first-order optimization procedure based on Nesterov accelerated gradient descent. We evaluate the proposed model on UCI, KEEL, AwA, and Corel5k benchmark datasets. Experimental results, together with statistical analyses and hyperparameter sensitivity analyses, show that XGRVFL-MV achieves competitive classification performance compared with the baseline methods across the evaluated benchmark datasets.

cs.LG

Graph-Embedded Intuitionistic Fuzzy Broad Learning System: A Multi-view Framework

The Broad Learning System (BLS) has been widely used for data classification and is based on a layer-by-layer feed-forward structure. However, it gives the same importance to all data points, which reduces its effectiveness on real-world datasets with noise and outliers. In addition, it does not consider the geometric structure of the data and has limitations in handling data from multiple sources. To address these challenges, we propose a Multi-View Graph-Embedded Intuitionistic Fuzzy Broad Learning System (MVGIFBLS) that integrates multi-view learning, graph embedding, and intuitionistic fuzzy theory into the BLS framework. This design enables the model to combine information from multiple sources and learn more discriminative representations. Graph embedding captures the geometric relationships among samples and improves class separation through intrinsic and penalty subspaces based on local Fisher discriminant analysis. Intuitionistic fuzzy theory enhances robustness to noise, while kernel-based neighborhood analysis captures local data structures. We evaluate the proposed framework on several UCI, KEEL, and AwA benchmark datasets using comparative evaluation, Gaussian feature noise analysis, ablation studies, and statistical analysis. The results demonstrate that each component contributes positively to the overall framework and that the proposed MVGIFBLS consistently achieves higher Area Under the Curve (AUC) scores and maintains robust performance under Gaussian feature noise.

cs.LG

Electric field controlled spin transport in a topological insulator interfaced with a ferroelectric antiferromagnet

Topological insulators have been explored extensively for spin-charge interconversion via magnetic interfaces, yet the true response of their spin-charge conversion, particularly in the absence of an external magnetic field, remains to be studied. Here, we report electric-field control of spin-charge conversion in the topological insulator Bi$_2$Te$_3$ with the antiferromagnetic multiferroic BiFeO$_3$, employing a nonlocal spin transport device. A systematic thickness dependence of the spin transport across the interface between Bi$_2$Te$_3$ and BiFeO$_3$ reveals a signature of topological surface-state-dominated spin transport in the bilayer system. The spin-charge conversion remains robust for thicknesses above 10 nm but falls rapidly with reducing thickness and vanishes at 5 nm. This is consistent with the hybridization-induced emergence of a trivial insulating phase, which is supported by the coherency factor estimated from the magnetoconductance of Bi$_2$Te$_3$. These results establish that spin-momentum-locked surface states dominate interfacial spin transport in the decoupled regime. Beyond presenting efficient spin-charge interconversion at an entirely insulating magnetic interface, this work also highlights sputter-deposited Bi$_2$Te$_3$ as a high-quality and scalable platform for integrating quantum materials into devices. The nonlocal spin transport approach presented here provides a simple and direct evidence of spin-charge conversion and opens an efficient and practical pathway toward designing energy-efficient spin-based devices.

cond-mat.mes-hall

Intuitionistic Fuzzy Graph Embedded Random Vector Functional Link with Multiview Learning

Random Vector Functional Link (RVFL) networks are popular due to their fast training and universal approximation capabilities. However, RVFL models face challenges in preserving geometric relationships and utilizing multiple feature views effectively. To address these limitations we propose the Intuitionistic Fuzzy Graph Embedded Random Vector Functional Link with Multiview Learning (IFGRVFL-MV) model. The proposed approach comprises three key components: intuitionistic fuzzy sets for uncertainty handling, graph embedding to capture intrinsic geometric structures, and multiview learning to use complementary information from multiple feature spaces. The model assigns intuitionistic fuzzy membership and non-membership values to data points making it robust to outliers. Also, the graph embedding framework preserves topological structures, increasing the generalization performance. We performed experiments on benchmark datasets from UCI and KEEL repositories which concludes that IFGRVFL-MV outperforms existing models in classification accuracy. Our results establish that IFGRVFL-MV is a promising advancement in the domain of uncertainty and multiview environments.

cs.LG

Thermodynamical analysis of QGP using effective PNJL model with Quasiparticle approach

We study the thermodynamics of the quark-gluon plasma using an effective Two flavor Polyakov Nambu Jona Lasinio (PNJL) model extended by a quasiparticle description for quarks and gluons, incorporating temperature dependent quark masses within the PNJL framework. Two variants, Quasiparticle Model-I and Quasiparticle Model-II, are implemented to investigate bulk thermodynamic observables such as pressure, energy density, entropy density, specific heat, and the speed of sound. The combined framework yields a robust baseline for the description of hot QGP dynamics in the high temperature regime at vanishing chemical potential and zero magnetic field. Systematic comparison with lattice QCD results shows an excellent agreement and clear improvement over conventional PNJL implementations. We observe that both variants complement each other, offering mutually consistent insight into quasiparticle mass effects and medium response in the deconfined phase. This mutual consistency validates the physical foundation of the overall quasiparticle mechanism, reinforcing the credibility of the calculated Equation of State. Finally, the quasiparticle model extension improves PNJL from a descriptive tool to a more qualitative phenomenological approach, enabling an improved description of the strong interacting quark-gluon plasma.

hep-ph

Graph-Guided Universum Learning in Generalized Eigenvalue Proximal SVMs for Alzheimer's Disease Classification

Early and accurate detection of Alzheimer's disease (AD) is important for timely intervention and disease management. Generalized Eigenvalue Proximal Support Vector Machine (GEPSVM) and its Universum-based variants have shown promising results for AD classification. However, existing methods treat Universum samples as independent points and do not consider the geometric relationships among them. This paper proposes two graph-guided Universum learning models, namely UG-GEPSVM and IUG-GEPSVM, for AD versus cognitively normal (CN) classification using structural MRI data. In the proposed framework, mild cognitive impairment (MCI) subjects are used as Universum data to provide intermediate information between AD and CN classes. A graph is constructed over the Universum samples using Gaussian similarity, Minimum Spanning Tree connectivity, and multi-hop propagation. From this graph, a Laplacian matrix is derived that captures the geometric structure of the MCI samples. This Laplacian-based regularization is incorporated into the learning process in place of the conventional independent Universum penalty term. UG-GEPSVM integrates this regularization into the generalized eigenvalue formulation, while IUG-GEPSVM extends the numerically stable improved GEPSVM framework using a standard eigenvalue formulation. Experiments on ADNI MRI dataset variants using ICA- and PCA-based features at five different noise levels show that both proposed models consistently outperform existing GEPSVM and Universum-based methods. UG-GEPSVM achieves the highest average AUC of 88.07% and maintains stable performance under increasing noise levels. Statistical tests further confirm the significance of the observed improvements.

cs.LG

Analyzing Linear Layers in Related-Differential Cryptanalysis

In AES-like ciphers, diffusion layers are commonly instantiated using MDS matrices, since their optimal branch number yields strong diffusion guarantees and underpins classical resistance arguments against differential and linear cryptanalysis. However, Daemen and Rijmen (2009) showed that linear layers may still exhibit related-differential structure beyond what the MDS criterion captures, and Bardeh and Rijmen (2022) demonstrated that this phenomenon can be exploited in attacks on reduced-round AES. In this work, we systematically investigate the conditions under which linear layers avoid or exhibit these differentials, identifying matrix classes for which such structure is unavoidable. We first prove that every non-MDS matrix admits a nontrivial pair of related differentials, showing that the MDS property is necessary for avoiding them. We then establish that every odd-order symmetric MDS matrix admits related differentials, which rules out broad families of Cauchy-based constructions. We also substantially strengthen the circulant case by proving that related differentials are unavoidable for every circulant matrix of order $n$ with $n \not\equiv \pm 2 \pmod{12}$. Finally, we revisit the characterization of $3 \times 3$ MDS matrices over $\mathbb{F}_{2^m}$ for the absence of related differentials, and derive an explicit necessary and sufficient criterion in terms of $15$ polynomial constraints.

cs.CR

MINT: Multimodal Imaging-to-Speech Knowledge Transfer for Early Alzheimer's Screening

Alzheimer's disease is a progressive neurodegenerative disorder in which mild cognitive impairment (MCI) marks a critical transition between aging and dementia. Neuroimaging modalities, such as structural MRI, provide biomarkers of this transition; however, their high costs and infrastructure needs limit their deployment at a population scale. Speech analysis offers a non-invasive alternative, but speech-only classifiers are developed independently of neuroimaging, leaving decision boundaries biologically ungrounded and limiting reliability on the subtle CN-versus-MCI distinction. We propose MINT (Multimodal Imaging-to-Speech Knowledge Transfer), a three-stage cross-modal framework that transfers biomarker structure from MRI into a speech encoder at training time. An MRI teacher, trained on 1,228 subjects, defines a compact neuroimaging embedding space for CN-versus-MCI classification. A residual projection head aligns speech representations to this frozen imaging manifold via a combined geometric loss, adapting speech to the learned biomarker space while preserving imaging encoder fidelity. The frozen MRI classifier, which is never exposed to speech, is applied to aligned embeddings at inference and requires no scanner. Evaluation on ADNI-4 shows aligned speech achieves performance comparable to speech-only baselines (AUC 0.720 vs 0.711) while requiring no imaging at inference, demonstrating that MRI-derived decision boundaries can ground speech representations. Multimodal fusion improves over MRI alone (0.973 vs 0.958). Ablation studies identify dropout regularization and self-supervised pretraining as critical design decisions. To our knowledge, this is the first demonstration of MRI-to-speech knowledge transfer for early Alzheimer's screening, establishing a biologically grounded pathway for population-level cognitive triage without neuroimaging at inference.

cs.LG

Anisotropic magnon transport in an antiferromagnetic trilayer heterostructure: is BiFeO$_3$ an altermagnet?

Magnons provide a route to ultra-fast transport and non-destructive readout of spin-based information transfer. Here, we report magnon transport and its emergent anisotropic nature in BiFeO$_3$ layers confined between ultrathin layers of the antiferromagnet LaFeO$_3$. Due to the confined state, BiFeO$_3$ serves as an efficient magnon transmission channel as well as a magnetoelectric knob by which to control the stack by means of an electric field. We discuss the mechanism of the anisotropic spin transport based on the interaction between the antiferromagnetic order and the electric field. This allows us to manipulate and amplify the spin transport in such a confined geometry. Furthermore, lower crystal symmetric and suppression of the spin cycloid in ultrathin BiFeO$_3$ stabilizes a non-trivial antiferromagnetic state exhibiting symmetry-protected spin-split bands that provide the non-trivial sign inversion of the spin current, which is a characteristic of an altermagnet. This work provides an understanding of the anisotropic spin transport in complex antiferromagnetic heterostructures where ferroelectricity and altermagnetism coexist, paving the way for a new route to realize electric-field control of a novel state of magnetism.

cond-mat.mes-hall

A Unified Framework for EEG Seizure Detection Using Universum-Integrated Generalized Eigenvalues Proximal Support Vector Machine

The paper presents novel Universum-enhanced classifiers: the Universum Generalized Eigenvalue Proximal Support Vector Machine (U-GEPSVM) and the Improved U-GEPSVM (IU-GEPSVM) for EEG signal classification. Using the computational efficiency of generalized eigenvalue decomposition and the generalization benefits of Universum learning, the proposed models address critical challenges in EEG analysis: non-stationarity, low signal-to-noise ratio, and limited labeled data. U-GEPSVM extends the GEPSVM framework by incorporating Universum constraints through a ratio-based objective function, while IU-GEPSVM enhances stability through a weighted difference-based formulation that provides independent control over class separation and Universum alignment. The models are evaluated on the Bonn University EEG dataset across two binary classification tasks: (O vs S)-healthy (eyes closed) vs seizure, and (Z vs S)-healthy (eyes open) vs seizure. IU-GEPSVM achieves peak accuracies of 85% (O vs S) and 80% (Z vs S), with mean accuracies of 81.29% and 77.57% respectively, outperforming baseline methods.

cs.LG

Temporal Object-Aware Vision Transformer for Few-Shot Video Object Detection

Few-shot Video Object Detection (FSVOD) addresses the challenge of detecting novel objects in videos with limited labeled examples, overcoming the constraints of traditional detection methods that require extensive training data. This task presents key challenges, including maintaining temporal consistency across frames affected by occlusion and appearance variations, and achieving novel object generalization without relying on complex region proposals, which are often computationally expensive and require task-specific training. Our novel object-aware temporal modeling approach addresses these challenges by incorporating a filtering mechanism that selectively propagates high-confidence object features across frames. This enables efficient feature progression, reduces noise accumulation, and enhances detection accuracy in a few-shot setting. By utilizing few-shot trained detection and classification heads with focused feature propagation, we achieve robust temporal consistency without depending on explicit object tube proposals. Our approach achieves performance gains, with AP improvements of 3.7% (FSVOD-500), 5.3% (FSYTV-40), 4.3% (VidOR), and 4.5 (VidVRD) in the 5-shot setting. Further results demonstrate improvements in 1-shot, 3-shot, and 10-shot configurations. We make the code public at: https://github.com/yogesh-iitj/fs-video-vit

cs.CV

Thermal gradient-driven skyrmion dynamics with near-zero skyrmion Hall angle

Thermal gradient driven skyrmion dynamics offers a promising route toward green spintronics, enabling the utilization of waste heat for information transport and processing. Using micromagnetic simulations, we investigate Neel skyrmions in a Co-Pt bilayer nanoracetrack and demonstrate that stochastic torques induced by a thermal gradient drive skyrmion motion toward the hotter region with a nearly vanishing Hall angle. The dynamics depends sensitively on intrinsic material parameters - the skyrmion velocity decreases with increasing damping constant, increases with stronger thermal gradients, and varies systematically with saturation magnetization, interfacial DMI strength, and uniaxial out of plane anisotropy. Importantly, we identify a specific range of material parameters within which the skyrmion velocity changes sharply while the Hall angle remains strongly suppressed, saturating near zero. This comprehensive parameter-dependent study establishes a universal design framework for minimizing the Hall effect in thermal gradient driven spintronic systems.

cond-mat.mes-hall

New Insights into Involutory and Orthogonal MDS Matrices

MDS matrices play a critical role in the design of diffusion layers for block ciphers and hash functions due to their optimal branch number. Involutory and orthogonal MDS matrices offer additional benefits by allowing identical or nearly identical circuitry for both encryption and decryption, leading to equivalent implementation costs for both processes. These properties have been further generalized through the notions of semi-involutory and semi-orthogonal matrices. While much of the existing literature focuses on identifying efficiently implementable MDS candidates or proposing new constructions for MDS matrices of various orders, this work takes a different direction. Rather than introducing novel constructions or prioritizing implementation efficiency, we investigate structural relationships between the generalized variants and their conventional counterparts. Specifically, we establish nontrivial interconnections between semi-involutory and involutory matrices, as well as between semi-orthogonal and orthogonal matrices. Exploiting these relationships, we show that the number of semi-involutory MDS matrices can be directly derived from the number of involutory MDS matrices, and vice versa. A similar correspondence holds for semi-orthogonal and orthogonal MDS matrices. We also examine the intersection of these classes and show that the number of $3 \times 3$ MDS matrices that are both semi-involutory and semi-orthogonal coincides with the number of semi-involutory MDS matrices over $\mathbb{F}_{2^m}$. Furthermore, we derive the general structure of orthogonal matrices of arbitrary order $n$ over $\mathbb{F}_{2^m}$. Finally, leveraging the aforementioned interconnections, we present an alternative and direct derivation of the explicit formulae for counting $3 \times 3$ semi-involutory MDS matrices and $3 \times 3$ semi-orthogonal MDS matrices.

cs.CR

Language-Guided Temporal Token Pruning for Efficient VideoLLM Processing

Vision Language Models (VLMs) struggle with long-form videos due to the quadratic complexity of attention mechanisms. We propose Language-Guided Temporal Token Pruning (LGTTP), which leverages temporal cues from queries to adaptively prune video tokens, preserving contextual continuity while reducing computational overhead. Unlike uniform pruning or keyframe selection, LGTTP retains higher token density in temporally relevant segments. Our model-agnostic framework integrates with TimeChat and LLaVA-Video, achieving a 65% reduction in computation while preserving 97-99% of the original performance. On QVHighlights, LGTTP improves HIT@1 by +9.5%, and on Charades-STA, it retains 99.6% of R@1. It excels on queries with explicit temporal markers and remains effective across general video understanding tasks.

cs.CV

Aligning Moments in Time using Video Queries

Video-to-video moment retrieval (Vid2VidMR) is the task of localizing unseen events or moments in a target video using a query video. This task poses several challenges, such as the need for semantic frame-level alignment and modeling complex dependencies between query and target videos. To tackle this challenging problem, we introduce MATR (Moment Alignment TRansformer), a transformer-based model designed to capture semantic context as well as the temporal details necessary for precise moment localization. MATR conditions target video representations on query video features using dual-stage sequence alignment that encodes the required correlations and dependencies. These representations are then used to guide foreground/background classification and boundary prediction heads, enabling the model to accurately identify moments in the target video that semantically match with the query video. Additionally, to provide a strong task-specific initialization for MATR, we propose a self-supervised pre-training technique that involves training the model to localize random clips within videos. Extensive experiments demonstrate that MATR achieves notable performance improvements of 13.1% in R@1 and 8.1% in mIoU on an absolute scale compared to state-of-the-art methods on the popular ActivityNet-VRL dataset. Additionally, on our newly proposed dataset, SportsMoments, MATR shows a 14.7% gain in R@1 and a 14.4% gain in mIoU on an absolute scale over strong baselines.

cs.CV