Search arXivSearch

arXiv subjects

Neeraj Kumar

Publications and source records attributed to Neeraj Kumar.

At least 19 recordsLinked to original sources

Noncommutative black holes: Topological bulk-boundary correspondence and Binary Merger Bounds

We investigate the thermodynamic topology of charged AdS black holes in a non-commutative spacetime sourced by Lorentzian-smeared matter distributions. Since exact analytical solutions for the critical thermodynamic quantities are not available, we employ a perturbative expansion in the non-commutative parameter and validate the resulting expressions through numerical analysis. Using the generalized off-shell free-energy framework, we explore the topological structure of the thermodynamic phase space and evaluate the corresponding winding number that characterizes the phase transitions. Our results reveal that non-commutative effects introduce qualitative modifications to the thermodynamic behavior compared with the standard Reissner-Nordstr\"om AdS black hole. Furthermore, we demonstrate that the bulk and boundary descriptions possess an identical global thermodynamic topology, providing strong evidence for the correspondence between their topological structures. We also investigate the lower bound on the remnant mass implied by the second law of black-hole thermodynamics and observe that non-commutative corrections modify key thermodynamic quantities, with particular emphasis on the entropy and the final black-hole mass.

gr-qc

Prompt Codebooks: Discrete Compositional Optimization for Language Model Instruction Refinement

Automatic prompt optimization (APO) has driven significant gains in LLM-based agentic workflows. However, most existing methods treat each task's prompt as a monolithic, instance-blind string optimized through global edits, producing brittle updates and preventing the reuse of learned sub-behaviors. We propose Prompt Codebook Optimization (PCO), a novel compositional prompt optimization framework that recasts APO as discrete learning over a finite vocabulary of natural-language instincts--atomic, reusable instruction units. PCO organizes prompt-construction knowledge in a discrete codebook and routes each input to a small subset of entries via an LLM-based encoder; a generator composes them into a prompt for the executor; a critic emits a structured verdict that decomposes by attribution into per-variable textual gradients, jointly training the encoder, generator, critic, and codebook under a language-valued min-max objective. The resulting routing is per-instance: different inputs in the same task receive different instinct compositions. Across six benchmarks, PCO improves aggregate performance over zero-shot by +13.50 points on Qwen3-8B and +11.80 points on LLaMA-3.1-8B. Crucially, PCO surpasses GEPA on HotpotQA by +6.34 points (Qwen3-8B) and +4.27 points (LLaMA-3.1-8B), while simultaneously reducing deployed prompt length by up to 14.1$\times$ vs. MIPROv2 and 3.0$\times$ vs. GEPA.

cs.AI

Free Resolutions of Symmetric Algebras of Ideals with Deviation Two

For a graded ideal I in a graded ring, the deviation of I is defined as the difference between the minimal number of generators of I and its grade. In this article, we provide bigraded free resolutions of the symmetric algebras for specific classes of ideals of deviation two. Additionally, we study the regularity of powers of deviation two ideals generated by d-sequences. In particular, we examine the powers of Huneke-Ulrich ideals and find bounds on their regularity.

math.AC

Revisiting Thermodynamics of the Hayward Black Holes and Exploring Binary Merger Bounds

In this article, we revisit the thermodynamics of Hayward black holes [1] in asymptotic flat spacetime and obtain the bounds on the final mass post merger after head-on collision event of two equal mass black holes. We revisit thermal properties of these black holes from a perspective that the laws of black hole thermodynamics remain valid. Under this condition a novel entropy formula appears naturally with a logarithmic correction term along with one extra term. We discuss phase structure of the black holes. Next, we obtain the bounds on the final black hole mass parameter using the validity of the second law of black hole thermodynamics with the new entropy formula. We discuss the impact of the Hayward parameter on thermal and merger properties of these black holes.

gr-qc

GOLDMARK: Governed Outcome-Linked Diagnostic Model Assessment Reference Kit

Computational biomarkers (CBs) are histopathology-derived patterns extracted from hematoxylin-eosin (H&E) whole-slide images (WSIs) using artificial intelligence (AI) to predict therapeutic response or prognosis. Recently, slide-level multiple-instance learning (MIL) with pathology foundation models (PFMs) has become the standard baseline for CB development. While these methods have improved predictive performance, computational pathology lacks standardized intermediate data formats, provenance tracking, checkpointing conventions, and reproducible evaluation metrics required for clinical-grade deployment. We introduce GOLDMARK (https://artificialintelligencepathology.org), a standardized benchmarking framework built on a curated TCGA cohort with clinically actionable OncoKB level 1-3 biomarker labels. GOLDMARK releases structured intermediate representations, including tile coordinate maps, per-slide feature embeddings from canonical PFMs, quality-control metadata, predefined patient-level splits, trained slide-level models, and evaluation outputs. Models are trained on TCGA and evaluated on an independent MSKCC cohort with reciprocal testing. Across 33 tumor-biomarker tasks, mean AUROC was 0.689 (TCGA) and 0.630 (MSKCC). Restricting to the eight highest-performing tasks yielded mean AUROCs of 0.831 and 0.801, respectively. These tasks correspond to established morphologic-genomic associations (e.g., LGG IDH1, COAD MSI/BRAF, THCA BRAF/NRAS, BLCA FGFR3, UCEC PTEN) and showed the most stable cross-site performance. Differences between canonical encoders were modest relative to task-specific variability. GOLDMARK establishes a shared experimental substrate for computational pathology, enabling reproducible benchmarking and direct comparison of methods across datasets and models.

cs.CV

Algorithms for numerical semigroups with fixed maximum primitive

We present an algorithm to explore various properties of the numerical semigroups with a given maximum primitive. In particular, we count the number of such numerical semigroups and verify that there is no counterexample to Wilf's conjecture among the numerical semigroups with maximum primitive up to \(60\).

math.CO

Simpson-Visser-AdS Black Holes: Thermodynamics and Binary Merger

In this article, we performed Simpson-Visser (SV)-regularization scheme to Anti-de Sitter (AdS) black holes and then studied thermal properties of the resulting spacetime geometry. We considered the validity of the first law of black hole thermodynamics in this case and derived an entropy formula consistent with this new regular geometry. Next, we carried out the free energy analysis and studied the phase structure of these black holes. We discovered non-trivial phase transition properties dependent on the SV-regularization parameter. We also considered the validity of the second law of black hole thermodynamics and analyzed a merger scenario of two equal mass SV-regular black holes. In particular, we investigated the impact of the SV-regularization parameter on the constraints on post-merger black hole mass. Intrestingly, we found that the bounds initially increase and then fall sharply with increasing the SV-regularization parameter. All results are compared with standard black holes for vanishing SV-regularization parameter.

gr-qc

Clustering with Set Outliers and Applications in Relational Clustering

We introduce and study the $k$-center clustering problem with set outliers, a natural and practical generalization of the classical $k$-center clustering with outliers. Instead of removing individual data points, our model allows discarding up to $z$ subsets from a given family of candidate outlier sets $\mathcal{H}$. Given a metric space $(P,\mathsf{dist})$, where $P$ is a set of elements and $\mathsf{dist}$ a distance metric, a family of sets $\mathcal{H}\subseteq 2^P$, and parameters $k, z$, the goal is to compute a set of $k$ centers $S\subseteq P$ and a family of $z$ sets $H\subseteq \mathcal{H}$ to minimize $\max_{p\in P\setminus(\bigcup_{h\in H} h)} \min_{s\in S}\mathsf{dist}(p,s)$. This abstraction captures structured noise common in database applications, such as faulty data sources or corrupted records in data integration and sensor systems. We present the first approximation algorithms for this problem in both general and geometric settings. Our methods provide tri-criteria approximations: selecting up to $2k$ centers and $2f z$ outlier sets (where $f$ is the maximum number of sets that a point belongs to), while achieving $O(1)$-approximation in clustering cost. In geometric settings, we leverage range and BBD trees to achieve near-linear time algorithms. In many real applications $f=1$. In this case we further improve the running time of our algorithms by constructing small \emph{coresets}. We also provide a hardness result for the general problem showing that it is unlikely to get any sublinear approximation on the clustering cost selecting less than $f\cdot z$ outlier sets. We demonstrate that this model naturally captures relational clustering with outliers: outliers are input tuples whose removal affects the join output. We provide approximation algorithms for both, establishing a tight connection between robust clustering and relational query evaluation.

cs.DS

R\'enyi Law Constraints on Gau{\ss}-Bonnet Black Hole Merger

In this article, we explore the R\'enyi law constraints on black hole merger in Gau{\ss}-Bonnet (GB) gravity. Specifically, we consider the case of static solutions in five-dimensional (5D) Anti-de-Sitter (AdS) spacetime and study the constraints on merger of two equal mass black holes. We calculate the general R\'enyi entropy expression and utilize it to study the bounds on the final black hole mass post-merger. We study its variation with the R\'enyi parameter. We also compare the results with those for black holes in General Relativity (GR). We find that the GB term has a significant impact on the bounds for black hole merger. The bounds for GB gravity become weaker for the zeroth order R\'enyi entropy and stronger for higher order R\'enyi entropies in comparison to GR.

gr-qc

The Role of Quantum Computing in Advancing Scientific High-Performance Computing: A perspective from the ADAC Institute

Quantum computing (QC) has gained significant attention over the past two decades due to its potential for speeding up classically demanding tasks. This transition from an academic focus to a thriving commercial sector is reflected in substantial global investments. While advancements in qubit counts and functionalities continues at a rapid pace, current quantum systems still lack the scalability for practical applications, facing challenges such as too high error rates and limited coherence times. This perspective paper examines the relationship between QC and high-performance computing (HPC), highlighting their complementary roles in enhancing computational efficiency. It is widely acknowledged that even fully error-corrected QCs will not be suited for all computational task. Rather, future compute infrastructures are anticipated to employ quantum acceleration within hybrid systems that integrate HPC and QC. While QCs can enhance classical computing, traditional HPC remains essential for maximizing quantum acceleration. This integration is a priority for supercomputing centers and companies, sparking innovation to address the challenges of merging these technologies. The Accelerated Data Analytics and Computing Institute (ADAC) is comprised of globally leading HPC centers. ADAC has established a Quantum Computing Working Group to promote and catalyze collaboration among its members. This paper synthesizes insights from the QC Working Group, supplemented by findings from a member survey detailing ongoing projects and strategic directions. By outlining the current landscape and challenges of QC integration into HPC ecosystems, this work aims to provide HPC specialists with a deeper understanding of QC and its future implications for computationally intensive endeavors.

quant-ph

Single GPU Task Adaptation of Pathology Foundation Models for Whole Slide Image Analysis

Pathology foundation models (PFMs) have emerged as powerful tools for analyzing whole slide images (WSIs). However, adapting these pretrained PFMs for specific clinical tasks presents considerable challenges, primarily due to the availability of only weak (WSI-level) labels for gigapixel images, necessitating multiple instance learning (MIL) paradigm for effective WSI analysis. This paper proposes a novel approach for single-GPU \textbf{T}ask \textbf{A}daptation of \textbf{PFM}s (TAPFM) that uses vision transformer (\vit) attention for MIL aggregation while optimizing both for feature representations and attention weights. The proposed approach maintains separate computational graphs for MIL aggregator and the PFM to create stable training dynamics that align with downstream task objectives during end-to-end adaptation. Evaluated on mutation prediction tasks for bladder cancer and lung adenocarcinoma across institutional and TCGA cohorts, TAPFM consistently outperforms conventional approaches, with H-Optimus-0 (TAPFM) outperforming the benchmarks. TAPFM effectively handles multi-label classification of actionable mutations as well. Thus, TAPFM makes adaptation of powerful pre-trained PFMs practical on standard hardware for various clinical applications.

cs.CV

On counting numerical semigroups by maximum primitive and Wilf's conjecture

We introduce a new way of counting numerical semigroups, namely by their maximum primitive, and show its relation with the counting of numerical semigroups by their Frobenius number. We show that these two ways of counting are M\"obius transforms of one another. We also establish that almost all numerical semigroups with large enough maximum primitive satisfy Wilf's conjecture. A crucial step in the proof is a result of independent interest: a numerical semigroup $S$ with multiplicity $\mathrm{m}$ such that $|S\cap (\mathrm{m},2 \mathrm{m})|\geq \sqrt{2\mathrm{m}}$ satisfies Wilf's conjecture.

math.CO

Comparative Analysis of Deep Learning Approaches for Harmful Brain Activity Detection Using EEG

The classification of harmful brain activities, such as seizures and periodic discharges, play a vital role in neurocritical care, enabling timely diagnosis and intervention. Electroencephalography (EEG) provides a non-invasive method for monitoring brain activity, but the manual interpretation of EEG signals are time-consuming and rely heavily on expert judgment. This study presents a comparative analysis of deep learning architectures, including Convolutional Neural Networks (CNNs), Vision Transformers (ViTs), and EEGNet, applied to the classification of harmful brain activities using both raw EEG data and time-frequency representations generated through Continuous Wavelet Transform (CWT). We evaluate the performance of these models use multimodal data representations, including high-resolution spectrograms and waveform data, and introduce a multi-stage training strategy to improve model robustness. Our results show that training strategies, data preprocessing, and augmentation techniques are as critical to model success as architecture choice, with multi-stage TinyViT and EfficientNet demonstrating superior performance. The findings underscore the importance of robust training regimes in achieving accurate and efficient EEG classification, providing valuable insights for deploying AI models in clinical practice.

cs.LG

Electrical Load Forecasting in Smart Grid: A Personalized Federated Learning Approach

Electric load forecasting is essential for power management and stability in smart grids. This is mainly achieved via advanced metering infrastructure, where smart meters (SMs) are used to record household energy consumption. Traditional machine learning (ML) methods are often employed for load forecasting but require data sharing which raises data privacy concerns. Federated learning (FL) can address this issue by running distributed ML models at local SMs without data exchange. However, current FL-based approaches struggle to achieve efficient load forecasting due to imbalanced data distribution across heterogeneous SMs. This paper presents a novel personalized federated learning (PFL) method to load prediction under non-independent and identically distributed (non-IID) metering data settings. Specifically, we introduce meta-learning, where the learning rates are manipulated using the meta-learning idea to maximize the gradient for each client in each global round. Clients with varying processing capacities, data sizes, and batch sizes can participate in global model aggregation and improve their local load forecasting via personalized learning. Simulation results show that our approach outperforms state-of-the-art ML and FL methods in terms of better load forecasting accuracy.

cs.LG

Assessing Reusability of Deep Learning-Based Monotherapy Drug Response Prediction Models Trained with Omics Data

Cancer drug response prediction (DRP) models present a promising approach towards precision oncology, tailoring treatments to individual patient profiles. While deep learning (DL) methods have shown great potential in this area, models that can be successfully translated into clinical practice and shed light on the molecular mechanisms underlying treatment response will likely emerge from collaborative research efforts. This highlights the need for reusable and adaptable models that can be improved and tested by the wider scientific community. In this study, we present a scoring system for assessing the reusability of prediction DRP models, and apply it to 17 peer-reviewed DL-based DRP models. As part of the IMPROVE (Innovative Methodologies and New Data for Predictive Oncology Model Evaluation) project, which aims to develop methods for systematic evaluation and comparison DL models across scientific domains, we analyzed these 17 DRP models focusing on three key categories: software environment, code modularity, and data availability and preprocessing. While not the primary focus, we also attempted to reproduce key performance metrics to verify model behavior and adaptability. Our assessment of 17 DRP models reveals both strengths and shortcomings in model reusability. To promote rigorous practices and open-source sharing, we offer recommendations for developing and sharing prediction models. Following these recommendations can address many of the issues identified in this study, improving model reusability without adding significant burdens on researchers. This work offers the first comprehensive assessment of reusability and reproducibility across diverse DRP models, providing insights into current model sharing practices and promoting standards within the DRP and broader AI-enabled scientific research community.

q-bio.BM

On Conjecture of Binomial Edge Ideals of Linear Type

An ideal $I$ of a commutative ring $R$ is said to be of linear type when its Rees algebra and symmetric algebra exhibit isomorphism. In this paper, we investigate the conjecture put forth by Jayanthan, Kumar, and Sarkar (2021) that if $G$ is a tree or a unicyclic graph, then the binomial edge ideal of $G$ is of linear type. Our investigation validates this conjecture for trees. However, our study reveals that not all unicyclic graphs adhere to this conjecture.

math.AC

CACTUS: Chemistry Agent Connecting Tool-Usage to Science

Large language models (LLMs) have shown remarkable potential in various domains, but they often lack the ability to access and reason over domain-specific knowledge and tools. In this paper, we introduced CACTUS (Chemistry Agent Connecting Tool-Usage to Science), an LLM-based agent that integrates cheminformatics tools to enable advanced reasoning and problem-solving in chemistry and molecular discovery. We evaluate the performance of CACTUS using a diverse set of open-source LLMs, including Gemma-7b, Falcon-7b, MPT-7b, Llama2-7b, and Mistral-7b, on a benchmark of thousands of chemistry questions. Our results demonstrate that CACTUS significantly outperforms baseline LLMs, with the Gemma-7b and Mistral-7b models achieving the highest accuracy regardless of the prompting strategy used. Moreover, we explore the impact of domain-specific prompting and hardware configurations on model performance, highlighting the importance of prompt engineering and the potential for deploying smaller models on consumer-grade hardware without significant loss in accuracy. By combining the cognitive capabilities of open-source LLMs with domain-specific tools, CACTUS can assist researchers in tasks such as molecular property prediction, similarity searching, and drug-likeness assessment. Furthermore, CACTUS represents a significant milestone in the field of cheminformatics, offering an adaptable tool for researchers engaged in chemistry and molecular discovery. By integrating the strengths of open-source LLMs with domain-specific tools, CACTUS has the potential to accelerate scientific advancement and unlock new frontiers in the exploration of novel, effective, and safe therapeutic candidates, catalysts, and materials. Moreover, CACTUS's ability to integrate with automated experimentation platforms and make data-driven decisions in real time opens up new possibilities for autonomous discovery.

cs.CL

Almost Complete Intersections: Regularity of powers and Rees algebra

An almost complete intersection ideal can be seen as a $d$-sequence ideal with the minimal number of generators being one more than its height. In this paper, we give exact formulas for the regularity of powers of graded almost complete intersection ideals satisfying certain conditions. We also present the minimal bigraded free resolutions of Rees algebras associated with linear type ideals having one generator more than their grade and study the properties of Cohen-Macaulayness, Koszulness associated to their diagonals.

math.AC