Search arXivSearch

arXiv subjects

Soham Ray

Publications and source records attributed to Soham Ray.

7 recordsLinked to original sources

Krylov complexity and fidelity susceptibility in two-band Hamiltonians

We investigate Krylov spread complexity for the ground state of two-band Hamiltonians, where the reference state is a generic state on the Bloch sphere. The spread complexity is obtained by using a purely geometric formulation in terms of Bloch sphere data without constructing the circuit Hamiltonian. For generic reference states, the derivative of the spread complexity is logarithmically divergent at the topological phase transition in the Su-Schrieffer-Heeger (SSH) model. We demonstrate that the derivative of the spread complexity is bounded by fidelity susceptibility for general two-band models, indicating the sensitivity of the spread complexity to any gap closing (topological or trivial). This is illustrated in the massive Dirac Hamiltonian with a trivial gap closing. Finally, we introduce a non-unitary duality in the SSH model between the topological and trivial phases, which manifests itself in the spread complexity and fidelity susceptibility.

quant-ph

$\tau$-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains

Full-duplex voice agents--systems that listen and speak simultaneously--are rapidly moving from research to production. However, existing evaluations address conversational dynamics and task completion in isolation. We introduce $\tau$-voice, a benchmark for evaluating voice agents on grounded tasks with real-world complexity: agents must navigate complex multi-turn conversations, adhere to domain policies, and interact with the environment. The framework extends $\tau^2$-bench into a novel voice agent benchmark combining verifiable completion of complex grounded tasks, full-duplex interaction, and realistic audio--enabling direct comparison between voice and text performance. A controllable and realistic voice user simulator provides diverse accents, realistic audio environments, and rich turn-taking dynamics; by decoupling simulation from wall-clock time, the user simulator can use the most capable LLM without real-time constraints. We evaluate task completion (pass@1) and voice interaction quality across 278 tasks: while GPT-5 (reasoning) achieves 85%, voice agents reach only 31--51% under clean conditions and 26--38% under realistic conditions with noise and diverse accents--retaining only 30--45% of text capability; qualitative analysis confirms 79--90% of failures stem from agent behavior, suggesting that observed failures primarily reflect agent behavior under our evaluation setup. $\tau$-voice provides a reproducible testbed for measuring progress toward voice agents that are natural, conversational, and reliable.

cs.SD

Symmetry and Topology in a Non-Hermitian Kitaev chain

We investigate the non-Hermitian Kitaev chain with non-reciprocal hopping amplitudes and asymmetric superconducting pairing. We work out the symmetry structure of the model and show that particle-hole symmetry (PHS) is preserved throughout the entire parameter regime. As a consequence of PHS, the topological phase transition point of a finite open chain coincides with that of the periodic (infinite) system. By explicitly constructing the zero-energy wave functions (Majorana modes), we show that Majorana modes necessarily occur as reciprocal localization pairs accumulating on opposite boundaries, whose combined probability density exhibits an exact cancellation of the non-Hermitian skin effect for the zero energy modes. Excited states, by contrast, generically display skin-effect localization, with particle and hole components accumulating at opposite ends of the system. At the level of bulk topology, we further construct a $\mathbb{Z}_2$ topological invariant in restricted parameter regimes that correctly distinguishes the topological and trivial phases. Finally, we present the topological phase diagram of the non-Hermitian Kitaev chain across a broad range of complex parameters and delineate the associated phase boundaries.

cond-mat.mes-hall

$\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Existing benchmarks for conversational AI agents simulate single-control environments, where only the AI agent can use tools to interact with the world, while the user remains a passive information provider. This differs from real-world scenarios like technical support, where users need to actively participate in modifying the state of the (shared) world. In order to address this gap, we introduce $\tau^2$-bench, with four key contributions: 1) A novel Telecom dual-control domain modeled as a Dec-POMDP, where both agent and user make use of tools to act in a shared, dynamic environment that tests both agent coordination and communication, 2) A compositional task generator that programmatically creates diverse, verifiable tasks from atomic components, ensuring domain coverage and controlled complexity, 3) A reliable user simulator tightly coupled with the environment, whose behavior is constrained by tools and observable states, improving simulation fidelity, 4) Fine-grained analysis of agent performance through multiple ablations including separating errors arising from reasoning vs communication/coordination. In particular, our experiments show significant performance drops when agents shift from no-user to dual-control, highlighting the challenges of guiding users. Overall, $\tau^2$-bench provides a controlled testbed for agents that must both reason effectively and guide user actions.

cs.AI

Sample-Efficient Diffusion for Text-To-Speech Synthesis

This work introduces Sample-Efficient Speech Diffusion (SESD), an algorithm for effective speech synthesis in modest data regimes through latent diffusion. It is based on a novel diffusion architecture, that we call U-Audio Transformer (U-AT), that efficiently scales to long sequences and operates in the latent space of a pre-trained audio autoencoder. Conditioned on character-aware language model representations, SESD achieves impressive results despite training on less than 1k hours of speech - far less than current state-of-the-art systems. In fact, it synthesizes more intelligible speech than the state-of-the-art auto-regressive model, VALL-E, while using less than 2% the training data.

cs.SD

Machine Learning ${\cal N}=8, D=5$ Gauged Supergravity

Type IIB string theory on a 5-sphere gives rise to ${\cal N}=8, SO(6)$ gauged supergravity in five dimensions. Motivated by the fact that this is the context of the most widely studied example of the AdS/CFT correspondence, we undertake an investigation of its critical points. The scalar manifold is an $E_{6(6)}/USp(8)$ coset, and the challenge is that it is 42-dimensional. We take a Machine Learning approach to the problem using TensorFlow, and this results in a substantial increase in the number of known critical points. Our list of 32 critical points contains all five of the previously known ones, including an ${\cal N}=2$ supersymmetric point identified by Khavaev, Pilch and Warner.

hep-th

Light cone OPE in a CFT with lowest twist scalar primary

We study the operator product expansion (OPE) of two identical scalar primary operators in the lightcone limit in a conformal field theory where a scalar is the operator with lowest twist. We see that in CFTs where both the stress tensor and a scalar are the lowest twist operators, the stress tensor contributes at the leading order in the lightcone OPE and the scalar contributes at the subleading order. We also see that there does not exist a scalar analogue of the average null energy condition (ANEC) for a CFT where a scalar is the lowest twist operator.

hep-th