Search arXiv⌕ Search

arXiv subjects

Xin Ma

Publications and source records attributed to Xin Ma.

At least 19 recordsLinked to original sources

Maximal Algebraic Ideals in Nonunital $C^*$-Algebras

Motivated by Ozawa's question of whether every maximal algebraic two-sided ideal in a $C^*$-algebra must be closed, we study the existence of maximal algebraic two-sided ideals in nonunital $C^*$-algebras. We formulate singular-distribution estimates intrinsically through lower semicontinuous 2-quasitraces and apply them to control algebraic ideal membership. This allows us to develop a novel criterionthe, the admissible quasitracial projection scale, for establishing the nonexistence of maximal ideals in nonunital $C^*$-algebras. This criterion applies to a wide class of simple $C^*$-algebras, including: (i) all $A\otimes K$ where $A$ is unital, simple, stably finite, $QT_2^1(A)$ nonempty, and the radius of comparison $rc(A)$ is finite, as well as all their hereditary $C^*$-subalgebras whenever $A$ satisfies further assumptions that A is of real rank zero and A has finitely many extreme quasitraces; (ii) all nonunital, simple, separable, stably finite, $Z$-stable $C^*$-algebras A that have an approximate identity consisting of increasing projections $(p_n)$ and for which the simplex $QT_2^1(A,p_1)$ of normalized traces at $p_1$ has finitely many extreme points. We also show that a large class of $C^*$-algebras E constructed from extensions of $C^*$-algebras above such that $E$ still has no maximal ideals. In particular, for these classes, Ozawa's question could be settled in an unexpected manner.

math.OA↗

An evidence-guided reinforcement learning method to improve psychiatric reasoning in small language models

Privacy and computational constraints limit the use of large language models in psychiatry, while adapting small language models (SLMs) often requires substantial data and expert annotation. We developed ClinMPO, an evidence-guided reinforcement-learning framework guided by the psychiatrist-defined Clinical Psychiatry Thinking Strategy (CPTS). ClinMPO uses ClinRM, a reward model trained on 18,569 question--answer pairs from 4,474 psychiatry articles. We evaluated four Qwen3 sizes on 1,737 model-screened questions. ClinMPO outperformed Base, supervised fine-tuning and standard group relative policy optimization across scales. From responses by 300 senior pre-licensure medical students, we established the human baseline, a medical-student reference. The 4B model approached this baseline, whereas the 8B model surpassed it and ranked first among 31 models and post-training variants. ClinMPO improved performance across two complementary schemes covering ICD-11 diagnostic categories and psychiatric practice competencies. Blinded assessment by three clinicians showed improved rationale quality across CPTS criteria. These findings highlight how existing clinical evidence and specialist knowledge can be incorporated into the development of medical AI systems through evidence-guided learning.

cs.CL↗

Spin triplet pairing by suppressing altermagnetism

The interplay between unconventional superconductivity and altermagnetic order has attracted much attention. In particular, whether spin-triplet superconductivity can be achieved by suppressing altermagnetism remains an open issue. We investigate this issue using a minimal single-orbital Hubbard model on a square lattice with a vacancy superstructure, in which both conventional antiferromagnetic and altermagnetic orders can emerge on an equal footing. We illustrate the existence of a metallic normal phase with altermagnetism even at half-filling due to geometric frustration and Coulomb interaction. Suppressing the altermagnetic long-range order using charge doping can lead to both conventional antiferromagnetic and altermagnetic spin fluctuations. Spin-singlet pairing is always favored when conventional antiferromagnetic spin fluctuations dominate. However, when altermagnetic spin fluctuations dominate, spin-triplet pairing will be induced. Implications of our results for possible material candidates are also briefly discussed.

cond-mat.supr-con↗

Tilting realizations of derived-equivalent matrix centralizer algebras

Let $A$ be the centralizer algebra of a matrix over an arbitrary field. We solve the fixed-source realization problem for matrix centralizers by proving that the matrix centralizer algebras derived equivalent to $A$ are precisely the opposite endomorphism algebras of tilting modules over $A$. We classify the basic tilting modules and determine their opposite endomorphism algebras. The tilting poset is a product of right weak orders on symmetric groups, with one factor for each primary block and degree equal to the number of distinct exponents in that block. Together with the center, this poset recovers the multiset of these numbers across all primary blocks, although it does not canonically match them with the local center factors. For each primary block, the target algebras are obtained by permuting the successive gaps between exponents, and their isomorphism classes are determined by the stabilizer of the gap word. Consequently, the quotient of the labeled mutation graph by target-algebra isomorphism is a Schreier multigraph. We also characterize when the weak-order orientation descends to its nonloop edges.

math.RT↗

StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows

Recent advances in Large Language Models(LLMs) and agents have substantially improved the ability of AI systems to execute complex tasks. Yet existing benchmarks largely rely on researcher-selected tasks, leaving uncertain whether such progress extends to the work that real-world users actually demand from AI systems. We introduce \textbf{StartupBench}, an E2E agent benchmark grounded in market-validated AI startup products. Rather than defining tasks from pre-defined assumptions about useful agent capabilities, we systematically study AI products with demonstrated adoption, together with their product workflows and users, to identify real-world tasks for which AI has established practical demand across diverse professional domains. We translate these workflows into complete deliverable-oriented tasks and evaluate them with fine-grained rubrics capturing their complex requirements. Across representative models evaluated under a unified agent harness, even the strongest model successfully completes only approximately 30\% of StartupBench, despite making substantial partial progress on many tasks. Further analysis identifies aspects like complex instruction following and domain-specific expertise as major sources of failure. Our results reveal that many market-validated workflows remain beyond the reliable capabilities of current general-purpose agents, establishing StartupBench as an empirical measure of progress toward E2E completions of real-world user tasks.

cs.AI↗

Almost elementary étale groupoids

Motivated by Matui and Kerr's work on almost finiteness, we introduce a new finite approximation property for (possibly non-ample) étale groupoids called {almost elementariness}, which unifies and generalizes both almost finiteness and pure infiniteness. This property serves as a dynamical analogue of regularity properties of $C^*$-algebras. In support of this view, we prove that minimal almost elementary groupoids yield tracially $\mathcal{Z}$-stable reduced groupoid $C^*$-algebras. Consequently, we obtain as a corollary that the reduced $C^*$-algebras of all minimal amenable second countable almost finite groupoids in Matui's sense are $\mathcal{Z}$-stable and thus classifiable by the Elliott invariants. Two basic ingredients underlying the definition of almost elementariness are castles in groupoids and groupoid subequivalence, both of which are developed extensively in this work. Notably, building on our flexible framework of castles, we introduce the technique of nesting of castles, which play a key role in the proof of our main theorem. Besides our main theorem on tracial $\mathcal{Z}$-stability, we also discuss in depth the relations between almost elementariness and other properties for groupoids such as effectiveness, the groupoid small boundary property, groupoid strict comparison and Matui's and Kerr's notions of almost finiteness.

math.OA↗

Almost elementary groupoid models for $C^*$-algebras

The notion of almost elementariness for a locally compact Hausdorff étale groupoid $\mathcal{G}$ with a compact unit space was introduced by the authors as a sufficient condition ensuring the reduced groupoid $C^*$-algebra $C^*_r(\mathcal{G})$ is (tracially) $\mathcal{Z}$-stable and thus classifiable under additional natural assumptions. In this paper, we explore the converse direction and show that many groupoids in the literature serving as models for classifiable $C^*$-algebras are almost elementary. In particular, for a large class $\mathcal{C}$ of Elliott invariants and a $C^*$-algebra $A$ with $\operatorname{Ell}(A)\in \mathcal{C}$, we show that $A$ is classifiable if and only if $A$ possesses a minimal, effective, amenable, second countable, almost elementary groupoid model, which leads to a groupoid-theoretic characterization of classifiability of $C^*$-algebras with certain Elliott invariants. In addition, we demonstrate obstructions to obtaining a transformation groupoid model for the Jiang-Su algebra $\mathcal{Z}$.

math.OA↗

Fiberwise amenability of étale groupoids

We introduce a new amenability property for étale groupoids, termed \textit{fiberwise amenability}, along with a stronger variant termed \emph{ubiquitous fiberwise amenability}. (Ubiquitous) fiberwise amenability emerges naturally from a coarse-geometric perspective on étale groupoids and, in the special case of transformation groupoids, it coincides precisely with the amenability of the acting group (rather than topological amenability of the action). It is also tightly linked to the existence of invariant measures on the unit space of the groupoid. The coarse-geometric framework for étale groupoids that we develop systematically in this work allows us to establish several foundational properties of (ubiquitous) fiberwise amenability. As an application, we prove a Følner--paradoxical dichotomy for minimal étale groupoids, which will serve as a key tool in a sequel on almost elementariness of étale groupoids.

math.DS↗

Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming

Real-time long-form avatar audio-video generation requires causal, continuous synthesis while maintaining audiovisual synchronization and visual consistency. Adapting a pretrained bidirectional model to this setting presents two key dilemmas. First, autoregressively reusing generated blocks as context creates exposure bias, causing errors and visual drift to accumulate over long rollouts. Second, a global speech utterance does not indicates a causal generator which portion should be spoken next when only limited local audio-video context is available. We present Vorch-Streamer, a post-training framework that addresses these challenges and enables real-time long-form Text-to-Audio-Video (T2AV) streaming. We construct a synthetic corpus of 80K avatar clips spanning 12-21 seconds and first train a causal generator with mixed Teacher Forcing and Diffusion Forcing. We then apply long-horizon Self Forcing with DMD distillation, exposing the model to its own rollout distribution while preserving the quality of the pretrained bidirectional teacher. To explicitly control speech progression, an external language model predicts discrete 25-Hz speech-planning tokens, whose continuous features condition the audio diffusion branch and align each causal block with the content it should speak. With bounded causal context and four-step denoising, Vorch-Streamer jointly generates audio and video from text at 27.12 FPS, exceeding the 24-FPS real-time playback rate while maintaining competitive audio-lip synchronization and strong identity preservation over long-form generation.

cs.CV↗

Bockstein operations and AD algebras with unbounded torsion in $K_1$

Eilers showed that for AD algebras of real rank zero with bounded torsion in $\mathrm{K}_1$, the coefficient transformations $κ$ are redundant in the classification by ordered scaled total $K$-theory. In this paper we treat the unbounded torsion case and prove that, in contrast, $κ$ becomes necessary. Specifically, we construct two non-isomorphic unital AD algebras of real rank zero, $E_0$ and $E_1$, such that their ordered scaled total $K$-theory invariants agree when the $κ$-maps are forgotten, i.e., \[ \bigl( \underline{\mathrm{K}}(E_0), \underline{\mathrm{K}}(E_0)_+, [1_{E_0}] \bigr)_{\underline{\mathrm{K}}_{\langleκ\rangle}} \cong \bigl( \underline{\mathrm{K}}(E_1), \underline{\mathrm{K}}(E_1)_+, [1_{E_1}] \bigr)_{\underline{\mathrm{K}}_{\langleκ\rangle}} \] but are not isomorphic under the full $Λ$-module structure: \[ \bigl( \underline{\mathrm{K}}(E_0), \underline{\mathrm{K}}(E_0)_+, [1_{E_0}] \bigr)_Λ \not\cong \bigl( \underline{\mathrm{K}}(E_1), \underline{\mathrm{K}}(E_1)_+, [1_{E_1}] \bigr)_Λ .\] This completes the picture for the necessity of all three operations $ρ$, $β$, and $κ$ in this context.

math.OA↗

Vorch-IR: Long-Form Unified Multimodal Identity Replacement Video Generation

Video identity replacement seeks to transfer the identities of one or more subjects while preserving the motion, expressions, and temporal structure of a driving video. Existing methods largely target single-person settings and often require task-specific structural controls, such as masks or pose representations, limiting their flexibility in general multimodal editing systems. Progress on multi-person replacement is further constrained by the scarcity of paired training data. We present Vorch-IR, a unified framework that supports single- and dual-person identity replacement, with optional background replacement, in a single model. Built on LTX2, Vorch-IR jointly conditions on a driving video, indexed reference images, and a textual editing instruction. The reference images need not match the pose, layout, or spatial configuration of the driving video: their roles as subject or background references are specified through the instruction. Dense visual conditions are fused through self-attention, while a vision-language context establishes semantic correspondence through cross-attention. We further develop an automatic data construction pipeline that synthesizes paired supervision for all four editing settings. Experiments using automatic metrics and pairwise human evaluation demonstrate strong identity preservation, motion fidelity, and temporal coherence across diverse scenarios. A temporal overlapping inference strategy additionally extends the short-clip model to minute-long generation without autoregressive continuation.

cs.CV↗

Vorch-Director: Interactive World Story Model via Noise-Aware Error Rectification

Autoregressive continuation provides a natural path toward minute-scale audio-visual generation by repeatedly extending a short-window generator conditioned on previously generated video and audio. However, models are trained on clean ground-truth histories, while inference relies on their own generated histories, where accumulated errors cause identity drift, over-smoothing, and audio-visual desynchronization. Recent methods reduce this mismatch by reusing prediction residuals as synthetic corruption, but we observe that the effectiveness of residual correction critically depends on the flow-matching noise level at which residuals are produced. We propose Vorch-Director, a noise-level-aware residual correction strategy that associates each residual with its originating noise level and injects residuals from matched noise regimes during training. By aligning injected errors with the denoising process, Vorch-Director produces more realistic autoregressive histories while retaining efficient teacher-forcing training. Built on the audio-visual LTX-2 diffusion transformer, Vorch-Director further introduces task embeddings to distinguish historical video, reference images, and target video, enabling unified conditioning for long-horizon generation. Together with a clean conditioning sink and mixed-task training, Vorch-Director supports multi-shot, multi-subject, reference-guided audio-visual long-video generation. We evaluate Vorch-Director on ST-Bench and introduce a new long-horizon audio-visual benchmark with metrics for quality drift and long-range consistency. Extensive experiments demonstrate improved stability and audio-visual fidelity over strong baselines.

cs.CV↗

Vorch-Omni: Multi-Task Orchestration of Sight and Sound

Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches often rely on fragmented task-specific models. A general model must distinguish heterogeneous target, source, and reference signals to determine what to generate, preserve, or use as guidance, while reducing interference among tasks. Joint audio-visual generation further increases this challenge by introducing diverse conditioning and output configurations across modalities. We present Vorch-Omni, a unified multi-task framework for audio-visual synthesis based on an arbitrary-condition-to-arbitrary-output formulation. It flexibly treats video and audio signals as either conditioning inputs or generation targets. Token-level conditioning masks and task identifiers distinguish targets, source content, and references, while position types separate temporal context from independent conditions. To capture semantic and structural information, Vorch-Omni employs complementary visual conditioning pathways: a vision-language model interprets sampled frames with text instructions, and a video VAE encodes conditions into latent tokens for direct guidance. We further build a distributed data pipeline to curate diverse temporally aligned audio-visual clips, generate structured captions and metadata, and balance heterogeneous task distributions. Built on a single flow-matching diffusion transformer without task-specific architectural changes, Vorch-Omni supports over 10 tasks, including text-to-video, text-to-audio-video, image- and reference-conditioned generation, temporal extension, audio-driven generation, video transformation, and audio-visual editing. This unified framework provides a scalable foundation for general-purpose audio-visual generation and manipulation.

cs.CV↗

Vorch-Human: Unified Multi-Task Human-Centric Generation via Long-Horizon Continuation

Human-centric audio-visual generation spans several closely related tasks: animating a person from driving speech, jointly generating speech and video from a voice reference, and synthesizing a scene from paired appearance and voice references. Existing systems commonly solve these tasks with separate models, even though they share the same target modalities and differ mainly in which observations are provided as conditions. We present Vorch-Human, a unified human-centric generation framework built on a dual-stream audio-video diffusion transformer. Vorch-Human augments the conventional noisy audio/noisy video interface with clean condition-audio and condition-video token groups. Per-token task embeddings, temporal position types, condition masks, and a shared multimodal prompt encoder allow driving speech, timbre examples, first frames, and subject images to be expressed within one model. To supply the supervision required by this interface, we develop a two-level data pipeline. Level 1 analyzes each clip with speech recognition, vocal separation, face detection and tracking, active-speaker and synchronization models, audio/visual speaker clustering, and multimodal caption correction; it produces subject-indexed speech, appearance, and timbre annotations. Level 2 links the same person across clips from a common source video and mines identity- and outfit-consistent reference images after face, body, quality, pose, and vision-language verification. Finally, we adapt Vorch-Human to long-form audio-driven generation by training with clean latent prefixes and using the same frozen-prefix recurrence at inference. Each segment contributes only its newly generated suffix, reducing boundary discontinuity and long-horizon identity drift. Experiments on short and five-minute generation demonstrate strong identity preservation, audio-visual synchronization, and temporal stability.

cs.CV↗

More Edits, More Stable: Understanding the Lifelong Normalization in Sequential Model Editing

Lifelong Model Editing aims to continuously update evolving facts in Large Language Models while preserving unrelated knowledge and general capabilities, yet it remains plagued by catastrophic forgetting and model collapse. Empirically, we find that recent editors resilient over long horizons share the same core strategy: Lifelong Normalization (LN), which normalizes value gradients using running statistics. Removing LN causes immediate performance collapse, and we observe a counter-intuitive positive cumulative effect where early edits can promote the success of future edits. Yet the mechanism of LN remains a "black box", leaving its precise role in lifelong stability poorly understood. In this work, we provide the first theoretical account of LN in the lifelong regime. Our analysis reveals a self-reinforcing stability loop and proves that, when combined with ridge-regularized regression, LN yields parameter updates with asymptotic orthogonality and bounded norms, directly mitigating forgetting and systemic collapse. Based on these insights, we derive StableEdit, which strengthens this stability loop via an explicit warm-up stage and full whitening, improving long-horizon stability at minimal overhead. Extensive experiments validate our theory and demonstrate competitive performance. Our code is available at https://github.com/MINE-USTC/StableEdit.

cs.LG↗

SutureFormer: Learning Surgical Trajectories via Goal-conditioned Offline RL in Pixel Space

Predicting surgical needle trajectories from endoscopic video is critical for robot-assisted suturing, enabling anticipatory planning, real-time guidance, and safer motion execution. Existing methods that directly learn motion distributions from visual observations tend to overlook the sequential dependency among adjacent motion steps. Moreover, sparse waypoint annotations often fail to provide sufficient supervision, further increasing the difficulty of supervised or imitation learning methods. To address these challenges, we formulate image-based needle trajectory prediction as a sequential decision-making problem, in which the needle tip is treated as an agent that moves step by step in pixel space. This formulation naturally captures the continuity of needle motion and enables the explicit modeling of physically plausible pixel-wise state transitions over time. From this perspective, we propose SutureFormer, a goal-conditioned offline reinforcement learning framework that leverages sparse annotations to dense reward signals via cubic spline interpolation, encouraging the policy to exploit limited expert guidance while exploring plausible future motion paths. SutureFormer encodes variable-length clips using an observation encoder to capture both local spatial cues and long-range temporal dynamics, and autoregressively predicts future waypoints through actions composed of discrete directions and continuous magnitudes. To enable stable offline policy optimization from expert demonstrations, we adopt Conservative Q-Learning with Behavioral Cloning regularization. Experiments on a new kidney wound suturing dataset containing 1,158 trajectories from 50 patients show that SutureFormer reduces Average Displacement Error by 58.6% compared with the strongest baseline, demonstrating the effectiveness of modeling needle trajectory prediction as pixel-level sequential action learning.

cs.RO↗

Knowledge-augmented Agentic AI for Mental Health Medication Information Seeking

Patients increasingly seek medication information online, yet safety knowledge for psychiatric drugs is split between regulatory adverse-event records, which are authoritative but abstract, and patient narratives, which are experience-near but unvalidated. Integrating them without conflating evidence and anecdote is especially consequential in psychiatry, where poorly contextualised information can amplify fear, nocebo responses, and non-adherence. Here we develop a provenance-aware, knowledge-graph-based multi-agent framework unifying 466,525 Reddit posts, 60,782 WebMD reviews, and twenty years of U.S. FDA Adverse Event Reporting System records for nine antidepressants. A large-language-model entity-recognition pipeline benchmarked against physician annotations reached highest F1 scores of 0.969 for medications and 0.973 for conditions. The two community platforms were far more concordant with each other (overlap up to a Jaccard similarity of 0.905) than with regulatory reports, indicating that patient-generated data form a partly independent safety signal. For sertraline, many adverse events appeared in community sources hundreds of days before the corresponding FDA date. A Neo4j knowledge graph grounded in ATC-N, ICD-10, and MedDRA vocabularies preserves provenance, keeping every claim traceable and regulatory facts distinct from patient experience. These results establish source-aware integration as a route to more auditable psychiatric medication information, with usefulness and patient benefit to be tested prospectively.

cs.AI↗

Full Gabor frames, its existence problem, and a non-uniform Balian-Low type theorem

For a broad class of Delone sets in $\mathbb{R}^n$ that are of significance in both mathematics and physics, we prove a non-uniform Balian-Low type theorem and settle the converse problem on the existence of Gabor frames, for arbitrary dimension $n$. To this end, we introduce a class of Gabor frames, termed full Gabor frames, and prove that the existence of such a frame on the Delone set with Schwartz window functions is equivalent to the condition that the lower Beurling density be strictly greater than one. In fact, the usual Balian-Low direction using window functions from the Feichtinger's algebra can be proven for arbitrary point sets, thereby improving an earlier density theorem by Christensen, Deng, and Heil. The corresponding dual result for Riesz sequences is also obtained. The main technical tools employed in this paper are tiling groupoid constructions and $C^*$-algebraic methods. As a byproduct, we resolve an open question from Ito's thesis concerning the bounded dynamical asymptotic dimension of tiling groupoids. Furthermore, this result allows us to extend the classification theorem of Ito, Whittaker, and Zacharias to the twisted case.

math.FA↗