Search arXiv⌕ Search

arXiv subjects

Ting Xu

Publications and source records attributed to Ting Xu.

At least 19 recordsLinked to original sources

Clinician-Friendly Foundation Models for Ophthalmic Image Diagnostics without Fine-Tuning or Technical Barriers

Artificial intelligence (AI) shows remarkable potential in medical imaging diagnostics, yet most current models require retraining when applied across different clinical settings, limiting their scalability. We developed GlobeReady, a deployment-oriented platform powered by the RetiGlobe foun- dation model and local feature augmentation. RetiGlobe was pretrained in two stages: 1) self-supervised learning using DINOv2 on 38 million synthetic ophthalmic images, and 2) contrastive learning using CLIP on 475,845 real image-text pairs spanning diverse ethnicities, imaging devices, and geographic regions worldwide. We evaluate GlobeReady on 488,448 ophthalmic images, including color fundus photographs (CFPs) and optical coherence tomography scans, from multi-centres in China, Singapore, Vietnam and the UK. Prospective testing included usability assessment with 31 ophthalmologists. Exploratory analyses evaluated domain generalisability, Bayesian uncertainty quantification, out-of-distribution (OOD) detection, and feature-based case retrieval.

cs.CV↗

The Ponderomotive Effects of Narrow-band, Superconducting Resonators in Open and Closed Loop

In this work, we present measurements of ponderomotive instabilities in narrow-band, coaxial resonators in open and closed control loop systems. We show an analytical scheme we will use in future studies to examine the dependency of open loop stability on mechanical parameters. Analytical and simulation models are used to predict the onset of the oscillatory instability in half wave resonators with active disturbance rejection control for amplitude and phase stabilization. We demonstrate that for high amplitude controller bandwidths, ponderomotive oscillations can couple with the controller frequency response via higher harmonics and lower the threshold for the oscillatory instability. We reaffirm the superiority of in-phase/quadrature $(I/Q)$ component control in regards to preventing the oscillatory instability, and show how the phase controller can amplify the crosstalk of disturbances in amplitude and phase control. In cases with large disturbances, this effect can lead to instabilities not present in $I/Q$ control.

physics.acc-ph↗

Modulating task outcome value to mitigate real-world procrastination via noninvasive brain stimulation

Procrastination represents one of the most prevalent behavioral problems associated with individual health and societal productivity. Despite its high prevalence and substantial impact on daily functioning, its underlying neurocognitive mechanisms remain poorly understood. A leading model posits that procrastination arises from imbalanced competing motivations: the avoidance of negative task aversiveness and the pursuit of positive task outcomes, yet this framework has not been fully validated in real-world settings and not applied effectively to guide interventions. Here, we addressed this gap with a double-blind, randomized controlled trial. We applied seven sessions of high-definition transcranial direct current stimulation (HD-tDCS) to the left dorsolateral prefrontal cortex (DLPFC) in chronic procrastinators. Using the intensive experience sampling method (iESM), we assessed the effect of anodal HD-tDCS on real-world procrastination at offline after-effect (2-day interval) and long-term after-effect (6-month follow-up). We found that this neuromodulation produced a lasting reduction in real-world procrastination, with effects sustained at a 6-month follow-up. While the intervention is significantly associated with both decreased task aversiveness and increased perceived task outcome value, a mediation analysis indicated a disassociable mechanism: the increase in task outcome value (but not task aversiveness) showed a statistical pattern consistent with accounting for the observed behavioral improvement. In conclusion, the findings are consistent with the hypothesis that enhancing DLPFC function may reduce procrastination by selectively amplifying the valuation of future rewards, not by simply reducing negative feelings about the task. These results align with established decision-theoretic frameworks and suggest a targeted, theory-informed avenue for future behavioral interventions.

q-bio.NC↗

Fed-Equilibrium Framework for Topological Pareto Control in Robust and Fair Clinical Federated Learning

The deployment of Federated Learning (FL) in multi-center clinical networks faces the challenge of "knowledge dominance," where high-volume hubs naturally overwhelm minority community nodes, implicitly treating the distinct clinical patterns of smaller cohorts as outliers. Existing geometric defenses provide a security baseline but leave this efficiency-fairness dilemma unresolved. To bridge this gap, we propose Fed-Equilibrium, a framework that advances the paradigm from simple defense to topological equilibrium. Unlike traditional aggregators, Fed-Equilibrium implements a sequential architectural synergy. It utilizes a two-stage gradient control cascade: Stage I (geometric quality assurance) enforces directional consistency via a cosine similarity funnel to filter malicious noise, creating a stabilized manifold; Stage II (topological Pareto control) then actively modulates verified contributions by identifying the optimal Pareto knee point. We validated this framework on a bi-national simulation integrating Canadian (CNODES) and U.S. (SyntheticMass) registries. Experimental results demonstrate that the system simultaneously secures the network against adversarial divergence while accommodating underrepresented signals. Notably, the minority U.S. spoke (representing less than 3% of data volume) achieved deep convergence comparable to the data-rich Canadian hub. This confirms that Fed-Equilibrium effectively counters "knowledge dominance," establishing a true "knowledge commons" where global generalizability does not come at the cost of local clinical representation.

cs.LG↗

DeepRHP: A Hybrid Variational Autoencoder for Designing Random Heteropolymers as Protein Mimics

Synthetic random heteropolymers (RHPs), consisting of a predefined set of monomers, offer an approach toward the design of protein-like materials. These RHPs, if designed appropriately, can mimic protein behavior and function. As such, there is a need for computational tools to efficiently guide RHP design. We bridge this gap by developing DeepRHP, a modified variational autoencoder (VAE) model under a semi-supervised framework. By equipping a classical VAE with an additional feature-based VAE, DeepRHP forces the latent space to capture structures of critical chemical features as well as individual RHP sequence patterns. In this sense, our method is versatile by allowing any relevant features to be incorporated in a hybrid manner. We demonstrate the effectiveness of DeepRHP by suggesting potential monomer compositions that stabilize membrane proteins (e.g. Aquaporin Z) in non-native environments and cross-validating our prediction with published results. The concordance between our model and true RHP function suggests strong potential in utilizing hybrid autoencoder architectures to guide RHP design for proteins and other biological compounds.

cs.LG↗

Unveiling the Entropy Dynamics of Chain-of-Thought Reasoning

This paper investigates the entropy dynamics of Chain-of-Thought (CoT) and uncovers a consistent two-phase structure: an Uncertainty Region of exploration transitioning sharply to a Confidence Region of convergence. We demonstrate that the Confidence Region possesses two critical properties: 1) High Reliability -- answers in the confidence region become highly accurate and stable, and 2) High Redundancy -- models generate unnecessary tokens long after reaching the correct answer. These properties unlock more efficient and reliable inference strategies: 1) Early Exit leverages reliability and redundancy to terminate computation safely when returns diminish, and 2)Test-Time Scaling uses the Confidence Region signal to prioritize converged trajectories. To operationalize these insights, we formulate Confidence Region detection as a sequential change-point detection problem, being the first to apply classical change-point methods to monitor CoT reasoning. Using the Cumulative Sum (CUSUM) algorithm, a statistically optimal change-point detector, we develop a training-free framework for real-time inference control. Experiments show our approach establishes a superior Pareto-frontier for early exit. CUSUM achieves 63.06% accuracy with 11.1% token reduction, outperforming DEER and Dynasor by 3.28% and 4.36% in accuracy respectively. For test-time scaling, CUSUM-weighted voting consistently outperforms self-consistency.

cs.CL↗

StepAudio 2.5 Technical Report

Unified audio-language modeling has emerged as a prominent trend in modern speech systems, promising to bring the reasoning capabilities of large language models to auditory tasks. However, existing unified foundations often struggle to match the depth of specialized systems across automatic speech recognition (ASR), text-to-speech synthesis (TTS), and realtime spoken interaction. Bridging this gap remains an open challenge. This report presents StepAudio 2.5, a unified audio-language foundation model that matches or exceeds specialized systems across all three capabilities. Rather than treating these tasks as architecturally distinct, we operate on the premise that once text and audio share a multimodal representational space, task specialization becomes a matter of operational regimes: data construction, optimization targets, and decoding constraints. Guided by this insight, we advance the post-training paradigm from standard supervised learning to task-tailored Reinforcement Learning from Human Feedback (RLHF), using it as the primary mechanism to define complex optimization targets. We leverage this RLHF-centric alignment, alongside specialized decoding, to shape a shared backbone into three distinct operational modes. Concretely, the ASR branch advances transcription efficiency via verifiable multi-token decoding; the TTS branch achieves controllable, expressive synthesis through preference-based RLHF and context-rich supervision; and the Realtime branch realizes low-latency, persona-consistent dialogue via generative reward modeling within an RLHF framework. On standard benchmarks, StepAudio 2.5 achieves state-of-the-art results across ASR, TTS, and Realtime, demonstrating that a singular audio-language foundation can successfully internalize the distinct deployment objectives of speech understanding, generation, and live interaction.

eess.AS↗

Behavioral Consistency and Transparency Analysis on Large Language Model API Gateways

Third-party Large Language Model (LLM) API gateways are rapidly emerging as unified access points to models offered by multiple vendors. However, the internal routing, caching, and billing policies of these gateways are largely undisclosed, leaving users with limited visibility into whether requests are served by the advertised models, whether responses remain faithful to upstream APIs, or whether invoices accurately reflect public pricing policies. To address this gap, we introduce GateScope, a lightweight black-box measurement framework for evaluating behavioral consistency and operational transparency in commercial LLM gateways. GateScope is designed to detect key misbehaviors, including model downgrading or switching, silent truncation, billing inaccuracies, and instability in latency by auditing gateways along four critical dimensions: response content analysis, multi-turn conversation performance, billing accuracy, and latency characteristics. Our measurements across 10 real-world commercial LLM API gateways reveal frequent gaps between expected and actual behaviors, including silent model substitutions, degraded memory retention, deviations from announced pricing, and substantial variation in latency stability across platforms.

cs.CR↗

Impact of heat treatments on the performance of low-frequency superconducting quarter-wave resonators at 4.3 K

We applied heat treatments to 80.5 MHz quarter-wave resonators made from bulk niobium and prepared with buffered chemical polishing BCP. We evaluated their performance at 4.3 K. We found that a 48 hour, 120 C bake-out ("low-temperature bake out") reduces the surface resistance by a factor of 2 to 3, stemming from a reduction in the Bardeen-Cooper-Schrieffer contribution, consistent with previous findings. This decrease leads to a 38% decrease on average in the medium-field Q-slope when compared to cavities which had only BCP. Mechanisms for the change in quality factor with low-temperature baking have been explored. We observed no improvement in cavity performance after a 3-hour bake-out at 350 C ("medium-temperature bake out"), in contrast to observations for higher-frequency cavities.

physics.acc-ph↗

Spectral radius and parity $[a,b]$-factors in graphs

Let $a$, $b$, and $n$ be three integers such that $1\leq a \leq b < n$, $a \equiv b$ (mod $2$), and $na$ is even. A parity $[a,b]$-factor of $G$ is a spanning subgraph $H$ such that for each vertex $v \in V(G)$, $a \leq d_H(v) \leq b$ and $d_H(v) \equiv a \equiv b$ (mod $2$). Recently, O [J. Graph Theory 100 (2022) 458-469] proved eigenvalue conditions for a regular graph to have a parity $[a,b]$-factor. In this paper, we prove a sharp lower bound on the spectral radius for an $n$-vertex graph $G$ to have a parity $[a,b]$-factor as follows: If $G$ is an $n$-vertex connected graph with $δ(G)\geq a$ and $ρ(G)\geqρ(G_{n}^{a})$, then $G$ contains a parity $[a,b]$-factor unless $G \cong G_{n}^{a}$, where $2\leq a<b$ and $G_{n}^{a}$ is the graph obtained from $K_{a-1}\vee(K_{n-2a-1}\cup(a+1)K_1)$ by adding a new vertex and adding all possible edges between the added vertex and each vertex in $(a+1)K_1$.

math.CO↗

Two-port CW measurements on RF cavities: Notes on self-consistency assessment and indirect methods

In the case of a radio-frequency (RF) cavity with a mismatched input coupler, a direct calculation of the power dissipation in the cavity and the intrinsic quality factor from continuous-wave (CW) measurements may have uncertainty due to systematic errors. Formulae for an indirect calculation of these quantities are derived for the case of a cavity with two couplers of fixed coupling strength. In this approach, the signal from the pickup coupler is used to infer the amplitude of the "emitted wave" from the input coupler. A graphical method for self-consistency assessment is evaluated. The impact of frequency offsets is considered. Applications of these methods are presented, drawing on cold tests of superconducting cavities produced for the Facility for Rare Isotope Beams.

physics.acc-ph↗

Visible Yet Unreadable: A Systematic Blind Spot of Vision Language Models Across Writing Systems

Writing is a universal cultural technology that reuses vision for symbolic communication. Humans display striking resilience: we readily recognize words even when characters are fragmented, fused, or partially occluded. This paper investigates whether advanced vision language models (VLMs) share this resilience. We construct two psychophysics inspired benchmarks across distinct writing systems, Chinese logographs and English alphabetic words, by splicing, recombining, and overlaying glyphs to yield ''visible but unreadable'' stimuli for models while remaining legible to humans. Despite strong performance on clean text, contemporary VLMs show a severe drop under these perturbations, frequently producing unrelated or incoherent outputs. The pattern suggests a structural limitation: models heavily leverage generic visual invariances but under rely on compositional priors needed for robust literacy. We release stimuli generation code, prompts, and evaluation protocols to facilitate transparent replication and follow up work. Our findings motivate architectures and training strategies that encode symbol segmentation, composition, and binding across scripts, and they delineate concrete challenges for deploying multimodal systems in education, accessibility, cultural heritage, and security.

cs.CV↗

Machine Learning Approaches to Clinical Risk Prediction: Multi-Scale Temporal Alignment in Electronic Health Records

This study proposes a risk prediction method based on a Multi-Scale Temporal Alignment Network (MSTAN) to address the challenges of temporal irregularity, sampling interval differences, and multi-scale dynamic dependencies in Electronic Health Records (EHR). The method focuses on temporal feature modeling by introducing a learnable temporal alignment mechanism and a multi-scale convolutional feature extraction structure to jointly model long-term trends and short-term fluctuations in EHR sequences. At the input level, the model maps multi-source clinical features into a unified high-dimensional semantic space and employs temporal embedding and alignment modules to dynamically weight irregularly sampled data, reducing the impact of temporal distribution differences on model performance. The multi-scale feature extraction module then captures key patterns across different temporal granularities through multi-layer convolution and hierarchical fusion, achieving a fine-grained representation of patient states. Finally, an attention-based aggregation mechanism integrates global temporal dependencies to generate individual-level risk representations for disease risk prediction and health status assessment. Experiments conducted on publicly available EHR datasets show that the proposed model outperforms mainstream baselines in accuracy, recall, precision, and F1-Score, demonstrating the effectiveness and robustness of multi-scale temporal alignment in complex medical time-series analysis. This study provides a new solution for intelligent representation of high-dimensional asynchronous medical sequences and offers important technical support for EHR-driven clinical risk prediction.

cs.LG↗

Tunable Nanostructures from Inverse Surfactants

Hierarchical materials in the natural world are often made through the self-assembly of amphiphilic molecules. Achieving similar structural complexity in synthetic materials requires understanding how various molecular parameters affect assembly behavior. In recent years, inverse surfactants -- molecules with hydrophobic head groups and hydrophilic macromolecular tails -- have been shown to self-assemble into supramolecular assemblies in aqueous solutions that show promise for a number of applications, including drug delivery. Here, we build an understanding of the morphological phase diagram of inverse surfactants using insights from scattering experiments, computer simulations, and statistical mechanics. The scattering and simulation results reveal that changing the head-group size is an important molecular knob in controlling morphological transitions. The molecular size ratio of the hydrophobic group to the hydrophilic emerges as a crucial dimensionless quantity in our theory and plays a determining role in setting the micelle structure and the transition from mesoscale to macroscale aggregates. Our minimal theory is able to qualitatively explain the key features of the morphological phase diagram, including the prevalence of fiber-like structures in comparison to spherical and planar micelles. Together, these findings provide a more complete picture for the molecular dependencies of assemblies of inverse surfactants, which we hope may aid in the de novo design of supramolecular structures.

cond-mat.soft↗

Report on first plasma processing trial for a FRIB quarter-wave resonator cryomodule

Plasma processing has been shown to help mitigate degradation of the performance of superconducting radio-frequency cavities, providing an alternative to removal of cryomodules from the accelerator for refurbishment. Studies of plasma processing for quarter-wave resonators (QWRs) and half-wave resonators (HWRs) are underway at the Facility for Rare Isotope Beams (FRIB), where a total of 324 such resonators are presently in operation. Plasma processing tests were done on several QWRs using the fundamental power coupler (FPC) to drive the plasma, with promising results. Driving the plasma with a higher-order mode allows for less mismatch at the FPC and higher plasma density. The first plasma processing trial for FRIB QWRs in a cryomodule was conducted in January 2024. Cold tests of the cryomodule showed a significant reduction in field emission X-rays after plasma processing.

physics.acc-ph↗

Benchmarking the Thinking Mode of Multimodal Large Language Models in Clinical Tasks

A recent advancement in Multimodal Large Language Models (MLLMs) research is the emergence of "reasoning MLLMs" that offer explicit control over their internal thinking processes (normally referred as the "thinking mode") alongside the standard "non-thinking mode". This capability allows these models to engage in a step-by-step process of internal deliberation before generating a final response. With the rapid transition to and adoption of these "dual-state" MLLMs, this work rigorously evaluated how the enhanced reasoning processes of these MLLMs impact model performance and reliability in clinical tasks. This paper evaluates the active "thinking mode" capabilities of two leading MLLMs, Seed1.5-VL and Gemini-2.5-Flash, for medical applications. We assessed their performance on four visual medical tasks using VQA-RAD and ROCOv2 datasets. Our findings reveal that the improvement from activating the thinking mode remains marginal compared to the standard non-thinking mode for the majority of the tasks. Their performance on complex medical tasks such as open-ended VQA and medical image interpretation remains suboptimal, highlighting the need for domain-specific medical data and more advanced methods for medical knowledge integration.

cs.CL↗

Improved high-gradient performance for medium-velocity superconducting half-wave resonators: Surface preparation and trapped flux mitigation

A development effort to improve the performance of superconducting radio-frequency half-wave resonators (SRF HWRs) is underway at the Facility for Rare Isotope Beams (FRIB), where 220 such resonators are in operation. Our goal was to achieve an intrinsic quality factor (Q0) of >= 2E10 at an accelerating gradient (Ea) of 12 MV/m. FRIB production resonators were prepared with buffered chemical polishing (BCP). First trials on electropolishing (EP) and post-EP low temperature baking (LTB) of FRIB HWRs allowed us to reach higher gradient (15 MV/m, limited by quench) with a higher quality factor at high gradient, but Q0 was still below our goal. Trapped magnetic flux during the Dewar test was found to be a source of Q0 reduction. Three strategies were used to reduce the trapped flux: (i) adding a local magnetic shield (LMGS) to supplement the ``global'' magnetic shield around the Dewar for reduction of the ambient magnetic field; (ii) performing a ``uniform cool-down'' (UC) to reduce the thermoelectric currents; and (iii) using a compensation coil to further reduce the ambient field with active field cancellation (AFC). The LMGS improved the Q0, but not enough to reach our goal. With UC and AFC, we exceeded our goal, reaching Q0 = 2.8E10 at Ea = 12 MV/m.

physics.acc-ph↗

Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel

Mixture-of-Experts (MoE) has become a cornerstone in recent state-of-the-art large language models (LLMs). Traditionally, MoE relies on $\mathrm{Softmax}$ as the router score function to aggregate expert output, a designed choice that has persisted from the earliest MoE models to modern LLMs, and is now widely regarded as standard practice. However, the necessity of using $\mathrm{Softmax}$ to project router weights into a probability simplex remains an unchallenged assumption rather than a principled design choice. In this work, we first revisit the classical Nadaraya-Watson regression and observe that MoE shares the same mathematical formulation as Nadaraya-Watson regression. Furthermore, we show that both feed-forward neural network (FFN) and MoE can be interpreted as a special case of Nadaraya-Watson regression, where the kernel function corresponds to the input neurons of the output layer. Motivated by these insights, we propose the \textbf{zero-additional-cost} Kernel Inspired Router with Normalization (KERN), an FFN-style router function, as an alternative to $\mathrm{Softmax}$. We demonstrate that this router generalizes both $\mathrm{Sigmoid}$- and $\mathrm{Softmax}$-based routers. \textbf{Based on empirical observations and established practices in FFN implementation, we recommend the use of $\mathrm{ReLU}$ activation and $\ell_2$-normalization in $\mathrm{KERN}$ router function.} Comprehensive experiments in MoE and LLM validate the effectiveness of the proposed FFN-style router function \methodNorm.

cs.CL↗