Search arXivSearch

arXiv subjects

Jing Zhou

Publications and source records attributed to Jing Zhou.

At least 19 recordsLinked to original sources

Continuous Token-Level Spatio-Temporal Context Modeling for Visual Object Tracking

Spatio-temporal context has become increasingly crucial for visual tracking. However, most existing approaches extract spatio-temporal cues via discrete sampling strategies, which inherently deviate from the continuity of spatio-temporal context, thereby deteriorating tracking performance. To address this challenge, we propose TLCTrack, a novel tracking framework that models token-level spatio-temporal context through continuously updated salient tokens, enabling more accurate target representation. Specifically, TLCTrack incorporates three components: Masked Unidirectional Attention (MUA), Spatial Salient Token Collection (SSTC), and Temporal Salient Token Bank (TSTB) modules. By explicitly integrating spatio-temporal context, MUA extracts discriminative targetaware spatial features in the search region. To avoid the negative impact of background on feature learning, SSTC progressively suppresses background interference, thereby enhancing target spatial representation. Finally, TSTB captures high-quality spatio-temporal information through continuous salient token updates. Extensive experiments on five benchmarks demonstrate that our method achieves superior performance over state-of-the-art trackers. Code and models are available at https://github.com/xiading123/TLCTrack.

cs.CV

A Unified Approach to Interpretable Causal Root Cause Attribution

Understanding why a target metric changes is a fundamental problem in data-driven decision making, beyond anomaly detection alone. We study root cause attribution for metric changes in complex e-commerce systems, focusing on trade-offs between interpretability, efficiency, and causal validity. As a starting point, we extend a metric-decomposition method into a recursive metric-tree framework for multi-level root cause analysis, but this relies on independence and decomposability assumptions that miss complex causal dependencies. In contrast, graphical causal models (GCMs) relax these assumptions and improve causal validity, at the cost of interpretability, higher computational and data demands, and potential attribution target misalignment. Through real-world applications, mathematical proofs, and simulations, we characterize the fundamental sources of these trade-offs. Guided by these insights, we propose a unified, causally informed attribution approach that integrates structural causal information into the metric-tree decomposition framework and corrects key sources of misalignment in GCM-based causal attributions, substantially improving causal validity while preserving interpretability and fast computation. Analytical proofs and simulations demonstrate that the proposed approach produces more accurate root cause attributions, and we also present a real-world application.

stat.ME

TELEVAL: A Benchmark Designed for Spoken Language Models in Chinese Interactive Scenarios

Spoken Language Models (SLMs) are expected to support natural spoken interaction beyond task completion. However, existing SLM benchmarks primarily evaluate semantic correctness in structured settings and provide limited assessment of interactional behavior grounded in acoustic context. To address this gap, we introduce TELEVAL, a large-scale SLM benchmark for Chinese spoken interaction in instruction-free, audio-conditioned settings. TELEVAL evaluates two complementary aspects: (1) Reliable Content Fulfillment, which measures semantic accuracy of SLMs under diverse acoustic and linguistic conditions, and (2) Interactional Appropriateness, which assesses whether models produce natural and appropriate responses by implicitly grounding behavior in auditory cues. Experiments show that while models perform competitively on semantic tasks, their performance degrades under acoustic variability and in interactional settings. We observe consistent degradation from perceptual instability to interactional errors, and further identify a recurring failure pattern, termed the "Caption Trap", where models tend to describe perceived audio signals rather than produce appropriate interactive responses. These results indicate that current SLMs remain insufficiently aligned with the requirements of natural spoken interaction. TELEVAL provides a targeted framework for evaluating and analyzing interactional behavior in SLMs.

cs.CL

A Self-Triggered Agentic Push Recommendation System

Push notification is a critical recommendation scenario on large-scale platforms, allowing the system to proactively reach users outside the application to improve long-term re-engagement. However, designing an optimal push system requires handling a complex action space for the "whether and when" delivery problem under strict system resource constraints. Existing solutions typically fall into two passive paradigms: pre-planned frequency methods that allocate delivery times via offline modeling, limiting real-time adaptability; and fixed-interval triggering methods that periodically poll the system, creating a strict dilemma between excessive computational overhead and diminished optimal timing capture. Furthermore, such multi-stage frameworks severely suffer from local optima. To overcome these limitations, in this paper, we propose STEPS, a proactive, Self-Triggered End-to-end Agentic Push Recommendation System, which is already fully deployed at Douyin with over 1 billion users. STEPS reformulates push recommendation as a self-triggered agentic process in which the system decides not only whether to send a push, but also when to invoke itself again, thereby forming a closed loop that balances real-time effectiveness and efficiency. Specifically, STEPS consists of two decision transformer-based agents: a planning agent that schedules the next system invocation using a gated ordinal regression method, and an execution agent that decides whether to send a push based on trajectory rewards. Furthermore, we introduce a lightweight filtering agent to both control computational overhead and act as a crucial safeguard against unreasonable planning behaviors. Online A/B testing demonstrates that STEPS significantly increases user active days by 0.2843% and reduces the push permission disablement rate by 1.9089%, while the filtering agent reduces computational overhead by 79.42%.

cs.IR

Free-Running Waveguide-Integrated Single-Photon Avalanche Detectors for Visible Light

Waveguide-integrated single-photon avalanche detectors (SPADs) are essential components of integrated photonics platforms for scalable extreme-low-light applications without the use of cryogenics. Here, we demonstrate an integrated SPAD for visible light operating at room temperature in a free-running mode without gating. The device is based on a doped silicon diode end-fire-coupled to a silicon nitride (SiN) photonic integrated circuit (PIC). We investigate a range of lateral and vertical doping profile designs, and operate the devices with a simple current-mode passive quenching circuit. The optimal device is a laterally-doped p-i-n+ SPAD with a maximum photon detection efficiency (PDE) of 1.95 +/- 0.32% for input light at 685 nm wavelength, when reverse-biased at an excess of 1.5 V beyond the breakdown voltage of 15.0 V. We identify promising avenues for improving device performance, which would enable such integrated SPADs to be an attractive choice for cutting-edge integrated photonics solutions in quantum technologies, low-light imaging, and high-speed communications at visible wavelengths.

physics.optics

DBHN-Net: Dual-Branch Hybrid Neural Network For Low-Complexity Monaural Speech Enhancement

Although artificial neural network (ANN) based speech enhancement (SE) methods demonstrate excellent performance, the high computational complexity and high energy consumption hinder their deployment in practical front-end processing tasks.} Currently, the spiking neural networks (SNNs) have shown potential in reducing power consumption. However, the discrete binary activation and complex spatio-temporal dynamics of SNNs often result in information loss. The current challenge therefore focuses on how to maintain performance and reduce computational complexity. To address this issue, this work propose a Dual-Branch Hybrid Neural (DBHN) Network. 1) In terms of network architecture: A dual-branch network integrating ANN and SNN was designed, where the SNN branch reduces power consumption while the ANN branch addresses information loss; The BandSplit and Time-Frequency (TF) -Mamba modules were developed to simultaneously compress energy consumption and enhance model performance; Spiking Feature Extraction Group (SFEG) and Information Transformation Block (ITB) components were implemented with residual connections to mitigate information loss while further refining feature representations. 2) To facilitate inter-branch information fusion: An Interaction module was designed to promote information exchange at various stages of the dual-branch network; A TF-Cross Attention-Fusion module was designed to perform time-frequency domain fusion of dual-branch information while data-adaptively guiding the SNN branch to retain more critical information. Results show that the proposed model maintains superior performance across three public datasets while achieving an average 7.5 fold reduction in computational complexity compared to baseline models.

cs.SD

GLEAM: A Multimodal Imaging Dataset and HAMM for Glaucoma Classification

We propose glaucoma lesion evaluation and analysis with multimodal imaging (GLEAM), the first publicly available tri-modal glaucoma dataset comprising scanning laser ophthalmoscopy fundus images, circumpapillary OCT images, and visual field pattern deviation maps, annotated with four disease stages, enabling effective exploitation of multimodal complementary information and facilitating accurate diagnosis and treatment across disease stages. To effectively integrate cross-modal information, we propose hierarchical attentive masked modeling (HAMM) for multimodal glaucoma classification. Our framework employs hierarchical attentive encoders and light decoders to focus cross-modal representation learning on the encoder.

eess.IV

TabCF: Distributional Control Function Estimation with Tabular Foundation Models

Instrumental variable (IV) and control function (CF) methods are powerful tools for causal effect estimation in the presence of unmeasured confounding, yet most existing approaches target only mean effects and/or demand substantial fitting and tuning effort. In this paper, we introduce a simple method, TabCF, for control function regression using tabular foundation models, which enables accurate, fast, identification-transparent, and tuning-light causal estimation of distributional quantities, such as interventional means and quantiles; we also propose a copula-based approximation for multivariate outcomes. TabCF performs favorably against representative methods across a broad range of small- to medium-sized synthetic and real data scenarios. The central message is two-fold: for practitioners, it highlights that TabCF is an effective tool for distributional causal inference; for researchers, it suggests that the proposed approach could be considered a strong baseline for future method development. Code is available at https://github.com/GepingChen/TabCF.

stat.ML

Detecting Breast Carcinoma Metastasis on Whole-Slide Images by Partially Subsampled Multiple Instance Learning

Breast cancer is the most prevalent cancer in women worldwide. Histopathology image analysis serves as the gold standard for cancer diagnosis. In this regard, whole-slide imaging (WSI), a revolutionary technology in digital pathology, allows for ultrahigh-resolution tissue analysis. Despite its promise, WSI analysis faces significant computational challenges due to its massive data size and tissue heterogeneity. To address this issue, we present a Gaussian mixture based multiple instance learning (MIL) framework for WSI analysis with partially subsampled instances. Our approach models a WSI as a bag of instances (i.e., randomly cropped sub-images), leveraging a bag-based maximum likelihood estimator (BMLE) to predict metastases. Furthermore, we introduce a subsampling-based maximum likelihood estimator (SMLE) to refine predictions by selectively labeling a subset of instances. Extensive evaluations of the breast carcinoma metastasis prediction demonstrate that BMLE surpasses state-of-the-art methods, while the SMLE further improves the prediction accuracy at both bag and instance levels. We find that our method is fairly robust against various plausible model mis-specifications. Theoretical analyses and simulation studies validate the performance and robustness of our methods.

stat.ME

Hypothesis Testing for Penalized Estimating Equations with Cross-Fitted Covariance Calibration

We study hypothesis testing for penalized estimators in settings where the full marginal distribution of a multivariate response is difficult to specify, such as longitudinal data with correlated measurements or high-dimensional heteroscedastic regression. Assuming that the conditional mean model is correctly specified, we establish that the penalized estimating equations admit a $\sqrt{n}$-consistent solution, even when the working covariance structure is misspecified. Our inferential target is a low-dimensional subvector of parameters associated with the mean model. We show that the resulting test statistic converges to a $χ^2$ distribution, and that its asymptotic power depends on the nuisance covariance function. To mitigate this dependence, we propose estimating the covariance function via cross-fitting, which provides a calibrated and robust procedure for inference.

stat.ME

A Bayes-Motivated Quadratic-Form Test for High-Dimensional Mean Testing

We propose a two-sample mean test based on the Bayes factor with non-informative priors, specifically designed for scenarios where the dimension $p$ grows with the sample size $n$ with a linear rate $p/n \to c_1 \in (0, \infty)$. We establish the asymptotic normality of the test statistic and the asymptotic power. Through extensive simulations, we demonstrate that the proposed test performs competitively against several existing methods, particularly when the marginal variances of the individual features are heterogeneous and when the sample size is small. Furthermore, our test remains robust under distribution misspecification. The proposed method not only effectively detects both sparse and non-sparse differences in mean vectors but also maintains a well-controlled type I error rate, even in small-sample scenarios. We also demonstrate the performance of our proposed test using the small round blue cell tumors (SRBCT) dataset.

stat.ME

Universal scaling laws for dynamical-thermal hysteresis

Dynamic hysteresis, the rate-dependent lagged response of materials to external fields, underpins applications from energy-efficient transformers to gas storage systems. A fundamental yet unresolved question is how the hysteresis loop area $A$ scales with the field sweep rate $R$. Here, we reveal that a competition between the field sweep and thermal fluctuations governs a universal crossover between two scaling regimes: $A - A_0 \propto R^{1/3}$ for $R < R^*$ and $A - A_0 \propto R^{2/3}$ for $R > R^*$, where $A_0$ is the quasi-static area and the crossover rate $R^* \propto T/T_c$ depends on the temperature $T$ and the material's critical temperature $T_c$. We demonstrate these scaling laws universally across experiments of magnetic materials, simulations of Ising and metal-organic framework models, and analytical solutions of a stochastic Langevin equation. This framework not only resolves the long-standing non-universality of reported scaling exponents but also provides a direct design principle for the application of dynamic hysteresis.

cond-mat.stat-mech

Discovery of a radio jet in the Cloverleaf Quasar at z = 2.56

The fast growth of supermassive black holes and their feedback to the host galaxies play an important role in regulating the evolution of galaxies, especially in the early Universe. However, due to cosmological dimming and the limited angular resolution of most observations, it is difficult to resolve the feedback from the active galactic nuclei (AGNs) to their host galaxies. Gravitational lensing, for its magnification, provides a powerful tool to spatially differentiate emission originating from AGN and host galaxy at high redshifts. Here we report a discovery of a jet-like radio structure in a strongly lensed starburst quasar, H1413+117 or Cloverleaf at redshift z= 2.56, based on observational data at optical, sub-millimetre, and radio wavelengths. With both parametric and non-parametric lens models and with reconstructed images in the source plane, we find a well-separated, kpc-scaled, single-sided radio jet located at projected ~1.2 kpc to the northwest of the host galaxy in the source plane. This could indicate the co-existence of feedback from the AGN by both wind and jet in the Cloverleaf quasar.

astro-ph.GA

HopChain: Multi-Hop Data Synthesis for Generalizable Vision-Language Reasoning

Vision-language models (VLMs) show strong multimodal capabilities but still struggle with fine-grained vision-language reasoning. We find that long chain-of-thought (CoT) reasoning exposes diverse failure modes, including perception, reasoning, knowledge, and hallucination errors, which can compound across intermediate steps. However, most existing vision-language data used for reinforcement learning with verifiable rewards (RLVR) does not involve complex reasoning chains that rely on visual evidence throughout, leaving these weaknesses largely unexposed. We therefore propose HopChain, a scalable framework for synthesizing multi-hop vision-language reasoning data for RLVR training of VLMs. Each synthesized multi-hop query forms a logically dependent chain of instance-grounded hops, where earlier hops establish the instances, sets, or conditions needed for later hops, while the final answer remains a specific, unambiguous number suitable for verifiable rewards. We train Qwen3.5-35B-A3B and Qwen3.5-397B-A17B under two RLVR settings: the original data alone, and the original data plus HopChain's multi-hop data, and compare them across 24 benchmarks spanning STEM and Puzzle, General VQA, Text Recognition and Document Understanding, and Video Understanding. Although this multi-hop data is not synthesized for any specific benchmark, it improves 20 of 24 benchmarks on both models, indicating broad and generalizable gains. Consistently, replacing full chained queries with half-multi-hop or single-hop variants reduces the average score across five representative benchmarks from 70.4 to 66.7 and 64.3, respectively. Notably, multi-hop gains peak in long-CoT vision-language reasoning, exceeding 50 points in the ultra-long-CoT regime. These experiments establish HopChain as an effective, scalable framework for synthesizing multi-hop data that improves generalizable vision-language reasoning.

cs.CV

Revealing Behavioral Plasticity in Large Language Models: A Token-Conditional Perspective

In this work, we reveal that Large Language Models (LLMs) possess intrinsic behavioral plasticity-akin to chameleons adapting their coloration to environmental cues-that can be exposed through token-conditional generation and stabilized via reinforcement learning. Specifically, by conditioning generation on carefully selected token prefixes sampled from responses exhibiting desired behaviors, LLMs seamlessly adapt their behavioral modes at inference time (e.g., switching from step-by-step reasoning to direct answering) without retraining. Based on this insight, we propose Token-Conditioned Reinforcement Learning (ToCoRL), a principled framework that leverages RL to internalize this chameleon-like plasticity, transforming transient inference-time adaptations into stable and learnable behavioral patterns. ToCoRL guides exploration with token-conditional generation and keep enhancing exploitation, enabling emergence of appropriate behaviors. Extensive experiments show that ToCoRL enables precise behavioral control without capability degradation. Notably, we show that large reasoning models, while performing strongly on complex mathematics, can be effectively adapted to excel at factual question answering, which was a capability previously hindered by their step-by-step reasoning patterns.

cs.CL

New reformulations for 0-1 quadratic programming problem using quadratic nonconvex reformulation techniques and valid inequalities

It is well-known that the quadratic convex reformulation (QCR) technique can speed up some general-purpose solvers such as CPLEX and Gurobi. Recently, the method of quadratic nonconvex reformulation (QNR) was proposed, which provides an alternative way for accelerating a solver via reformulation technique. This paper proposes several new reformulations for 0-1 quadratic programming problems using the QNR technique. Such a technique provides more flexibility in adding nonconvex quadratic constraints into the problem formulation, so that some valid inequalities, such as the triangle inequalities, can be incorporated into the formulation to tighten the lower bound of the problem. We analyze the effects of the proposed reformulations on the lower bounds implemented in the solver, and propose some methods to maximize the McCormick relaxation bounds of the reformulations. Our numerical experiments compare the proposed reformulations with the existing quadratic convex reformulations, showing the effectiveness of the proposed reformulations on 0-1 quadratic programming problems.

math.OC

An Information-Theoretic Framework for Receiver Quantization in Communication

We investigate information-theoretic limits and design of communication under receiver quantization. Unlike most existing studies, this work is more focused on the impact of resolution reduction from high to low. We consider a standard transceiver architecture, which includes i.i.d. complex Gaussian codebook at the transmitter, and a symmetric quantizer cascaded with a nearest neighbor decoder at the receiver. Employing the generalized mutual information (GMI), an achievable rate under general quantization rules is obtained in an analytical form, which shows that the rate loss due to quantization is $\log\left(1+γ\mathsf{SNR}\right)$, where $γ$ is determined by thresholds and levels of the quantizer. Based on this result, the performance under uniform receiver quantization is analyzed comprehensively. We show that the front-end gain control, which determines the loading factor of quantization, has an increasing impact on performance as the resolution decreases. In particular, we prove that the unique loading factor that minimizes the MSE also maximizes the GMI, and the corresponding irreducible rate loss is given by $\log\left(1+\mathsf {mmse}\cdot\mathsf{SNR}\right)$, where mmse is the minimum MSE normalized by the variance of quantizer input, and is equal to the minimum of $γ$. A geometrical interpretation for the optimal uniform quantization at the receiver is further established. Moreover, by asymptotic analysis, we characterize the impact of biased gain control, showing how small rate losses decay to zero and providing rate approximations under large bias. From asymptotic expressions of the optimal loading factor and mmse, approximations and several per-bit rules for performance are also provided. Finally we discuss more types of receiver quantization and show that the consistency between achievable rate maximization and MSE minimization does not hold in general.

cs.IT

High-dimensional Newey-Powell Test Via Approximate Message Passing

We propose a high-dimensional extension of the heteroscedasticity test proposed in Newey and Powell (1987). Our test is based on expectile regression in the proportional asymptotic regime where n/p \to δ\in (0,1]. The asymptotic analysis of the test statistic uses the Approximate Message Passing (AMP) algorithm, from which we obtain the limiting distribution of the test and establish its asymptotic power. The numerical performance of the test is validated through an extensive simulation study. As real-data applications, we present the analysis based on ``international economic growth" data (Belloni et al., 2011), which is found to be homoscedastic, and ``supermarket" data (Lan et al., 2016), which is found to be heteroscedastic.

stat.ME