Search arXiv⌕ Search

arXiv subjects

Xiyuan Zhang

Publications and source records attributed to Xiyuan Zhang.

At least 19 recordsLinked to original sources

TCSAlgBench: Benchmarking Automated Proving for Research-Level Theoretical Computer Science

Large language models perform strongly on competition mathematics, but their research-level reasoning remains difficult to evaluate systematically. Theoretical computer science (TCS) connects algorithm design to explicit guarantees and fundamental limits, providing a setting for evaluating whether models can justify computational improvements with arguments humans can inspect. We introduce TCSAlgBench, a benchmark and reusable pipeline for natural-language proof discovery, comprising 398 theorem-level challenges from 138 STOC and COLT 2026 papers. Expert-designed rules complete paper-specific context, preserve computational assumptions and quantitative guarantees, and withhold constructions when discovering an algorithm is part of the task. For each task, prover systems receive theorem statements and access to cited prior work. The pipeline supports fresh, versioned challenge batches from newly released papers. We evaluate ten model configurations from four families under direct inference and prover-verifier discussion, and compare four agent workflows under matched model-call opportunities. All evaluations use the full benchmark. In the model comparison, GPT-5.6 Sol max achieves the highest five-run verifier-accepted coverage at 23.6% after 10-round discussion. Discussion and repeated sampling improve coverage. In the separate agent comparison using GPT-5.5 xhigh, decomposition improves coverage over discussion, and agentic planning achieves the highest five-run verifier-accepted coverage at 25.4%. TCSAlgBench provides a refreshable testbed for measuring progress in model reasoning and studying how agent workflows support research-level proof discovery.

cs.AI↗

Mitra-v2 Technical Report

We introduce Mitra-v2, a tabular foundation model that delivers state-of-the-art performance on real-world classification and regression problems, from credit-risk scoring and clinical prediction to equipment-failure detection and house-price estimation. Mitra-v2 is trained only on synthetic data, with a pretraining distribution that is much larger and more diverse than Mitra-v1's. Built on a small 2D Transformer backbone, Mitra-v2 supports longer contexts and larger feature spaces. Improved optimization lets it learn from this larger task distribution. We evaluate Mitra-v2 on the TabArena and TALENT benchmarks, comprising more than 300 real-world datasets under two evaluation protocols. On the full TabArena benchmark, Mitra-v2 delivers state-of-the-art performance at the level of the industry-scale TabFM and EXAONE Tabular models, while surpassing TabPFN-3 by a wide margin in both classification and regression. Mitra-v2 matches the 1.6B-parameter TabFM with only 5% of its size (77M parameters), delivering frontier performance at a fraction of the cost. On TALENT, Mitra-v2 remains among the leading models, clearly outperforming TabPFN-3 and TabICLv2. It also ranks first on classification tasks with more than ten classes, even though it was pretrained only on tasks with at most ten classes. These results make Mitra-v2 one of the strongest and most broadly applicable open tabular foundation models released to date. We release the model weights, the inference and fine-tuning code, and our evaluation results under the Apache-2.0 license.

cs.LG↗

Rise Time and Charge Collection Efficiency of Graphene-Optimized 4H-SiC p-i-n Detector

Silicon carbide detectors exhibit good detection performance and are being considered for detection applications. However, the presence of surface electrode of detector limits the application of low-penetration particle detectors, photodetectors and heavy-ion detection. A graphene-optimized 4H-SiC detector has been fabricated to expand the application of SiC detectors.Its electrical properties and the charge collection performance of α particles are reported. The effective doping concentration of lightly doped 4H-SiC epitaxial layer is about 4.5\times10^{13}cm^{-3}, approaching the limit of the lowest doping level by the SiC epitaxial growth technique. The rise time of the graphene-optimized ring electrode detector is reduced by 24% at 200 V, compared to ring electrode detector. The charge collection efficiency (CCE) of graphene-optimized 4H-SiC PIN is 99.22%. When the irradiation dose is 2\times10^{11} n_{eq}/cm^2, the irradiation has no significant impact on the rise time and uniformity of the rise time for the graphene-optimized 4H-SiC detectors. This study proves that graphene has a certain radiation resistance. Graphene-optimized 4H-SiC detectors can not only reduce the signal rise time, but also improve uniformity of signal rise time and stability of charge collection. This research will expand the application of graphene-based 4H-SiC detectors in fields such as low energy ions, X-ray, UV light detection, particle physics, medical dosimetry and heavy-ion detection.

physics.ins-det↗

Stable Charge Collection and Sub-45 ps Time Resolution in a 4H-SiC PIN Detector Irradiated With Low Fluence 16.5 MeV/u Ta Ions

A silicon carbide PIN detector was fabricated and its radiation tolerance under Ta heavy ion irradiation of 2370 MeV was evaluated. Its electrical properties, charge collection performance and time resolution of $β$-particles ($^{90}$Sr) are reported. The leakage currents for unirradiated and irradiated 4H-SiC PIN detectors are $1.47 \times 10^{-10}$~A @ 300 V and 1.49~$\times$ 10$^{-10}$A@ 300 V. The effective doping concentrations for unirradiated and irradiated 4H-SiC PIN detectors are $6.23\times 10^{13}$~cm$^{-3}$ and $6.13\times 10^{13}$~cm$^{-3}$. The irradiated detector exhibits good electrical performance and stable device architecture. The 4H-SiC PIN detector exhibits a charge collection efficiency (CCE) of 99.24\% under Ta Heavy Ion Irradiation. The time resolutions of the detector before and after irradiation are 40 ps and 45 ps, respectively. Experimental results indicate that the CCE and time resolution performance exhibit good stability before and after irradiation. These results demonstrate stable performance under Ta heavy ion irradiation, highlighting the detectors potential for radiation-hard applications in high-energy physics, space missions, and nuclear reactor monitoring.

physics.ins-det↗

Stability of Charge Collection Efficiency in a Novel Graphene-Optimized Silicon Carbide Detector Under 160 keV X-Ray Irradiation

A novel graphene-optimized silicon carbide PIN detector was fabricated. Its electrical properties, charge collection performance and signal rise time were evaluated under non-irradiated conditions and under X-ray irradiation with an energy of 160 keV at doses of 0.1 MGy and 1 MGy. The leakage currents of the detectors under non-irradiated, 0.1 MGy, and 1 MGy irradiation conditions are approximately 1.45e-10 A, 1.51e-10 A, and 1.57e-10 A, respectively. The effective doping concentration of the detector is approximately 8.08e13 cm^-3 before and after irradiation, with no significant change. The rise times of the signals from alpha particles signal detected by the detector under unirradiated, 0.1 MGy, and 1 MGy X-ray irradiation conditions are 336 ps, 368 ps, and 387 ps, respectively. The rise times of the beta particles signal detected by the detector under unirradiated, 0.1 MGy, and 1 MGy X-ray irradiation conditions are 342 ps, 375 ps, and 398 ps, respectively. After 0.1 MGy and 1 MGy X-ray irradiation, the charge collection efficiencies (CCEs) of the detector for alpha particles are 97.2% and 90.0%, respectively; for beta particles, they are 100.0% and 97.0%, respectively. Experiments confirm that 160 keV X-ray irradiation may not cause significant displacement damage in the 4H-SiC, and the minor performance degradation may be attributed to ionization induced changes in the graphene electrode. The detector exhibits excellent charge collection performance and fast time response. These results demonstrate stable performance under extreme X-ray exposure, highlighting the detector's potential for radiation-hard applications in high-energy physics, space missions, and nuclear reactor monitoring.

physics.ins-det↗

Stability of Charge Collection Efficiency and Time Resolution in a Novel Ultra-fast Graphene-Optimized Silicon Carbide Detector Under X-ray Irradiation

A graphene-optimized silicon carbide PIN detector was fabricated and its radiation tolerance under X-ray irradiation of 160 keV was evaluated. Its electrical properties, charge collection performance and time resolution of beta-particles (90Sr) are reported. After 1 MGy irradiation, the detector maintains an ultralow leakage current of approximately 2.2e-10 A @ 300 V and the C-V characteristics are basically consistent with full depletion at 120V. The time resolution of the graphene-optimized silicon carbide detector is 58.0 ps. The time resolution is comparable to that of state-of-the-art 4H-SiC low-gain avalanche detectors (LGADs). The G/RE 4H-SiC PIN detector exhibits outstanding time resolution performance. Compared with the time resolution of the RE 4H-SiC PIN detector, the time resolution of the G/RE 4H-SiC PIN detector has decreased by 39.6%. This demonstrates the significance of the graphene electrode design. The graphene detector exhibits a charge collection efficiency (CCE) of 99.24% after X-ray irradiation, along with excellent stability. The graphene-optimized silicon carbide detector maintains good timing resolution: 58.0ps before and 64.0ps after X-ray irradiation. Experimental results indicate that the CCE and time resolution performance exhibit good stability before and after irradiation. These results demonstrate stable performance under extreme X-ray exposure, highlighting the detectors potential for radiation-hard applications in high-energy physics, space missions, and nuclear reactor monitoring.

physics.ins-det↗

MCGA: A Multi-task Classical Chinese Literary Genre Audio Corpus

With the rapid advancement of Multimodal Large Language Models (MLLMs), their potential has gained significant attention in Chinese Classical Studies (CCS). While existing research primarily focuses on text and visual modalities, the audio corpus within this domain remains largely underexplored. To bridge this gap, we introduce the Multi-task Classical Chinese Literary Genre Audio Corpus (MCGA), a 119-hour corpus comprising 22,000 audio samples. It encompasses a diverse range of literary genres across six tasks: Automatic Speech Recognition (ASR), Speech-to-Text Translation (S2TT), Speech Emotion Captioning (SEC), Spoken Question Answering (SQA), Speech Understanding (SU), and Speech Reasoning (SR). Through the evaluation of ten MLLMs, our experimental results demonstrate that current MLLMs still face substantial challenges on the MCGA test set. Furthermore, we introduce a domain-specific metric for SEC and a metric to measure the consistency between speech and text capabilities. We release MCGA to the public to facilitate the development of more robust MLLMs. MCGA Corpus: https://github.com/yxduir/MCGA

cs.CL↗

HetRL: Efficient Reinforcement Learning for LLMs in Heterogeneous Environments

As large language models (LLMs) continue to scale and new GPUs are released even more frequently, there is an increasing demand for LLM post-training in heterogeneous environments to fully leverage underutilized mid-range or previous-generation GPUs and alleviate the shortage of homogeneous high-end GPUs within a single availability zone. However, achieving high-performance reinforcement learning (RL) training for LLMs on such computing resources remains challenging because the workflow involves multiple models and tasks with complex computation and data dependencies. In this paper, we present HetRL, a distributed system for efficient RL training in infrastructures with heterogeneous GPUs and networks. HetRL formulates the scheduling of RL training in heterogeneous environments as a constrained joint optimization problem and provides two complementary approaches for addressing this problem: (1) a hybrid scheduling algorithm that efficiently identifies near-optimal solutions, and (2) an integer linear programming (ILP)-based scheduling algorithm that obtains optimal solutions, enabling flexible trade-offs between solution optimality and efficiency. Our extensive evaluation, consuming 20,000 GPU-hours, shows that HetRL achieves up to 9.17x the throughput of state-of-the-art systems, and 3.17x on average, across a wide range of workloads and settings.

cs.DC↗

A Telescope System for Charge and Position Measurement of High Energy Nuclei

A high-granularity telescope system with a large sensitive area and low material budget has been developed for high-energy heavy ion beam tests. The telescope consists of nine layers of silicon microstrip detectors (SSDs), whose performance was validated through a heavy ion beam test at the CERN SPS. A hybrid machine learning algorithm is proposed to address the challenges of nuclear charge measurement with SSDs. The system achieves a spatial resolution of $\mathcal{O}(1) \,$\SI{}{\micro\metre} and a charge resolution better than 0.16 charge units for nuclei from $Z = 1$ to $Z = 29$, with a sensitive area of $8 \times 8 \, \mathrm{cm}^2$. To the best of our knowledge, this represents the most precise charge and spatial resolution simultaneously achieved by a silicon telescope to date.

physics.ins-det↗

Beam Test Characterization of Silicon Microstrip Detector Flight-Model Ladders for the AMS-02 Upgrade

The AMS-02 experiment plans to install a new silicon microstrip tracker layer (Layer-0) on top of the existing detector, increasing the cosmic-ray acceptance by a factor of 3. Layer-0 employs a design in which multiple silicon microstrip detectors (SSDs) are connected in series to form long detector ladders. We present a detailed performance study of the flight-model ladders using a 350~GeV mixed hadron beam at the CERN SPS. The study focuses on the following aspects: (i) the performance of ladders with different numbers of SSDs, for which the intrinsic spatial resolution at normal incidence varies from $9.5~μ\mathrm{m}$ to $11.4~μ\mathrm{m}$ for ladders composed of 8 to 12 SSDs; (ii) the response consistency for particles impacting on the \emph{Head} and \emph{Tail} regions of the ladder; and (iii) the dependence of the detector performance on the particle incidence angle.

physics.ins-det↗

Stability of Charge Collection Efficiency and Time Resolution in 4H-SiC PIN Diodes Under X-ray Irradiation

This study evaluates the radiation tolerance of a 4H-SiC PIN detector under X-ray irradiation up to \SI{2}{MGy} (Si) at \SI{160}{keV}. The detector features a fully epitaxial vertical PIN structure with mesa terminations and field plates. Comprehensive pre- and post-irradiation characterization includes I-V/C-V measurements, charge collection efficiency (CCE) and timing resolution tests using $β$-particles ($^{90}$Sr). After \SI{2}{MGy} irradiation, the reverse leakage current remains at an ultralow level of $\sim 10^{-11}$ \si{A/cm^2} at \SI{-300}{V} with negligible degradation. C-V characteristics are basically consistent, with full depletion at \SI{~130}{V}. CCE for $β$-particles decreases by less than 5\%. The detector maintains good timing resolution: \SI{21}{ps} before and \SI{31}{ps} after irradiation, with jitter increasing moderately. These results demonstrate stable performance under extreme X-ray exposure, highlighting the detector's potential for radiation-hard applications in high-energy physics, space missions, and nuclear reactor monitoring.

physics.ins-det↗

FIRE: A Comprehensive Benchmark for Financial Intelligence and Reasoning Evaluation

We introduce FIRE, a comprehensive benchmark designed to evaluate both the theoretical financial knowledge of LLMs and their ability to handle practical business scenarios. For theoretical assessment, we curate a diverse set of examination questions drawn from widely recognized financial qualification exams, enabling evaluation of LLMs deep understanding and application of financial knowledge. In addition, to assess the practical value of LLMs in real-world financial tasks, we propose a systematic evaluation matrix that categorizes complex financial domains and ensures coverage of essential subdomains and business activities. Based on this evaluation matrix, we collect 3,000 financial scenario questions, consisting of closed-form decision questions with reference answers and open-ended questions evaluated by predefined rubrics. We conduct comprehensive evaluations of state-of-the-art LLMs on the FIRE benchmark, including XuanYuan 4.0, our latest financial-domain model, as a strong in-domain baseline. These results enable a systematic analysis of the capability boundaries of current LLMs in financial applications. We publicly release the benchmark questions and evaluation code to facilitate future research.

cs.AI↗

SenTSR-Bench: Thinking with Injected Knowledge for Time-Series Reasoning

Time-series diagnostic reasoning is essential for many applications, yet existing solutions face a persistent gap: general reasoning large language models (GRLMs) possess strong reasoning skills but lack the domain-specific knowledge to understand complex time-series patterns. Conversely, fine-tuned time-series LLMs (TSLMs) understand these patterns but lack the capacity to generalize reasoning for more complicated questions. To bridge this gap, we propose a hybrid knowledge-injection framework that injects TSLM-generated insights directly into GRLM's reasoning trace, thereby achieving strong time-series reasoning with in-domain knowledge. As collecting data for knowledge injection fine-tuning is costly, we further leverage a reinforcement learning-based approach with verifiable rewards (RLVR) to elicit knowledge-rich traces without human supervision, then transfer such an in-domain thinking trace into GRLM for efficient knowledge injection. We further release SenTSR-Bench, a multivariate time-series-based diagnostic reasoning benchmark collected from real-world industrial operations. Across SenTSR-Bench and other public datasets, our method consistently surpasses TSLMs by 9.1%-26.1% and GRLMs by 7.9%-22.4%, delivering robust, context-aware time-series diagnostic insights.

cs.LG↗

Mitra: Mixed Synthetic Priors for Enhancing Tabular Foundation Models

Since the seminal work of TabPFN, research on tabular foundation models (TFMs) based on in-context learning (ICL) has challenged long-standing paradigms in machine learning. Without seeing any real-world data, models pretrained on purely synthetic datasets generalize remarkably well across diverse datasets, often using only a moderate number of in-context examples. This shifts the focus in tabular machine learning from model architecture design to the design of synthetic datasets, or, more precisely, to the prior distributions that generate them. Yet the guiding principles for prior design remain poorly understood. This work marks the first attempt to address the gap. We systematically investigate and identify key properties of synthetic priors that allow pretrained TFMs to generalize well. Based on these insights, we introduce Mitra, a TFM trained on a curated mixture of synthetic priors selected for their diversity, distinctiveness, and performance on real-world tabular data. Mitra consistently outperforms state-of-the-art TFMs, such as TabPFNv2 and TabICL, across both classification and regression benchmarks, with better sample efficiency.

cs.LG↗

Understanding the Implicit Biases of Design Choices for Time Series Foundation Models

Time series foundation models (TSFMs) are a class of potentially powerful, general-purpose tools for time series forecasting and related temporal tasks, but their behavior is strongly shaped by subtle inductive biases in their design. Rather than developing a new model and claiming that it is better than existing TSFMs, e.g., by winning on existing well-established benchmarks, our objective is to understand how the various ``knobs'' of the training process affect model quality. Using a mix of theory and controlled empirical evaluation, we identify several design choices (patch size, embedding choice, training objective, etc.) and show how they lead to implicit biases in fundamental model properties (temporal behavior, geometric structure, how aggressively or not the model regresses to the mean, etc.); and we show how these biases can be intuitive or very counterintuitive, depending on properties of the model and data. We also illustrate in a case study on outlier handling how multiple biases can interact in complex ways; and we discuss implications of our results for learning the bitter lesson and building TSFMs.

cs.LG↗

Chronos-2: From Univariate to Universal Forecasting

Pretrained time series models have enabled inference-only forecasting systems that produce accurate predictions without task-specific training. However, existing approaches largely focus on univariate forecasting, limiting their applicability in real-world scenarios where multivariate data and covariates play a crucial role. We present Chronos-2, a pretrained model capable of handling univariate, multivariate, and covariate-informed forecasting tasks in a zero-shot manner. Chronos-2 employs a group attention mechanism that facilitates in-context learning (ICL) through efficient information sharing across multiple time series within a group, which may represent sets of related series, variates of a multivariate series, or targets and covariates in a forecasting task. These general capabilities are achieved through training on synthetic datasets that impose diverse multivariate structures on univariate series. Chronos-2 delivers state-of-the-art performance across three comprehensive benchmarks: fev-bench, GIFT-Eval, and Chronos Benchmark II. On fev-bench, which emphasizes multivariate and covariate-informed forecasting, Chronos-2's universal ICL capabilities lead to substantial improvements over existing models. On tasks involving covariates, it consistently outperforms baselines by a wide margin. Case studies in the energy and retail domains further highlight its practical advantages. The in-context learning capabilities of Chronos-2 establish it as a general-purpose forecasting model that can be used "as is" in real-world forecasting pipelines.

cs.LG↗

Understanding Transformers for Time Series: Rank Structure, Flow-of-ranks, and Compressibility

Transformers are widely used across data modalities, and yet the principles distilled from text models often transfer imperfectly to models trained to other modalities. In this paper, we analyze Transformers through the lens of rank structure. Our focus is on the time series setting, where the structural properties of the data differ remarkably from those of text or vision. We show that time-series embeddings, unlike text or vision, exhibit sharply decaying singular value spectra: small patch sizes and smooth continuous mappings concentrate the data into low-rank subspaces. From this, we prove that the associated $Q/K/V$ projections admit accurate low-rank approximations, and that attention layers become compressible in proportion to the decay of the embedding spectrum. We introduce the concept of flow-of-ranks, a phenomenon by which nonlinear mixing across depth inflates the rank, explaining why early layers are most amenable to compression and why ranks grow with depth. Guided by these theoretical and empirical results, we use these insights to compress Chronos, a large time series foundation model, achieving a reduction of $65\%$ in inference time and $81\%$ in memory, without loss of accuracy. Our findings provide principled guidance for allocating width, depth, and heads in time series foundation models, and for exploiting their inherent compressibility.

cs.LG↗

When Does Multimodality Lead to Better Time Series Forecasting?

Recently, there has been growing interest in incorporating textual information into foundation models for time series forecasting. However, it remains unclear whether and under what conditions such multimodal integration consistently yields gains. We systematically investigate these questions across a diverse benchmark of 16 forecasting tasks spanning 7 domains, including health, environment, and economics. We evaluate two popular multimodal forecasting paradigms: aligning-based methods, which align time series and text representations; and prompting-based methods, which directly prompt large language models for forecasting. Our findings reveal that the benefits of multimodality are highly condition-dependent. While we confirm reported gains in some settings, these improvements are not universal across datasets or models. To move beyond empirical observations, we disentangle the effects of model architectural properties and data characteristics, drawing data-agnostic insights that generalize across domains. Our findings highlight that on the modeling side, incorporating text information is most helpful given (1) high-capacity text models, (2) comparatively weaker time series models, and (3) appropriate aligning strategies. On the data side, performance gains are more likely when (4) sufficient training data is available and (5) the text offers complementary predictive signal beyond what is already captured from the time series alone. Our study offers a rigorous, quantitative foundation for understanding when multimodality can be expected to aid forecasting tasks, and reveals that its benefits are neither universal nor always aligned with intuition.

cs.CL↗