Search arXivSearch

arXiv subjects

Yuting Wang

Publications and source records attributed to Yuting Wang.

At least 19 recordsLinked to original sources

A Gaussian Covariance Matrix for Joint Pre- and Post-Reconstruction Full-Shape Power Spectrum Analysis

We apply the Gaussian covariance formalism to develop a semi-analytical covariance model for the joint analysis of pre-reconstruction, post-reconstruction, and cross full-shape galaxy power spectra. We model the reconstruction-reduced, scale-dependent cross shot noise using displacement-field statistics and introduce a new estimator that directly measures this term. Using the measured power spectra and the modeled shot-noise predictions as inputs, we construct the Gaussian covariance while accounting for correlations between the pre- and post-reconstruction density fields. We validate the resulting semi-analytical Gaussian covariance against mock catalogues. Using emulator-based parameter inference, we demonstrate that the semi-analytical Gaussian covariance adequately captures the dominant contribution to the covariance structure of the full data vector ($P_{\ell}^{\rm pre}, P_{\ell}^{\rm post}, P_{\ell}^{\rm cross}$). For the joint fit to these three power spectra, it yields cosmological constraints consistent with those obtained using the mock-based numerical covariance over the adopted fitting ranges: $k_{\rm max}=0.18\,h\,{\rm Mpc}^{-1}$ for $P_{\rm pre}$ and $P_{\rm post}$, and $k_{\rm max}=0.12\,h\,{\rm Mpc}^{-1}$ for $P_{\rm cross}$.

astro-ph.CO

Chinese Competitive Debating Dataset and Benchmark

Debate adjudication requires tracking how arguments develop through interaction, yet existing datasets rarely combine fine-grained debate transcripts with professional judgments collected during real competitions under a shared rubric. We introduce a dataset and benchmark for evaluating large language models' understanding of competitive Chinese-language debate at the match, stage, and speaker levels. We organized 182 matches and recruited 120 professional judges, with each match independently adjudicated by three judges using a predefined rubric. After excluding matches with incomplete records, the dataset contains 148 matches, 2,698 stages, and 20,542 exchange units, with manually verified transcripts and segmentation. It preserves original stage scores, match votes, best-debater ballots, and adjudication rationales. We define three tasks: winner-tendency prediction, stage-score prediction, and best-debater prediction. Zero-shot evaluation of multiple large language models yields a highest winner-prediction accuracy of 66.2%, a highest Pearson correlation of 0.250 between model stage scores and mean human ratings, and a highest best-debater prediction accuracy of 56.8%. The dataset and benchmark provide a testbed for studying large language models' understanding of interactive argumentation and their agreement with professional judges.

cs.CL

Aggregating many estimators using estimated weights

Consider an increasing number of consistent estimators to be averaged when only estimated weights are available. The underlying parameter of interest can be identical across estimators (homogeneity) or not (heterogeneity). The contribution of the paper is threefold. First, it is shown that the interaction of the estimated weights with the estimators can generate specific bias terms. This constrains the number of estimators that can be aggregated when weight estimation is ignored in inference. Second, the paper proposes estimated adaptive weights, which allow for standard Gaussian inference in a uniform manner and are asymptotically optimal both under homogeneity and heterogeneity. Third, conditions ensuring the validity of the Cochran (1937) Q test of homogeneity with estimated variance are given.

econ.EM

Periodicity-driven revision of the phase diagram of the generalized Baxter-Wu model with asymmetric complex couplings

The conventional self-dual lines of the generalized Baxter-Wu (GBW) model with asymmetric complex couplings are known to be $\sinh(2K)=\pm \cos(2ϕ)$, where $K$ and $ϕ$ are the real and imaginary parts of the coupling. We demonstrate that these lines are incomplete: the periodicity of the partition function, encoded in the cosine factor of the bundled Boltzmann weight, generates additional self-dual lines $\sinh(2K)=\pm \sin(2ϕ)$. Guided by the complete set of self-dual candidates, we perform Monte Carlo simulations using brute-force reweighting (Metropolis) and the Wang-Landau methods. Simulations indicate that the self-dual lines at the partition-function minima $ϕ_{\mathcal{Z}_{\min}}=(2n+1)π/8$ constitute a critical threshold. They are genuine critical boundaries for $|K| \ge \frac{1}{2}\operatorname{arsinh}(\cos(π/4)) \approx 0.32924$, while for smaller $|K|$ they are not. At $ϕ_{\mathcal{Z}_{\min}}$, the sign problem is most severe and finite-size scaling corrections are largest; the local peak observed below the phase boundary in the temperature scan is thus a finite-size artifact, not a genuine new phase. We further clarify the capability and limitations of the average sign and its derivatives for detecting phase transitions. In particular, the negative peak of the average sign at $ϕ_{\mathcal{Z}_{\min}}$ does not correspond to a genuine phase transition. We also evaluate the Wang-Landau method, which, despite formally circumventing the sign problem, still faces the exponential barrier.

cond-mat.str-el

HERALD: Counterfactual Audits and Minimal Repairs for Proof-of-Retrieval Rewards

Search-agent rewards mix answer quality, citation grounding, tool cost, and anti-hacking terms; a high score therefore need not imply that cited evidence was retrieved, and added penalties can cancel. We introduce HERALD, an offline audit that applies exact same-question interventions, separates candidate-visible from oracle information, and enumerates detector contracts before policy optimization. On four Qwen3-8B pools from HotpotQA, 2WikiMultiHopQA, and MuSiQue, $R_0$ rejects search deletion and fake IDs, but a label-free citation-laundering attack succeeds. A complete $2^3$ ablation identifies targeted strengthening of $L$---citing a corpus passage absent from the retrieved evidence---as the observed inclusion-minimal repair: $R[L]$ has zero empirical ASR with a 0.50% one-sided cluster upper bound. The gap persists across pool rules, a visible BM25 attacker, and four models; broader hardening remains vulnerable when the attack removes an oracle support-ID penalty. Under strict 5M-token matched training evaluated on 256 paired questions per benchmark, $R[L]$ meets the EM non-inferiority gate on HotpotQA and 2Wiki but not MuSiQue. Equal-suite citation precision and support recall improve by 2.02 and 1.46 points, unsupported citations fall by 1.69, and laundering attackability falls on 2Wiki and MuSiQue. Natural $L$ is not reduced, and the detector appears in only 18 of 58,368 training trajectories. HERALD thus separates robust scoring, sparse learning signal, and policy transfer.

cs.AI

Cosmological inference from the eBOSS QSO full-shape analysis with optimal redshift weights

We present a full-shape power-spectrum analysis of the eBOSS DR16 quasar sample with optimal redshift weights. The DR16 QSO catalog contains 343,708 quasars over $0.8<z<2.2$, a redshift interval broad enough to contain useful light-cone evolution but not naturally captured by a single effective-redshift measurement. We construct Karhunen--Loève weights for the parameters of interest and measure the resulting monopole and quadrupole with a cross-correlation estimator, which remains well defined for sign-changing weights. The theoretical spectra are convolved with the measured Fourier-space survey-window kernels for each Galactic cap and weighting scheme, and both the covariance matrix and the end-to-end validation are based on 1000 EZ light-cone mock catalogs. In $Λ$CDM, the redshift-weighted and standard analyses give consistent constraints, as expected from the near-standard effective redshifts of the weights targeting $h$, $Ω_{\rm m}$, and $A_s$. In the Chevallier--Polarski--Linder (CPL) model, the redshift-weighted DR16 analysis reduces the marginalized uncertainties on $H_0$, $σ_8$, and $w_0$ by $43.3\%$, $19.7\%$, and $20.5\%$, respectively, and turns the standard one-sided constraint on $w_a$ into a bounded posterior, $w_a=-0.98^{+1.0}_{-1.3}$. The gain is therefore concentrated where the model contains genuine redshift evolution, demonstrating that optimal redshift weighting can recover tomographic information from a wide QSO light cone while keeping the full-shape data vector compact.

astro-ph.CO

A multitracer analysis for the eBOSS galaxy sample based on the effective field theory of large-scale structure

We perform a multitracer full-shape analysis in Fourier space based on the effective field theory of large-scale structure (EFTofLSS) using the complete Sloan Digital Sky Survey IV (SDSS-IV) extended Baryon Oscillation Spectroscopic Survey (eBOSS) DR16 luminous red galaxy (LRG) and emission line galaxy (ELG) samples. We study in detail the impact of the volume projection effect and different prior choices when doing the full-shape analysis based on the EFTofLSS model. We show that adopting a combination of Jeffreys prior and Gaussian prior can mitigate the volume effect and avoid exploring unphysical regions in the parameter space at the same time, which is crucial when jointly analysing the eBOSS LRG and ELG samples. We validate our pipeline using 1000 eBOSS EZmocks. By performing a multitracer analysis on mocks with comparable footprints, we find that cosmological constraints can be improved by $\sim10-35$ per cent depending on whether we assume zero stochastic terms in the cross power spectrum, which breaks the degeneracy and boosts the constraints on the standard deviation of matter density fluctuation $σ_8$. Combining with the Big Bang Nucleosynthesis (BBN) prior and fixing the spectral tilt $n_s$ to Planck value, our multitracer full-shape analysis measures $H_0=70.0\pm2.3~{\mathrm{km}}~{\mathrm{s}}^{-1}{\mathrm{Mpc}}^{-1}$, $Ω_m=0.317^{+0.017}_{-0.021}$, $σ_8=0.787_{-0.062}^{+0.055}$ and $S_8=0.809_{-0.078}^{+0.064}$, consistent with the Planck~2018 results. In particular, the constraint on $σ_8$ is improved beyond that obtained from the single tracer analysis by $18$ per cent, or by $27$ per cent when assuming zero stochastic terms in the cross power spectrum.

astro-ph.CO

Towards Long-Horizon Vessel Trajectory and Destination Forecasting with Reasoning Large Language Models

Long-horizon maritime trajectory prediction is important for shipping management, logistics planning, and maritime risk analysis, yet month-level forecasting remains insufficiently studied. Existing deep learning methods mainly focus on short- and mid-term coordinate extrapolation and often struggle to preserve route feasibility and destination correctness over extended horizons. This paper investigates joint long-horizon vessel trajectory and destination forecasting with reasoning-capable large language models, and develops a Maritime LLM post-training framework based on Reinforcement Learning with Verifiable Reward (RLVR). An AIS-based benchmark is constructed with 60-day historical trajectories and 30-day forecasting horizons, where trajectories are converted into semantic textual representations for RL prompt construction. RLVR aligns LLMs with maritime forecasting objectives by enforcing physical validity, providing early-weighted trajectory supervision, and evaluating destination correctness through hierarchical matching and curriculum learning. Experimental results show that RLVR-trained LLMs substantially improve over zero-shot LLMs and representative deep learning baselines, especially on destination-related metrics. Among the evaluated RLVR-trained variants, 4B LLMs achieve the best overall performance, suggesting that reward-compatible optimization and task-specific capacity matching are more important than simply using larger 8B or 14B LLMs. The results also show that LSTM remains a strong deep learning baseline under limited fine-tuning data, while Transformer-style spatio-temporal models typically require larger datasets and richer structured inputs. Overall, this work advances semantic, verifier-aligned maritime forecasting for operational decision support.

cs.AI

Tropical Cartan's second main theorem for hyperplanes in general position

We prove a tropical analogue of Cartan's second main theorem for holomorphic curves intersecting hyperplanes in general position--a setting that was not fully resolved by previous tropical Nevanlinna theory. Two versions are obtained. The first (Theorem 1.7) requires subnormal growth and involves the tropical Casorati determinant. The second and main version (Theorem 1.9) is completely free of growth conditions and exceptional sets; it replaces the Casorati term by the sum of the counting functions of the curve's components, yielding an inequality valid for every r. The proof uses a tropical Cramer theorem, bypassing the logarithmic derivative lemma. This improves upon previous results by Korhonen-Tohge and Cao-Zheng, where the coefficient could be suboptimal even under the general position hypothesis. We also clarify the relation between different notions of linear independence, and present the first counterexample to the truncated second main theorem in the tropical setting (Example 5.4).

math.AG

GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents

We present GLM-5V-Turbo, a step toward native foundation models for multimodal agents. As foundation models are increasingly deployed in real environments, agentic capability depends not only on language reasoning, but also on the ability to perceive, interpret, and act over heterogeneous contexts such as images, videos, webpages, documents, GUIs. GLM-5V-Turbo is built around this objective: multimodal perception is integrated as a core component of reasoning, planning, tool use, and execution, rather than as an auxiliary interface to a language model. This report summarizes the main improvements behind GLM-5V-Turbo across model design, multimodal training, reinforcement learning, toolchain expansion, and integration with agent frameworks. These developments lead to strong performance in multimodal coding, visual tool use, and framework-based agentic tasks, while preserving competitive text-only coding capability. More importantly, our development process offers practical insights for building multimodal agents, highlighting the central role of multimodal perception, hierarchical optimization, and reliable end-to-end verification.

cs.CV

MUltiplexed Survey Telescope (MUST) Science White Paper I: Overview of Large-Scale Structure Cosmology in the Era of Stage-V Spectroscopic Surveys

The MUltiplexed Survey Telescope (MUST) is a 6.5-meter telescope under development. Dedicated to highly-multiplexed, wide-field spectroscopic surveys, MUST observes over 20,000 targets simultaneously using 6.2-mm pitch positioning robots within a ~5 deg$^2$ field of view. MUST aims to conduct the first Stage-V spectroscopic survey in the 2030s, mapping the 3D Universe with over 100 million galaxies and quasars, spanning from the nearby Universe to a redshift of z ~ 5.5, corresponding to approximately 1 billion years after the Big Bang. To cover this extensive redshift range, we present an initial conceptual target selection algorithm for different types of galaxies, ranging from local bright galaxies and luminous red galaxies to emission-line galaxies, and high-redshift (2 < z < 5.5) Lyman-break galaxies. Using Fisher forecasts, we demonstrate that MUST can address fundamental questions in cosmology, including the nature of dark energy, tests of gravity theories, and investigations into primordial physics. This is the first paper in the series of science white papers for MUST, with subsequent developments focusing on additional scientific cases such as galaxy and quasar evolution, Milky Way physics, and dynamic phenomena in the time-domain Universe.

astro-ph.CO

Revisiting Uncertainty: On Evidential Learning for Partially Relevant Video Retrieval

Partially relevant video retrieval aims to retrieve untrimmed videos using text queries that describe only partial content. However, the inherent asymmetry between brief queries and rich video content inevitably introduces uncertainty into the retrieval process. In this setting, vague queries often induce semantic ambiguity across videos, a challenge that is further exacerbated by the sparse temporal supervision within videos, which fails to provide sufficient matching evidence. To address this, we propose Holmes, a hierarchical evidential learning framework that aggregates multi-granular cross-modal evidence to quantify and model uncertainty explicitly. At the inter-video level, similarity scores are interpreted as evidential support and modeled via a Dirichlet distribution. Based on the proposed three-fold principle, we perform fine-grained query identification, which then guides query-adaptive calibrated learning. At the intra-video level, to accumulate denser evidence, we formulate a soft query-clip alignment via flexible optimal transport with an adaptive dustbin, which alleviates sparse temporal supervision while suppressing spurious local responses. Extensive experiments demonstrate that Holmes outperforms state-of-the-art methods. Code is released at https://github.com/lijun2005/ICML26-Holmes.

cs.CV

Efficient estimators for power spectrum and bispectrum multipole measurements

Large galaxy surveys demand fast and scalable estimators for anisotropic clustering statistics beyond the monopole. We present a suite of efficient FFT-based estimators for power-spectrum and bispectrum multipoles, built upon exact conjugation and parity symmetries of spherical-harmonic--weighted Fourier transforms of real fields. These symmetries eliminate redundant magnetic sub-configurations, thereby reducing the computational cost by a factor of 2. For the Yamamoto power-spectrum multipoles, we further decrease the cost of high-order even multipoles by algebraically expressing ${L}_{2n}$ in terms of lower-order Legendre polynomials, thereby measuring modified high-order multipoles using only low-$\ell$ fields with a small and controlled deviation from the traditional definition. We introduce a new TripoSH bispectrum estimator obtained by compressing the Scoccimarro bispectrum along an alternative triangle side, which substantially reduces the FFT scaling for commonly used quadrupole configurations in the large-$k$-bin limit. We also derive an analytic treatment of bispectrum shot noise by integrating spherical-harmonic kernels over the triangle-constrained $k$-space volumes, avoiding additional FFTs or costly spherical-Bessel evaluations and enabling fast and accurate shot-noise subtraction. Based on these optimizations, we also introduce CosmoNPC, an open-source Python package for large-scale-structure clustering measurements.

astro-ph.CO

From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents

While multimodal large language models have demonstrated impressive short-term reasoning, they struggle with long-horizon video understanding due to limited context windows and static memory mechanisms that fail to mirror human cognitive efficiency. Existing paradigms typically fall into two extremes: vision-centric methods that incur high latency and redundancy through dense visual accumulation, or text-centric approaches that suffer from detail loss and hallucination via aggressive captioning. To bridge this gap, we propose MM-Mem, a pyramidal multimodal memory architecture grounded in Fuzzy-Trace Theory. MM-Mem structures memory hierarchically into a Sensory Buffer, Episodic Stream, and Symbolic Schema, enabling the progressive distillation of fine-grained perceptual traces (verbatim) into high-level semantic schemas (gist). Furthermore, to govern the dynamic construction of memory, we derive a Semantic Information Bottleneck objective and introduce SIB-GRPO to optimize the trade-off between memory compression and task-relevant information retention. In inference, we design an entropy-driven top-down memory retrieval strategy. Extensive experiments across 4 benchmarks confirm that MM-Mem achieves state-of-the-art performance on both offline and streaming tasks, demonstrating robust generalization and validating the effectiveness of cognition-inspired memory organization. Code and associated configurations are publicly available at https://github.com/EliSpectre/MM-Mem.

cs.CV

Imagine Before Concentration: Diffusion-Guided Registers Enhance Partially Relevant Video Retrieval

Partially Relevant Video Retrieval (PRVR) aims to retrieve untrimmed videos based on text queries that describe only partial events. Existing methods suffer from incomplete global contextual perception, struggling with query ambiguity and local noise induced by spurious responses. To address these issues, we propose DreamPRVR, which adopts a coarse-to-fine representation learning paradigm. The model first generates global contextual semantic registers as coarse-grained highlights spanning the entire video and then concentrates on fine-grained similarity optimization for precise cross-modal matching. Concretely, these registers are generated by initializing from the video-centric distribution produced by a probabilistic variational sampler and then iteratively refined via a text-supervised truncated diffusion model. During this process, textual semantic structure learning constructs a well-formed textual latent space, enhancing the reliability of global perception. The registers are then adaptively fused with video tokens through register-augmented Gaussian attention blocks, enabling context-aware feature learning. Extensive experiments show that DreamPRVR outperforms state-of-the-art methods. Code is released at https://github.com/lijun2005/CVPR26-DreamPRVR.

cs.CV

TacSIm: A Dataset and Benchmark for Football Tactical Style Imitation

Current football imitation research primarily aims to opti mize reward-based objectives, such as goals scored or win rate proxies, paying less attention to accurately replicat ing real-world team tactical behaviors. We introduce Tac SIm, a large-scale dataset and benchmark for Tactical Style Imitation in football. TacSIm imitates the acitons of all 11 players in one team in the given broadcast footage of Pre mier League matches under a single broadcast view. Under a offensive or defensive broadcast footage, TacSIm projects the beginning positions and actions of all 22 players from both sides onto a standard pitch coordinate system. Tac SIm offers an explicit style imitation task and evaluation protocols. Tactics style imitation is measured by using spatial occupancy similarity and movement vector similarity in defined time, supporting the evaluation of spatial and tem poral similarities for one team. We run multiple baseline methods in a unified virtual environment to generate full team behaviors, enabling both quantitative and visual as sessment of tactical coordination. By using unified data and metrics from broadcast to simulation, TacSIm estab lishes a rigorous benchmark for measuring and modeling style-aligned tactical imitation task in football.

cs.CV

Nonlinear Information from DESI Luminous Red Galaxies: An Emulator-Based Analysis of Pre- and Post-Reconstruction Power Spectra

We present joint measurements of the pre- and post-reconstruction power spectra, $P_{\rm pre}$ and $P_{\rm post}$, together with their cross-power spectrum, $P_{\rm cross}$, for the Luminous Red Galaxies (LRGs) in the DESI Data Release 1 (DR1). We jointly analyse these observables with an emulator-based full-shape modeling framework, thereby, for the first time, we extract complementary nonlinear information from the galaxy density field before and after reconstruction in real survey data. Specifically, including $P_{\rm post}$ and $P_{\rm cross}$ in addition to $P_{\rm pre}$ (hereafter $P_{\rm all}$) yields an improvement of approximately $18$-$27\%$ in the $σ_8$ constraint in both $Λ$CDM and $w$CDM, depending on the redshift bin, relative to the $P_{\rm pre}$-only analysis with the cosmic microwave background distance priors (hereafter CMB). In $w$CDM, the joint CMB+$P_{\rm all}$ analysis can tighten the constraints on $w$ by approximately $5$-$15\%$ across the two LRG redshift bins, compared to the CMB+$P_{\rm pre}$ combination. Further incorporating the Type Ia supernova dataset and comparing the cosmological constraints in $w$CDM from each individual power-spectrum component with those from the full combination, we find that $P_{\rm all}$ consistently provides the tightest constraints. From the joint CMB+$P_{\rm all}$+DES-Dovekie dataset, we obtain $Ω_m = 0.314 \pm 0.0048$ and $w = -0.988 \pm 0.023$ for the \texttt{LRG1} sample, and $Ω_m = 0.318 \pm 0.0046$ and $w = -0.988 \pm 0.025$ for \texttt{LRG2}. These results demonstrate that combining pre- and post-reconstruction power spectra with their cross-correlation enables DESI to harvest additional nonlinear information, leading to tighter constraints on cosmological parameters.

astro-ph.CO

Probing the Critical Behavior of a Sign-Problematic Model with Monte Carlo Simulations

The sign-problematic generalized Baxter-Wu (GBW) model with asymmetric complex couplings is mapped onto a one-dimensional quantum model. Utilizing the model's exactly known critical properties, we study the relation between the conventional and the modified average signs and the phase transitions in the GBW model. We find that the average sign develops a negative peak near the critical point, but it is not a unique indicator of phase transition, as similar features can appear in non-critical regions. While the average modified sign provides a viable probe for the phase transition, the practical effectiveness of this method is limited by the exponential scaling of computational cost with the system's volume. We propose that the universal properties of the original model can be investigated through simulating the related reference model, based on the universality assumption. Using finite-size scaling analysis based on Monte Carlo simulations, we confirm the validity of this method, which thereby provides a novel framework for investigating phase transitions in systems plagued by the sign problem.

cond-mat.str-el