Search arXiv⌕ Search

arXiv subjects

Jiahui Chen

Publications and source records attributed to Jiahui Chen.

At least 19 recordsLinked to original sources

WaveTLM: Reliable Time-Series Language Modeling through Task Compilation

Time-series language models provide a shared natural-language interface across temporal tasks, but plausible text does not guarantee reliable task outputs. Responses may appear reasonable while hallucinating the required object: numerical sequences can violate shape, scale, channel order, or temporal alignment, and textual decisions can fall outside the legal label space. We formulate reliable time-series language modeling, separating task-object reliability from predictive quality. We introduce ExecTS-QA, a contract-grounded benchmark spanning forecasting, imputation, classification, anomaly detection, and waveform analysis. We further propose WaveTLM, a unified compiler-executor model whose task compiler transforms user requests, visible arguments, and wave-grounded evidence into typed task states, while task-native executors construct numerical tensors, legal decisions, or structured records. On ExecTS-QA, a single WaveTLM checkpoint achieves 99.40% contract-valid coverage, compared with 37.83% for the strongest evaluated string-first baseline, while retaining balanced predictive performance across all five task families. Evaluations on SciTS, TSQA, IRTS-ToolBench, and ARFBench provide additional evidence of transfer. The code, construction scripts, and ExecTS-QA dataset will be publicly released upon publication. These results show that task compilation can convert plausible language generation into reliable time-series outputs.

cs.LG↗

Unified Text-Image Generation with Weakness-Targeted Post-Training

Unified multimodal generation architectures that jointly produce text and images have recently emerged as a promising direction for text-to-image (T2I) synthesis. However, many existing systems rely on explicit modality switching, generating reasoning text before switching manually to image generation. This separate, sequential inference process limits cross-modal coupling and prohibits automatic multimodal generation. This work explores post-training to achieve fully unified text-image generation, where a model autonomously transitions from textual reasoning to visual synthesis within a single inference process. We study this on BAGEL, a 14B mixture-of-transformers model that pairs autoregressive text generation with flow-matching image synthesis. We examine the impact of joint text-image generation on T2I performance and the relative importance of each modality during post-training. We additionally explore different post-training data strategies, showing that a targeted dataset addressing specific limitations achieves superior results compared to broad image-caption corpora or benchmark-aligned data. Using offline, reward-weighted post-training with fully self-generated synthetic data, our approach enables improvements in multimodal image generation across four diverse, independent T2I benchmarks, demonstrating the effectiveness of reward-weighting both modalities and strategically designed post-training data.

cs.CV↗

Multimodal Language Models as Text-to-Image Model Evaluators

The steady improvements of text-to-image (T2I) generative models lead to slow deprecation of automatic evaluation benchmarks that rely on static datasets, motivating researchers to seek alternative ways to evaluate T2I progress. We present Multimodal Text-to-Image Eval (MT2IE), an evaluation framework in which a single multimodal large language model (MLLM) acts as an evaluator agent, iteratively generating the evaluation prompts and scoring the resulting images. We show that MT2IE's image-text consistency scores have higher correlation with human judgment than metrics previously introduced in the literature. MT2IE generates prompts that are efficient at probing T2I model performance: closely recovering the official T2I model rankings of three structurally distinct benchmarks from just 20 generated evaluation prompts, 28-105x fewer than the benchmarks' own prompt sets. When compared to existing evaluation metrics such as CLIPScore, VIEScore, and VQAScore, MT2IE's T2I model rankings are more faithful and far more consistent across multiple evaluation seeds when using the same number of prompts. MT2IE can also adapt evaluation to the model being tested: rewriting each prompt based on the model's own measured performance to produce a bespoke per-model benchmark that still recovers the official rankings and keeps the evaluated model in an informative scoring range. We hope that these results will encourage the development of dynamic and interactive evaluation frameworks, and mitigate the deprecation of automatic evaluation benchmarks.

cs.CV↗

Test-time Scaling over Perception: Resolving the Grounding Paradox in Thinking with Images

Recent multimodal large language models (MLLMs) support Thinking with Images, invoking visual tools such as zooming and cropping to inspect image regions during inference. Yet these systems remain brittle in fine-grained reasoning: to acquire a decisive detail, a model must ground its attention on the correct region, but knowing which region is correct presupposes having already observed that detail. We identify this circular dependency as the grounding paradox, show that grounding errors are rarely self-corrected within a single trajectory---once a misleading region is inspected, all subsequent reasoning conditions on that observation and the error propagates to the final answer---and observe that because each trajectory constructs its own evidence, answer-level aggregation discards the very information that distinguishes trajectories. We propose Test-Time Scaling over Perception (TTSP), a closed-loop framework that treats perception as the unit of scalable inference and allocates compute along two axes: Entropy-Gated Perceptual Exploration samples diverse trajectories and uses critical-token entropy to withhold evidence the model cannot commit to, while Evidence-Guided Iterative Refinement distills validated observations into a correctable Evidence Ledger that steers later rounds to re-inspect unresolved regions. Across high-resolution and general multimodal benchmarks, TTSP consistently outperforms strong test-time scaling baselines, while improving grounding quality with favorable token efficiency.

cs.CV↗

FasTac: A Curved Multispectral Vision-Based Tactile Sensor for High-Speed High-Precision 3D Shape and Force Perception

Curved tactile fingertips for dexterous manipulation must resolve fine contact geometry, distinguish normal and tangential loads, and capture transient signals. Existing curved vision-based tactile sensors struggle to combine accurate 3D reconstruction, three-axis force estimation, and high-speed processing in a compact form. This article presents FasTac, a curved vision-based tactile sensor integrating multispectral photometric stereo, dynamic-convolution force estimation, and hardware acceleration on a field-programmable gate array (FPGA). Single-image-sensor simultaneous multispectral imaging provides spatially aligned observations for robust surface normal estimation, followed by boundary-prior fast Poisson depth reconstruction. HyperForce uses position-aware dynamic convolution to model the spatially nonuniform mechanical response of curved elastomers and estimate three-axis forces. The complete image-to-normal-force pipeline is deployed on an FPGA. Experiments show that near-infrared (NIR) illumination and the boundary prior decrease depth mean absolute error (MAE) from 0.2730 mm to 0.0415 mm; HyperForce achieves normalized mean absolute error (NMAE) values of 2.74% and 2.39% for normal and shear forces, respectively; and FPGA deployment shortens processing latency from 3.26 ms on the GPU to 1.09 ms. Multi-object reconstruction, feedback grasping, and vibration measurement validate fine geometric perception, stable force feedback, and dynamic contact sensing.

cs.RO↗

Geometry-Induced Hodge Stars on Rips and Dowker--Rips Complexes

The Vietoris--Rips complex $\mathrm{VR}_ε(X)$, the Dowker complex $\mathrm{D}_R(X,Y)$, and its flagified Dowker--Rips variant $\mathrm{DR}_R(X,Y)=\mathrm{F}(\mathrm{D}_R(X,Y))$ are simplicial complexes constructed from metric data or witness relations. They are useful in topological data analysis because they encode topology through combinatorial data derived from pairwise information, but at a fixed scale they retain little of the underlying geometry. Unlike alpha complexes or mesh-based discretizations, Rips-type complexes carry no canonical primal--dual cell structure, which is the ingredient used by the discrete exterior calculus Hodge star to encode metric information. We address this gap by equipping a Rips-type complex $K$ with diagonal geometry-induced Hodge stars represented by positive simplex weights $W_k=\operatorname{diag}\{w_k(σ):σ\in K_k\}$, which define weighted inner products on $k$-cochains. The resulting weighted discrete Hodge Laplacian $Δ_k^W$ has kernel dimension equal to the $k$th Betti number of the underlying complex, while its nonzero spectrum is governed by the chosen geometric weights. The central issue is therefore not the existence of a weighted Laplacian, since any positive diagonal weights define one, but the design of weights that encode meaningful metric or witness geometry. We focus on two computable choices: simplex-volume weights, based on Euclidean simplex volumes, and soft witness weights, based on a Dowker-style support function $s_t(σ;Y)$ that quantifies higher-order witness support lost under flagification. We prove positivity, weighted self-adjointness, and Betti-number preservation for arbitrary positive diagonal weights, establish an asymptotic decay-rate characterization for soft witness support, and describe spectral descriptors derived from $Δ_k^W$ for comparing geometry-aware Hodge spectra on Rips complexes.

math.AT↗

Final assessment of radioactive impurities in the JUNO detector

The Jiangmen Underground Neutrino Observatory (JUNO) collaboration has completed the construction of the 20,000-ton liquid scintillator detector and the associated muon veto detector system. To meet the physics objectives, the materials used in the detector must exhibit low radioactive contamination. The single-event rate in the fiducial volume (R $<$ 17.2 m) of the scintillator is required to be approximately 7 Hz for energies above 0.7 MeV, resulting in an accidental coincidence background of about 1 event per day for reactor neutrino physics analyses. Since the beginning of the construction phase, we have screened the natural radioactivity content of thousands of materials, to select those that meet the design background budget. The radioactive impurity concentrations of the materials ultimately used in the JUNO detector are summarized in this paper. The construction of the entire detector and the subsequent filling of the liquid scintillator were completed in August 2025. From the initial data, the total count rate of natural radioactivity within the detector's fiducial volume has met the requirements and is sufficient to support the reactor antineutrino analysis.

physics.ins-det↗

A Low-energy Threshold and Multi-messenger Trigger System for the JUNO Experiment

The Jiangmen Underground Neutrino Observatory (JUNO) is a 20-kiloton liquid scintillator neutrino detector, located 650 meters (1800 m.w.e.) underground in Jiangmen, Guangdong, China. JUNO is primarily designed for reactor neutrino measurements and has been taking data since 2025. With the largest mass of its kind and an excellent energy resolution, JUNO is a leading observatory for high-precision measurements of MeV neutrinos. The standard global trigger system serves as the primary trigger for JUNO. We present a newly developed multi-messenger trigger system that extends the capabilities of the global trigger by providing a lower energy threshold and an independent monitoring capability. During the 2025 operation, it achieved an effective energy threshold of approximately 110 +/- 10 keV, providing a lower threshold configuration suitable for low-energy event analysis. The system shows the potential to further reduce the threshold to well below 100 keV. Based on the multi-messenger trigger system, an astrophysical monitor has been developed to receive and process external alerts from other messengers, such as gravitational-wave observations. A Transient Neutrino Burst Monitor is integrated to detect short-time-scale neutrino burst events and enables real-time monitoring of transient astrophysical phenomena. The system is sensitive to neutrino bursts from core-collapse supernovae within a distance of about 250 kpc.

hep-ex↗

HiMem-WAM: Hierarchical Memory-Gated World Action Models for Robotic Manipulation

World Action Models (WAMs) have emerged as a new powerful paradigm for embodied intelligence, learning action-relevant visual dynamics that significantly enhance generalization and robustness. However, existing WAMs still struggle with task-relevant memory in long-horizon robotic manipulation. To address this, we present HiMem-WAM, a Hierarchical Memory-Gated WAM that integrates motion-centric latent actions, high-level skill latents, and boundary-triggered memory updates. Specifically, we develop a hierarchical latent action framework that jointly learns low-level motion and high-level skill latents, providing structured temporal abstraction. Meanwhile, a boundary-aware memory gate writes compact task states at predicted skill transitions, enabling causal inference without test-time generation of future video or optical flow estimation. Evaluated on LIBERO, LIBERO-PLUS, RMBench and real-world tasks, HiMem-WAM shows that hierarchical latents improve robustness under deployment perturbations, and the memory module substantially benefits memory-dependent long-horizon manipulation.

cs.RO↗

Mamba-Enhanced Implicit Motion Learning for Audio-Driven Portrait Animation

Audio-driven human motion video generation aims to synthesize realistic and temporally coherent human animations from a single static image, with applications in talking-head synthesis, co-speech gesture generation, and dynamic presentations. Moving beyond conventional keypoint-based methods that often struggle to capture subtle motion dynamics, We propose a novel implicit-motion framework for generating realistic and temporally coherent human motion videos from a single static image and audio. Our approach uses a two-stage pipeline that decouples motion prediction from rendering. The first stage integrates appearance priors and hierarchical depth cues into a region-aware attention mechanism to model latent motion features. The second stage employs a Mamba-enhanced diffusion model to directly predict these features from audio and the source image, enabling unsupervised learning of fine-grained motion patterns. This decoupled architecture enhances flexibility and efficiency. Trained on a new 380-hour high-quality dataset, our method outperforms prior work across multiple public benchmarks and our collected data in accuracy, naturalness, and temporal coherence, setting a new state-of-the-art.

cs.CV↗

Embedded underwater front-end electronics for the 3-inch photomultipliers in the JUNO experiment

The Jiangmen Underground Neutrino Observatory (JUNO) is a 20-kton liquid scintillator-based, low-radioactivity, multi-purpose neutrino detector located 693 meters (1800 m.w.e.) underground in the Guangdong province, China. To detect scintillation light produced in the target, the detector is equipped with 17,612 20-inch photomultipliers (PMTs), forming the Large PMT system (LPMT). In addition, 25,600 3-inch photomultipliers (the Small Photomultiplier System or SPMT) are deployed in the gaps between the LPMTs. This paper presents the design and performance of the underwater front-end electronics developed for the SPMT system. It details the individual electronics boards and their key components, the inter-board interfaces, the system-level design, and the firmware architecture that supports data acquisition and control. It also outlines mechanical and thermal integration, board validation procedures, and system performance metrics. The readout chain includes digitization of 128 PMT channels per unit, synchronized time-stamping, charge measurement, event packaging, and bandwidth management. Comprehensive validation confirms the system's readiness to meet JUNO's stringent physics goals. The underwater electronics achieve noise levels as low as 0.04 photoelectrons with minimal crosstalk (below 0.4%) and a bandwidth of 57 MB/s, ensuring reliable single photo-electron detection and operation under high-rate conditions. The SPMT system has now been fully integrated and installed in JUNO. Its commissioning and physics performance will be reported in a future publication.

physics.ins-det↗

SpectraLLM: Uncovering the Ability of LLMs for Molecular Structure Elucidation from Multi-Spectral Data

Automated molecular structure elucidation remains challenging, as existing approaches often depend on pre-compiled databases or restrict themselves to single spectroscopic modalities. Here we introduce SpectraLLM, a large language model that performs end-to-end structure prediction by reasoning over one or multiple spectra. Unlike conventional spectrum-to-structure pipelines, SpectraLLM represents both continuous (IR, Raman, UV-Vis, NMR) and discrete (MS) modalities in a shared language space, enabling it to capture substructural patterns that are complementary across different spectral types. We pretrain and fine-tune the model on small-molecule domains and evaluate it on four public benchmark datasets. SpectraLLM achieves state-of-the-art performance, substantially surpassing single-modality baselines. Moreover, it demonstrates strong robustness in unimodal settings and further improves prediction accuracy when jointly reasoning over diverse spectra, establishing a scalable paradigm for language-based spectroscopic analysis. Code is available at https://github.com/OPilgrim/SpectraLLM.

q-bio.QM↗

ViTok-v2: Scaling Native Resolution Auto-Encoders to 5 Billion Parameters

Vision Transformer (ViT) autoencoders have emerged as compelling tokenizers for images, offering improved reconstruction over convolutional tokenizers. However, existing ViT tokenizers cannot explore this landscape as performance degrades outside training resolutions, and reliance on adversarial losses prevents stable scaling. ViTok (Hansen-Estruch et al., 2025) found that the compression ratio r mediates a reconstruction-generation trade-off where lower r means better reconstructions but harder generations, so improving tokenizer reconstruction is key to more Pareto-optimal tokenizers. We introduce ViTok-v2, which addresses these limitations with native resolution support via NaFlex for generalization across resolutions and aspect ratios, and a novel DINOv3 perceptual loss that replaces both LPIPS and GAN objectives for stable training at any scale. ViTok-v2 is trained on about 2B images and scaled to 5B parameters, the largest image autoencoder to date. ViTok-v2 matches or exceeds state-of-the-art reconstruction at 256p and outperforms all baselines at 512p and above. In joint scaling experiments with flow matching generators, we show that scaling both the autoencoder and the generator advances the Pareto frontier of this trade-off.

cs.CV↗

Engineering Precise and Robust Effective Hamiltonians

Engineering effective Hamiltonians is essential for advancing quantum technologies including quantum simulation, sensing, and computing. This paper presents a general framework for effective Hamiltonian engineering, enabling robust, precise, and efficient quantum control strategies. To achieve efficiency, we focus on creating target zeroth-order effective Hamiltonians while minimizing higher-order contributions and enhancing robustness against systematic errors. The control design identifies the minimal subspace of the toggling-frame Hamiltonian and the full set of achievable, zeroth-order, effective Hamiltonians. The framework also enables robust state transfer, characterization of achievable density matrices, and extension to stochastic parameter fluctuations via a cumulant expansion. Examples are included to illustrate the process flow and resultant precision and robustness.

quant-ph↗

Engineering Higher-order Effective Hamiltonians

Advancing quantum technologies requires precise and robust coherent control of quantum systems. Robust higher-order Hamiltonian engineering is essential for high-precision control and for accessing effective dynamics absent at zeroth order. Here, we introduce a systematic methodology for achieving the precision, robustness, and complexity required for quantum control through the engineering of higher-order processes and effective Hamiltonians. We identify the minimal subspace of achievable effective Hamiltonian at each order and provide universal cost functions for achieving desired targets. Examples include robust sequences for decoupling, three-body interactions and detuning/interaction correlations.

quant-ph↗

Towards Practical Benchmarking of Data Cleaning Techniques: On Generating Authentic Errors via Large Language Models

Data quality remains an important challenge in data-driven systems, as errors in tabular data can severely compromise downstream analytics and machine learning performance. Although numerous error detection algorithms have been proposed, the lack of diverse, real-world error datasets limits comprehensive evaluation. Manual error annotation is both time-consuming and inconsistent, motivating the exploration of synthetic error generation as an alternative. In this work, we introduce TableEG, a framework that leverages large language models (LLMs) to generate authentic errors. By employing a table fine-tuning strategy and a triplet representation $(I, T, O)$ to model error generation, detection, and correction tasks, TableEG captures the complex dependencies inherent in two-dimensional tables. Trained on 12 real-world datasets spanning 10 diverse domains, TableEG ensures that the synthesized errors faithfully reflect authentic error distributions. Experimental results indicate that errors generated by TableEG exhibit superior pattern and distribution similarity compared to both rule-based methods and LLM-generated errors without fine-tuning. Furthermore, performance metrics on TableEG-generated errors closely align with those on real-world errors across nearly all datasets and detection algorithms, particularly for machine learning based detection techniques. Overall, TableEG not only bridges the gap between synthetic and real-world errors but also establishes a robust benchmark for subsequent error detection and correction tasks.

cs.DB↗

Decoupling Defense Strategies for Robust Image Watermarking

Deep learning-based image watermarking, while robust against conventional distortions, remains vulnerable to advanced adversarial and regeneration attacks. Conventional countermeasures, which jointly optimize the encoder and decoder via a noise layer, face 2 inevitable challenges: (1) decrease of clean accuracy due to decoder adversarial training and (2) limited robustness due to simultaneous training of all three advanced attacks. To overcome these issues, we propose AdvMark, a novel two-stage fine-tuning framework that decouples the defense strategies. In stage 1, we address adversarial vulnerability via a tailored adversarial training paradigm that primarily fine-tunes the encoder while only conditionally updating the decoder. This approach learns to move the image into a non-attackable region, rather than modifying the decision boundary, thus preserving clean accuracy. In stage 2, we tackle distortion and regeneration attacks via direct image optimization. To preserve the adversarial robustness gained in stage 1, we formulate a principled, constrained image loss with theoretical guarantees, which balances the deviation from cover and previous encoded images. We also propose a quality-aware early-stop to further guarantee the lower bound of visual quality. Extensive experiments demonstrate AdvMark outperforms with the highest image quality and comprehensive robustness, i.e. up to 29\%, 33\% and 46\% accuracy improvement for distortion, regeneration and adversarial attacks, respectively.

cs.CV↗

HiVid: LLM-Guided Video Saliency For Content-Aware VOD And Live Streaming

Content-aware streaming requires dynamic, chunk-level importance weights to optimize subjective quality of experience (QoE). However, direct human annotation is prohibitively expensive while vision-saliency models generalize poorly. We introduce HiVid, the first framework to leverage Large Language Models (LLMs) as a scalable human proxy to generate high-fidelity weights for both Video-on-Demand (VOD) and live streaming. We address 3 non-trivial challenges: (1) To extend LLMs' limited modality and circumvent token limits, we propose a perception module to assess frames in a local context window, autoregressively building a coherent understanding of the video. (2) For VOD with rating inconsistency across local windows, we propose a ranking module to perform global re-ranking with a novel LLM-guided merge-sort algorithm. (3) For live streaming which requires low-latency, online inference without future knowledge, we propose a prediction module to predict future weights with a multi-modal time series model, which comprises a content-aware attention and adaptive horizon to accommodate asynchronous LLM inference. Extensive experiments show HiVid improves weight prediction accuracy by up to 11.5\% for VOD and 26\% for live streaming over SOTA baselines. Real-world user study validates HiVid boosts streaming QoE correlation by 14.7\%.

cs.CV↗