Search arXiv⌕ Search

arXiv subjects

Bo Peng

Publications and source records attributed to Bo Peng.

At least 37 records · Page 2Linked to original sources

Geometry-Dependent Nonlocal Valence Screening Following Core Ionization of the Water Dimer

Core ionization is spatially localized, but the correlated valence response that screens the resulting hole need not be. Using the water dimer as a minimal hydrogen-bonded system, we ask whether acceptor-site O~1s ionization recruits valence channels on the neighboring donor molecule and how proton displacement redistributes that response. We introduce correlated shifted-start real-time $Λ$-coupled-cluster theory to connect the core-hole spectrum with orbital- and fragment-resolved screening pathways. At equilibrium, the satellite response contains donor-local and intermolecular charge-transfer-like contributions. Their intermolecular origin is supported by the strong suppression of both contributions when the monomers are separated. In a fixed-nuclei geometry scan, elongating the hydrogen-bonded donor O--H bond nearly doubles their combined share of the satellite response, while analysis of the first approximately 10~fs of the fixed-nuclei electronic response reveals stronger modulation involving a donor-localized valence orbital. More broadly, this work shows how localized core ionization can serve as a site-selective probe of electronic communication, with satellite structure revealing how nuclear geometry redirects many-electron screening through molecular environments.

physics.chem-ph↗

Difficulty-Aware Semantic-ID Optimization for Generative Recommendation

Semantic-ID-based generative recommendation casts retrieval and ranking as autoregressive generation over hierarchical item identifiers. A common recipe is SFT followed by GRPO, yet vanilla GRPO is poorly matched to this tree-structured task. Under the frozen SFT checkpoint, the exact target is absent from the first 16 candidates of the 50-beam constrained ranking for many prompts, and in harder cases none of these candidates enters the target SID branch. This prompt-level diagnostic motivates a training concern: when on-policy GRPO groups are similarly target-missing, item-level rewards may produce weak or degenerate reward variation even if some candidates follow part of the target path. We propose Difficulty-Aware Semantic-ID Optimization (DASO), a tree-aware post-training method that addresses this failure mode as an online rollout-allocation problem. Instead of using fixed difficulty buckets or uniformly injecting ground-truth completions, DASO profiles each current rollout group by prefix-match depth, locates the bottleneck SID levels where candidates leave the target path, and reallocates a bounded portion of the group to prefix-guided completions while retaining raw rollouts for contrast. A SID-prefix reward provides graded credit, while an auxiliary SFT anchor mitigates regression on examples already solved by the SFT checkpoint. On the public benchmarks, DASO improves over MiniOneRec-style GRPO on 11 of 12 metrics and achieves the best result on 9 of 12 metrics; it also improves most level-wise recall metrics on the internal recommendation task.

cs.AI↗

A Two-Dimensional Counterexample to Radical Equality in Primitive Axial Algebras

Let $R(A,X)$ denote the largest ideal of a primitive axial algebra $(A,X)$ that contains no axis from the specified generating set $X$, and let $J(A)$ be the intersection of the maximal ideals of $A$. Mamontov, Shpectorov, and Zhelyabin asked whether $R(A,X)=J(A)$ always holds. We give a negative answer. Over every field of characteristic different from $2$, the two-dimensional commutative algebra with basis $a,b$ and multiplication $a^2=a$, $ab=2b$, and $b^2=b$ is a primitive axial algebra for an explicit fusion law and the generating set $X={a,b}$. Its complete ideal lattice is $0<\mathbb{F}b<A$, whence $R(A,X)=0$ and $J(A)=\mathbb{F}b$; in particular, the primitive axis $b$ lies in $J(A)$. Over $\mathbb{C}$, this axial presentation is equivalent to the previously classified $D(-1),{e_2,a_6}$ presentation with fusion law $F_{D3}$. Thus the algebra and axial structure are known; the new point is the computation of its Jacobson radical and the resulting counterexample to the radical-equality question.

math.RA↗

M3TR: Temporal Retrieval Enhanced Multi-Modal Micro-video Popularity Prediction

Accurately predicting the popularity of micro-videos is a critical but challenging task, characterized by volatile, `rollercoaster-like' engagement dynamics. Existing methods often fail to capture these complex temporal patterns, leading to inaccurate long-term forecasts. This failure stems from two fundamental limitations: \ding{172} a superficial understanding of user feedback dynamics, which overlooks the mutually exciting and decaying nature of interactions such as likes, comments, and shares; and~\ding{173} retrieval mechanisms that rely solely on static content similarity, ignoring the crucial patterns of how a video's popularity evolves over time. To address these limitations, we propose \textbf{M$^3$TR}, a \textbf{T}emporal \textbf{R}etrieval enhanced \textbf{M}ulti-\textbf{M}odal framework that uniquely synergizes fine-grained temporal modeling with a novel temporal-aware retrieval process for \textbf{M}icro-video popularity prediction. At its core, M$^3$TR introduces a Mamba-Hawkes Process (MHP) module to explicitly model user feedback as a sequence of self-exciting events, capturing the intricate, long-range dependencies within user interactions (for \textbf{limitation} \ding{172}). This rich temporal representation then powers a temporal-aware retrieval engine that identifies historically relevant videos based on a combined similarity of both their multi-modal content (visual, audio, text) and their popularity trajectories (for \textbf{limitation} \ding{173}). By augmenting the target video's features with this retrieved knowledge, M$^3$TR achieves a comprehensive understanding of prediction. Extensive experiments on two real-world datasets demonstrate the superiority of our framework. M$^3$TR achieves state-of-the-art performance, outperforming previous methods by up to \textbf{19.3}\% in nMSE and showing significant gains in addressing long-term prediction challenges.

cs.MM↗

Focal-point scanning for dose delivery and optimization with focused laser-accelerated very-high-energy electron beams

Focused very-high-energy electron (VHEE) beams can produce localized dose enhancement at selected depths, but irradiation of a finite target requires coordinated control of multiple focal positions, incidence directions, and beam weights while limiting exposure of nearby organs at risk (OARs). We present Focal-Point Scanning (FPS), a dose delivery and optimization method developed for laser wakefield accelerator (LWFA)-driven VHEE beams. The method is based on a two-dipole focusing system that produces single-plane beam convergence and allows the focal position to be varied by changing the magnetic field strength. FPS distributes focal points throughout the planning target volume and determines focal-point-specific incidence sectors according to the geometry of nearby critical OARs. The method was evaluated using the AAPM TG119 C-shape benchmark and one previously treated lung radiotherapy case. At matched target coverage, FPS reduced the TG119 Core mean dose by approximately one half relative to parallel VHEE and intensity-modulated x-ray plans, approaching the single-field proton pencil-beam-scanning reference. In the lung case, FPS maintained target coverage comparable to the clinical volumetric modulated arc therapy reference while reducing the mean dose to every evaluated OAR; spinal-cord mean and maximum doses decreased by 93.2% and 87.2%, respectively. The evaluated OAR mean doses varied little across rms energy spreads of 0 to 10% and for a flat-top electron spectrum spanning 150 to 250 MeV. These results demonstrate that focal-point-specific angular selection can translate focused-beam physics into effective OAR sparing and support FPS as a planning strategy for broadband LWFA-VHEE radiotherapy.

physics.med-ph↗

ALMA visits the QSO MUSEUM: Connecting molecular gas and the cool circumgalactic medium around 37 z~3 quasars

Extended Ly$α$ emission is ubiquitous around quasars and traces the cool circumgalactic medium, providing insights into halo gas dynamics and active galactic nucleus feedback. However, its connection to the cold molecular gas of the host galaxies remains largely unexplored. We characterize the molecular gas reservoirs of quasars at cosmic noon and investigate their connection to extended Ly$α$ emission using ALMA CO(4-3) observations of 37 quasars at $z\sim3$ from the QSO MUSEUM survey, previously mapped in Ly$α$ with VLT/MUSE. We derive molecular gas masses and gas fractions, explore correlations with Ly$α$ nebula and quasar properties, and search for CO-emitting companions. We detect 21/37 quasars in CO(4-3), with gas masses of $M_\mathrm{gas}\approx(3-40) \times10^9\,\mathrm{M_\odot}$. Quasars with the most massive molecular gas reservoirs are associated with the centrally dimmest Ly$α$ nebulae, while those hosting the centrally brightest Ly$α$ nebulae are generally not detected in CO. This suggests that gas and dust in the hosts regulate Ly$α$ escape and consequently affect the emission from halo gas. We find evidence that lower-Eddington-ratio quasars harbor more massive gas reservoirs, while strongly accreting quasars ($λ_\mathrm{Edd} \gtrapprox 0.9$) likely deplete their gas; for example, through powerful quasar-driven outflows. Despite their higher molecular gas masses within the sample, CO-detected low-Eddington quasars exhibit low gas fractions with a median $M_\mathrm{gas}/M_* \sim 0.10$, below what is typically found for inactive star-forming galaxies. Six quasars are marginally resolved in CO, with effective radii up to $\sim 8\,\mathrm{kpc}$. In addition, we detect 14 high-fidelity companion galaxies, indicating overdense quasar environments with a quasar-galaxy cross-correlation length of $9.81^{+2.22}_{-2.05}\,h^{-1}\mathrm{cMpc}$.

astro-ph.GA↗

Harmonic Ranking for Edge-Weighted Oblivious Matching

We study edge-weighted oblivious bipartite matching. The weight of every potential edge is known, but its existence is revealed only when the edge is probed, and a successful probe between two free vertices must be accepted immediately. We give an explicit randomized algorithm with certified competitive ratio $0.698$, improving the previous best guarantee of $0.659$ (Huang, Sun, Wu, and Zhao, FOCS 2025). The result is computer-assisted and verified by a reproducible exact-integer computation. The same algorithm has a $0.698$-competitive online implementation for the vertex-weighted random-arrival model, improving the previous $0.696$ unweighted guarantee of Mahdian and Yan (STOC 2011) and the $0.686$ vertex-weighted guarantee of Peng and Tang (EC 2025). Our algorithm, Harmonic Ranking, is a role-symmetric generalization of \textsc{Ranking}. It assigns an independent random rank $x_z$ to each vertex and probes a potential edge $uv$ in decreasing order of \[ w_{uv}\frac{h(x_u)h(x_v)}{h(x_u)+h(x_v)}. \] This harmonic priority arises from a budget-balanced gain split and a mutual-proposal interpretation. The analysis lifts two cutoff curves into indicators, reducing the exponential-size factor-revealing problem to a polynomial-size directed minimum-cut instance. A maximum-flow computation with rounded-down integer capacities gives a rigorous certificate. Independently, we observe that the finite-grid unweighted relaxation of our factor-revealing program coincides exactly with a Mahdian--Yan program.

cs.DS↗

Breaking the 4-Approximation Barrier in Strategyproof Two-Facility Location

We study strategyproof mechanism design without transfers for the two-facility location problem in metric spaces. A mechanism selects two facility locations based on agents' reported locations; each agent incurs her distance to the nearer facility, and the objective is to minimize the expected social cost. A mechanism is strategyproof if no agent ever benefits from misreporting her location. The best approximation ratio achieved by a randomized strategyproof mechanism has been $4$, attained by the Proportional mechanism of Lu, Sun, Wang, and Zhu (EC 2010), and the best lower bound has been $1.045$, due to Lu, Wang, and Zhou (WINE 2009). Neither bound has moved since then, even on the line $\mathbb{R}$. We improve both bounds. Our main result is a randomized strategyproof mechanism with approximation ratio $11/3 \approx 3.667$ on every Ptolemaic metric space, a rich class containing all Euclidean spaces. The mechanism randomizes between the Proportional mechanism and a new mechanism that we call Global Pair. Global Pair draws an unordered pair of agents with probability proportional to their distance and opens facilities at their reported locations. Although Global Pair and Proportional each have approximation ratio $4$, the two mechanisms attain their worst-case approximation ratios on complementary instances. Randomizing between them balances these complementary weaknesses and breaks the $4$-approximation barrier. On the lower-bound side, we construct a new two-profile instance that yields a lower bound of $(1+\sqrt{2})/2 \approx 1.207$, improving upon the previous lower bound of $1.045$.

cs.GT↗

Alpha as an Efficiency Signal: Visibility-Routed RGBA Image-to-Video Generation

RGBA videos combine RGB appearance with an alpha channel, enabling animated assets to be applied across arbitrary backgrounds, which are heavily used in gaming industry. However, generating high-quality RGBA animations for games remains challenging for two reasons. First, most existing RGBA video datasets are dominated by photorealistic content, with limited coverage of game assets. Second, the traditional generate-then-matte pipelines estimate alpha only after RGB synthesis, so semi-transparent regions are often blurred by background, resulting in unstable matting outputs. More recently, many methods have begun to model RGB and alpha jointly, but existing approaches are mostly text-conditioned, and still have unresolved issues in efficiency and quality. To address these challenges, we introduce GameAlpha-2.4K, a 2.4K-clip game-style RGBA video dataset built with matte-friendly synthesis, multi-hypothesis alpha recovery, and compositing-based quality gates. Using this dataset, we train a reference-conditioned RGBA video generator that jointly produces RGB frames and alpha mattes in a single pass. To improve efficiency, we propose a visibility router that identifies transparent tokens in an early stage and bypasses their later DiT updates, while x_0-lock guides them along the original flow-matching schedule toward self-predicted endpoints. Our model obtains lower FVD than traditional two-stage pipelines, and the visibility router skips 35% of token evaluations in the final two DiT denoising steps, providing a 1.2x backbone speedup with negligible quality degradation compared to dense inference.

cs.CV↗

Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucination

Large Vision-Language Models (LVLMs) often suffer from object hallucination, generating objects that are absent from the image. Prior work largely attributes this to insufficient visual attention. However, we find that both real and hallucinated objects receive equally strong visual attention in the model's mid-to-late layers, suggesting that the key issue may not be how much the model attends, but what it attends to and why. To this end, we decode the visual features of high-attention regions using Logit Lens, and observe that regions corresponding to real objects can be correctly decoded to the target object tokens, whereas those for hallucinated objects cannot. Building on this, we identify two hallucination mechanisms: (i) visual uncertainty, triggered by semantically similar or confusable regions; masking these regions eliminates the hallucination. (ii) contextual prior, triggered by strong co-occurrence priors; even when the initially attended region is masked, the hallucination persists and attention drifts to other regions. Based on these findings, we propose a simple yet effective training-free Detect-Mitigate framework comprising a Logit-Lens Consistency Check to detect hallucination and targeted remedies: High-Attention Regions Masking (HARM) for visual uncertainty hallucination, and Visual Evidence Enhanced Decoding (VEED) for contextual prior hallucination. Our approach achieves state-of-the-art results on multiple hallucination benchmarks. Code will be available.

cs.CV↗

FAST Ultra-Deep Survey: the baryonic Tully-Fisher relation in FUDS0 field

The Baryonic Tully-Fisher relation (BTFR) is one of the tightest scaling relations for disk galaxies in the local Universe, and therefore is an important tool for studying the fomation and evolution of galaxies. However, the evolution of the BTFR over cosmic time is poorly understood due to the limited sample of HI galaxies beyond the local Universe, limitations of optically-derived rotation curves, and selection effects. In this work, we explore the BTFR at redshifts up to $z=0.42$ from galaxies detected in the pilot FAST Ultra-Deep Survey (FUDS) field, FUDS0. As found in previous work, we identify two components in the plane of baryonic mass versus rotational velocity, $C_{\rm BTFR}$ (tight) and $C_{\rm Outlier}$ (dispersed). A Gaussian mixture model is employed to recover the BTFR, yielding the best fit parameters for the slope $k=3.32_{-0.11}^{+0.12}$, zero point $b=10.07_{-0.03}^{+0.03}$, and intrinsic scatter $σ_{\rm BTFR}=0.036_{-0.009}^{+0.010}$. A random forest classifier is used to investigate the origin of the outlier component. We find that low signal significance and inaccurate inclinations are the key factors that contribute to the outlier population, indicating that observational effects are the dominant origin. Evolutionary trends are examined in three different redshift bins. Both the slope and zero point show consistency within 1-$σ$ uncertainty in the two low redshift bins, indicating no significant evolution. The indirectly inferred BTFR parameters from the $C_{\rm Outlier}$ component in the highest redshift bin aligns with the conclusion. The ongoing full FUDS survey will provide a larger sample to enable more accurate constraints on BTFR evolution.

astro-ph.GA↗

Structured High-Angular-Momentum Coulomb Tensors from Real and Complex Solid-Harmonic Integral Engines: A Perspective

Electron-repulsion integrals describe the Coulomb interaction between charge distributions built from orbital basis functions. Most integral algorithms generate these quantities through Cartesian Gaussian functions, whose angular shapes are written as powers of $x$, $y$, and $z$, and then transform the result to spherical functions. This route is effective, but from $d$ shells onward the Cartesian representation contains more functions than the spherical space required by the calculation. Direct real or complex solid-harmonic engines work in that target space from the beginning. They therefore produce a smaller final Coulomb tensor while preserving the ordering, phase, and magnetic-quantum-number labels that describe its angular structure. Following this structure beyond integral evaluation reveals direct connections to the algorithms that use the tensor. Simple analytical counts quantify tensor size, angular blocks, radial Slater--Condon parameters, and pair-space work. These quantities guide low-rank factorization, local Hamiltonian construction, quantum simulation, and transformations to spinor or effective-model bases. In this way, solid-harmonic integral engines provide a direct bridge between efficient integral generation and structured many-electron computation.

physics.chem-ph↗

Random-Order Online Facility Location Beyond Uniform Opening Costs

We study online metric facility location in the random-order model with arbitrary positive opening costs. A finite set of candidate facilities and their costs is known in advance, while an adversary fixes a multiset of demand points that arrives in a uniformly random order. This setting includes both prescribed candidate sites and the classical finite full-space node-cost model. For a known horizon, we give a deterministic $4.2674$-competitive algorithm, improving the previous factor $33$ for nonuniform opening costs. At rank $t$, the algorithm uses the positive normalized rank $q_t=t/n$, chooses a candidate minimizing $d(x,y)+λ_t f_y$, where $λ_t=\min\{1,q_t/μ\}$, and opens it when the current connection distance covers this penalized objective. The analysis uses a monotone one-round charge and an upper-envelope decomposition to control later points and the first point of each optimal cluster. With unit opening costs, the rule reduces exactly to a cutoff on the distance improvement attainable from a nearest candidate. A supplementary appendix gives the sharper analysis of the closely related zero-start rank cutoff and obtains a ratio below $3.2805$. We also prove a $3-o(1)$ lower bound for arbitrary randomized online algorithms. The lower bound already holds with uniform costs on a prescribed candidate set and transfers, without loss, to the finite full-space model with nonuniform opening costs. Together with the recent competitive ratio below $2.42$ for full-space uniform costs, this yields a strict separation between the full-space uniform- and nonuniform-cost models.

cs.DS↗

Absolute charge calibration of DRZ phosphor screens for relativistic electron bunches

Laser-plasma accelerators have been the subject of extensive research in recent years. The electron beams they generate exhibit a broad energy spread. To conveniently characterize beams from laser wakefield acceleration (LWFA), electron spectrometers employing scintillating screens coupled with CCD cameras are typically used. In this work, we calibrate a series of DRZ phosphor screens and measure the spectra of the light they emit. The calibration was performed using the radio-frequency linear electron accelerator at Tsinghua University, which provided monoenergetic electron beams with peak energy of approximately 30 MeV.

physics.acc-ph↗

Unified Uncertainty Quantification Framework Bridging Noisy Quantum Backends Across Variational Quantum Algorithms and Quantum Signal Processing

We present an uncertainty quantification (UQ) framework for application level benchmarking and characterization of noisy quantum backends. The framework compares two workload classes under one statistical pipeline: noisy intermediate scale quantum (NISQ) variational quantum algorithms (VQAs) and Quantum Singular Value Transformation (QSVT) based Green's function reconstruction. For the VQA branch, we evaluate ten benchmark families spanning chemistry, optimization, simulation, compiling, linear solving, partial differential equations, metrology, error correction, tomography, and channel fidelity estimation. For the QSVT branch, we reconstruct orbital resolved Green's functions and spectral peaks from a block encoded real time propagator. The workflow combines Bayesian optimization, posterior distribution refinement, sensitivity analysis, robust parameter density estimation, backend ranking, noise correlation, and resource estimation analysis. Instead of reporting only one best parameter vector, the framework identifies robust parameter regions, residual gaps to ideal behavior, backend specific failure modes, and calibration sensitive uncertainty. The result is a common benchmark for variational and non-variational workloads that measures how reliably each backend reaches useful task level behavior.

cs.ET↗

The Stack Search Tests on FAST Data: Discovery of Six Faint Isolated Millisecond Pulsars in NGC 6517 and NGC 7078 (M15)

We report the discovery of six faint millisecond pulsars (MSPs) in the globular clusters NGC 6517 and NGC 7078 (M15) using the Five-hundred-meter Aperture Spherical radio Telescope (FAST). These discoveries were enabled by stacking power spectra from multiple observations, a method that effectively boosts the signal-to-noise ratio of faint sources. In NGC 6517, we identified four new MSPs (NGC 6517S-V) with spin periods ranging from 3.68 to 6.02 ms and dispersion measures (DMs) between 182.45 and 182.85 pc cm^-3. In M15, two additional MSPs (M15M and M15N) were discovered, with spin periods of 4.83 and 9.28 ms, and DMs of 67.89 and 66.65 pc cm^-3, respectively. A phase-coherent timing solution has been obtained for M15M; however, sparse detection rates currently preclude phase-connected solutions for the remaining five pulsars. Current timing parameters suggest all six MSPs are isolated, which is consistent with the expected pulsar populations in core-collapsed globular clusters. Notably, pulsars M15N, NGC 6517U, and NGC 6517V eluded detection by standard frequency-domain searches (e.g., PRESTO-based) and the Fast Folding Algorithm, demonstrating that the stack search technique significantly enhances detection sensitivity to inherently faint pulsar signals.

astro-ph.HE↗

Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents

In long-horizon tasks, decision-relevant state is often scattered across an expanding trajectory, while the action agent must surface it and act. As trajectories grow, task requirements, environment facts, prior attempts, diagnoses, and open subgoals can be buried in the context window or pushed beyond it, failing to influence decisions when needed. We call this failure mode "behavioral state decay". We study memory as an active intervention mechanism rather than passive retrieval. A separate memory agent runs alongside an unmodified action agent, updating a structured memory bank from the recent trajectory and deciding whether to inject a memory-grounded reminder or remain silent. The module is plug-and-play with frontier action agents and existing agent harnesses. Across Terminal-Bench 2.0 and $τ^2$-Bench, it improves pass@1 for both weaker and stronger action agents, with gains of +8.3 pp on Terminal-Bench and +6.8 pp on $τ^2$-Bench. Ablations show that selective intervention outperforms passive bank exposure, always-on injection, advisor-only guidance, and general retrieval. As an early step toward open-weight memory policies, we train Qwen3.5-27B on SETA using SFT and GRPO, improving validation reward and achieving partial transfer to Terminal-Bench.

cs.AI↗

Measurement-Access Risk Frontiers for Autonomous Scientific Control

Rapidly scaling autonomous science is limited not only by algorithms, compute or data volume, but by which physical records a platform exposes before action. We formulate physically accessible decision-making (PADM) and a measurement-access risk frontier: the Bayes-optimal target risk minimized over records realizable under cost, bandwidth, latency, disturbance, memory and actuation constraints. The frontier gives a no-free-autonomy limit: automation cannot collapse decision uncertainty by computation alone; an optimal controller cannot remove target components absent from its record, and closing that gap requires expanded access, auditing, tolerated disturbance, slower operation or restricted deployment. In monitored feedback, displacement-only control remains exposed to a hidden switching force, whereas a finite-bandwidth cue recovers part of the missing projection before action. A chemistry-aware candidate-ranking audit with a 1000-target stress panel, Gaussian sensing, hidden-regime decisions and cost-aware/thermodynamic channel selection provide reproducible checks. PADM identifies target-specific audit value and residual oracle gaps before deployment.

math-ph↗