Search arXiv⌕ Search

arXiv subjects

Bo Yang

Publications and source records attributed to Bo Yang.

At least 37 records · Page 2Linked to original sources

RelightFormer: Feed-forward Generative Transformer for Multiview Object Relighting

Image relighting is traditionally tackled via complex inverse rendering pipelines, which suffer from ill-posed optimization, or single-image generative models that ignore crucial multi-view cues necessary for understanding 3D geometry and material interactions. To address these limitations, we introduce a feed-forward generative Transformer for direct single- and multi-view image relighting that entirely bypasses explicit intrinsic property estimation. Adapted from a video foundation model, our architecture features a latent illumination module that dynamically injects target environment maps into spatial features via cross-attention. Furthermore, we employ permutation-invariant positional encodings to symmetrically process unordered multi-view inputs without sequential bias. To train this robust data-driven model, we construct the massive Laval Objaverse Dataset (LOD), comprising 90K objects and 39K unique illuminations. Extensive experiments demonstrate state-of-the-art visual quality, photorealistic relighting quality, and strong zero-shot generalization across single-view, multi-view, and novel-view relighting tasks.

cs.CV↗

ReliableRAG: Combating Misinformation in Retrieval-Augmented Generation via Reliability-Guided Reasoning Chains

Retrieval-Augmented Generation (RAG) has emerged as a powerful architecture for Question Answering (QA) by integrating external information into Large Language Models (LLMs). However, false, inaccurate, and misleading information in news and social media poses a serious challenge to real-world RAG systems, especially in multi-hop QA, where complex multi-step reasoning can be misled by even a single deceptive misinformation segment in the retrieved documents. Existing approaches mainly rely on implicit alignment or explicit regulation, but their limited ability to assess fine-grained information reliability makes them vulnerable to deceptive misinformation that is semantically relevant to the question yet factually incorrect, leading to erroneous answers. To address this limitation, we propose ReliableRAG, which, to the best of our knowledge, is the first reliability-driven framework that mitigates deceptive misinformation in multi-hop QA through fine-grained evaluation of individual triples. ReliableRAG first extracts information segments from source documents and represents them as structured triples. It then quantifies triple reliability by combining query-triple semantic relevance with triple credibility, retaining only the top-$K$ reliable and non-redundant triples. Based on these refined triples, ReliableRAG autoregressively constructs robust reasoning chains to consolidate trustworthy evidence and filter deceptive misinformation, producing accurate answers faithful to reliable information. Experiments on three multi-hop QA datasets show that ReliableRAG outperforms existing methods, substantially improving the factual reliability and robustness of RAG systems under deceptive misinformation injection.

cs.CL↗

Sequential Magnetic Reconnections in a Fishbone-like Structure Leading to Recurrent Brightenings

The fine-scale release of magnetic free energy in the solar atmosphere is a fundamental open question in solar physics. Multi-wavelength observations at high spatiotemporal resolution now offer a direct window into this process. Using data from NVST, SDO, IRIS, and Hinode, we reveal the energy release process in a fishbone-like magnetic structure within active region 12297. The fishbone-like structure consists of a spine along a narrow, elongated positive-polarity field, with herringbone branches rooted in negative-polarity sunspots. Persistent photospheric magnetic flux emergence and shearing motions are observed beneath the fishbone-like structure, which may play a key role in maintaining its topology and producing the recurrent brightenings. During brightenings, compact bright features propagate sequentially from west to east along the spine, accompanied by bidirectional flows along the branches and plasma blobs ejected toward the distant positive sunspot. These propagating features could be interpreted as signatures of sequential magnetic reconnection events occurring in chronological order at nodes along the spine, predominantly in the chromosphere and transition region. Our observations may provide evidence that recurrent brightenings could arise from repeated sequential reconnection events organized by a coherent magnetic structure, deepening our understanding of how such recurrent brightenings are generated and how magnetic free energy is dissipated at fine scales in solar active regions.

astro-ph.SR↗

Spatiotemporal organization in bike-sharing systems using gravity model parameters

Gravity models describe bike-sharing origin-destination (OD) flows as increasing with origin and destination activity and decreasing with distance. Yet how these relationships change within a day and across observation areas remains unclear. We examine temporal and spatial variation in origin, destination and distance exponents $α$, $β$, and $γ$, together with $R^2$ and mean absolute error (MAE), across eight bike-sharing systems. Temporal modeling uses sliding windows, while spatial modeling expands a circular area and separates intra-zonal, cross-zonal outflow, cross-zonal inflow and extra-zonal trips. We find recurring patterns across cities: morning and evening peaks differ in origin and destination dependence, while cross-zonal outflow and inflow across the same boundary show opposite changes in their relative origin and destination dependence as radius increases. Model performance and distance dependence also vary with time and radius. A full-day, citywide fit therefore combines time periods and flow types with different gravity relationships.

physics.soc-ph↗

SkyNative: A Native Multimodal Architecture for Remote Sensing Vision-Language Understanding

Remote sensing vision-language models (RS-VLMs) commonly employ a pretrained vision encoder and a projection module to map image features into the token space of a large language model. Although effective, this modular RS-VLMs separates visual representation from language reasoning, potentially limiting the direct involvement of fine-grained visual evidence in complex spatial inference. This challenge is particularly relevant to remote sensing imagery, which often covers broad geographic areas and contains multi-scale objects, dense target distributions, and intricate spatial layouts. In this paper, we propose SkyNative, the first study to explore a native multimodal architecture for remote sensing vision-language tasks. SkyNative converts remote sensing images into visual tokens through a lightweight patch embedding module and places them together with text tokens in a shared autoregressive sequence, allowing textual tokens to directly access the preceding visual context. To accommodate the heterogeneous characteristics of the two modalities, we further adopt a modality-aware decoupling mechanism that applies modality-specific projections, normalization, and feed-forward transformations while processing both modalities through shared causal self-attention. Extensive experiments demonstrate SkyNative's strong capabilities in dense small-object perception, large-format contextual understanding, complex reasoning, and robustness, with scores of 68.93%, 47.40%, and 63.43% on HRRSD, RSHR reasoning, and OmniEarth, respectively. These results suggest that the native VLM architecture explored in SkyNative represents a promising approach to RS vision-language modeling.

cs.CV↗

Scalar curvature of blow-ups of compact Kähler manifolds along complex submanifolds

Let $(M,ω)$ be a compact Kähler manifold with its scalar curvature $S(ω)$, Assume that $M$ has complex dimension at least $3$ and contains a complex submanifold $X$ of complex codimension at least $2$. Let $σ: Bl_{X} M \rightarrow M$ denote the blow-up of $M$ along $X$. We show that $Bl_{X} M$ admits a sequence of Kähler metrics $\{\widetildeω_{i}\}_{i \geq 1}$ whose scalar curvatures $S(\widetildeω_i)$ converge to $σ^{\ast} (S(ω))$ in the $C^0(Bl_{X} M)$ norm. Our work is motivated by a recent result of Brown, who established the corresponding result for blow-ups at a point. The proof is based on the gluing method for constructing extremal Kähler metrics on blow-ups, together with new analytic tools and several modifications needed in our setting.

math.DG↗

Non-Abelian braiding in Abelian Fractional Quantum Hall Phases from realistic interactions

We propose a method of realizing non-Abelian braiding of fractionalized quasiholes in the Laughlin fractional quantum Hall phase at $ν=1/3$ with realistic two-body interactions within the lowest Landau level. It is numerically shown that low-lying gapped excitations near $ν=1/3$ are contained almost entirely within the null space of the three-body Moore-Read model Hamiltonian. They are thus quantum fluids of non-Abelian quasiholes that are in principle physically accessible. In particular, Laughlin ground state can be described as a fluid of ``$ψ$-type" quasiholes formed by binding a magnetic flux with a Majorana fermion (MF), and the Laughlin quasiholes are described by the ``$1$-type'' quasiholes, which are magnetic fluxes without a MF attached. Within the Laughlin phase, Laughlin quasiholes can be locally fractionalized into non-Abelian quasiholes, when the strong attraction between them is overcome by properly designed one-body electronstatic trapping potentials. Extensive numerics with proper finite-size scaling corroborate this physical picture, and our study points to the possibility of realizing non-Abelian braiding within an Abelian topological phase in experiment without the need for fine-tuning realistic electron-electron interaction.

cond-mat.str-el↗

Holomorphic functions on complete Hermitian manifolds with flat Chern connection

A classical result of Boothby states that any compact Hermitian manifold with flat Chern connection is covered by a complex Lie group. In this work, we prove a sharp generalization: the universal cover of any complete Hermitian manifold with vanishing Chern curvature and torsion of sublinear growth is holomorphically isometric to a complex Lie group equipped with a left-invariant metric. The proof relies on a new gradient estimate for holomorphic functions on complete Hermitian manifolds with nonnegative second Ricci curvature. Combining this estimate with methods from sub-Riemannian geometry, we establish quantitative characterizations of function theory on these manifolds.

math.DG↗

ContestTrade: A Multi-Agent Trading System Based on Internal Contest Mechanism

In financial trading, large language model (LLM)-based agents demonstrate significant potential, but their decisions can be sensitive to noisy and non-stationary market information. We propose ContestTrade, a multi-agent trading system with an internal competitive mechanism inspired by institutional investment workflows. The system consists of two specialized teams: (1) a Data Team that processes and condenses massive market data into diversified textual factors optimized for constrained LLM context windows, and (2) a Research Team that produces parallelized multipath trading decisions via tool-augmented deep research. The core design is a "Quantify-Predict-Allocate" contest mechanism within each team: agent outputs are scored only after market outcomes become observable, future utility is predicted from historical scores, and resources are allocated to agents with positive predicted utility. In a post-2024 A-share backtest, ContestTrade achieves higher backtested return and risk-adjusted performance than the evaluated baselines. We further describe the temporal protocol, implementation choices, and limitations to clarify the scope of these results.

q-fin.TR↗

Many-Anyon Braiding in Non-Abelian Fractional Quantum Hall Effect with Hybrid Monte Carlo Simulation

We employ the hybrid Monte Carlo method to efficiently compute the many-anyon non-Abelian braiding matrices associated with different braiding schemes of the Moore-Read quasiholes. A novel proposal in this work is that anyon braiding schemes based on a global rotation are robust against finite-size effects, as demonstrated by benchmarking their errors in the braiding matrix against those of a simple two-anyon exchange. Moreover, we investigate how electron-electron interactions and local electrostatic trapping potentials influence the energetic preference of different fusion channels. Their effect on the non-Abelian braiding matrices has been verified, a surprising phenomenon that demonstrates long-range entanglement of non-Abelian states. Our results are relevant to the experimental realization of non-Abelian physics in fractional quantum Hall and other analogous systems, including the fast-growing field of fractional quantum anomalous Hall states in moiré materials.

cond-mat.str-el↗

Parameter- and Bandwidth-Efficient Edge--cloud Many-to-Many Speech-to-Text Translation

Multimodal large language models (MLLMs) have demonstrated significant potential for speech-to-text translation (S2TT). However, existing deployment paradigms face critical challenges: pure on-device models suffer from resource constraints, while centralized cloud systems incur bandwidth bottlenecks and privacy risks by transmitting raw voice data. In this paper, we propose Edge--cloud Speech Recognition and Translation (ESRT), a parameter-efficient, bandwidth-efficient, and privacy-aware collaborative Edge--cloud MLLM framework. First, we introduce a multi-task weighted curriculum learning strategy to mitigate catastrophic forgetting, improve multilingual balance, and train parameter-efficient ESRT-1B, ESRT-4B, and ESRT-12B models. Second, we enable bandwidth-efficient Edge--cloud inference by retaining a lightweight speech encoder and adapter on the device and transmitting only a compressed tensor to the cloud. Extensive experiments on FLEURS demonstrate that ESRT models achieve state-of-the-art S2TT performance across 45 languages ($45 \times 44$ directions). Relative to raw audio, ESRT and ESRT-Lite reduce the transmitted tensor size by $5.1\times$ and $10.2\times$, respectively, while keeping raw speech on-device and avoiding its direct exposure to the cloud. The code and models are released to facilitate reproducible, privacy-aware S2TT research.

cs.AI↗

Rogue-wave and lump patterns associated with the third Painlevé equation

We report rogue-wave and lump patterns associated with Umemura polynomials, which arise in rational solutions of the third Painlevé equation. We first show that in many integrable equations such as the nonlinear Schrödinger equation and the Boussinesq equation, when internal parameters of their rogue wave solutions are large and of certain form, then their rogue patterns in the spatial-temporal plane can be asymptotically predicted by root distributions of Umemura polynomials (or equivalently, pole distributions of rational solutions to the third Painlevé equation). Specifically, every simple root of the Umemura polynomial would induce a fundamental rogue wave whose spatial-temporal location is linearly related to that simple root, while a multiple root of the Umemura polynomial would induce a non-fundamental rogue wave in the $O(1)$ neighborhood of the spatial-temporal origin. Next, we show that in a certain class of higher-order lump solutions of the Kadomtsev-Petviashvili-I (KPI) equation, when their internal parameters are large and of certain form, then their lump patterns at $O(1)$ time can also be predicted asymptotically by root distributions of Umemura polynomials, where simple and multiple roots of the polynomial would give rise to fundamental and non-fundamental lumps in the spatial plane, respectively. These results reveal the importance of the third Painlevé equation in studies of nonlinear wave patterns. We also report a new transformation which turns bilinear rogue-wave solutions of the nonlinear Schrödinger equation to higher-order lump solutions of the KPI equation.

nlin.SI↗

Breaking the Curse of Multilinguality in Many-to-Many Speech-to-Text Translation via a Resource-Aware Mixture of Speech Encoders

Multimodal large language models (MLLMs) have achieved significant success in speech-to-text translation (S2TT). However, when processing multilingual speech inputs, a single speech encoder shared across all languages suffers from the curse of multilinguality: languages at different resource levels compete for limited representation capacity, leading to strong high-resource performance but substantial degradation on low-resource speech. To address this problem and improve multilingual consistency, we propose MSRT, a novel framework built around a resource-aware Mixture of Speech Encoders (MoSE). MoSE uses an explicit language router to assign each utterance to an appropriate expert encoder. A frozen expert preserves high-resource language capabilities, while a trainable expert adapts to and specializes in medium- and low-resource languages. We further introduce a five-stage curriculum learning strategy that substantially reduces data dependence, requiring only 10 hours of paired S2TT data per language for effective alignment. We conduct extensive experiments on 45 languages, systematically evaluating all $45 \times 44$ translation directions. Our 4B-parameter model achieves state-of-the-art performance, outperforming substantially larger baselines. Empirical analyses show that MoSE improves high-, medium-, and low-resource languages simultaneously, with the largest gains on low-resource speech, thereby breaking the curse of multilinguality without compromising high-resource performance. To support future multilingual S2TT research, we release our code and models.

cs.CL↗

Fairness in Augmented Graph Learning: A Survey

Graph learning has evolved into Augmented Graph Learning (AGL) by integrating specialized machine learning (ML) techniques. Examples include federated learning, graph transformers, and graph condensation. While enhancing model utility, AGL introduces unique intersectional fairness challenges that traditional GNN debiasing frameworks, which primarily focus on message-passing regulations, fail to address. This paper provides a systematic investigation into this emerging field, termed FairGX. We first delineate the shift from conventional fairness-aware graph learning to the FairGX paradigm, identifying novel bias sources inherent in ML augmentations, such as dual-side disparities in federated aggregation and attention-head skewness. A structured taxonomy is established to categorize existing literature based on their technical integration and fairness objectives. Furthermore, we analyze the impact of diverse ML paradigms on algorithmic equity, emphasizing the unique challenges in human-centered applications and the absence of a unified framework. We conclude by identifying five critical future directions, including novel metrics for AGL, fairness-privacy synergy, and Fairness-aware LLM4Graph/Graph4LLM. This survey serves as a foundational roadmap for developing robust and equitable graph systems in complex ML environments.

cs.LG↗

Verifiable blind probabilistic error cancellation

Quantum error mitigation (QEM) is an essential tool for mitigating hardware noise without incurring space overhead. Yet, its reliability depends on modeling, calibration, and implementation, leaving end-to-end security on untrusted quantum hardware unresolved. We address this problem by introducing verifiable blind probabilistic error cancellation (VBPEC), the first secure verification protocol against a fully malicious adversary that integrates QEM. VBPEC brings probabilistic error cancellation (PEC), a widely studied QEM technique, within the scope of composable security by formalizing delegated mitigation as a cryptographic resource in the abstract cryptography framework. The protocol performs PEC with perfect blindness and an exponentially small security error. VBPEC retains the absence of quantum-space overhead from recent statistically-secure verified quantum computation protocols and from PEC. The only overhead takes the form of additional repetitions due to the QEM procedure. To achieve this, we extend trap-based verification from deterministic pass/fail checks to statistical tests that benefit from QEM and develop a new proof technique that integrates the corresponding additional deviation sources. Rather than merely tolerating honest noise below a fixed threshold, VBPEC actively cancels it, enabling correctly mitigated estimates to be accepted with high probability without compromising security. Our framework thus establishes an essential route towards secure, reliable, and practical delegated quantum computation on near-future quantum hardware: VBPEC fundamentally improves the practicality of verification.

quant-ph↗

Three-term Recurrence Relation with Arbitrary Degree Steps for Orthogonal Polynomials

An approach to generate three-term recurrence relations with arbitrary degree steps is proposed for orthogonal polynomials. Specifically, given any class of orthogonal polynomials $\{Q_{p}(x)\}_{p=0}^{\infty}$ defined by Favard's theorem, we employ the adjacent members $Q_{p}(x)$ and $Q_{p-1}(x)$ to compute $Q_{p+s}(x)$ of high degree and the one of low degree $Q_{p-t}(x)$, where $(s,t)$ are parameters for degree step adjustment. The coefficients of both relations are analyzed, revealing novel properties that enable the derivation of three-term recurrence relations with respect to $Q_{p+s}(x)$, $Q_{p}(x)$ and $Q_{p-t}(x)$ by eliminating $Q_{p-1}(x)$. Furthermore, in addition to the standard recursive formula, which is characterized by degree increase, the formulas for degree decrease and end-to-middle directions are also formulated. Moreover, explicit recurrence relations with 2-degree steps are presented for Hermite, Gegenbauer and Legendre polynomials. The computation precision of the proposed recurrence relations is also compared with that of the standard ones.

math.NA↗

On Success and Simplicity: A Second Look at Transferable Vision-Language Attack Pipeline

Vision-Language Pre-training Models (VLPMs) are known to be vulnerable to adversarial attacks. Recent transferable attacks on VLPMs have followed a common pipeline with complicated loss functions or multi-stage text/image attacks. However, in this paper, we demonstrate that such a sophisticated attack pipeline can be simpler yet more successful. Specifically, we identify three previously overlooked issues caused by inappropriate cross-modal interactions and excessive operations. To address them, we propose the Simple Vision-Language Attack (SimVLA) pipeline, which observably improves transferability and efficiency. Experiments on four datasets and three downstream tasks validate the superiority of our pipeline. For instance, on Flickr30k text-image retrieval dataset, our SimVLA outperforms the SOTA baseline in R@1 transferability by 8.01\%-14.71\%, while consuming only about 35.73\% of the time and 46.26\% of the max VRAM. Overall, the superiority of our SimVLA highlights the importance of leveraging domain knowledge (e.g., our proposed cross-modal word identification), while blindly pursuing intricate operations (e.g, complex loss functions and redundant multi-stage designs) may even be harmful. We hope our SimVLA can serve as a simple yet effective backbone for future extensions. Code is available at https://github.com/RYC-98/SimVLA.

cs.CV↗

Task-Oriented Sensing and Covert Transmissions for Collaborative Multi-AUV Systems

In underwater covert cooperative missions, autonomous underwater vehicles (AUVs) often cannot rely on active sonar to continuously obtain complete information, since active sensing and frequent communications increase the risk of exposure. As a result, AUVs primarily rely on passive observation, an approach that yields incomplete local perception and limited task efficiency. Although underwater acoustic communications can mitigate this limitation through information sharing, they are simultaneously constrained by long delays, severe interference, low reliability, and the risk of covert exposure. Existing communications-oriented multi-agent reinforcement learning (MARL) studies often model communication as an ideal information flow, whereas traditional communication optimization primarily focuses on link-level performance. However, both are insufficient to characterize the actual contribution of perceptual information to cooperative tasks under realistic conditions of covert physical communications. This paper proposes a Sensed Information Value Realization Multi-Agent Reinforcement Learning (SVR-MARL) framework that leverages practical information to characterize the utility of information for cooperative tasks and learns distributed cooperative policies under realistic communication and covert constraints. Through a case study of covert multi-AUV cooperative localization and tracking, the potential of the proposed framework to improve collaborative task efficiency while reducing unnecessary communication and exposure risks is demonstrated.

cs.LG↗