Search arXiv⌕ Search

arXiv subjects

Yang Yu

Publications and source records attributed to Yang Yu.

At least 181 records · Page 10Linked to original sources

POINTS1.5: Building a Vision-Language Model towards Real World Applications

Vision-language models have made significant strides recently, demonstrating superior performance across a range of tasks, e.g. optical character recognition and complex diagram analysis. Building on this trend, we introduce a new vision-language model, POINTS1.5, designed to excel in various real-world applications. POINTS1.5 is an enhancement of POINTS1.0 and incorporates several key innovations: i) We replace the original CLIP vision encoder, which had a fixed image resolution, with a NaViT-style vision encoder that supports native dynamic high resolution. This allows POINTS1.5 to process images of any resolution without needing to split them into tiles. ii) We add bilingual support to POINTS1.5, significantly enhancing its capability in Chinese. Due to the scarcity of open-source Chinese datasets for vision-language models, we collect numerous images from the Internet and annotate them using a combination of manual and automatic methods. iii) We propose a set of rigorous filtering methods for visual instruction tuning datasets. We comprehensively evaluate all these filtering methods, and choose the most effective ones to obtain the final visual instruction tuning set. Thanks to these innovations, POINTS1.5 significantly outperforms POINTS1.0 and demonstrates strong performance across a range of real-world applications. Notably, POINTS1.5-7B is trained on fewer than 4 billion tokens and ranks first on the OpenCompass leaderboard among models with fewer than 10 billion parameters

cs.CV↗

Universal and Context-Independent Triggers for Precise Control of LLM Outputs

Large language models (LLMs) have been widely adopted in applications such as automated content generation and even critical decision-making systems. However, the risk of prompt injection allows for potential manipulation of LLM outputs. While numerous attack methods have been documented, achieving full control over these outputs remains challenging, often requiring experienced attackers to make multiple attempts and depending heavily on the prompt context. Recent advancements in gradient-based white-box attack techniques have shown promise in tasks like jailbreaks and system prompt leaks. Our research generalizes gradient-based attacks to find a trigger that is (1) Universal: effective irrespective of the target output; (2) Context-Independent: robust across diverse prompt contexts; and (3) Precise Output: capable of manipulating LLM inputs to yield any specified output with high accuracy. We propose a novel method to efficiently discover such triggers and assess the effectiveness of the proposed attack. Furthermore, we discuss the substantial threats posed by such attacks to LLM-based applications, highlighting the potential for adversaries to taking over the decisions and actions made by AI agents.

cs.CL↗

Stable Continual Reinforcement Learning via Diffusion-based Trajectory Replay

Given the inherent non-stationarity prevalent in real-world applications, continual Reinforcement Learning (RL) aims to equip the agent with the capability to address a series of sequentially presented decision-making tasks. Within this problem setting, a pivotal challenge revolves around \textit{catastrophic forgetting} issue, wherein the agent is prone to effortlessly erode the decisional knowledge associated with past encountered tasks when learning the new one. In recent progresses, the \textit{generative replay} methods have showcased substantial potential by employing generative models to replay data distribution of past tasks. Compared to storing the data from past tasks directly, this category of methods circumvents the growing storage overhead and possible data privacy concerns. However, constrained by the expressive capacity of generative models, existing \textit{generative replay} methods face challenges in faithfully reconstructing the data distribution of past tasks, particularly in scenarios with a myriad of tasks or high-dimensional data. Inspired by the success of diffusion models in various generative tasks, this paper introduces a novel continual RL algorithm DISTR (Diffusion-based Trajectory Replay) that employs a diffusion model to memorize the high-return trajectory distribution of each encountered task and wakeups these distributions during the policy learning on new tasks. Besides, considering the impracticality of replaying all past data each time, a prioritization mechanism is proposed to prioritize the trajectory replay of pivotal tasks in our method. Empirical experiments on the popular continual RL benchmark \texttt{Continual World} demonstrate that our proposed method obtains a favorable balance between \textit{stability} and \textit{plasticity}, surpassing various existing continual RL baselines in average success rate.

cs.LG↗

PhyTracker: An Online Tracker for Phytoplankton

Phytoplankton, a crucial component of aquatic ecosystems, requires efficient monitoring to understand marine ecological processes and environmental conditions. Traditional phytoplankton monitoring methods, relying on non-in situ observations, are time-consuming and resource-intensive, limiting timely analysis. To address these limitations, we introduce PhyTracker, an intelligent in situ tracking framework designed for automatic tracking of phytoplankton. PhyTracker overcomes significant challenges unique to phytoplankton monitoring, such as constrained mobility within water flow, inconspicuous appearance, and the presence of impurities. Our method incorporates three innovative modules: a Texture-enhanced Feature Extraction (TFE) module, an Attention-enhanced Temporal Association (ATA) module, and a Flow-agnostic Movement Refinement (FMR) module. These modules enhance feature capture, differentiate between phytoplankton and impurities, and refine movement characteristics, respectively. Extensive experiments on the PMOT dataset validate the superiority of PhyTracker in phytoplankton tracking, and additional tests on the MOT dataset demonstrate its general applicability, outperforming conventional tracking methods. This work highlights key differences between phytoplankton and traditional objects, offering an effective solution for phytoplankton monitoring.

cs.CV↗

Impressions of Parton Distribution Functions

Parton distribution functions (DFs) have long been recognised as key measures of hadron structure. Today, theoretical prediction of such DFs is becoming possible using diverse methods for continuum and lattice analyses of strong interaction (QCD) matrix elements. Recent developments include a demonstration that continuum and lattice analyses yield mutually consistent results for all pion DFs, with behaviour on the far valence domain of light-front momentum fraction that matches QCD expectations. Theory is also delivering an understanding of the distributions of proton mass and spin amongst its constituents, which varies, of course, with the resolving scale of the measuring probe. Aspects of the pion DF and proton spin developments are sketched herein, along with some novel perspectives on gluon and quark orbital angular momentum.

hep-ph↗

WHALE: Towards Generalizable and Scalable World Models for Embodied Decision-making

World models play a crucial role in decision-making within embodied environments, enabling cost-free explorations that would otherwise be expensive in the real world. To facilitate effective decision-making, world models must be equipped with strong generalizability to support faithful imagination in out-of-distribution (OOD) regions and provide reliable uncertainty estimation to assess the credibility of the simulated experiences, both of which present significant challenges for prior scalable approaches. This paper introduces WHALE, a framework for learning generalizable world models, consisting of two key techniques: behavior-conditioning and retracing-rollout. Behavior-conditioning addresses the policy distribution shift, one of the primary sources of the world model generalization error, while retracing-rollout enables efficient uncertainty estimation without the necessity of model ensembles. These techniques are universal and can be combined with any neural network architecture for world model learning. Incorporating these two techniques, we present Whale-ST, a scalable spatial-temporal transformer-based world model with enhanced generalizability. We demonstrate the superiority of Whale-ST in simulation tasks by evaluating both value estimation accuracy and video generation fidelity. Additionally, we examine the effectiveness of our uncertainty estimation technique, which enhances model-based policy optimization in fully offline scenarios. Furthermore, we propose Whale-X, a 414M parameter world model trained on 970K trajectories from Open X-Embodiment datasets. We show that Whale-X exhibits promising scalability and strong generalizability in real-world manipulation scenarios using minimal demonstrations.

cs.LG↗

Provably and Practically Efficient Adversarial Imitation Learning with General Function Approximation

As a prominent category of imitation learning methods, adversarial imitation learning (AIL) has garnered significant practical success powered by neural network approximation. However, existing theoretical studies on AIL are primarily limited to simplified scenarios such as tabular and linear function approximation and involve complex algorithmic designs that hinder practical implementation, highlighting a gap between theory and practice. In this paper, we explore the theoretical underpinnings of online AIL with general function approximation. We introduce a new method called optimization-based AIL (OPT-AIL), which centers on performing online optimization for reward functions and optimism-regularized Bellman error minimization for Q-value functions. Theoretically, we prove that OPT-AIL achieves polynomial expert sample complexity and interaction complexity for learning near-expert policies. To our best knowledge, OPT-AIL is the first provably efficient AIL method with general function approximation. Practically, OPT-AIL only requires the approximate optimization of two objectives, thereby facilitating practical implementation. Empirical studies demonstrate that OPT-AIL outperforms previous state-of-the-art deep AIL methods in several challenging tasks.

cs.LG↗

Interacting hypersurfaces and multiple scalar-tensor theories

We propose a novel method to construct ghost-free multiple scalar-tensor theories. The key idea is to use the geometric quantities of hypersurfaces defined by the scalar fields, rather than the covariant derivatives of scalar fields or spacetime curvature, to build the theory. This approach has proven effective in developing ghost-free scalar-tensor theories in the single-field case. When multiple scalar fields are present, each field specifies a foliation of spacelike hypersurfaces, on which we can define the normal vector, induced metric, extrinsic and intrinsic curvatures, as well as extrinsic (Lie) and intrinsic (spatial) derivatives, respectively. By employing these hypersurface geometric quantities as foundational elements, we construct the Lagrangian for interacting hypersurfaces that describes a multiple scalar-tensor theory. Given that temporal (Lie) and spatial derivatives are separated, it becomes relatively easier to control the order of time derivatives, thus helping to avoid ghost-like or unwanted degrees of freedom. In this work, we use bi-scalar-field theory as an example, focusing on polynomial-type Lagrangians. We construct monomials of hypersurface geometric quantities up to $d=3$, where $d$ denotes the number of derivatives in each monomial. Additionally, we present the correspondence between expressions in terms of hypersurface quantities and those in covariant bi-scalar-tensor theory. Through a cosmological perturbation analysis of a simple model, we demonstrate that the theory propagates two tensor and two scalar degrees of freedom at the linear order in perturbations, thereby remaining free from any extra degrees of freedom.

gr-qc↗

Magnetism measurements of two-dimensional van der Waals antiferromagnet CrPS4 using dynamic cantilever magnetometry

Recent experimental and theoretical work has focused on two-dimensional van der Waals (2D vdW) magnets due to their potential applications in sensing and spintronics devises. In measurements of these emerging materials, conventional magnetometry often encounters challenges in characterizing the magnetic properties of small-sized vdW materials, especially for antiferromagnets with nearly compensated magnetic moments. Here, we investigate the magnetism of 2D antiferromagnet CrPS4 with a thickness of 8nm by using dynamic cantilever magnetometry (DCM). Through a combination of DCM experiment and the calculation based on a Stoner--Wohlfarth-type model, we unravel the magnetization states in 2D CrPS4 antiferromagnet. In the case of H parallel with c, a two-stage phase transition is observed. For H perpendicular to c, a hump in the effective magnetic restoring force is noted, which implies the presence of spin reorientation as temperature increases. These results demonstrate the benefits of DCM for studying magnetism of 2D magnets.

cond-mat.mes-hall↗

Dual-AEB: Synergizing Rule-Based and Multimodal Large Language Models for Effective Emergency Braking

Automatic Emergency Braking (AEB) systems are a crucial component in ensuring the safety of passengers in autonomous vehicles. Conventional AEB systems primarily rely on closed-set perception modules to recognize traffic conditions and assess collision risks. To enhance the adaptability of AEB systems in open scenarios, we propose Dual-AEB, a system combines an advanced multimodal large language model (MLLM) for comprehensive scene understanding and a conventional rule-based rapid AEB to ensure quick response times. To the best of our knowledge, Dual-AEB is the first method to incorporate MLLMs within AEB systems. Through extensive experimentation, we have validated the effectiveness of our method. The source code will be available at https://github.com/ChipsICU/Dual-AEB.

cs.RO↗

Hybrid Centralized-Distributed Resource Allocation Based on Deep Reinforcement Learning for Cooperative D2D Communications

Device-to-device (D2D) technology enables direct communication between adjacent devices within cellular networks. Due to its high data rate, low latency, and performance improvement in spectrum and energy efficiency, it has been widely investigated and applied as a critical technology in 5G New Radio (NR). In addition to conventional overlay and underlay D2D communications, cooperative D2D communication, which can achieve a win-win situation between cellular users (CUs) and D2D users (DUs) through cooperative relaying technique, has attracted extensive attention from academic and industrial circles in the past decade. This paper delves into optimizing joint spectrum allocation, power control, and link-matching between multiple CUs and DUs for cooperative D2D communications, using weighted sum energy efficiency (WSEE) as the performance metric to address the challenges of green communication and sustainable development. This integer programming problem can be decomposed into a classic weighted bipartite graph matching and a series of nonconvex spectrum allocation and power control problems between potentially matched cellular and D2D link pairs. To address this issue, we propose a hybrid centralized-distributed scheme based on deep reinforcement learning (DRL) and the Kuhn-Munkres (KM) algorithm. Leveraging the latter, the CUs and DUs autonomously optimize spectrum allocation and power control by only utilizing local information. Then, the base station (BS) determines the link matching. Simulation results reveal that it achieves near-optimal performance and significantly enhances the network convergence speed with low signaling overheads. In addition, we also propose and utilize cooperative link sets for corresponding D2D links to accelerate the proposed scheme and reduce signaling exchange further.

cs.NI↗

Lowering the strong coupling mode of modified teleparallel gravity theories

We investigate the strong coupling problem in modified teleparallel gravity theories using the effective field theory (EFT) approach, demonstrating that it is possible to shift the emergence of new degrees of freedom (DoFs) to lower orders in perturbation theory. We first focus on the case of $f(T)$ gravity, and we show that in its conformally equivalent form the scalar perturbations are non-dynamical up to the cubic action. We then propose a simple modification of the theory, which lowers the appearance of new DoFs to cubic order, compared to the quartic order in standard $f(T)$ gravity. Our work opens a new avenue to address the issue of strong coupling in modified teleparallel gravity, and suggests a new classification scheme of these theories based on the perturbative order at which new DoFs appear.

gr-qc↗

Weighted Sum Power Minimization for Cooperative Spectrum Sharing in Cognitive Radio Networks

This letter introduces weighted sum power (WSP), a new performance metric for wireless resource allocation during cooperative spectrum sharing in cognitive radio networks, where the primary and secondary nodes have different priorities and quality of service (QoS) requirements. Compared to using energy efficiency (EE) and weighted sum energy efficiency (WSEE) as performance metrics and optimization objectives of wireless resource allocation towards green communication, the linear character of WSP can reduce the complexity of optimization problems. Meanwhile, the weights assigned to different nodes are beneficial for managing their power budget. Using WSP as the optimization objective, a suboptimal resource allocation scheme is proposed, leveraging linear programming and Newton's method. Simulations verify that the proposed scheme provides near-optimal performance with low computation time. Furthermore, the initial approximate value selection in Newton's method is also optimized to accelerate the proposed scheme.

cs.NI↗

Single-photon detectors on arbitrary photonic substrates

Detecting non-classical light is a central requirement for photonics-based quantum technologies. Unrivaled high efficiencies and low dark counts have positioned superconducting nanowire single photon detectors (SNSPDs) as the leading detector technology for fiber and integrated photonic applications. However, a central challenge lies in their integration within photonic integrated circuits regardless of material platform or surface topography. Here, we introduce a method based on transfer printing that overcomes these constraints and allows for the integration of SNSPDs onto arbitrary photonic substrates. We prove this by integrating SNSPDs and showing through-waveguide single-photon detection in commercially manufactured silicon and lithium niobate on insulator integrated photonic circuits. Our method eliminates bottlenecks to the integration of high-quality single-photon detectors, turning them into a versatile and accessible building block for scalable quantum information processing.

quant-ph↗

Suppression of static ZZ interaction in an all-transmon quantum processor

The superconducting transmon qubit is currently a leading qubit modality for quantum computing, but gate performance in quantum processor with transmons is often insufficient to support running complex algorithms for practical applications. It is thus highly desirable to further improve gate performance. Due to the weak anharmonicity of transmon, a static ZZ interaction between coupled transmons commonly exists, undermining the gate performance, and in long term, it can become performance limiting. Here we theoretically explore a previously unexplored parameter region in an all-transmon system to address this issue. We show that an feasible parameter region, where the ZZ interaction is heavily suppressed while leaving XY interaction with an adequate strength to implement two-qubit gates, can be found for all-transmon systems. Thus, two-qubit gates, such as cross-resonance gate or iSWAP gate, can be realized without the detrimental effect from static ZZ interaction. To illustrate this, we demonstrate that an iSWAP gate with fast gate speed and dramatically lower conditional phase error can be achieved. Scaling up to large-scale transmon quantum processor, especially the cases with fixed coupling, addressing error, idling error, and crosstalk that arises from static ZZ interaction could also be strongly suppressed.

quant-ph↗

High-contrast ZZ interaction using superconducting qubits with opposite-sign anharmonicity

For building a scalable quantum processor with superconducting qubits, ZZ interaction is of great concern because its residual has a crucial impact to two-qubit gate fidelity. Two-qubit gates with fidelity meeting the criterion of fault-tolerant quantum computationhave been demonstrated using ZZ interaction. However, as the performance of quantum processors improves, the residual static-ZZ can become a performance-limiting factor for quantum gate operation and quantum error correction. Here, we introduce a superconducting architecture using qubits with opposite-sign anharmonicity, a transmon qubit and a C-shunt flux qubit, to address this issue. We theoretically demonstrate that by coupling the two types of qubits, the high-contrast ZZ interaction can be realized. Thus, we can control the interaction with a high on/off ratio to implement two-qubit CZ gates, or suppress it during two-qubit gate operation using XY interaction (e.g., an iSWAP gate). The proposed architecture can also be scaled up to multi-qubit cases. In a fixed coupled system, ZZ crosstalk related to neighboring spectator qubits could also be heavily suppressed.

quant-ph↗

Long-Range $ZZ$ Interaction via Resonator-Induced Phase in Superconducting Qubits

Superconducting quantum computing emerges as one of leading candidates for achieving quantum advantage. However, a prevailing challenge is the coding overhead due to limited quantum connectivity, constrained by nearest-neighbor coupling among superconducting qubits. Here, we propose a novel multimode coupling scheme using three resonators driven by two microwaves, based on the resonator-induced phase gate, to extend the $ZZ$ interaction distance between qubits. We demonstrate a CZ gate fidelity exceeding 99.9\% within 160 ns at free spectral range (FSR) of 1.4 GHz, and by optimizing driving pulses, we further reduce the residual photon to nearly $10^{-3}$ within 100 ns at FSR of 0.2 GHz. These facilitate the long-range CZ gate over separations reaching sub-meters, thus significantly enhancing qubit connectivity and making a practical step towards the scalable integration and modularization of quantum processors. Specifically, our approach supports the implementation of quantum error correction codes requiring high connectivity, such as low-density parity check codes that paves the way to achieving fault-tolerant quantum computing.

quant-ph↗

Energy-Guided Diffusion Sampling for Offline-to-Online Reinforcement Learning

Combining offline and online reinforcement learning (RL) techniques is indeed crucial for achieving efficient and safe learning where data acquisition is expensive. Existing methods replay offline data directly in the online phase, resulting in a significant challenge of data distribution shift and subsequently causing inefficiency in online fine-tuning. To address this issue, we introduce an innovative approach, \textbf{E}nergy-guided \textbf{DI}ffusion \textbf{S}ampling (EDIS), which utilizes a diffusion model to extract prior knowledge from the offline dataset and employs energy functions to distill this knowledge for enhanced data generation in the online phase. The theoretical analysis demonstrates that EDIS exhibits reduced suboptimality compared to solely utilizing online data or directly reusing offline data. EDIS is a plug-in approach and can be combined with existing methods in offline-to-online RL setting. By implementing EDIS to off-the-shelf methods Cal-QL and IQL, we observe a notable 20% average improvement in empirical performance on MuJoCo, AntMaze, and Adroit environments. Code is available at \url{https://github.com/liuxhym/EDIS}.

cs.LG↗