Search arXiv⌕ Search

arXiv subjects

Hang Yu

Publications and source records attributed to Hang Yu.

At least 127 records · Page 7Linked to original sources

Focus On What Matters: Separated Models For Visual-Based RL Generalization

A primary challenge for visual-based Reinforcement Learning (RL) is to generalize effectively across unseen environments. Although previous studies have explored different auxiliary tasks to enhance generalization, few adopt image reconstruction due to concerns about exacerbating overfitting to task-irrelevant features during training. Perceiving the pre-eminence of image reconstruction in representation learning, we propose SMG (Separated Models for Generalization), a novel approach that exploits image reconstruction for generalization. SMG introduces two model branches to extract task-relevant and task-irrelevant representations separately from visual observations via cooperatively reconstruction. Built upon this architecture, we further emphasize the importance of task-relevant features for generalization. Specifically, SMG incorporates two additional consistency losses to guide the agent's focus toward task-relevant areas across different scenarios, thereby achieving free from overfitting. Extensive experiments in DMC demonstrate the SOTA performance of SMG in generalization, particularly excelling in video-background settings. Evaluations on robotic manipulation tasks further confirm the robustness of SMG in real-world applications.

cs.CV↗

SDP: Spiking Diffusion Policy for Robotic Manipulation with Learnable Channel-Wise Membrane Thresholds

This paper introduces a Spiking Diffusion Policy (SDP) learning method for robotic manipulation by integrating Spiking Neurons and Learnable Channel-wise Membrane Thresholds (LCMT) into the diffusion policy model, thereby enhancing computational efficiency and achieving high performance in evaluated tasks. Specifically, the proposed SDP model employs the U-Net architecture as the backbone for diffusion learning within the Spiking Neural Network (SNN). It strategically places residual connections between the spike convolution operations and the Leaky Integrate-and-Fire (LIF) nodes, thereby preventing disruptions to the spiking states. Additionally, we introduce a temporal encoding block and a temporal decoding block to transform static and dynamic data with timestep $T_S$ into each other, enabling the transmission of data within the SNN in spike format. Furthermore, we propose LCMT to enable the adaptive acquisition of membrane potential thresholds, thereby matching the conditions of varying membrane potentials and firing rates across channels and avoiding the cumbersome process of manually setting and tuning hyperparameters. Evaluating the SDP model on seven distinct tasks with SNN timestep $T_S=4$, we achieve results comparable to those of the ANN counterparts, along with faster convergence speeds than the baseline SNN method. This improvement is accompanied by a reduction of 94.3\% in dynamic energy consumption estimated on 45nm hardware.

cs.RO↗

Uncovering Stealth Bias in LISA observations of Double White Dwarf Binaries due to Tidal Coupling

Double white dwarfs are important gravitational wave sources for LISA, as they are some of the most numerous compact systems in our universe. Here we consider finite-sized effects due to tidal interactions, as they are expected to have a measurable impact on these systems. Previous studies suggested that tidal effects would allow the individual masses to be measured, but there was a subtle error in those analyses. Using a fully Bayesian analysis we find that while tidal effects do not allow us to constrain the individual masses, they do yield informative lower bounds on the total mass of the system. Including tidal effects is crucial to the accuracy of our estimation of the chirp and total mass. Neglecting tidal effects leads to significant biases towards higher chirp masses, and we see that the lower bound of the total masses is biased towards a higher value as well. For many systems observed by LISA, tidal effects can lead to a "stealth" bias, since only the first derivative of the frequency can be measured. To separate tidal effects from the usual point particle decay we need to be able to measure the change in the second derivative of the frequency cause by the tides. This can only be done for high frequency systems observed with high signal-to-noise. The bias, if not accounted for, can have significant astrophysical implications; for example, it could lead to an incorrect estimation of the population of potential Type IA supernovae progenitors.

astro-ph.HE↗

DUPLEX: Dual GAT for Complex Embedding of Directed Graphs

Current directed graph embedding methods build upon undirected techniques but often inadequately capture directed edge information, leading to challenges such as: (1) Suboptimal representations for nodes with low in/out-degrees, due to the insufficient neighbor interactions; (2) Limited inductive ability for representing new nodes post-training; (3) Narrow generalizability, as training is overly coupled with specific tasks. In response, we propose DUPLEX, an inductive framework for complex embeddings of directed graphs. It (1) leverages Hermitian adjacency matrix decomposition for comprehensive neighbor integration, (2) employs a dual GAT encoder for directional neighbor modeling, and (3) features two parameter-free decoders to decouple training from particular tasks. DUPLEX outperforms state-of-the-art models, especially for nodes with sparse connectivity, and demonstrates robust inductive capability and adaptability across various tasks. The code is available at https://github.com/alipay/DUPLEX.

cs.LG↗

SQLfuse: Enhancing Text-to-SQL Performance through Comprehensive LLM Synergy

Text-to-SQL conversion is a critical innovation, simplifying the transition from complex SQL to intuitive natural language queries, especially significant given SQL's prevalence in the job market across various roles. The rise of Large Language Models (LLMs) like GPT-3.5 and GPT-4 has greatly advanced this field, offering improved natural language understanding and the ability to generate nuanced SQL statements. However, the potential of open-source LLMs in Text-to-SQL applications remains underexplored, with many frameworks failing to leverage their full capabilities, particularly in handling complex database queries and incorporating feedback for iterative refinement. Addressing these limitations, this paper introduces SQLfuse, a robust system integrating open-source LLMs with a suite of tools to enhance Text-to-SQL translation's accuracy and usability. SQLfuse features four modules: schema mining, schema linking, SQL generation, and a SQL critic module, to not only generate but also continuously enhance SQL query quality. Demonstrated by its leading performance on the Spider Leaderboard and deployment by Ant Group, SQLfuse showcases the practical merits of open-source LLMs in diverse business contexts.

cs.CL↗

Dynamical tides during the inspiral of rapidly spinning neutron stars: Solutions beyond mode resonance

We investigate the dynamical tide in a gravitational wave (GW)-driven coalescing binary involving a neutron star (NS). The NS is assumed to spin rapidly, with its spin axis anti-aligned with the orbit. Such an NS may exist if the binary forms dynamically in a dense environment, and it can lead to a strong tide because the f-mode can be resonantly excited during the inspiral. We present a new analytical solution for the f-mode resonance by decomposing the tide into a resummed equilibrium component and a dynamical component that is excited only around resonance. This solution simplifies numerical implementations by avoiding the subtraction of two diverging terms. It also extends the solution's validity to frequencies beyond mode resonance. When the dynamical tide back reacts on the orbit, the commonly adopted effective Love number is insufficient because it does not capture the tidal torque on the orbit that dominates the back reaction during mode resonance. An additional dressing factor originating from the imaginary part of the Love number is introduced to model the torque. The dissipative interaction between the NS and the orbital mass multipoles is computed including the dynamical tide. Orbital phase shifts caused by the $l=3$ and $l=2$ f-modes can reach 0.5 and 10 radians at their respective resonances if the NS has a spin rate of 850 Hz. Because of the large impact of the dynamical tide, a linearized analytical description becomes insufficient. After mode excitation, the orbit cannot remain quasi-circular, and the eccentricity excited by the dynamical tide can approach $e\simeq 0.1$, leading to non-monotonic frequency evolution which breaks the stationary phase approximation commonly adopted by frequency-domain waveform constructions. The GW radiation from the excited f-mode alone can be detected with a signal-to-noise ratio exceeding unity with the next-generation detectors.

gr-qc↗

Approaching Outside: Scaling Unsupervised 3D Object Detection from 2D Scene

The unsupervised 3D object detection is to accurately detect objects in unstructured environments with no explicit supervisory signals. This task, given sparse LiDAR point clouds, often results in compromised performance for detecting distant or small objects due to the inherent sparsity and limited spatial resolution. In this paper, we are among the early attempts to integrate LiDAR data with 2D images for unsupervised 3D detection and introduce a new method, dubbed LiDAR-2D Self-paced Learning (LiSe). We argue that RGB images serve as a valuable complement to LiDAR data, offering precise 2D localization cues, particularly when scarce LiDAR points are available for certain objects. Considering the unique characteristics of both modalities, our framework devises a self-paced learning pipeline that incorporates adaptive sampling and weak model aggregation strategies. The adaptive sampling strategy dynamically tunes the distribution of pseudo labels during training, countering the tendency of models to overfit easily detected samples, such as nearby and large-sized objects. By doing so, it ensures a balanced learning trajectory across varying object scales and distances. The weak model aggregation component consolidates the strengths of models trained under different pseudo label distributions, culminating in a robust and powerful final model. Experimental evaluations validate the efficacy of our proposed LiSe method, manifesting significant improvements of +7.1% AP$_{BEV}$ and +3.4% AP$_{3D}$ on nuScenes, and +8.3% AP$_{BEV}$ and +7.4% AP$_{3D}$ on Lyft compared to existing techniques.

cs.CV↗

Are WASP-107-like Systems Consistent with High-eccentricity Migration?

WASP-107 b seems to be a poster child of the long-suspected high-eccentricity migration scenario. It is on a 5.7-day, polar orbit. The planet is Jupiter-like in radius but Neptune-like in mass with exceptionally low density. WASP-107 c is on a 1100-day, $e=0.28$ orbit with at least Saturn mass. Planet b may still have a residual eccentricity of $0.06\pm 0.04$: the ongoing tidal dissipation leads to the observed internally heated atmosphere and hydrodynamic atmospheric erosion. We present a population synthesis study coupling octopole Lidov-Kozai oscillations with various short-range forces, while simultaneously accounting for the radius inflation and tidal disruption of the planet. We find that a high-eccentricity migration scenario can successfully explain nearly all observed system properties. Our simulations further suggest that the initial location of WASP-107 b at the onset of migration is likely within the snowline ($<0.5\,{\rm AU}$). More distant initial orbits usually lead to tidal disruption or orbit crossing. WASP-107 b most likely lost no more than 20% of its mass during the high-eccentricity migration, i.e. it did not form as a Jupiter-mass object. More vigorous tidally-induced mass loss leads to disruption of the planet during migration. We predict that the current-day mutual inclination between the planets b and c is substantial: at least 25-55$^\circ$ which may be tested with future Gaia astrometric observations. Knowing the current-day mutual inclination may further constrain the initial orbit of planet b. We suggest that the proposed high-eccentricity migration scenario of WASP-107 may be applicable to HAT-P-11, GJ-3470, HAT-P-18, and GJ-436 which have similar orbital architectures.

astro-ph.EP↗

Unifying the Perspectives of NLP and Software Engineering: A Survey on Language Models for Code

In this work we systematically review the recent advancements in software engineering with language models, covering 70+ models, 40+ evaluation tasks, 180+ datasets, and 900 related works. Unlike previous works, we integrate software engineering (SE) with natural language processing (NLP) by discussing the perspectives of both sides: SE applies language models for development automation, while NLP adopts SE tasks for language model evaluation. We break down code processing models into general language models represented by the GPT family and specialized models that are specifically pretrained on code, often with tailored objectives. We discuss the relations and differences between these models, and highlight the historical transition of code modeling from statistical models and RNNs to pretrained Transformers and LLMs, which is exactly the same course that had been taken by NLP. We also go beyond programming and review LLMs' application in other software engineering activities including requirement engineering, testing, deployment, and operations in an endeavor to provide a global view of NLP in SE, and identify key challenges and potential future directions in this domain. We keep the survey open and updated on GitHub at https://github.com/codefuse-ai/Awesome-Code-LLM.

cs.CL↗

D2LLM: Decomposed and Distilled Large Language Models for Semantic Search

The key challenge in semantic search is to create models that are both accurate and efficient in pinpointing relevant sentences for queries. While BERT-style bi-encoders excel in efficiency with pre-computed embeddings, they often miss subtle nuances in search tasks. Conversely, GPT-style LLMs with cross-encoder designs capture these nuances but are computationally intensive, hindering real-time applications. In this paper, we present D2LLMs-Decomposed and Distilled LLMs for semantic search-that combines the best of both worlds. We decompose a cross-encoder into an efficient bi-encoder integrated with Pooling by Multihead Attention and an Interaction Emulation Module, achieving nuanced understanding and pre-computability. Knowledge from the LLM is distilled into this model using contrastive, rank, and feature imitation techniques. Our experiments show that D2LLM surpasses five leading baselines in terms of all metrics across three tasks, particularly improving NLI task performance by at least 6.45%. The source code is available at https://github.com/codefuse-ai/D2LLM.

cs.CL↗

Complex scaling in finite volume

Quantum resonances, i.e., metastable states with a finite lifetime, play an important role in nuclear physics and other domains. Describing this phenomenon theoretically is generally a challenging task. In this work, we combine two established techniques to address this challenge. Complex scaling makes it possible to calculate resonances with bound-state-like methods. Finite-volume simulations exploit the fact that the infinite-volume properties of quantum systems are encoded in how discrete energy levels change as one varies the size of the volume. We apply complex scaling to systems in finite periodic boxes and derive the volume dependence of states in this scenario, demonstrating with explicit examples how one can use these relations to infer infinite-volume resonance energies and lifetimes.

nucl-th↗

Channel Access Methods for RF-Powered IoT Networks: A Survey

Many Internet of Things (IoT) networks with Radio Frequency (RF) powered devices operate over a shared medium. They thus require a channel access protocol. Unlike conventional networks where devices have unlimited energy, in an RF-powered IoT network, devices must first harvest RF energy in order to transmit or/and receive data. To this end, this survey presents the {\em first} comprehensive review of prior works that employ contention-based and contention-free protocols in IoT networks with one or more {\em dedicated} energy sources. Specifically, these protocols work in conjunction with RF-energy sources to deliver energy delivery or/and data. In this respect, this survey covers protocols based on Aloha, Carrier Sense Multiple Access (CSMA), polling, and dynamic Time Division Multiple Access (TDMA). Further, it covers successive interference cancellation protocols. It highlights key issues and challenges addressed by prior works, and provides a qualitative comparison of these works. Lastly, it identifies gaps in the literature and presents a list of future research directions.

cs.NI↗

MAVRL: Learn to Fly in Cluttered Environments with Varying Speed

Many existing obstacle avoidance algorithms overlook the crucial balance between safety and agility, especially in environments of varying complexity. In our study, we introduce an obstacle avoidance pipeline based on reinforcement learning. This pipeline enables drones to adapt their flying speed according to the environmental complexity. Moreover, to improve the obstacle avoidance performance in cluttered environments, we propose a novel latent space. The latent space in this representation is explicitly trained to retain memory of previous depth map observations. Our findings confirm that varying speed leads to a superior balance of success rate and agility in cluttered environments. Additionally, our memory-augmented latent representation outperforms the latent representation commonly used in reinforcement learning. Finally, after minimal fine-tuning, we successfully deployed our network on a real drone for enhanced obstacle avoidance.

cs.RO↗

BasisFormer: Attention-based Time Series Forecasting with Learnable and Interpretable Basis

Bases have become an integral part of modern deep learning-based models for time series forecasting due to their ability to act as feature extractors or future references. To be effective, a basis must be tailored to the specific set of time series data and exhibit distinct correlation with each time series within the set. However, current state-of-the-art methods are limited in their ability to satisfy both of these requirements simultaneously. To address this challenge, we propose BasisFormer, an end-to-end time series forecasting architecture that leverages learnable and interpretable bases. This architecture comprises three components: First, we acquire bases through adaptive self-supervised learning, which treats the historical and future sections of the time series as two distinct views and employs contrastive learning. Next, we design a Coef module that calculates the similarity coefficients between the time series and bases in the historical view via bidirectional cross-attention. Finally, we present a Forecast module that selects and consolidates the bases in the future view based on the similarity coefficients, resulting in accurate future predictions. Through extensive experiments on six datasets, we demonstrate that BasisFormer outperforms previous state-of-the-art methods by 11.04\% and 15.78\% respectively for univariate and multivariate forecasting tasks. Code is available at: \url{https://github.com/nzl5116190/Basisformer}

cs.LG↗

CodeFuse-13B: A Pretrained Multi-lingual Code Large Language Model

Code Large Language Models (Code LLMs) have gained significant attention in the industry due to their wide applications in the full lifecycle of software engineering. However, the effectiveness of existing models in understanding non-English inputs for multi-lingual code-related tasks is still far from well studied. This paper introduces CodeFuse-13B, an open-sourced pre-trained code LLM. It is specifically designed for code-related tasks with both English and Chinese prompts and supports over 40 programming languages. CodeFuse achieves its effectiveness by utilizing a high quality pre-training dataset that is carefully filtered by program analyzers and optimized during the training process. Extensive experiments are conducted using real-world usage scenarios, the industry-standard benchmark HumanEval-x, and the specially designed CodeFuseEval for Chinese prompts. To assess the effectiveness of CodeFuse, we actively collected valuable human feedback from the AntGroup's software development process where CodeFuse has been successfully deployed. The results demonstrate that CodeFuse-13B achieves a HumanEval pass@1 score of 37.10%, positioning it as one of the top multi-lingual code LLMs with similar parameter sizes. In practical scenarios, such as code generation, code translation, code comments, and testcase generation, CodeFuse performs better than other models when confronted with Chinese prompts.

cs.SE↗

Detecting gravitational lensing in hierarchical triples in galactic nuclei with space-borne gravitational-wave observatories

Stellar-mass binary black holes (BBHs) may merge in the vicinity of a supermassive black hole (SMBH). It is suggested that the gravitational-wave (GW) emitted by a BBH has a high probability to be lensed by the SMBH if the BBH's orbit around the SMBH (i.e., the outer orbit) has a period of less than a year and is less than the duration of observation of the BBH by a space-borne GW observatory. For such a BBH + SMBH triple system, the de Sitter precession of the BBH's orbital plane is also significant. In this work, we thus study GW waveforms emitted by the BBH and then modulated by the SMBH due to effects including Doppler shift, de Sitter precession, and gravitational lensing. We show specifically that for an outer orbital period of 0.1 yr and an SMBH mass of $10^7 M_\odot$, there is a 3\%-10\% chance for the standard, strong lensing signatures to be detectable by space-borne GW detectors such as LISA and/or TianGO. For more massive lenses ($\gtrsim 10^8 M_\odot$) and more compact outer orbits with periods <0.1 yr, retro-lensing of the SMBH might also have a 1%-level chance of detection. Furthermore, by combining the lensing effects and the dynamics of the outer orbit, we find the mass of the central SMBH can be accurately determined with a fraction error of $\sim 10^{-4}$. This is much better than the case of static lensing because the degeneracy between the lens' mass and the source's angular position is lifted by the outer orbital motion. Including lensing effects also allows the de Sitter precession to be detectable at a precession period 3 times longer than the case without lensing. Lastly, we demonstrate that one can check the consistency between the SMBH's mass determined from the orbital dynamics and the one inferred from gravitational lensing, which serves as a test on theories behind both phenomena. The statistical error on the deviation can be constrained to a 1% level.

gr-qc↗

Accurate and Efficient Waveform Model for Precessing Binary Black Holes

We present IMRPhenomXODE, a new phenomenological frequency-domain waveform approximant for gravitational wave (GW) signals from precessing binary black holes (BBHs) with generic spin configurations. We build upon the success of IMRPhenomXPHM [G. Pratten et al., Phys. Rev. D 103, 104056 (2021), which is one of the most widely adopted waveform approximants in GW data analyses that include spin precession, and introduce two additional significant improvements. First, we employ an efficient technique to numerically solve the (next-to)$^4$-leading-order post-Newtonian precession equations, which allows us to accurately determine the evolution of the orientation of the orbital angular momentum $\boldsymbol{\hat{L}}_{\rm N}$ even in cases with complicated precession dynamics, such as transitional precession. Second, we recalibrate the phase of GW modes in the frame coprecessing with $\boldsymbol{\hat{L}}_{\rm N}$ against SEOBNRv4PHM [S. Ossokine et al., Phys. Rev. D 102, 044055 (2020)] to capture effects due to precession such as variations in the spin components aligned with $\boldsymbol{\hat{L}}_{\rm N}$. By incorporating these new features, IMRPhenomXODE achieves matches with SEOBNRv4PHM that are better than 99% for most BBHs with mass ratios $q \geq 1/6$ and with arbitrary spin configurations. In contrast, the mismatch between IMRPhenomXPHM and SEOBNRv4PHM often exceeds 10% for a BBH with $q\lesssim 1/2$ and large in-plane or antialigned spin components. Our implementation is also computationally efficient, with waveform evaluation times that can even be shorter than those of IMRPhenomXPHM for BBH signals with long durations and hence high frequency resolutions. The accuracy and efficiency of IMRPhenomXODE position it as a valuable tool for GW event searches, parameter estimation analyses, and the inference of underlying population properties.

gr-qc↗

Charged-particle bound states in periodic boxes

We consider the binding energy of a two-body system with a repulsive Coulomb interaction in a finite periodic volume. We define the finite-volume Coulomb potential as the usual Coulomb potential, except that the distance is defined as the shortest separation between the two bodies in the periodic volume. We investigate this problem in one and three-dimensional periodic boxes and derive the asymptotic behavior of the volume dependence for bound states with zero angular momentum in terms of Whittaker functions. We benchmark our results against numerical calculations and show how the method can be used to extract asymptotic normalization coefficients for charged-particle bound states. The results we derive here have immediate applications for calculations of atomic nuclei in finite periodic volumes for the case where the leading finite-volume correction is associated with two charged clusters.

nucl-th↗