Search arXiv⌕ Search

arXiv subjects

Yang Long

Publications and source records attributed to Yang Long.

At least 37 records · Page 2Linked to original sources

Unsupervised Classification of Non-Hermitian Topological Phases under Symmetries

The integration of artificial intelligence (AI) into fundamental science has opened new possibilities to address long-standing scientific challenges rooted in mathematical limitations. For example, topological invariants are used to characterize topology, but there is no universally applicable one. This limitation explains why, in the past decades-long classification of topological phases of matter -- mainly focused on Hermitian systems -- many phases initially classified ``trivial" were later identified as topological. Recently, the discovery of non-Hermitian band topology has spurred substantial efforts in non-Hermitian topological classification, including the development of new topological invariants. However, such classifications similarly risk overlooking key topological features. Here, without relying on any topological invariant, we develop an AI-based unsupervised classification of symmetry-protected non-Hermitian topological phases. This algorithm distinguishes topological differences among non-Hermitian Hamiltonians with symmetries, and constructs, in an unsupervised manner, a topological periodic table for non-Hermitian systems. Additionally, it can account for the boundary effects, enabling the exploration of open-boundary effects on the topological phase diagram. These results introduce an unsupervised approach for classifying symmetry-protected non-Hermitian topological phases without omission and provide valuable guidance for the development of theories and experiments.

cond-mat.mes-hall↗

Learning Global Representation from Queries for Vectorized HD Map Construction

The online construction of vectorized high-definition (HD) maps is a cornerstone of modern autonomous driving systems. State-of-the-art approaches, particularly those based on the DETR framework, formulate this as an instance detection problem. However, their reliance on independent, learnable object queries results in a predominantly local query perspective, neglecting the inherent global representation within HD maps. In this work, we propose \textbf{MapGR} (\textbf{G}lobal \textbf{R}epresentation learning for HD \textbf{Map} construction), an architecture designed to learn and utilize a global representations from queries. Our method introduces two synergistic modules: a Global Representation Learning (GRL) module, which encourages the distribution of all queries to better align with the global map through a carefully designed holistic segmentation task, and a Global Representation Guidance (GRG) module, which endows each individual query with explicit, global-level contextual information to facilitate its optimization. Evaluations on the nuScenes and Argoverse2 datasets validate the efficacy of our approach, demonstrating substantial improvements in mean Average Precision (mAP) compared to leading baselines.

cs.CV↗

Decoding Visual Neural Representations by Multimodal with Dynamic Balancing

In this work, we propose an innovative framework that integrates EEG, image, and text data, aiming to decode visual neural representations from low signal-to-noise ratio EEG signals. Specifically, we introduce text modality to enhance the semantic correspondence between EEG signals and visual content. With the explicit semantic labels provided by text, image and EEG features of the same category can be more closely aligned with the corresponding text representations in a shared multimodal space. To fully utilize pre-trained visual and textual representations, we propose an adapter module that alleviates the instability of high-dimensional representation while facilitating the alignment and fusion of cross-modal features. Additionally, to alleviate the imbalance in multimodal feature contributions introduced by the textual representations, we propose a Modal Consistency Dynamic Balance (MCDB) strategy that dynamically adjusts the contribution weights of each modality. We further propose a stochastic perturbation regularization (SPR) term to enhance the generalization ability of semantic perturbation-based models by introducing dynamic Gaussian noise in the modality optimization process. The evaluation results on the ThingsEEG dataset show that our method surpasses previous state-of-the-art methods in both Top-1 and Top-5 accuracy metrics, improving by 2.0\% and 4.7\% respectively.

cs.CV↗

Observation of Embedded Topology in a Trivial Bulk via Projective Crystal Symmetry

Bulk-boundary correspondence is the foundational principle of topological physics, first established in the quantum Hall effect, where a $D$-dimensional topologically nontrivial bulk gives rise to $(D-1)$-dimensional boundary states. The advent of higher-order topology has generalized this principle to a hierarchical chain, enabling topological states to appear at $(D-2)$ or even lower-dimensional boundaries. To date, all known realizations of topological systems must require a topologically nontrivial bulk to initiate the chain of action for bulk-boundary correspondence. Here, in an acoustic crystal platform, we experimentally demonstrate an exception to this paradigm--embedded topology in a trivial bulk--where the bulk-boundary correspondence originates from a trivial bulk. Rather than relying on global symmetries, we employ projective crystal symmetry, which induces nontrivial topology not at the outset in the $D$-dimensional bulk, but midway through the correspondence hierarchy in lower-dimensional boundaries. We further realize a three-dimensional system exhibiting embedded topology that supports zero-dimensional topological states, achieving the longest possible chain of action for such an unconventional bulk-boundary correspondence in physical space. Our work experimentally establishes a new form of bulk-boundary correspondence initiated from a trivial bulk, opening additional degrees of freedom for the design of robust topological devices.

cond-mat.mes-hall↗

Towards Scalable Spatial Intelligence via 2D-to-3D Data Lifting

Spatial intelligence is emerging as a transformative frontier in AI, yet it remains constrained by the scarcity of large-scale 3D datasets. Unlike the abundant 2D imagery, acquiring 3D data typically requires specialized sensors and laborious annotation. In this work, we present a scalable pipeline that converts single-view images into comprehensive, scale- and appearance-realistic 3D representations - including point clouds, camera poses, depth maps, and pseudo-RGBD - via integrated depth estimation, camera calibration, and scale calibration. Our method bridges the gap between the vast repository of imagery and the increasing demand for spatial scene understanding. By automatically generating authentic, scale-aware 3D data from images, we significantly reduce data collection costs and open new avenues for advancing spatial intelligence. We release two generated spatial datasets, i.e., COCO-3D and Objects365-v2-3D, and demonstrate through extensive experiments that our generated data can benefit various 3D tasks, ranging from fundamental perception to MLLM-based reasoning. These results validate our pipeline as an effective solution for developing AI systems capable of perceiving, understanding, and interacting with physical environments.

cs.CV↗

Observation of wave amplification and temporal topological state in a genuine photonic time crystal

Photonic time crystals (PTCs) are materials whose dielectric permittivity is periodically modulated in time, giving rise to bandgaps not in energy-as in conventional photonic crystals-but in momentum, known as k-gaps. These k-gaps enable wave amplification by extracting energy from temporal modulation, offering a mechanism for coherent light generation that bypasses traditional optical gain. PTCs also extend the concept of topological insulators to the time domain, inducing a temporal topological state at the mid-gap of the k-gap, characterized by the Zak phase-a topological invariant originally defined for spatial lattices. Here, we experimentally demonstrate the properties of a k gap in a genuine PTC, realized in a dynamically modulated transmission-line metamaterial. Wave amplification within the k-gap is observed, with an initial power spectrum narrowing and shifting toward the gap. To probe the mid-gaptopological state, we introduce a temporal interface separating two PTCs with distinct topological phases. The measured phase shift between time-reflected and time-refracted waves, together with the temporal confinement of the topological state, provides direct evidence of nontrivial temporal topology. By integrating kgap amplification with time-domain topological features, our work opens new avenues for light generation and manipulation in time-varying photonic materials.

physics.optics↗

Realization of Weyl elastic metamaterials with spin skyrmions

Topological elastic metamaterials provide a topologically robust way to manipulate the phononic energy and information beyond the conventional approaches. Among various topological elastic metamaterials, Weyl elastic metamaterials stand out, as they are unique to three dimensions and exhibit numerous intriguing phenomena and potential applications. To date, however, the realization of Weyl elastic metamaterials remains elusive, primarily due to the full-vectoral nature of elastic waves and the complicated couplings between polarizations, leading to complicated and tangled three-dimensional (3D) bandstructures that unfavorable for experimental demonstration. Here, we overcome the challenge and realize an ideal, 3D printed, all-metallic Weyl elastic metamaterial with low dissipation losses. Notably, the elastic spin of the excitations around the Weyl points exhibits skyrmion textures, a topologically stable structure in real space. Utilizing 3D laser vibrometry, we reveal the projection of the Weyl points, the Fermi arcs and the unique spin characteristics of the topological surface states. Our work extends the Weyl metamaterials to elastic waves and paves a topological way to robust manipulation of elastic waves in 3D space.

physics.app-ph↗

Rethinking Brain Tumor Segmentation from the Frequency Domain Perspective

Precise segmentation of brain tumors, particularly contrast-enhancing regions visible in post-contrast MRI (areas highlighted by contrast agent injection), is crucial for accurate clinical diagnosis and treatment planning but remains challenging. However, current methods exhibit notable performance degradation in segmenting these enhancing brain tumor areas, largely due to insufficient consideration of MRI-specific tumor features such as complex textures and directional variations. To address this, we propose the Harmonized Frequency Fusion Network (HFF-Net), which rethinks brain tumor segmentation from a frequency-domain perspective. To comprehensively characterize tumor regions, we develop a Frequency Domain Decomposition (FDD) module that separates MRI images into low-frequency components, capturing smooth tumor contours and high-frequency components, highlighting detailed textures and directional edges. To further enhance sensitivity to tumor boundaries, we introduce an Adaptive Laplacian Convolution (ALC) module that adaptively emphasizes critical high-frequency details using dynamically updated convolution kernels. To effectively fuse tumor features across multiple scales, we design a Frequency Domain Cross-Attention (FDCA) integrating semantic, positional, and slice-specific information. We further validate and interpret frequency-domain improvements through visualization, theoretical reasoning, and experimental analyses. Extensive experiments on four public datasets demonstrate that HFF-Net achieves an average relative improvement of 4.48\% (ranging from 2.39\% to 7.72\%) in the mean Dice scores across the three major subregions, and an average relative improvement of 7.33% (ranging from 5.96% to 8.64%) in the segmentation of contrast-enhancing tumor regions, while maintaining favorable computational efficiency and clinical applicability. Code: https://github.com/VinyehShaw/HFF.

eess.IV↗

Three-dimensional flat Landau levels in an inhomogeneous acoustic crystal

When electrons moving in two-dimensions (2D) are subjected to a strong uniform magnetic field, they form flat bands called Landau levels, which are the basis for the quantum Hall effect. Landau levels can also arise from pseudomagnetic fields (PMFs) induced by lattice distortions; for example, mechanically straining graphene causes its Dirac quasiparticles to form a characteristic set of unequally-spaced Landau levels, including a zeroth Landau level. In three-dimensional (3D) systems, there has thus far been no experimental demonstration of Landau levels or any other type of flat band. For instance, applying a uniform magnetic field to materials hosting Weyl quasiparticles, the 3D generalizations of Dirac quasiparticles, yields bands that are non-flat in the direction of the field. Here, we report on the experimental realization of a flat 3D Landau level in an acoustic crystal. Starting from a lattice whose bandstructure exhibits a nodal ring, we design an inhomogeneous distortion corresponding to a specific pseudomagnetic vector potential (PVP) that causes the nodal ring states to break up into Landau levels, with a zeroth Landau level that is flat along all three directions. These findings point to the possibility of using nodal ring materials to generate 3D flat bands, to access strong interactions and other interesting physical regimes in 3D.

cond-mat.mes-hall↗

Observation of returning Thouless pumping

Introduced by David Thouless in 1983, Thouless pumping exemplifies topological properties in topological systems, where the transported charge is quantized by the Chern number. Recently, returning Thouless pumping was theoretically proposed, in which quantized charge is pumped during the first half of the cycle but returns to zero in the second half. This mechanism leads to crystalline symmetry-protected delicate topological insulators. Unlike conventional topological bands, a delicate topological band is Wannierizable but not atomically obstructed, which features multicellular Wannier functions extending beyond a single unit cell. Here, by replacing the second dimension with a synthetic dimension, we realize a two-dimensional delicate topological insulator via a set of one-dimensional acoustic crystals with fine-tuned geometric parameters. Through acoustic bands and wavefunction measurements, we directly observe returning Thouless pumping and symmetric multicellular Wannier functions, followed by establishing the bulk-boundary correspondence between sub-Brillouin zone Chern numbers and gapless boundary modes. As enriched by crystalline symmetries, our experimental demonstration of returning Thouless pumping expands the current understanding of topological phases of matter.

cond-mat.mes-hall↗

Rethinking Score Distilling Sampling for 3D Editing and Generation

Score Distillation Sampling (SDS) has emerged as a prominent method for text-to-3D generation by leveraging the strengths of 2D diffusion models. However, SDS is limited to generation tasks and lacks the capability to edit existing 3D assets. Conversely, variants of SDS that introduce editing capabilities often can not generate new 3D assets effectively. In this work, we observe that the processes of generation and editing within SDS and its variants have unified underlying gradient terms. Building on this insight, we propose Unified Distillation Sampling (UDS), a method that seamlessly integrates both the generation and editing of 3D assets. Essentially, UDS refines the gradient terms used in vanilla SDS methods, unifying them to support both tasks. Extensive experiments demonstrate that UDS not only outperforms baseline methods in generating 3D assets with richer details but also excels in editing tasks, thereby bridging the gap between 3D generation and editing. The code is available on: https://github.com/xingy038/UDS.

cs.CV↗

Attention in Diffusion Model: A Survey

Attention mechanisms have become a foundational component in diffusion models, significantly influencing their capacity across a wide range of generative and discriminative tasks. This paper presents a comprehensive survey of attention within diffusion models, systematically analysing its roles, design patterns, and operations across different modalities and tasks. We propose a unified taxonomy that categorises attention-related modifications into parts according to the structural components they affect, offering a clear lens through which to understand their functional diversity. In addition to reviewing architectural innovations, we examine how attention mechanisms contribute to performance improvements in diverse applications. We also identify current limitations and underexplored areas, and outline potential directions for future research. Our study provides valuable insights into the evolving landscape of diffusion models, with a particular focus on the integrative and ubiquitous role of attention.

cs.LG↗

FMDConv: Fast Multi-Attention Dynamic Convolution via Speed-Accuracy Trade-off

Spatial convolution is fundamental in constructing deep Convolutional Neural Networks (CNNs) for visual recognition. While dynamic convolution enhances model accuracy by adaptively combining static kernels, it incurs significant computational overhead, limiting its deployment in resource-constrained environments such as federated edge computing. To address this, we propose Fast Multi-Attention Dynamic Convolution (FMDConv), which integrates input attention, temperature-degraded kernel attention, and output attention to optimize the speed-accuracy trade-off. FMDConv achieves a better balance between accuracy and efficiency by selectively enhancing feature extraction with lower complexity. Furthermore, we introduce two novel quantitative metrics, the Inverse Efficiency Score and Rate-Correct Score, to systematically evaluate this trade-off. Extensive experiments on CIFAR-10, CIFAR-100, and ImageNet demonstrate that FMDConv reduces the computational cost by up to 49.8\% on ResNet-18 and 42.2\% on ResNet-50 compared to prior multi-attention dynamic convolution methods while maintaining competitive accuracy. These advantages make FMDConv highly suitable for real-world, resource-constrained applications.

cs.CV↗

Asynchronous Personalized Federated Learning through Global Memorization

The proliferation of Internet of Things devices and advances in communication technology have unleashed an explosion of personal data, amplifying privacy concerns amid stringent regulations like GDPR and CCPA. Federated Learning offers a privacy preserving solution by enabling collaborative model training across decentralized devices without centralizing sensitive data. However, statistical heterogeneity from non-independent and identically distributed datasets and system heterogeneity due to client dropouts particularly those with monopolistic classes severely degrade the global model's performance. To address these challenges, we propose the Asynchronous Personalized Federated Learning framework, which empowers clients to develop personalized models using a server side semantic generator. This generator, trained via data free knowledge transfer under global model supervision, enhances client data diversity by producing both seen and unseen samples, the latter enabled by Zero-Shot Learning to mitigate dropout-induced data loss. To counter the risks of synthetic data impairing training, we introduce a decoupled model interpolation method, ensuring robust personalization. Extensive experiments demonstrate that AP FL significantly outperforms state of the art FL methods in tackling non-IID distributions and client dropouts, achieving superior accuracy and resilience across diverse real-world scenarios.

cs.LG↗

Observation of non-Hermitian topological disclination states and charge fractionalization

There has been significant interest in exploring topological disclination states, which effectively probe the band topology of the host material beyond the conventional bulk-edge correspondence. While most studies in this area have primarily focused on Hermitian systems, recent theoretical work predicts that non-Hermiticity can drive topological phase transitions and host topological disclination states associated with fractional charge. However, no experimental observations have been reported to date. Here, we report the first experimental observation of topological disclination states in electric circuits, induced solely by gain and loss. Through admittance matrix measurements and eigenstate analysis, we confirm their emergence and compute the corresponding fractional charge. Moreover, the disclination mode profile and localization effect can be directly visualized via monochromatic field excitation. Additionally, we demonstrate the emergence of degenerate zero-energy topological disclination states, devoid of fractional charge, in distinct non-Hermitian geometries. Our findings open the possibility of non-Hermiticity-induced fractional charges in two-dimensional non-Hermitian lattices, which may pave the way for advancements in active topological photonic devices.

cond-mat.mes-hall↗

Laser: Efficient Language-Guided Segmentation in Neural Radiance Fields

In this work, we propose a method that leverages CLIP feature distillation, achieving efficient 3D segmentation through language guidance. Unlike previous methods that rely on multi-scale CLIP features and are limited by processing speed and storage requirements, our approach aims to streamline the workflow by directly and effectively distilling dense CLIP features, thereby achieving precise segmentation of 3D scenes using text. To achieve this, we introduce an adapter module and mitigate the noise issue in the dense CLIP feature distillation process through a self-cross-training strategy. Moreover, to enhance the accuracy of segmentation edges, this work presents a low-rank transient query attention mechanism. To ensure the consistency of segmentation for similar colors under different viewpoints, we convert the segmentation task into a classification task through label volume, which significantly improves the consistency of segmentation in color-similar areas. We also propose a simplified text augmentation strategy to alleviate the issue of ambiguity in the correspondence between CLIP features and text. Extensive experimental results show that our method surpasses current state-of-the-art technologies in both training speed and performance. Our code is available on: https://github.com/xingy038/Laser.git.

cs.CV↗

Topologically protected edge states in time photonic crystals with chiral symmetry

Time photonic crystals are media in which their electromagnetic parameters are modulated periodically in time, showing promising applications in non-resonant lasers and particle accelerators, among others. Traditionally utilized to study space photonic crystals, topological band theory has also been translated recently to analyze time photonic crystals with time inversion symmetry, enabling the construction of the temporal version of topological edge states. However, temporal disorder can readily break time inversion symmetry in practice, hence likely destroying the edge states associated with this type of time photonic crystals. To overcome this limitation, here we propose a new class of time photonic crystals presenting chiral symmetry instead, whose edge states exhibit superior robustness over the time-reversal-symmetry-protected counterparts. Our time photonic crystal is equivalent to a temporal version of the Su-Schrieffer-Heeger model, and the chiral symmetry of this type of time photonic crystals quantizes the winding number defined in the Bloch frequency band. Remarkably, random temporal disorders do not impact the eigenfrequencies of these chiral-symmetry-protected edge states, while instead enhancing their temporal localizations. Our findings thus provide a promising paradigm to control field amplification with exceptional robustness as well as being a feasible platform to investigate various topological phases in time-varying media.

physics.optics↗

Momentum flatband and superluminal propagation in a photonic time Moiré superlattice

Flat bands typically describe energy bands whose energy dispersion is entirely or almost entirely degenerate. One effective method to form flat bands is by constructing Moiré superlattices. Recently, there has been a shift in perspective regarding the roles of space (momentum) and time (energy) in a lattice, with the concept of photonic time crystals that has sparked discussions on momentum dispersion such as the presence of a bandgap in momentum. Here we propose a photonic time moiré superlattice achieved by overlaying two photonic time crystals with different periods. The resulting momentum bandgap of this superlattice supports isolated momentum bands that are nearly independent of energy, which we refer to as momentum flat bands. Unlike energy flat bands, which have zero group velocity, momentum flat bands exhibit infinitely large group velocity across a broad frequency range. Unlike previous optical media supporting broadband superluminal propagation based on gain, the effective refractive index of the momentum flat bands is real-valued, leading to more stabilized superluminal pulse propagation.

physics.optics↗