Search arXivSearch

arXiv subjects

Hang Yang

Publications and source records attributed to Hang Yang.

At least 19 recordsLinked to original sources

VADER: Adaptive Debiasing for Hallucination Mitigation in Video Large Language Models

Large vision-language models (LVLMs) have demonstrated strong performance in open-ended video understanding, yet they remain prone to fluent responses unsupported by video evidence. Existing training-free methods typically apply a globally fixed visual intervention or construct a contrastive branch through input perturbation. The former cannot accommodate video-dependent fusion paths, while the latter can be compensated by cross-frame redundancy. We therefore propose Video-Adaptive Debiasing via Evidence Reweighting (VADER), a training-free framework with two complementary modules. Visual Focus Reallocation (VFR) automatically instantiates an intervention policy for each video-question input: it diagnoses layer-wise visual-to-text evidence flow, determines where to intervene, and derives how strongly to reallocate pre-softmax attention from system-token to video-token blocks. Selective Evidence Erasure (SEE) independently masks high-importance visual tokens in every frame, constructing a prior-biased branch that is difficult to compensate through neighboring frames. Contrastive decoding then down-weights predictions that remain confident after selective evidence erasure. Across multiple VideoLLMs, VADER yields substantial improvements on event-level grounding and temporal consistency; on LLaVA-Video-7B, it reaches 72.60% accuracy on EventHallusion.

cs.CV

GIFT: Geometry-Invariant Fine-Tuning for Non-Lambertian Monocular Depth Estimation

Monocular depth foundation models, benefiting from large-scale synthetic training data, have demonstrated strong generalization. However, they often hallucinate depth on non-Lambertian surfaces, estimating reflected content in mirrors or transmitted content behind glass rather than the physical surface itself. Adapting these models with real-world data is challenging because conventional depth sensors are also unreliable in such regions. We observe that while the appearance of a non-Lambertian surface varies with its reflected or transmitted environment, its underlying geometry remains unchanged. Based on this observation, we propose GIFT (Geometry-Invariant Fine-Tuning), a parameter-efficient post-training framework that requires no measured depth labels. We collect groups of RGB images under controlled appearance changes while keeping the camera and target geometry fixed. GIFT exploits geometric invariance across these observations to suppress non-Lambertian depth hallucinations while retaining general depth estimation capability. We further construct a controlled benchmark that evaluates non-Lambertian depth recovery, robustness to appearance changes, and performance retention in other regions. Experiments on our benchmark and an independent real-world dataset demonstrate that GIFT improves depth prediction for mirrors and transparent objects while largely preserving the base model's performance, providing a practical and low-cost approach for adapting monocular depth foundation models to non-Lambertian scenes.

cs.CV

d-Spectral Bitopological Spaces

We introduce and study the category of \emph{d-spectral spaces}, a bitopological analogue of the classical spectral spaces of Stone and Hochster. A d-spectral space is a compact, d-sober bitopological space such that both open set lattices are coherent frames, where d-sobriety is the bitopological notion of sobriety due to Jung and Moshier. We show that the category of spectral spaces embeds into the category of d-spectral spaces as a simultaneously reflective and coreflective full subcategory. Moreover, we prove that d-spectral spaces are precisely the spectra of d-lattices. Key to this result is the d-lattice of compact open sets associated to a d-spectral space and the spectrum construction for d-lattices. We also show that the patch space of a d-spectral space is d-Boolean and that the de Groot dual of a d-spectral space is again d-spectral, mirroring the corresponding classical properties of spectral spaces. Our results demonstrate that d-spectral spaces form a natural and well-behaved bitopological extension of the spectral space framework.

math.GN

Diophantine analysis and the Braid group ${\bf B}_3$

Given a finite dimensional representation $\pi$ of a finitely generated group $G=\langle g_1, \ldots, g_n\rangle$, the associated characteristic polynomial is defined as $Q_\pi(z):=\det(z_0I+z_1\pi(g_1)+\cdots +z_n\pi(g_n))$, and it is known to contain a good amount of structural information about $G$ and $\pi$. This paper is a part of an ongoing project to investigate the number-theoretic properties of the algebraic varieties (called {\em eigensurfaces}) $\{z\in \mathbb{C}^{n+1}: Q_\pi(z)=0\}$. Its focus is the distribution of prime triples in the eigensurface $S:=\{z\in \mathbb{C}^3: (z_0+z_1+z_2)^2+z_0z_1=0\}$ associated with the braid group ${\bf B}_3$ and its reduced Burau representation. We prove that such triples occur with higher frequency on $S$ than in the ambient lattice, revealing an unexpected connection between group representation theory and analytic number theory.

math.NT

Serrin's Problem under Dirichlet Perturbations: Geometric Compactness and Sharp Planar Stability

In earlier work [21], we posed a stability question for Serrin's overdetermined problem under Dirichlet perturbations and proved that the answer is negative in dimensions $n\ge3$. Here we resolve the question in the planar convex class and obtain a sharp quantitative theory without any a priori geometric nondegeneracy. Let $u_\Omega$ solve \[ -\Delta u_\Omega=1\ \text{in }\Omega,\qquad \partial_\nu u_\Omega=-\frac{|\Omega|}{P(\Omega)}\ \text{on }\partial\Omega, \qquad \int_{\partial\Omega}u_\Omega\,d\sigma=0, \] and set $O(\Omega):=\text{osc}_{\partial \Omega}u_\Omega$. We construct fixed-area annuli with $O(\Omega_k)\to0$ that remain far from every disk, showing that convexity is essential in dimension two. By contrast, if $\Omega_k\subset\mathbb R^2$ are convex, $|\Omega_k|=\pi$, and $O(\Omega_k)\to0$, then, up to translations, $\Omega_k$ converges in Hausdorff distance to the unit disk. Moreover, \[ R_\Omega-r_\Omega+\inf_{z\in\mathbb R^2}d_H(\Omega,B_1(z)) \le C\,O(\Omega) \] for all planar convex $\Omega$ with $|\Omega|=\pi$ and sufficiently small $O(\Omega)$, and the linear order is optimal. The proof combines a new mechanism excluding long-thin degeneration, the rough-domain Serrin rigidity theorem of Figalli--Zhang, new tangential-gradient and linear boundary-growth estimates, a boundary $P$-function estimate, and the reverse-Serrin identity of Magnanini--Molinarolo--Poggesi. We also study the weaker deficit \[ A(\Omega):=\frac1{P(\Omega)}\int_{\partial\Omega}u_\Omega,d\sigma-\min_{\partial\Omega}u_\Omega. \] In the planar convex class, $A(\Omega_k)\to0$ still forces convergence to a disk, and \[ R_\Omega-r_\Omega+\inf_z d_H(\Omega,B_1(z)) \le C A(\Omega)^{2/3} \] for $|\Omega|=\pi$ and sufficiently small $A(\Omega)$.

math.AP

Autonomous End-to-End SOH Prediction Services for Battery Systems via Temporal-Contrastive Representation Learning

Accurate state of health (SOH) estimation is a critical diagnostic service for lithium-ion battery management. However, reliance on labor-intensive manual feature engineering and opaque black-box models hinders scalable industrial deployment. To address this, we introduce TC-SOH: a modular, plug-and-play service architecture for autonomous, end-to-end SOH prediction. TC-SOH employs a temporal-contrastive mechanism and a cross-window prediction pretext task to extract degradation-relevant representations directly from raw operational data. To improve transparency, we connect model efficacy with representation diagnostics: visualization, sensitivity analysis, redundancy analysis, bidirectional probing, future-SOH probing, and temporal shuffling show that learned features overlap with selected expert descriptors while retaining additional SOH-relevant variation, and that ordered temporal context improves subsequent-SOH prediction. Across four public datasets, TC-SOH outperforms the considered physics-informed and data-driven baselines, reducing MAPE by 1.91 times and RMSE by 2.13 times.

cs.LG

Large Language Model-Powered Query-Driven Event Timeline Summarization in Industrial Search

Understanding how events evolve over time is essential for search engines handling queries about trending news. We present QDET (Query-Driven Event Timeline Summarization), a production system deployed on Baidu Search that constructs focused event timelines to explain specific query events. Unlike traditional topic-centric approaches that aim for comprehensive coverage, QDET identifies and organizes sub-events closely relevant to the query from noisy candidate sets formed by millions of documents retrieved daily. QDET incorporates two key innovations: (1) multi-task supervised fine-tuning with three auxiliary tasks-temporal ordering, causal judgment, and timeline completion-that enable compact models to match the performance of much larger general-purpose models in specialized domains; (2) reinforcement learning-based event concise summarization that enforces strict length constraints while maintaining semantic quality, achieving 88.2% length compliance and outperforming 671B-scale models by 7.7 points in constraint satisfaction. Our fine-tuned 7B parameter model achieves 76.2% F1 score on timeline summarization, slightly surpassing the zero-shot performance of DeepSeek-R1-671B (76.1% F1) while using only 1% of its parameters-demonstrating that domain-specific optimization enables production-ready models with comparable quality at drastically reduced computational costs. Online A/B tests on Baidu Search validate real-world effectiveness, showing 5.5% CTR improvement, 4.6% longer dwell time, and 4.4% deeper exploration compared to single-task baselines. We further demonstrate that timeline understanding transfers to heat prediction, confirming effective knowledge transfer to downstream tasks.

cs.CL

Microwave photonic radar jamming and target detection integration based on advanced waveform editing, forwarding, and self-squaring reception

The integrated radar and jamming (IRAJ) system provides a promising solution that meets the demands for miniaturization, integration, and multifunctionality in complex warfare environments. However, traditional electronic-domain IRAJ systems face limitations in operating frequency and bandwidth. In this paper, we propose and experimentally demonstrate a microwave photonic IRAJ system based on pseudo-random binary phase modulation and segmented frequency shifting. By modulating pseudo-random binary coding sequence and frequency-shifting signals onto linearly frequency-modulated (LFM) pulses, an IRAJ waveform is generated to achieve noise-like jamming against the adversary radar. To overcome the random {\pi}-phase jumps introduced by pseudo-random binary modulation in the de-chirped signal, a time-domain squaring operation is implemented during de-chirped reception, restoring the radar detection ability of our system and enabling accurate target sensing without prior knowledge of the coding sequence. Experimental results demonstrate that the system can generate IRAJ waveforms with a bandwidth of up to 4 GHz, covering both 10-28 GHz. The proposed system achieves effective jamming against adversary radars employing either de-chirped reception or pulse compression, with the generated jamming results exhibiting an irregular and random distribution of false targets. Meanwhile, the system maintains radar performance with a ranging error of around 5 cm and a radial velocity measurement error below 4 cm/s.

physics.optics

Cross-Modal RGB-D Fusion Transformer for 6D Pose Estimation of Non-Cooperative Spacecraft with Stereo-Derived Depth

On-orbit servicing and active debris removal involving non-cooperative spacecraft require reliable pose estimation to supply accurate position and orientation data for autonomous visual navigation. Learning-based monocular methods have seen widespread adoption in spacecraft pose estimation, yet they suffer from an intrinsic depth ambiguity problem and tend to fail under the harsh illumination conditions routinely encountered in orbit. Active depth sensors could in principle address the geometric ambiguity, but their power and mass requirements make them poorly suited to most spacecraft platforms. This work addresses these issues through a passive stereo vision framework for six-degree-of-freedom (6-DOF) pose estimation of non-cooperative spacecraft. A binocular stereo matching network called TSCA-Stereo is developed to cope with weak-texture surfaces, specular highlights, and severe lighting variations typical of space imagery. A cross-modal fusion Transformer is introduced to combine RGB appearance information with stereo depth features in an adaptive manner, supporting reliable pose recovery. A synthetic binocular multimodal dataset is also built for the experiments, covering stereo disparity maps and 6-DOF pose annotations across a range of illumination scenarios, attitude configurations, and noise levels. Experimental results show that TSCA-Stereo outperforms the baseline across every evaluated metric on this space-specific dataset. The full pose estimation pipeline achieves a mean translation error of 0.0419 m and a mean orientation error of 0.8632{\deg} under varied imaging conditions, confirming that the passive stereo approach is both effective and resilient when operating under the demanding visual conditions of the space environment.

cs.CV

Enhanced Phase Sensitive SD-OCT for flow imaging using ultrasonically sculpted optical waveguides

Phase sensitive detection in spectral domain optical coherence tomography (SD-OCT) is a powerful method for functional imaging of biological events with high spatiotemporal resolution. The depth-dependent signal-to-noise ratio (SNR) is a limiting factor on the minimum detectable phase changes of phase in shot noise-limited SD-OCT systems. The SNR over a depth is constrained by the terminal optics, usually using a focusing lens to project light into the tissue and collect the backscattered light. In situ ultrasonically sculpted optical waveguides have been used to improve SNR roll-off over depth compared to conventional SD-OCT systems. In this paper, we extend this feature to demonstrate phase sensitive detection at depth using ultrasonically enhanced OCT (ue-OCT). Our experimental results show that ultrasonically sculpted optical waveguides are phase stable and follow near shot-noise limited behavior. We measured milk flow velocity changes to demonstrate a phase sensitivity of 5.25 mrad at 10 dB SNR and dynamic range of 0.8 mm/s to 14.7 cm/s using ue-OCT. Our results show flow detection with ue-OCT at extended depths (i.e., 3.5 mm) otherwise not possible with conventional SD-OCT systems with matched focal lengths. The results in this paper show the potential of ue-OCT for phase-sensitive flow measurement from the depth of tissue for a gamut of applications such as cerebral blood flow imaging as a proxy to neural activity mapping.

physics.optics

Dynamically cold discs in high-redshift galaxies: comparison between ALMA observations and TNG50

Observations of highly rotationally supported gas discs in high redshift ($z$ > 3) star-forming galaxies challenge our understanding of galaxy formation, as the prevailing view holds that galaxies in the early universe are dynamically hot due to frequent mergers, gas accretion, and strong stellar feedback. We examined the kinematic properties of massive ($M_{\star} \geq 10^{10}\,M_{\odot}$) star-forming galaxies in the TNG50 cosmological hydrodynamical simulation in the redshift range $3\leq z \leq 5$. Mock emission line datacubes were constructed and analysed using the same methodology as for [CII] observations with ALMA. We measured the ratio of the gas rotation velocity ($V$) to velocity dispersion ($\sigma$) finding that most galaxies have $V/\sigma\sim$ $2-3$, lower than observed. However, a few simulated galaxies show $V/\sigma$ > 5. Such "cold" discs, selected at $z=4$, remain dynamically colder than most of the TNG population across $z=3-5$. A galaxy with $V/\sigma\gtrsim10$ appears in a transient phase that lasts $\leq200$ Myr. Dynamically cold disc formation in TNG50 is promoted by gas accretion with angular momentum aligned with the pre-existing disc, while most galaxies undergo misaligned accretion. Dynamically cold discs also show lower mass accretion rates and better aligned stellar and dark-matter angular momentum vectors. By tracing their evolution to $z = 0$, we find that one-third become massive disc galaxies and two-thirds become ETGs.

astro-ph.GA

A Flow Matching Framework for Soft-Robot Inverse Dynamics

Learning the inverse dynamics of soft continuum robots remains challenging due to high-dimensional nonlinearities and complex actuation coupling. Conventional feedback-based controllers often suffer from control chattering due to corrective oscillations, while deterministic regression-based learners struggle to capture the complex nonlinear mappings required for accurate dynamic tracking. Motivated by these limitations, we propose an inverse-dynamics framework for open-loop feedforward control that learns the system's differential dynamics as a generative transport map. Specifically, inverse dynamics is reformulated as a conditional flow-matching problem, and Rectified Flow (RF) is adopted as a lightweight instance to generate physically consistent control inputs rather than conditional averages. Two variants are introduced to further enhance physical consistency: RF-Physical, utilizing a physics-based prior for residual modeling; and RF-FWD, integrating a forward-dynamics consistency loss during flow matching. Extensive evaluations demonstrate that our framework reduces trajectory tracking RMSE by over 50% compared to standard regression baselines (MLP, LSTM, Transformer). The system sustains stable open-loop execution at a peak end-effector velocity of 1.14 m/s with sub-millisecond inference latency (0.995 ms). This work demonstrates flow matching as a robust, high-performance paradigm for learning differential inverse dynamics in soft robotic systems.

cs.RO

APOSTLE vs. AURIGA Simulations: How Subgrid Models Shape Milky Way Analogs

Despite significant progress in cosmological simulations of galaxy formation, the role of subgrid physics in shaping the detailed properties of galaxies remains incompletely understood. In this work, we analyze two sets of zoom-in simulations that share identical initial conditions but adopt distinct implementations of baryonic physics, enabling a controlled comparison of their predictions. We examine the stellar properties, morphological structures, and satellite populations of the simulated galaxies at $z=0$. We find that AURIGA galaxies systematically exhibit higher stellar masses and surface densities than their APOSTLE counterparts. These differences are primarily driven by variations in the efficiency of gas cooling from the circumgalactic medium (CGM) into the star-forming gas. Both simulations form well-defined disk galaxies; however, AURIGA systems generally display higher disk-to-total mass ratios, earlier disk formation, and more prominent dynamical structures such as bars and spiral arms. Nevertheless, strongly disk-dominated systems are present in both simulations, although they do not arise in the same host haloes. The vertical disk structure in both simulations is well described by a sech density profile, with scale heights below ~ 1 kpc in the inner regions. The satellite populations also differ, with AURIGA producing systematically more massive satellites, including a ~ 0.3 dex increase in the most massive system, while the number of satellites above $10^6 M_{\odot}$ remains comparable in most halo pairs. Both simulations reproduce similar satellite stellar mass--metallicity relations, albeit ~ 0.25 dex higher than observation. This comparative study therefore provides useful benchmarks for future efforts to better constrain galaxy formation models.

astro-ph.GA

The capture of halo material by orbiting subhaloes

When a dark matter halo falls into a more massive object and becomes a subhalo, it typically loses much of its mass through tidal stripping. The reverse process is also possible in principle. The subhalo may gravitationally capture material from its host. If sufficiently efficient, this process could make an initially starless subhalo visible. We use high-resolution N-body simulations to estimate the efficiency of capture. We find that after an extended period orbiting within its host, at most $\sim 10^{-4}$ of a subhalo's remaining mass has been acquired since infall. This captured material is less concentrated to subhalo centre than material retained from before infall. It is also very much less abundant than host material that is instantaneously passing through the subhalo on almost unperturbed orbits. Captured stars are not sufficiently spatially concentrated to be distinguished from the dominant background of "field" stars, and their concentration in velocity space is no greater than that of typical stellar streams in the halo. Unfortunately, stellar capture is not efficient enough to allow initially starless low-mass subhaloes to be detected.

astro-ph.GA

Generative Recommendation for Large-Scale Advertising

Generative recommendation has recently attracted widespread attention in industry due to its potential for scaling and stronger model capacity. However, deploying real-time generative recommendation in large-scale advertising requires designs beyond large-language-model (LLM)-style training and serving recipes. We present a production-oriented generative recommender co-designed across architecture, learning, and serving, named GR4AD (Generative Recommendation for ADdvertising). As for tokenization, GR4AD proposes UA-SID (Unified Advertisement Semantic ID) to capture complicated business information. Furthermore, GR4AD introduces LazyAR, a lazy autoregressive decoder that relaxes layer-wise dependencies for short, multi-candidate generation, preserving effectiveness while reducing inference cost, which facilitates scaling under fixed serving budgets. To align optimization with business value, GR4AD employs VSL (Value-Aware Supervised Learning) and proposes RSPO (Ranking-Guided Softmax Preference Optimization), a ranking-aware, list-wise reinforcement learning algorithm that optimizes value-based rewards under list-level metrics for continual online updates. For online inference, we further propose dynamic beam serving, which adapts beam width across generation levels and online load to control compute. Large-scale online A/B tests show up to 4.2% ad revenue improvement over an existing DLRM-based stack, with consistent gains from both model scaling and inference-time scaling. GR4AD has been fully deployed in Kuaishou advertising system with over 400 million users and achieves high-throughput real-time serving.

cs.IR

Dynamic Dispersion Accumulation in Fiber Loops for Realizing Record-High Frequency Resolution or Ultra-Low Signal Sampling Rate in Dispersion-Based Photonics-Assisted Wideband Microwave Measurement Systems

Dispersion-based photonics-assisted microwave measurement systems provide immense potential for real-time analysis of wideband and dynamic signals. However, they face two critical challenges: a difficulty in achieving high frequency resolution over a wideband analysis bandwidth, and a reliance on large-bandwidth-and-high-sampling-rate oscilloscopes to capture the resulting ultra-narrow pulses. We introduce a dynamic dispersion accumulation technique to overcome these limitations. By circulating the optical signal in fiber loops containing a dispersion-compensating fiber, we achieve a high accumulated dispersion of -215700 ps/nm. This high dispersion relaxes the required chirp rate of the chirped optical signal, enabling two distinct advantages: When the analysis bandwidth is fixed, a lower chirp rate enables a longer temporal period, yielding a record-high frequency resolution of 27.9 MHz; When the temporal period is fixed, a lower chirp rate enables a smaller bandwidth, generating a wider pulse and thus relaxing pulse sampling requirements at the expense of analysis bandwidth. This sacrifice in analysis bandwidth can be compensated by a duty-cycle-enabling technique, which holds the potential to extend the analysis bandwidth beyond 100 GHz. This work breaks the performance and hardware limitations in dispersion-based systems, paving the way for high frequency resolution, wideband microwave measurement systems that are both real-time and cost-effective.

physics.optics

ScholarGym: Benchmarking Large Language Model Capabilities in the Information-Gathering Stage of Deep Research

Large language models have advanced from single-turn question answering to deep research systems that iteratively decompose research questions, invoke retrieval tools, and synthesize information across multiple rounds. Evaluating such systems typically involves scoring their final research reports holistically, but this end-to-end paradigm tightly couples the language model's decision-making, workflow design, and environmental feedback, precluding decomposable analysis of individual components. We introduce ScholarGym, an evaluation environment that isolates the information-gathering stage of deep research on academic literature. Under a unified workflow, ScholarGym decomposes the research process into three explicit stages -- Query Planning, Tool Invocation, and Relevance Assessment -- and evaluates each against 2,536 expert-annotated queries over a static corpus of 570K papers with deterministic retrieval. Systematic experiments reveal that iterative query decomposition yields 2.9--3.3$\times$ F1 gains over single-query retrieval, models with extended thinking trade recall for precision, and Query Planning quality together with Relevance Assessment constitute dual bottlenecks that separate proprietary from open-source model performance.

cs.AI

Mass distribution of ultralight boson in binary black hole systems

Ultralight bosons are compelling dark-matter candidates. Both scalar and vector bosons can be produced through black hole superradiance, forming a boson cloud surrounding a rotating black hole. Self-interaction of bosons, together with transition mixing in binary black hole systems, give rise to dynamical phenomena that could be potentially observable with future gravitational wave observations. In this work, we investigate the dynamics of bosons in binary black hole systems. In particular, we focus on boson mass transfer in unequal-mass binary black hole systems with arbitrary spin-orientation of the companion. Our results show that the mass ratio between the companion and the primary black holes significantly affects cloud absorption through mass transfer. Moreover, when the companion's spin is not aligned with that of the primary, the efficiency of cloud depletion is further modified.

gr-qc