Search arXiv⌕ Search

arXiv subjects

Yuhan Zhang

Publications and source records attributed to Yuhan Zhang.

At least 19 recordsLinked to original sources

Label-Efficient Generative Inverse Design and Design-Rule Discovery in Freeform Topological Photonics

Topological photonic crystals support robust, disorder-resilient light transport protected by their band-structure topology, yet their design remains confined to a small set of symmetry-defined templates. The surrounding freeform design space, where optimal performance and new functionality may reside, is difficult to access because of its high dimensionality and costly full-wave simulation needed for evaluating each candidate. % Here we report generative inverse design of freeform valley photonic crystals. A label-free diffusion prior learns the distribution of three-fold rotationally symmetric geometries from 7,689 unlabeled designs that cost 0.15~s each to produce, while a convolutional surrogate is trained on 1,666 full-wave band-gap simulations costing 4.28~min each, exploiting a 1,700-fold cost asymmetry. Across nine target band gaps up to $225~$meV, we generate and verify by full-wave simulation 270 valley photonic crystal designs, reaching a mean absolute error of $1.4$--$6.7$~meV within the labeled range and retaining below $5.1\%$ fractional error at targets $25\%$ beyond its upper bound, while a label-conditioned diffusion model trained on the same label data performs substantially worse. Beyond inverse design, the trained model functions as an instrument for discovering structure--property relations. At a target band gap where the labeled set contains only five designs, it generates hundreds, resolving geometric trends that are statistically inaccessible in the training data and yielding interpretable rules for large-gap valley photonic crystals. Our results establish surrogate-guided diffusion as a label-efficient route to freeform topological-photonic design, and as a means of extracting design principles where physical intuition and simulation labels are both scarce.

physics.optics↗

Stealthy in Semantics, Antagonistic in Space: Attacking Visible-Infrared Object Detectors via Object-Level Misalignment

Visible-infrared object detectors are used for robust perception under challenging illumination and weather conditions. Current physical attacks apply conspicuous patches to spatially aligned target regions, which are noticeable to human observers. Meanwhile, most of these methods only perturb the appearance within the aligned region, without explicitly targeting the correspondence between modalities or the fusion process. In this paper, we propose CamoShift, an adversarial framework for visible-infrared object detection. By combining visual camouflage with object-level infrared shifting, CamoShift breaks cross-modal spatial alignment and disrupts fusion. Specifically, the Semantic Camouflage Module (SCM) generates a stealthy camouflaged patch that can be attached to the host object and maintains its effectiveness in the infrared branch through an RGB-IR adapter. The Object-level Spatial Decoupling Module (OSDM) shifts the infrared target evidence in a scale-aware manner, so as to break object-level correspondence and disrupt cross-modal fusion. Then, the Harmonic Adversarial loss (HarAdv loss) further balances attack strength and visual stealth during optimization. To the best of our knowledge, we are the first to target both visual stealthiness and attack success in visible-infrared object detection. Extensive experimental results show that CamoShift achieves a superior balance between attack effectiveness and visual stealth. Code and models will be available on GitHub.

cs.CV↗

Multi-View Mixture-of-Experts with Vision-Language Reranking for Cross-View Object Geo-Localization

Cross-view object geo-localization (CVOGL) locates a target in satellite imagery using drone or street-view queries. Existing methods train separate detectors for each viewpoint, leading to parameter redundancy and impeding cross-view knowledge sharing. Moreover, top-ranked satellite candidates are often visually similar, so visual appearance and categorical labels alone are insufficient to resolve such ambiguity. To address these, we propose MVLGeo, an efficient framework designed to unify multiple viewpoints and reduce model redundancy. First, we introduce environmental contextual text from the query view as cues to distinguish visually similar candidates via Vision-Language Reranking (VL-Rerank). Second, we design a multi-view Mixture-of-Experts architecture (MV-MoE) with a shared encoder and view-specific experts to reduce redundancy and promote knowledge sharing, while cross-view contrastive learning aligns their representations for consistency. Third, we introduce an adaptive elliptical prior (ESAM-Prior) as auxiliary positional encoding for anisotropic geometric perception. Extensive experiments on the CVOGL benchmarks confirm that MVLGeo, as a unified model for multiple query viewpoints, achieves state-of-the-art performance, demonstrating robustness to input degradation and generalization across viewpoints. Code and models will be available on GitHub to facilitate future work.

cs.CV↗

ReUnit: Multi-Granularity Visual Unitization for Long Video Understanding

Long-video understanding is constrained by the limited visual input capacity of video multimodal large language models (Video-MLLMs). Existing methods mainly optimize which content to retain, while the presentation of retained content often remains fixed. As a result, the same balance between spatial detail and content coverage is imposed across the entire visual input. We propose ReUnit, a training-free and query-aware framework that jointly determines which content to retain and how it should be presented. Guided by frame-level query relevance, ReUnit constructs and allocates visual units at multiple presentation granularities. More relevant content receives finer presentation, while broader context is represented more compactly. It realizes these granularities using visual units that carry one, four, or nine source frames and are each rendered as a standard image. Across four benchmarks and visual input budgets from 4 to 32 units, ReUnit achieves the highest average score at every tested budget. It improves over Uniform Sampling by 8.3--10.1 points on average, and the gains transfer across three additional Video-MLLM families. Project resources are available at https://github.com/charon525/ReUnit.

cs.CV↗

Tight instances of the Lonely Runner Conjecture: complete classification of one-entry modifications, a new infinite family, and the growth bound

For a set V of n-1 distinct positive integers write LR(V) = max_t min_{v in V} ||vt||, where ||x|| is the distance from x to the nearest integer; V is tight if LR(V) = 1/n, the value predicted by the Lonely Runner Conjecture. The baseline [n-1] = {1,...,n-1} is tight for every n, and Perarnau and Serra list the characterization of tight instances as Problem 1 of their survey, noting no further progress since Goddyn and Wong, who classified the multiple case and proved finiteness for each fixed deleted speed. We settle the one-entry case completely: ([n-1] minus {r}) union {w} with w > n-1 is tight if and only if either 2r > n-1, r divides w and the Goddyn-Wong gcd criterion holds, or (n,r,w) = (5,2,7) or (6,2,9). In particular no tight one-entry modification exists for 3 <= r <= (n-1)/2, and the only tight cases beyond the Goddyn-Wong multiples are the two sporadic sets of Wills, {1,3,4,7} and {1,3,4,5,9}. The proof is purely theoretical: the connected components of the uncovered region U(n,r) are determined exactly in both regimes 2r > n-1 and 2r <= n-1, yielding the effective bound w <= 4rI/(2s-I), with s = n-r and I the least integer of [s, n-1] coprime to r; a single inequality, proved via the Jacobsthal function, closes the mid-range without computation. Tight one-entry modifications also satisfy max V <= 0.60 n log n + 52 n unconditionally, with sharp leading constant 1/2 along n = p#+2 (p# the primorial), so no linear bound confines tight instances. We further isolate an explicit CRT doubling subfamily of the Goddyn-Wong multi-acceleration theorem and prove for two-speed replacements that no tight instance contains a removed speed q with 2 <= q <= n/10 when n >= max(40,10q), whatever the second removal and the inserted speeds. An exact census over all n <= 140, w <= 8n finds only two tight two-swaps, the Wills set {1,4,5,6,7,11,13} and the Goddyn-Wong doubling at n = 74.

math.CO↗

Dense-core approach to the Brualdi--Hoffman--Turán problem on odd wheels

We present a unified presentation of the fixed-size adjacency-spectral extremal problem for odd wheels $W_{2k+1}$, where $k\geq2$ and $W_{2k+1}=K_1\vee C_{2k}$. The exceptional case $W_5$ and the general case $W_{2k+1}$, $k\ge3$, share the same dense-core reduction and edge-spectral stability, but have different rigidity structures. We prove that every $W_5$-free graph of sufficiently large size $m$ satisfies $ρ(G)^2-ρ(G)\le m,$ with equality precisely for $K_{n,n}$ with a perfect matching embedded in each part, where $n$ is even and $m=n^2+n$. For any fixed $k\ge3$, every $W_{2k+1}$-free graph of sufficiently large size $m$ satisfies $ρ(G)^2-(k-1)ρ(G)\le m-\binom{k}{2},$ with equality precisely for $K_k\vee qK_1$ when $m=\binom{k}{2}+kq$. Our results completely settle a conjecture proposed by Yu, Li and Peng and, via a distinct approach, further strengthen known results concerning odd cycles, friendship graphs and odd fan graphs for sufficiently large $m.$ The proof combines the edge-spectral stability theorem, residual functions and the dense-core method.

math.CO↗

Focal-point scanning for dose delivery and optimization with focused laser-accelerated very-high-energy electron beams

Focused very-high-energy electron (VHEE) beams can produce localized dose enhancement at selected depths, but irradiation of a finite target requires coordinated control of multiple focal positions, incidence directions, and beam weights while limiting exposure of nearby organs at risk (OARs). We present Focal-Point Scanning (FPS), a dose delivery and optimization method developed for laser wakefield accelerator (LWFA)-driven VHEE beams. The method is based on a two-dipole focusing system that produces single-plane beam convergence and allows the focal position to be varied by changing the magnetic field strength. FPS distributes focal points throughout the planning target volume and determines focal-point-specific incidence sectors according to the geometry of nearby critical OARs. The method was evaluated using the AAPM TG119 C-shape benchmark and one previously treated lung radiotherapy case. At matched target coverage, FPS reduced the TG119 Core mean dose by approximately one half relative to parallel VHEE and intensity-modulated x-ray plans, approaching the single-field proton pencil-beam-scanning reference. In the lung case, FPS maintained target coverage comparable to the clinical volumetric modulated arc therapy reference while reducing the mean dose to every evaluated OAR; spinal-cord mean and maximum doses decreased by 93.2% and 87.2%, respectively. The evaluated OAR mean doses varied little across rms energy spreads of 0 to 10% and for a flat-top electron spectrum spanning 150 to 250 MeV. These results demonstrate that focal-point-specific angular selection can translate focused-beam physics into effective OAR sparing and support FPS as a planning strategy for broadband LWFA-VHEE radiotherapy.

physics.med-ph↗

AnomExpert: Identifying and Selecting Anatomical Planes for Prenatal Ultrasound Anomaly Diagnosis

Life-limiting congenital anomalies require accurate prenatal diagnosis for appropriate clinical decision-making. Prenatal ultrasound (US) examinations involve multiple anatomical planes, and diagnosis depends on identifying anatomical planes and selecting diagnostically relevant planes for each anomaly. Existing automated methods either rely on plane-level annotations or aggregate heterogeneous images without explicitly modeling these diagnostic capabilities. We propose AnomExpert, a prototype-driven framework for prenatal US anomaly diagnosis using only case-level supervision. AnomExpert introduces learnable plane prototypes to organize unordered images into latent representations corresponding to anatomical planes without requiring plane annotations. A disease-aware sparse selection mechanism further selects diagnostically relevant planes for each anomaly. Experiments on a multi-center dataset of 3,654 cases show that AnomExpert consistently outperforms nine representative multi-instance learning methods. Using a ViT-small backbone, it achieves 86.9% accuracy and 84.2% F1-score while maintaining parameter efficiency. These findings indicate that modeling anatomical plane identification and disease-specific plane selection improves weakly supervised multi-plane prenatal US anomaly classification. The code is available at https://github.com/TIanCat/AnomExpert.

cs.CV↗

Graded strength of comparative illusions is explained by Bayesian inference

Like visual processing, language processing is susceptible to illusions in which people systematically misperceive stimuli. In one such case--the comparative illusion (CI), e.g., More students have been to Russia than I have--comprehenders tend to judge the sentence as acceptable despite its underlying nonsensical comparison. Prior research has argued that this phenomenon can be explained as Bayesian inference over a noisy channel: the posterior probability of an interpretation of a sentence is proportional to both the prior probability of that interpretation and the likelihood of corruption into the observed (CI) sentence. Initial behavioral work has supported this claim by evaluating a narrow set of alternative interpretations of CI sentences and showing that comprehenders favor interpretations that are more likely to have been corrupted into the illusory sentence. In this study, we replicate and go substantially beyond this earlier work by directly predicting the strength of illusion with a quantitative model of the posterior probability of plausible interpretations, which we derive through a novel synthesis of statistical language models with human behavioral data. Our model explains not only the fine gradations in the strength of CI effects, but also a previously unexplained effect caused by pronominal vs. full noun phrase than-clause subjects. These findings support a noisy-channel theory of sentence comprehension by demonstrating that the theory makes novel predictions about the comparative illusion that bear out empirically. This outcome joins related evidence of noisy channel processing in both illusory and non-illusory contexts to support noisy channel inference as a unified computational-level theory of diverse language processing phenomena.

cs.CL↗

Prototype Memory-Guided Training-Free Anomaly Classification and Localization in Prenatal Ultrasound

Prenatal anomaly classification and localization is of critical importance for fetal health and pregnancy management. Although ultrasound (US) is the primary modality for prenatal screening, accurate diagnosis remains challenging due to the low prevalence and high heterogeneity of anomalies. Existing deep learning methods for prenatal tasks rely on large-scale annotated datasets, which are difficult to obtain in practice. Although few-shot learning alleviates data scarcity, it typically requires fine-tuning for new categories, limiting its practicality in resource-limited clinical settings. To address these challenges, we propose a training-free framework for multi-class prenatal US anomaly classification and localization that operates with only a few reference images per class, representing the first exploration of this setting. Our framework comprises three key components: (1) a memory bank with multi-granular prototypes that explicitly models both class-level semantics and anomaly characteristics; (2) a prototype-driven soft merging mechanism that aggregates discriminative features to detect the anomaly region; and (3) a class-aware refinement strategy that leverages prototype consistency to improve category prediction. Extensively validated on a multi-center prenatal US dataset containing 1,149 cases, with a total of 2,357 images and 9 categories, our proposed method outperforms the competitors.

cs.CV↗

FrameONE: Hierarchical Motion Modeling for Universal Multi-View Echocardiographic Keyframe Detection

Accurate detection of end-systole (ES) and end-diastole (ED) frames is fundamental to echocardiographic assessment. Existing methods are typically developed in a view-specific manner, depend on auxiliary annotations or intensive visual modeling, which limits their generalizability. In multi-view modeling, keyframe detection is driven by shared cardiac motion, yet large appearance differences and motion patterns make unified modeling challenging. To address these issues, we propose FrameONE, a unified end-to-end framework for multi-view echocardiographic keyframe detection. FrameONE introduces a Hierarchical Motion Modeling strategy: an intra-view multi-task learning reduces appearance bias and promotes motion-focused representations within each view; an inter-view general motion learning module further separates view-agnostic dynamics from view-specific patterns, enabling shared yet flexible motion representation learning across views. Extensive experiments on 25,872 videos spanning four standard views demonstrate that FrameONE achieves state-of-the-art keyframe detection accuracy with strong cross-view generalization. Code is available at https://github.com/szuboy/FrameONE.

cs.CV↗

Mandol: An Agglomerative Agent Memory System for Long-Term Conversations

Long-term conversational agents need to remember and query cross-session, multi-typed information with complex correlations. Existing agent memory systems rely on heterogeneous vector and graph databases, which fragment memory information and cause high cross-database I/O latency. For retrieval, common RAG-style methods tend to introduce noise, miss correlated clues, and lack token budget control, degrading LLM accuracy and efficiency. We propose Mandol, an agglomerative memory system that consolidates fragmented memory representations and storage into a unified memory-native architecture. Its core components include: (1) a hierarchical memory model that organizes memory into a basic layer representing raw memory information and a high-level abstract layer that agglomerates basic memories into traceable abstract memories, both uniformly represented as structured semantic graphs; (2) an agglomerative semantic data structure combining SemanticMap and SemanticGraph, which natively fuses key-value, vector, and graph structures and provides unified hybrid retrieval operators to eliminate cross-database I/O; and (3) a quantitative query mechanism with query-adaptive routing, quantitative denoising and conflict resolution, and token-constrained context generation, all without involving LLMs during retrieval. Experiments on two widely used long-term conversation benchmarks, LoCoMo and LongMemEval, show that Mandol achieves the best overall accuracy among representative agent memory systems. For performance comparison, Mandol also obtains a 5.4x retrieval speedup and a 4.8x insertion speedup under 10 QPS concurrent load, while still maintaining low latency on consumer-grade hardware.

cs.DB↗

A spectral threshold for triangle counting

The 1970 spectral extension of Mantel's theorem, proved by Nosal, states that every graph with $m$ edges and spectral radius $ρ_1>\sqrt{m}$ contains at least one triangle. Its quantitative refinement by Ning and Zhai later established that any graph $G$ with $m$ edges and spectral radius $ρ_1\geq\sqrt{m}$ contains at least $\lfloor\frac{\sqrt{m}-1}{2}\rfloor$ triangles, unless $G$ is a complete bipartite graph. In this paper, we further investigate the minimum number of triangles guaranteed under the strengthened spectral condition $ρ_1\geq\sqrt{m}+c$, where $c$ is a positive constant. We prove that for any constant $c\in (0,\frac{1}{2}]$ and all sufficiently large $m$, if $s=s(m)$ is a real-valued function satisfying $\lim_{m\to\infty} \frac{s}{m}=c$, then every $m$-edge graph $G$ with spectral radius $ρ_1$ satisfying $ρ_1^2\geq m-1+\frac{2s}{ρ_1-1}$ contains at least $s$ triangles. Moreover, we characterize the extremal graph achieving the minimal number of triangles. In particular, when $s=\frac{m-1}2$, our result settles a conjecture proposed by Li, Feng, and Peng.

math.CO↗

Noisy memory encoding explains negative polarity illusions

A sentence like "The authors that no critics recommended have ever received acknowledgment for a best-selling novel" is sometimes rated as acceptable even though, strictly speaking, it is ungrammatical because the negative polarity word "ever" is not licensed where it is. This behavioral effect is sometimes called a "negative polarity illusion". Here we propose that the lossy context surprisal theory of Hahn et al. (2022) -- whereby people have an imperfect encoding of complex sentences -- might explain this effect. We hypothesize that people have poor memory representation of the determiners in the main-clause and embedded-clause subjects and could entertain a determiner exchange that licenses ever. We propose that more similar determiners in those positions would trigger stronger illusion effects. Acceptability judgment tasks with six novel determiner pairs (e.g., "few" and "many", "few" and "most") support our proposal, showing, specifically, that a novel sentence, "Many authors that few critics recommended have ever received acknowledgment for a best-selling novel", triggered a much stronger illusion than the canonical one even without time pressure. These results offer further support for the suggestion that human language processing is imperfect and resource-rational: in face of working memory limitations, humans rationally reconstruct what is most likely from noisy linguistic input to facilitate downstream processing.

cs.CL↗

Efficient Dataset Distillation for Pre-Trained Self-Supervised Models via Statistical Flow Matching

Dataset distillation seeks to synthesize a highly compact dataset that achieves performance comparable to the original dataset on downstream tasks. For the classification task that use pre-trained self-supervised models as backbones, previous linear gradient matching optimizes synthetic images by encouraging them to mimic the gradient updates induced by real images on the linear classifier. However, this batch-level formulation requires loading thousands of real images and applying multiple rounds of differentiable augmentations to synthetic images at each distillation step, leading to substantial computational and memory overhead. In this paper, we introduce statistical flow matching , a stable and efficient supervised learning framework that optimizes synthetic images by aligning constant statistical flows from target class centers to non-target class centers in the original data. Our approach loads raw statistics only once and performs a single augmentation pass on the synthetic data, achieving performance comparable to or better than the state-of-the-art methods with 10x lower GPU memory usage and 4x shorter runtime. Furthermore, we propose a classifier inheritance strategy that reuses the classifier trained on the original dataset for inference, requiring only an extremely lightweight linear projector and marginal storage while achieving substantial performance gains.

cs.CV↗

DSPE: An Energy-Efficient Edge Processor for DeepSeek Inference with MerkleTree-based Incremental Pruning, Multi-Stage Boothing Lookup and Dynamic Adaptive Posit Processing

In recent years, DeepSeek has achieved strong inference performance but remains hard to deploy on energy-constrained edge devices. This paper presents the DeepSeek Processing Element (DSPE), an edge-oriented architecture that alleviates the model's heavy computational and energy demands. DSPE introduces three techniques: the MerkleTree-based Incremental Pruning Scheme (MIPS) for secure redundant-vector reduction, the Multi-Stage Boothing Lookup Method (MBLM) for bit-flip-aware approximate multiplication, and the Dynamic Adaptive Posit Processing Mechanism (DAPPM), which introduces a new DA-Posit format and its corresponding hardware multiplication architecture. Implemented in TSMC 28nm CMOS, DSPE achieves 109.4 TFLOPS/W energy efficiency compared with state-of-the-art designs and offers a scalable foundation for edge deployment.

cs.AR↗

Artificial Intelligence for Detecting Fetal Orofacial Clefts and Advancing Medical Education

Orofacial clefts are among the most common congenital craniofacial abnormalities, yet accurate prenatal detection remains challenging due to the scarcity of experienced specialists and the relative rarity of the condition. Early and reliable diagnosis is essential to enable timely clinical intervention and reduce associated morbidity. Here we show that an artificial intelligence system, trained on over 45,139 ultrasound images from 9,215 fetuses across 22 hospitals, can diagnose fetal orofacial clefts with sensitivity and specificity exceeding 93% and 95% respectively, matching the performance of senior radiologists and substantially outperforming junior radiologists. When used as a medical copilot, the system raises junior radiologists' sensitivity by more than 6%. Beyond direct diagnostic assistance, the system also accelerates the development of clinical expertise. A pilot study involving 24 radiologists and trainees demonstrated that the model can improve the expertise development for rare conditions. This dual-purpose approach offers a scalable solution for improving both diagnostic accuracy and specialist training in settings where experienced radiologists are scarce.

cs.CV↗

Cooperative Double IRS aided Secure Communication for MIMO-OFDM Systems

Cooperative double intelligent reflecting surface (double-IRS) has emerged as a promising approach for enhancing physical layer security (PLS) in MIMO systems. However, existing studies are limited to narrowband scenarios and fail to address wideband MIMO-OFDM. In this regime, frequency-flat IRS phases and cascaded IRS links cause severe coupling, rendering narrowband designs inapplicable. To overcome this challenge, we introduce cooperative double-IRS-assisted wideband MIMO-OFDM and propose an efficient manifold-based solution. By regarding the power and constant modulus constraints as Riemannian manifolds, we reformulate the non-convex secrecy sum rate maximization as an unconstrained optimization on a product manifold. Building on this formulation, we further develop a product Riemannian gradient descent (PRGD) algorithm with guaranteed stationary convergence. Simulation results demonstrate that the proposed scheme effectively resolves the OFDM coupling issue and achieves significant secrecy rate gains, outperforming single-IRS and distributed multi-IRS benchmarks by 32.0% and 22.3%, respectively.

eess.SP↗