Search arXivSearch

arXiv subjects

Yiming Chen

Publications and source records attributed to Yiming Chen.

At least 19 recordsLinked to original sources

CMAMBADEPTH: Self-supervised Monocular Depth Estimation with Channel Mamba and Hybrid Attention

Accurate monocular depth estimation serves as a core enabler for single camera scene understanding. However, existing self-supervised monocular depth estimation methods generally suffer from the bottleneck of inefficient cross-scale information interaction and difficulty in balancing local and global spatial modeling. In this paper, we propose CMambaDepth, a self-supervised framework that achieves efficient multi-scale feature fusion and fine-grained contextual modeling via channel-wise selective state propagation. Specifically, Bidirectional Channel Mamba (Bi-CMamba) aligns encoder features across scales and enables bidirectional information exchange among ordered scale groups. Unidirectional Channel Mamba (Uni-CMamba) progressively aggregates decoder features and retains fine-grained scale groups through a group selection mechanism for subsequent fusion. Furthermore, a Hybrid Attention Module (HAM) is introduced to combine large-kernel local context and Manhattan self-attention for complementary spatial modeling. Experimental results demonstrate that our method achieves highly competitive performance. Specifically, our model achieves an AbsRel of 0.094 and an RMSE of 4.156 on KITTI, and an AbsRel of 0.140 on DDAD. In the zero-shot cross-dataset generalization test on NYUv2, it attains an AbsRel of 0.232, outperforming the baseline RA-Depth by 7.2%.

cs.CV

Mixing time under monotone censoring

We prove that the lazy random walk on the discrete cube, censored to any increasing set of fixed positive density, mixes in time $O(n\log n)$, answering a question of Ding and Mossel. More precisely, for every nonempty increasing set $A\subseteq \{0,1\}^n$, \begin{equation} t_{\mathrm{mix}}(P) \le Kμ(A)^{-3}n\log(en). \label{eq:mixing-time-bound} \end{equation} where $μ$ is uniform on the cube and $K$ is an absolute constant. The proof uses hypercontractivity on the ambient cube to strengthen Poincaré inequality on coordinate sections. A stopping-time occupation inequality for increasing sets converts the resulting local bound into a uniform bound on hitting times of large sets. See the appendix for a better estimates of the constant and the dependence on \(μ(A)\).

math.PR

High-Moment Stability and Error Analysis of a Fully Discrete LDG-IMEX Method for High Dimensional Nonlinear Stochastic Convection-Diffusion Equations

A fully discrete local discontinuous Galerkin (LDG) method coupled with an implicit-explicit (IMEX) Euler time discretization is presented and analyzed for a class of high dimensional nonlinear stochastic convection-diffusion equations driven by multiplicative $\mathcal Q$-Wiener noise. The model allows nonlinear leading coefficients, nonlinear convection terms, dissipative source terms, and gradient-dependent noise. The diffusion operator is treated implicitly through the LDG formulation, while the nonlinear convection, lower-order drift, and stochastic terms are evaluated explicitly. The main contribution is a high-moment stability and error analysis for the fully discrete scheme. A central difficulty is that the nonlinear terms lead to pathwise growth factors that cannot be controlled uniformly on the full sample space. To provide the stability and error estimate, we introduce recursively defined nested subsets adapted to the numerical solution. Under the stated stochastic parabolicity and refinement conditions, we prove that these subsets have probabilities converging to one. On these subsets, the numerical solution satisfies high-moment stability, and the fully discrete error converges with order arbitrarily close to $r+1$ in space and $1/2$ in time. We also derive a pathwise error estimate by combining the high-moment error bound with a discrete Kolmogorov argument. Numerical experiments for stochastic Burgers' and Allen-Cahn equations confirm the theoretical rates and demonstrate the robustness of the proposed method for nonlinear stochastic models.

math.NA

Chaos of Berry curvature for BPS microstates

We expect black hole microstates to differ in their chaotic properties from states associated with other geometries. For supersymmetric black holes, ordinary level statistics cannot diagnose this distinction, since their energy levels are exactly degenerate. We propose that there is an intrinsic probe of chaos, encoded in the mixing of the microstates under changes in the couplings of the theory, as determined by the non-Abelian Berry curvature of the BPS states under certain deformations. For states dual to horizonless geometries in holographic systems, such as 1/2-BPS states in the D1/D5 CFT and 1/4-BPS states in $\mathcal{N}=4$ SYM, we find that the Berry curvature for marginal deformations is non-random and often exactly zero at generic couplings. By contrast, for states dual to supersymmetric black holes, we show through computations in $\mathcal{N}=2$ super-JT gravity and explicit numerics in the $\mathcal{N}=2$ SYK model that the Berry curvature resembles a random matrix. We also uncover interesting topological features of the $\mathcal{N}=2$ SYK moduli space, as probed by Chern numbers. These results suggest that the Berry curvature sharply distinguishes black hole microstates from smooth horizonless states and provides a robust diagnostic of chaos in supersymmetric sectors.

hep-th

PSMP-CLIP: Patch-Prompt SAM and Multi-Semantic Prompting for CLIP-Based Zero-Shot Anomaly Detection

Zero-shot anomaly detection aims to localize anomalies without target-domain samples. Existing CLIP-based methods suffer from coarse anomaly maps and limited semantic prompts. We propose PSMP-CLIP, integrating patch-prompt SAM2 segmentation (PPSS) and multi-semantic guided prompt regularization (MSGPR). PPSS samples prompts directly from intermediate patch features, avoiding threshold drift and guiding SAM2 to produce precise masks. MSGPR uses multiple learnable prompts constrained by semantic anchors to preserve generalization. Experiments on 14 datasets show highly competitive performance, achieving the best pixel-level AUROC on MVTec AD, BTAD, DTD-Synthetic, CVC-ClinicDB, TN3K, Endo, and Kvasir.

cs.CV

Understanding the Limits of Agentic ICD Coding

ICD-10-CM codes are alphanumeric codes used in the US to classify diagnoses and injuries for medical billing and epidemiological reporting. Standard ICD-10-CM benchmarks report aggregate metrics that obscure performance on complex coding scenarios. We evaluate neural, workflow, and agentic systems on a rarity-stratified set of MIMIC-IV discharge summaries and identify two orthogonal failure modes. Neural classifiers exhibit a 0.43 micro-F1 gap between rare and common codes. Workflow systems handle rare codes well but score near zero on injury and external cause codes that require multi-step guideline following. A tool-augmented agentic configuration with structured access to official ICD-10-CM reference materials recovers up to 0.34 micro-F1 on this subset. No single system dominates across all conditions.

cs.CL

A Fully Discrete Local Discontinuous Galerkin Method for Quasilinear Stochastic Convection-Diffusion-Type Equations

In this paper, we develop and analyze a fully discrete local discontinuous Galerkin (LDG) method with IMEX-Euler time discretization for a class of multi-dimensional quasilinear stochastic convection-diffusion-type equations driven by multiplicative $\mathcal Q$-Wiener noise. The leading diffusion matrix may depend on the solution as well as the spatial and temporal variables, while the lower-order drift and noise coefficients may depend on both the solution and its gradient. Under a suitable stochastic parabolicity condition, we establish unconditional high-moment stability estimates for the fully discrete scheme in the quasilinear setting. In the semilinear setting, where the leading diffusion matrix is independent of the solution but may vary in space and time, we further prove optimal high-moment strong error estimates of order $\mathcal O(h^{r+1})$ in space and $\mathcal O(k^{1/2})$ in time. A pathwise error estimate is then derived by combining the high-moment error bound with a discrete Kolmogorov argument. Numerical experiments are presented to illustrate the stability and convergence properties of the proposed method.

math.NA

An Exponential Lower Bound for the Permanent of Random Bernoulli Matrix

Let $M_n$ be an $n\times n$ matrix with independent uniform sign entries. We prove that there exist absolute constants $C,c>0$ such that, for all sufficiently large $n$, \[ \mathbb{P}\!\left( \left|\operatorname{Per}(M_n)\right| \ge e^{-Cn}\sqrt{n!} \right) \ge 1-n^{-c}. \] This confirms, up to the exponential scale, the lower bound suggested by Tao and Vu.

math.PR

A semicircle law for the normalized Laplacian of sparse random graphs

We study the limiting spectral distribution of the normalized Laplacian $\mathcal L$ of an Erdős-Rényi graph $G(n,p)$. To account for the presence of isolated vertices in the sparse regime, we define $\mathcal L$ using the Moore-Penrose pseudoinverse of the degree matrix. Under this convention, we show that the empirical spectral distribution of a suitably normalized $\mathcal L$ converges weakly in probability to the semicircle law whenever $np\to\infty$, thereby providing a rigorous justification of a prediction made in (Akara-pipattana and Evnin, 2023). Moreover, if $np>\log n+ω(1)$, so that $G(n,p)$ has no isolated vertices with high probability, the same conclusion holds for the standard definition of $\mathcal L$. We further strengthen this result to almost sure convergence when $np=Ω(\log n)$. Finally, we extend our approach to the Chung-Lu random graph model, where we establish a semicircle law for $\mathcal L$ itself, improving upon (Chung, Lu, and Vu 2003), which obtained the semicircle law only for a proxy matrix.

math.PR

Stringent Constraints on Spin-Spin-Velocity-Dependent Exotic Interactions with a Levitated Magnet Force Sensor

Exotic spin-spin-velocity-dependent interactions, predicted in extensions of the Standard Model involving new bosonic fields, could resolve fundamental puzzles from dark matter to cosmic asymmetry. However, exploring these weak potential interactions at centimeter scales presents formidable challenges, primarily due to the overwhelming dominance of electromagnetic backgrounds that can easily obscure the weak exotic signals. Here, we utilize a levitated magnet force sensor with ultrahigh electron spin density to probe these interactions. We constrain two interactions individually through a designed spin source and a multi-layer magnetic shielding system that suppresses electromagnetic backgrounds. In this study, we constrain two types of interactions: the V_6 potential at force ranges from $10^{-3}$ m to $6 \times 10^{-2}$ m and the V_{14} potential at ranges greater than $10^{-3}$ m. Our measurements establish 95% confidence-level bounds of $|f_6| \leq 2.12 \times 10^{-13}$ and $|f_{14}| \leq 2.34 \times 10^{-23}$ at $λ= 1.6 \times 10^{-2}$ m, improving prior limits by up to 12 and 13 orders of magnitude, respectively. Our result demonstrates the levitated magnet as a highly sensitive probe for detecting new bosonic fields in extensions of the Standard Model.

physics.app-ph

Visual General Intelligence: A White Paper

This paper reconsiders intelligence from a vision-centered perspective and examines whether intelligence emerging from visual experience and learning may provide a pathway toward AGI. In the language domain, beginning with the introduction of the Transformer architecture, the GPT series has demonstrated transfer to unseen tasks through autoregressive language modeling on web-scale text combined with aggressive scaling. This raises a natural question, namely, what capabilities and forms of intelligence can emerge from visual modalities such as images, videos, and geometry? In this paper, we discuss whether visual intelligence can serve as a pathway toward AGI, referred to in this paper as visual general intelligence (VGI), by bringing together contributors from diverse standpoints and affiliations. Our aim is not to offer a single definition of visual intelligence, but to clarify the principles that computer vision should pursue in the AGI era, the visual input modalities, the benchmarks, the learning paradigms, and the relationship between vision, when taken as the core, and other modalities such as language.

cs.CV

Low-Degree Fourier Threshold for Random Boolean Functions

We study whether a uniformly random Boolean function $f : \{-1,1\}^p \to \{-1,1\}$ is determined by its Walsh--Fourier coefficients of degree at most $d$. We show that the threshold lies at $p/2$ up to an $O(\sqrt{p \log p})$ window: if \[ d \le \frac{p}{2} - \sqrt{\frac{p}{2}\bigl(\log p + ω(1)\bigr)}, \] then with probability $1-o(1)$ there exists another Boolean function $g \ne f$ with the same degree-$\le d$ coefficients. Conversely, for every fixed $η\in (0,1)$, if \[ d \ge \frac{p}{2} + \sqrt{\frac{p}{2}\log\frac{6p}{η^2}}, \] then with probability at least $1-2^{-p}$, the function $f$ is uniquely determined by its degree-$\le d$ coefficients, even among all bounded functions $g : \{-1,1\}^p \to [-1,1]$. This resolves a question of Vershynin.

math.PR

Non-vanishing of Single, Double, and Triple Schubert Structure Constants

The Schubert vanishing problem asks whether the single Schubert coefficients $c_{u,v}^w$ are zero. In this paper, we consider the non-vanishing problems of double Schubert coefficients $c_{u,v}^w(t)$ and triple Schubert coefficients $c_{u,v}^w(t;y)$. We show that the non-vanishing of $c_{u,v}^w(t;y)$ is completely determined by the non-vanishing of single Schubert coefficients. As a byproduct, we obtain the saturation property of the triple Littlewood--Richardson coefficients $c_{λ,μ}^ν(t;y)$. Moreover, we pose a conjecture asserting that the non-vanishing of $c_{u,v}^w(t)$ is also determined by the non-vanishing of single or triple Schubert coefficients. We prove a one-side inclusion of the conjecture. For the reverse inclusion, we show that the conjecture holds for the following three cases: the Pieri case, the separated descents case, and the inverse Grassmannian case.

math.CO

Generics-Aware Fuzz Target Generation for Rust Libraries via Structured API Analysis

Fuzzing Rust library APIs requires constructing well-typed, compilable call sequences that satisfy ownership rules, generic parameters, and trait bounds; existing tools ignore these constraints or use shallow heuristics, yielding low coverage. We present GRAFT, which extracts structured API information from Rust documentation, builds an API dependency graph via recursive generics-aware type matching, and uses topology-guided traversal plus LLM synthesis with compiler-error feedback to produce compilable fuzz targets. On 13 crates from crates.io, GRAFT achieves 80.75% macro-average API coverage at 96.19% compilation success, outperforming RULF and RPG by 4.76x and 2.43x, and reaching 1.41x the average API coverage of deepSURF on crates with unsafe-reaching APIs.

cs.SE

Levitated Milligram-scale Ferromagnetic Magnetometer at Room Temperature

Levitated mechanical oscillators are emerging ultrasensitive sensors with tremendous potential in both applied and fundamental physics. Levitated ferromagnets, with internal spin noises rapidly averaged, promise ultrahigh magnetic sensitivity. Here, we demonstrate a milligram-scale diamagnetically levitated ferromagnet system operating at room temperature. Through optimized geometry and multi-channel dissipation control, we achieve a magnetic sensitivity of 23~fT$/\sqrt{\text{Hz}}$ at frequency of 100-Hz level. We anticipate that a ferromagnetic magnetometer with subfemtotesla sensitivity is within reach, after modest technical improvements. This platform establishes a high-performance magnetometer for biomagnetic field detection and beyond-standard-model force searches.

quant-ph

SEED: Simple ViT and Evolving Harness for Explainable Text Forgery Detection

AI-assisted image editing threatens trust in financial, legal, and identity records. The GenText-Forensics Challenge at ACM MM 2026 addresses this by requiring structured forensic reports, in which integrating detection, pixel-level localization, and natural language explanation for multilingual text-centric forgery images. We present SEED, a modular system with three components. First, a similarity-guided pipeline augments training with diverse synthetic forgeries. Second, a single ViT, built on DINOv3 with LoRA adaptation, jointly performs detection and pixel-level localization while preserving pre-trained priors with minimal trainable parameters. Third, an evolving harness takes the detector's predictions and generates a complete forensic report via an MLLM, iteratively improved through a proposer-evaluator loop optimizing report quality. SEED ranked 3rd in the GenText-Forensics Challenge. Code and data are available at https://github.com/KahimWong/GenText-Forensics-3rd-Place.

cs.CV

When Modalities Fail to Tango: Conformal Backdoor Detection in Multimodal Contrastive Learning

Backdoor attacks in multimodal contrastive learning (MCL) have garnered growing attention in recent years, as many downstream tasks critically depend on pre-trained MCL models. Existing detection-based defenses predominantly rely on the CLIPScore metric, under the assumption that poisoned pairs exhibit lower semantic similarity between the image and the caption. However, we identify two critical flaws remaining in existing methods: (1) the substantial overlap between CLIPScore distributions of benign and poisoned pairs undermines the reliability of this metric, and (2) fixed-threshold detection cannot provide statistical guarantees for ambiguous samples within overlapping regions. To overcome these limitations, we propose integrating conformal prediction (CP), a statistical framework that quantifies uncertainty through nonconformity scores (NCSs), to establish provable confidence bounds for detecting poisoned image-caption pairs. Building on CP, we introduce CASCADE, a novel two-stage Coarse-to-Fine Conformal Backdoor Detection framework. The coarse-grained stage uses cross-modality consistency to identify high-confidence benign and poisoned pairs. In the fine-grained stage, a reference set is constructed from high-confidence poisoned pairs, and instance-level NCSs based on text-space similarity are computed for each sample in the unidentified subset. These NCSs measure conformity to the poisoning distribution and enable precise identification of latent poisoned pairs within the unidentified subset. Extensive experiments on the large-scale CC3M dataset demonstrate that CASCADE achieves an average FPR of 5.79% at 100% TPR and an average AUROC of 0.9867 across diverse attacks, while remaining effective against adaptive attacks.

cs.CR

Test-time Scaling over Perception: Resolving the Grounding Paradox in Thinking with Images

Recent multimodal large language models (MLLMs) support Thinking with Images, invoking visual tools such as zooming and cropping to inspect image regions during inference. Yet these systems remain brittle in fine-grained reasoning: to acquire a decisive detail, a model must ground its attention on the correct region, but knowing which region is correct presupposes having already observed that detail. We identify this circular dependency as the grounding paradox, show that grounding errors are rarely self-corrected within a single trajectory---once a misleading region is inspected, all subsequent reasoning conditions on that observation and the error propagates to the final answer---and observe that because each trajectory constructs its own evidence, answer-level aggregation discards the very information that distinguishes trajectories. We propose Test-Time Scaling over Perception (TTSP), a closed-loop framework that treats perception as the unit of scalable inference and allocates compute along two axes: Entropy-Gated Perceptual Exploration samples diverse trajectories and uses critical-token entropy to withhold evidence the model cannot commit to, while Evidence-Guided Iterative Refinement distills validated observations into a correctable Evidence Ledger that steers later rounds to re-inspect unresolved regions. Across high-resolution and general multimodal benchmarks, TTSP consistently outperforms strong test-time scaling baselines, while improving grounding quality with favorable token efficiency.

cs.CV