Search arXivSearch

arXiv subjects

Fan Cheng

Publications and source records attributed to Fan Cheng.

At least 19 recordsLinked to original sources

Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models

Modern interactive video world models have achieved impressive visual fidelity, yet lack fine-grained multi-entity control and cross-entity, cross-world generalization. We trace this gap to the action interface: standard control protocols (e.g. animation IDs, device inputs, scene-level captions) bind action semantics to specific entities or engines at design time. We propose natural language as the interface to unlock expressiveness that no prior interface can achieve, and we present Incantation, the first interactive video world model with per-latent-frame (0.25 s) natural-language conditioning that supports simultaneous multi-entity control and concept-level cross-entity transfer beyond any fixed rendering pipeline. We pair a pretrained bidirectional video backbone with frame-local text cross-attention, and enable real-time long-horizon streaming through ODE-initialized Self-Forcing distillation with a RoPE-decoupled sliding KV-cache. We surpass the Action-Index baseline on cross-entity transfer (89% vs. 43%) and out-of-vocabulary prompts (90% vs. 0%), and our 2-step student sustains 19.7 FPS at 480p with stable FVD over 2-hour rollouts. We further apply the same architecture and training recipe to The King of Fighters, changing only the per-entity action vocabulary slots. We have released a preview subset of the Incantation dataset at https://huggingface.co/datasets/zhush/incantation-elden-ring-scenes, containing manually collected Elden Ring player-boss combat clips with structured action-oriented metadata. Larger-scale Elden Ring and KOF data will be released with the full project.

cs.CV

TIE: Time Interval Encoding for Video Generation over Events

Director-style prompting, robotic action prediction, and interactive video agents demand temporal grounding over concurrent events -- a regime in which 68% of general clips and over 99% of robotics/gameplay clips contain overlapping events, yet existing multi-event generators rest on a single-active-prompt assumption. However, modern video generators, such as Diffusion Transformers (DiT), represent time as discrete points through point-wise positional encodings. This formulation creates a fundamental dimension mismatch: temporally extended intervals and overlapping events are mathematically unrepresentable to the attention mechanism. In this paper, we propose Time Interval Encoding (TIE), a principled, plug-and-play interval-aware generalization of rotary embeddings that elevates time intervals to first-class primitives inside DiT cross-attention. Rather than introducing another heuristic interval embedding, we show that, within RoPE-compatible bilinear attention, TIE is characterized by two basic principles: Temporal Integrability, which requires an event to aggregate positional evidence over its full duration, and Duration Invariance, which removes the trivial bias toward longer intervals. Under a uniform kernel, this characterization yields an efficient closed-form sinc-based solution that preserves the standard attention interface and naturally attenuates boundary noise through interval integration. Empirically, TIE preserves the visual quality of the base DiT model while substantially improving temporal controllability. In our experiments on the OmniEvents dataset, it improves human-verified Temporal Constraint Satisfaction Rate from 77.34% to 96.03% and reduces temporal boundary error from 0.261s to 0.073s, while also improving trajectory-level temporal alignment metrics. The code and dataset are available at https://github.com/MatrixTeam-AI/TIE.

cs.CV

Energy-variational solutions for geodynamical two-phase flows -- From logarithmic to double-obstacle potentials by variational convergence

In [Cheng, Lasarzik, Thomas 2025 ARXIV-Preprint 2509.25508], we studied a Cahn--Hilliard two-phase model describing the flow of two viscoelastoplastic fluids in the framework of dissipative solutions using a logarithmic potential for the phase-field variable. This choice of potential has the effect that the fluid mixture cannot fully separate into two pure phases. The notion of dissipative solutions is based on a relative energy-dissipation inequality featuring a suitable regularity weight. In this way, this is a very weak solution concept. In the present work, we study the well-posedness of the geodynamical two-phase flow in the notion of energy-variational solutions. They feature an additional scalar energy variable that majorizes the system energy along solutions and they are further characterized by a variational inequality that combines an energy-dissipation estimate with the weak formulation of the system adding an error term that accounts for the mismatch between the energy variable and the system energy multiplied by a suitable regularity weight. We give a comparison of these two concepts. We further study different phase-field potentials for the geodynamical two-phase flow model. In particular, we address the variational limit from a potential with a logarithmic contribution to a double-obstacle potential, then also allowing for the emergence of pure phases. This study underlines that, thanks to its structure, the energy-variational solution is better suited for variational convergence methods than the dissipative solution.

math.AP

Coding for Fading Channels with Imperfect CSI at the Transmitter and Quantized Feedback

The classical Schalkwijk-Kailath (SK) scheme for the additive Gaussian noise channel with noiseless feedback is highly efficient since its coding complexity is extremely low and the decoding error doubly exponentially decays as the coding blocklength tends to infinity. However, how to extend the SK scheme to channel models with memory has yet to be solved. In this paper, we first investigate how to design SK-type scheme for the 2-path quasi-static fading channel with noiseless feedback. By viewing the signal of the second path as a relay and adopting an amplify-and-forward (AF) relay strategy, we show that the interference path signal can help to enhance the transmission rate. Besides this, for arbitrary multi-path fading channel with feedback, we also present an SK-type scheme for such a model, which transforms the time domain channel into a frequency domain MIMO channel.

cs.IT

Addressing the ID-Matching Challenge in Long Video Captioning

Generating captions for long and complex videos is both critical and challenging, with significant implications for the growing fields of text-to-video generation and multi-modal understanding. One key challenge in long video captioning is accurately recognizing the same individuals who appear in different frames, which we refer to as the ID-Matching problem. Few prior works have focused on this important issue. Those that have, usually suffer from limited generalization and depend on point-wise matching, which limits their overall effectiveness. In this paper, unlike previous approaches, we build upon LVLMs to leverage their powerful priors. We aim to unlock the inherent ID-Matching capabilities within LVLMs themselves to enhance the ID-Matching performance of captions. Specifically, we first introduce a new benchmark for assessing the ID-Matching capabilities of video captions. Using this benchmark, we investigate LVLMs containing GPT-4o, revealing key insights that the performance of ID-Matching can be improved through two methods: 1) enhancing the usage of image information and 2) increasing the quantity of information of individual descriptions. Based on these insights, we propose a novel video captioning method called Recognizing Identities for Captioning Effectively (RICE). Extensive experiments including assessments of caption quality and ID-Matching performance, demonstrate the superiority of our approach. Notably, when implemented on GPT-4o, our RICE improves the precision of ID-Matching from 50% to 90% and improves the recall of ID-Matching from 15% to 80% compared to baseline. RICE makes it possible to continuously track different individuals in the captions of long videos.

cs.CV

Analysis of a Cahn--Hilliard model for viscoelastoplastic two-phase flows

We study a Cahn--Hilliard two-phase model describing the flow of two viscoelastoplastic fluids, which arises in geodynamics. A phase-field variable indicates the proportional distribution of the two fluids in the mixture. The motion of the incompressible mixture is described in terms of the volume-averaged velocity. Besides a volume-averaged Stokes-like viscous contribution, the Cauchy stress tensor in the momentum balance contains an additional volume-averaged internal stress tensor to model the elastoplastic behavior. This internal stress has its own evolution law featuring the nonlinear Zaremba-Jaumann time-derivative and the subdifferential of a non-smooth plastic potential. The well-posedness of this system is studied in two cases: Based on a regularization by stress-diffusion we obtain the existence of Leray-Hopf-type weak solutions. In order to deduce existence results also in the absence of the regularization, we introduce the concept of dissipative solutions, which is based on an estimate for the relative energy. We discuss general properties of dissipative solutions and show their existence for the viscoelastoplastic two-phase model in the setting of stress-diffusion. By a limit passage in the relative energy inequality for vanishing stress-diffusion, we conclude an existence result for the non-regularized model.

math.AP

Towards Interpretable Visual Decoding with Attention to Brain Representations

Recent work has demonstrated that complex visual stimuli can be decoded from human brain activity using deep generative models, offering new ways to probe how the brain represents real-world scenes. However, many existing approaches first map brain signals into intermediate image or text feature spaces before guiding the generative process, which obscures the contributions of different brain areas to the final reconstruction output. In this work, we propose NeuroAdapter, a visual decoding framework that directly conditions a latent diffusion model on brain representations, bypassing the need for intermediate feature spaces. Our method demonstrates competitive visual reconstruction quality on public fMRI datasets compared to prior work, while providing greater transparency into how brain signals drive visual reconstruction. To this end, we introduce an Image-Brain BI-directional interpretability framework (IBBI) that analyzes cross-attention patterns across diffusion denoising steps to reveal how different cortical areas influence the unfolding generative trajectory. Our work highlights the potential of end-to-end brain-to-image reconstruction and establishes a path for interpretable neural decoding.

cs.CV

MAMBO-G: Magnitude-Aware Mitigation for Boosted Guidance

High-fidelity text-to-image and text-to-video generation typically relies on Classifier-Free Guidance (CFG), but achieving optimal results often demands computationally expensive sampling schedules. In this work, we propose MAMBO-G, a training-free acceleration framework that significantly reduces computational cost by dynamically optimizing guidance magnitudes. We observe that standard CFG schedules are inefficient, applying disproportionately large updates in early steps that hinder convergence speed. MAMBO-G mitigates this by modulating the guidance scale based on the update-to-prediction magnitude ratio, effectively stabilizing the trajectory and enabling rapid convergence. This efficiency is particularly vital for resource-intensive tasks like video generation. Our method serves as a universal plug-and-play accelerator, achieving up to 3x speedup on Stable Diffusion v3.5 (SD3.5) and 4x on Lumina. Most notably, MAMBO-G accelerates the 14B-parameter Wan2.1 video model by 2x while preserving visual fidelity, offering a practical solution for efficient large-scale video synthesis. Our implementation follows a mainstream open-source diffusion framework and is plug-and-play with existing pipelines.

cs.CV

Photonic Origami of Silica on a Silicon Chip with Microresonators and Concave Mirrors

3D printing of high-quality silica photonic structures is particularly challenging, as surface roughness at the nanoscale can severely degrade optical performance through scattering losses. Here, we develop a technique to fold ultrasmooth silica on silicon chips into such desired 3D structures. A laser-induced, surface-tension-driven method achieves folding with 20 nm alignment accuracy, enabling origami-like polylines and helices with integrated 0.5 nm-smooth photonic devices. The technique allows for the fabrication of record length-to-thickness ratio structures, incorporating concave micromirrors with numerical aperture of 0.41, and microresonators with quality factors exceeding ${8 \times 10^6}$. This on-chip silica origami approach offers a pathway to transform planar electro-opto-mechanical circuits into high-quality 3D configurations.

physics.optics

Coding for Quasi-Static Fading Channel with Imperfect CSI at the Transmitter and Quantized Feedback

The classical Schalkwijk-Kailath (SK) scheme for the additive Gaussian noise channel with noiseless feedback is highly efficient since its coding complexity is extremely low and the decoding error doubly exponentially decays as the coding blocklength tends to infinity. However, its application to the fading channel with imperfect CSI at the transmitter (I-CSIT) is challenging since the SK scheme is sensitive to the CSI. In this paper, we investigate how to design SK-type scheme for the quasi-static fading channel with I-CSIT and quantized feedback. By introducing modulo lattice function and an auxiliary signal into the SK-type encoder-decoder of the transceiver, we show that the decoding error caused by the I-CSIT can be perfectly eliminated, resulting in the success of designing SK-type scheme for such a case. The study of this paper provides a way to design efficient coding scheme for fading channels in the presence of imperfect CSI and quantized feedback.

cs.IT

Optimal Feedback Schemes for Dirty Paper Channels With State Estimation at the Receiver

In the literature, it has been shown that feedback does not increase the optimal rate-distortion region of the dirty paper channel with state estimation at the receiver (SE-R). On the other hand, it is well-known that feedback helps to construct low-complexity coding schemes in Gaussian channels, such as the elegant Schalkwijk-Kailath (SK) feedback scheme. This motivates us to explore capacity-achieving SK-type schemes in dirty paper channels with SE-R and feedback. In this paper, we first propose a capacity-achieving feedback scheme for the dirty paper channel with SE-R (DPC-SE-R), which combines the superposition coding and the classical SK-type scheme. Then, we extend this scheme to the dirty paper multiple-access channel with SE-R and feedback, and also show the extended scheme is capacity-achieving. Finally, we discuss how to extend our scheme to a noisy state observation case of the DPC-SE-R. However, the capacity-achieving SK-type scheme for such a case remains unknown.

cs.IT

Instability in Diffusion ODEs: An Explanation for Inaccurate Image Reconstruction

Diffusion reconstruction plays a critical role in various applications such as image editing, restoration, and style transfer. In theory, the reconstruction should be simple - it just inverts and regenerates images by numerically solving the Probability Flow-Ordinary Differential Equation (PF-ODE). Yet in practice, noticeable reconstruction errors have been observed, which cannot be well explained by numerical errors. In this work, we identify a deeper intrinsic property in the PF-ODE generation process, the instability, that can further amplify the reconstruction errors. The root of this instability lies in the sparsity inherent in the generation distribution, which means that the probability is concentrated on scattered and small regions while the vast majority remains almost empty. To demonstrate the existence of instability and its amplification on reconstruction error, we conduct experiments on both toy numerical examples and popular open-sourced diffusion models. Furthermore, based on the characteristics of image data, we theoretically prove that the instability's probability converges to one as the data dimensionality increases. Our findings highlight the inherent challenges in diffusion-based reconstruction and can offer insights for future improvements.

cs.LG

Accelerating Diffusion Sampling via Exploiting Local Transition Coherence

Text-based diffusion models have made significant breakthroughs in generating high-quality images and videos from textual descriptions. However, the lengthy sampling time of the denoising process remains a significant bottleneck in practical applications. Previous methods either ignore the statistical relationships between adjacent steps or rely on attention or feature similarity between them, which often only works with specific network structures. To address this issue, we discover a new statistical relationship in the transition operator between adjacent steps, focusing on the relationship of the outputs from the network. This relationship does not impose any requirements on the network structure. Based on this observation, we propose a novel training-free acceleration method called LTC-Accel, which uses the identified relationship to estimate the current transition operator based on adjacent steps. Due to no specific assumptions regarding the network structure, LTC-Accel is applicable to almost all diffusion-based methods and orthogonal to almost all existing acceleration techniques, making it easy to combine with them. Experimental results demonstrate that LTC-Accel significantly speeds up sampling in text-to-image and text-to-video synthesis while maintaining competitive sample quality. Specifically, LTC-Accel achieves a speedup of 1.67-fold in Stable Diffusion v2 and a speedup of 1.55-fold in video generation models. When combined with distillation models, LTC-Accel achieves a remarkable 10-fold speedup in video generation, allowing real-time generation of more than 16FPS.

cs.CV

Assessing the electronic excitation spectra of chromium, palladium and samarium from their stopping quantities

The electronic excitation spectrum of a material characterises the response to external electromagnetic perturbations through its energy loss function (ELF), which is obtained from several experimental sources that usually do not completely agree among them. In this work, we assess the available ELF of three metals, namely chromium, palladium, and samarium, by using the dielectric formalism to calculate relevant stopping quantities, such as the stopping cross sections for protons and alpha particles, as well as the corresponding electron inelastic mean free paths. The comparison of these quantities (as calculated from different sets of ELF) with the available experimental data for each of the analyzed metals highlights the promising capability of the recently proposed reverse Monte Carlo method for the determination of the ELF. This work also analyzes the contribution of different electronic shells to the electronic excitation spectra of these materials, and reveals the important role that the excitation of "semi-core" bands plays on the energy loss mechanism for these metals.

cond-mat.mtrl-sci

BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs

Advancements in large Vision-Language Models have brought precise, accurate image captioning, vital for advancing multi-modal image understanding and processing. Yet these captions often carry lengthy, intertwined contexts that are difficult to parse and frequently overlook essential cues, posing a great barrier for models like GroundingDINO and SDXL, which lack the strong text encoding and syntax analysis needed to fully leverage dense captions. To address this, we propose BACON, a prompting method that breaks down VLM-generated captions into disentangled, structured elements such as objects, relationships, styles, and themes. This approach not only minimizes confusion from handling complex contexts but also allows for efficient transfer into a JSON dictionary, enabling models without linguistic processing capabilities to easily access key information. We annotated 100,000 image-caption pairs using BACON with GPT-4V and trained an LLaVA captioner on this dataset, enabling it to produce BACON-style captions without relying on costly GPT-4V. Evaluations of overall quality, precision, and recall-as well as user studies-demonstrate that the resulting caption model consistently outperforms other SOTA VLM models in generating high-quality captions. Besides, we show that BACON-style captions exhibit better clarity when applied to various models, enabling them to accomplish previously unattainable tasks or surpass existing SOTA solutions without training. For example, BACON-style captions help GroundingDINO achieve 1.51x higher recall scores on open-vocabulary object detection tasks compared to leading methods.

cs.CV

Radiation Pressure Induced Oscillations of an Optically Levitating Mirror

Optical Fabry-Perot cavity with a movable mirror is a paradigmatic optomechanical systems. While usually the mirror is supported by a mechanical spring, it has been shown that it is possible to keep one of the mirrors in a stable equilibrium purely by optical levitation without any mechanical support. In this work we expand previous studies of nonlinear dynamics of such a system by demonstrating a possibility for mechanical parametric instability and emergence of the ``phonon laser'' phenomenon.

physics.optics

Cavity Continuum

We experimentally demonstrate and numerically analyze large arrays of whispering gallery resonators. Using fluorescent mapping, we measure the spatial distribution of the cavity-ensemble's resonances, revealing that light reaches distant resonators in various ways, including while passing through dark gaps, resonator groups, or resonator lines. Energy spatially decays exponentially in the cavities. Our practically infinite periodic array of resonators, with a quality factor [Q] exceeding 10^7, might impact a new type of photonic ensembles for nonlinear optics and lasers using our cavity continuum that is distributed, while having high-Q resonators as unit cells.

physics.optics

Lipschitz Singularities in Diffusion Models

Diffusion models, which employ stochastic differential equations to sample images through integrals, have emerged as a dominant class of generative models. However, the rationality of the diffusion process itself receives limited attention, leaving the question of whether the problem is well-posed and well-conditioned. In this paper, we explore a perplexing tendency of diffusion models: they often display the infinite Lipschitz property of the network with respect to time variable near the zero point. We provide theoretical proofs to illustrate the presence of infinite Lipschitz constants and empirical results to confirm it. The Lipschitz singularities pose a threat to the stability and accuracy during both the training and inference processes of diffusion models. Therefore, the mitigation of Lipschitz singularities holds great potential for enhancing the performance of diffusion models. To address this challenge, we propose a novel approach, dubbed E-TSDM, which alleviates the Lipschitz singularities of the diffusion model near the zero point of timesteps. Remarkably, our technique yields a substantial improvement in performance. Moreover, as a byproduct of our method, we achieve a dramatic reduction in the Fr\'echet Inception Distance of acceleration methods relying on network Lipschitz, including DDIM and DPM-Solver, by over 33%. Extensive experiments on diverse datasets validate our theory and method. Our work may advance the understanding of the general diffusion process, and also provide insights for the design of diffusion models.

cs.CV