Search arXiv⌕ Search

arXiv subjects

Jie Xiao

Publications and source records attributed to Jie Xiao.

At least 19 recordsLinked to original sources

WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation

Video generation begins in text space by authoring a cinematic screenplay, then materializes into pixels. As contemporary video generators scale to 30 seconds and faithfully follow complex conditions, the textual prompt largely directs the production, planning how actions, camera trajectories, lighting, and sound unfold across multi-shot sequences. In this paper, we present WanPE, a 397B-parameter prompt enhancement model trained on 1.05M real-world videos to master director-level cinematic planning. WanPE formulates shot-level cinematic plans via video-grounded reverse construction and employs Semantic-Consistency GRPO (SC-GRPO) to faithfully preserve user requirements across shots and over time. To benchmark this capability, we curate WanPEval, a human-annotated testbed covering durations from 5 to 30 seconds across varying intent granularities, supported by approximately 11K blind pairwise assessments. When powering Wan3.0's video generator, WanPE-397B boosts human preference over raw user prompts by 10.66-18.84 points at 5-15 seconds and by a dramatic 50.86 points in the 30-second arena. Ablation studies show that reverse construction demonstrates clear superiority over forward rewriting, while SC-GRPO robustly preserves semantic fidelity across model scales. Ultimately, WanPE leads all evaluated commercial offerings at 5-15 seconds and remains competitive with Seedance 2.5 at 30 seconds.

cs.CV↗

Code Plans, Diffusion Renders: Open-Ended Generative World Modeling

We introduce \textbf{CoDeR}, a new paradigm for world modeling. Unlike existing video world models that implicitly represent world dynamics through visual observations, our system explicitly constructs an executable world with code and employs video generation models for visual realization. Specifically, we coordinate five complementary roles to translate high-level concepts into structured world rules, executable dynamics, and perceptual observations. This design enables \textit{long-term memory}, \textit{open-ended interactions}, \textit{autonomous world evolution}, and \textit{multi-agent scenarios}, where multiple entities can act, interact, and evolve persistently beyond the current observation. Extensive experiments demonstrate that our framework substantially extends the capabilities of existing world models, enabling long-term memory, open-ended interactions, autonomous evolution, and persistent multi-agent dynamics, while achieving state-of-the-art performance across multiple evaluation settings. Code and model weights will be made publicly available. Project Page: \href{https://becauseimbatman0.github.io/CoDeR}{CoDeR}.

cs.CV↗

Back on Track: Aligning Rewards and States for Reasoning in Diffusion Large Language Models

Reinforcement learning (RL) holds immense promise for enhancing the reasoning capabilities of diffusion large language models (dLLMs). However, progress is fundamentally constrained by a dual misalignment between authentic generation trajectory and the gradient update process: (i) Process-reward misalignment. Sparse, terminal rewards are indiscriminately assigned to all intermediate steps of the generation process, failing to provide discriminative credit assignment. (ii) State-trajectory misalignment. Policy updates are often diverted toward artificial, out-of-trajectory states, squandering gradients on less informative samples. To address these limitations, we introduce Process Aligned Policy Optimization (PAPO), a novel framework that holistically aligns the RL update with the dLLM's generative trajectory via Step-Aware Process Rewards (SPR) that transform sparse terminal rewards into dense, step-wise credit, and Entropy-Guided Historical Re-enactment (EHR) that replays authentic trajectories at high-uncertainty steps. Extensive experiments on four benchmarks demonstrate that PAPO significantly outperforms baselines, achieving gains of 4.5% on GSM8K, 4.8% on MATH500, 42.2% on Countdown and 16.1% on Sudoku.

cs.CL↗

Capacitary-Distance Hardy Inequality

Let $n\ge3$, $Ω\subset\mathbb R^n$ be an open set, $F:=\mathbb R^n\setminusΩ$, and $α\in(0,\infty)$. For any $x\inΩ$, we define the capacitary distance \begin{align*} d_α(x) := \inf\left\{ r>0: \operatorname{cap}(\overline{F\cap B(x,r)}) \ge α\operatorname{cap}(B(\mathbf0,r)) \right\}. \end{align*} In this article, we prove that there exists a positive constant $C_n$, depending only on $n$, such that, for any $α\in(0,1]$ and any $u\in C_{\rm{c}}^\infty(Ω)$, \begin{align*} \int_Ω\frac{|u(x)|^2}{d_α(x)^2}\,d x \le \frac{C_n}{α^{2}} \int_Ω|\nabla u(x)|^2\,d x. \end{align*} This gives an affirmative answer to Problem 8 of Maz'ya [25]. Moreover, this dependence on $α$ is sharp: there exists a positive constant $c_n$, depending only on $n$, such that, for every $α\in(0,1]$, we are able to construct a bounded connected domain $Ω_α$ on which the optimal constant in the above Hardy inequality is at least $\frac{c_n}{α^{2}}$. The proof combines a variable-time semigroup estimate for the killed Brownian motion with finite-time exit estimates derived from capacity.

math.CA↗

Fuzzy-MoE: Interpretable Regime-Conditioned Expert Routing for Non-Stationary Multivariate Time Series Forecasting

In non-stationary multivariate time series, different variables and samples often exhibit heterogeneous latent dynamic states, while existing deep forecasting models usually compress them into a unified end-to-end mapping, leading to suboptimal modeling of time-varying dynamics and limited interpretability regarding which forecasting mechanism is activated under different latent states. To overcome these limitations, we reformulate time series forecasting as a unified framework of latent temporal state identification and interpretable expert routing, and propose Fuzzy-MoE, a fuzzy logic-based dynamic Mixture-of-Experts model. Fuzzy-MoE consists of multiple parallel expert mapping networks and a dual-view fuzzy router. By jointly exploiting local convolutional dynamics and global segmented statistics, the router infers latent temporal states and computes expert activation strengths through learnable Gaussian membership functions, enabling explicit IF-THEN rule-based expert selection. This fine-grained routing strategy allows different variables within the same sequence to activate different experts, effectively capturing heterogeneous temporal dynamics while improving model interpretability. Experimental results on multiple public time series benchmark datasets show that Fuzzy-MoE significantly outperforms mainstream forecasting methods in forecasting accuracy. Moreover, fuzzy memberships and rule activations provide interpretable routing diagnostics, demonstrating the effectiveness of the proposed framework in both forecasting performance and mechanism transparency. Unlike traditional MoE models that use black-box routing, Fuzzy-MoE`s routing is based on clear, interpretable fuzzy rules. This makes the expert selection transparent and traceable.

cs.LG↗

Sharp $p$-Capacity Estimates via Quermassintegrals in Hyperbolic Space

This paper establishes sharp upper bounds for $p$-capacities $\mathrm{Cap}_{1 2m+1$, an interpolating radius combines the $2m$-th curvature-excess radius with the $L^\infty$ curvature scale, thereby linking the finite-moment and supremum regimes. Equality in the sharp comparisons characterizes geodesic balls.

math.DG↗

A Unified Quermassintegral Approach to Quasilinear Heat Dispersion and Loss

This paper establishes a fundamental connection between quasilinear potential theory and convex geometric analysis by investigating the interplay between the quasilinear Laplace operator and quermassintegrals. We introduce a quasilinear heat dispersion law for convex conductors and prove that, among all convex conductors of a fixed mean width, the closed ball is a unique maximizer of this dispersion. By characterizing the quasilinear heat loss of a convex conductor explicitly in terms of its quermassintegrals, we demonstrate not only a formal equivalence between the isocapacitary and isoperimetric inequalities in the setting of mathematical physics but also that, among all convex conductors of a fixed mean width, the closed ball is a unique maximizer of this loss. These results provide a novel bridge between the metric properties of convex conductors and the variational analysis of quasilinear elliptic operators, offering a unified perspective on sharp geometric inequalities and their extremal cases.

math.AP↗

MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale

LLM agents execute heterogeneous sequences of model calls within a single task: some invocations require careful reasoning, while others are structured steps such as formatting or tool-argument construction. Prior routing methods exploit this asymmetry by assigning easy invocations to a cheaper small model and difficult ones to a large model. Such policies reduce inference cost, but they leave the small model's capability unchanged, so attainable savings remain bounded by the work the student can already solve. MERA instead improves the small model itself, using a single model invocation as the unit of adaptation. In each cycle, MERA replays failed student invocations to obtain execution-verified teacher demonstrations, distills recurring procedures into an iteratively updated SkillBook, and fine-tunes a student LoRA adapter via supervised learning and optional GRPO. Routing serves as supporting machinery for deployment: the improved student is served behind a cost-calibrated router with verifier-backed fallback, and a candidate SkillBook, adapter, or router is admitted only when joint replay preserves task quality. Empirically, four-cycle adaptation raises Qwen2.5-Coder-1.5B from 28.7% to 49.7% pass on held-out HumanEval+MBPP. Under verifier-backed fallback, the deployed policy retains 88.3% pass at 60.8% of always-Luna cost. On TAU-2, a fine-tuned Qwen3.5-2B improves from 14/35 to 18/35 and matches an unadapted 4B model. These results indicate that verifier-backed multi-cycle adaptation can increase small-model capability, rather than only routing around a fixed student.

cs.LG↗

A Counterexample to the Liu--Lou--Zhu $\mathcal Q_p$--Carleson Embedding Conjecture

In this paper, we disprove a conjecture of Liu, Lou, and Zhu concerning Carleson embeddings of $\mathcal Q_p$ spaces into tent spaces for $0<p<1$. More precisely, we construct a finite positive $p$-Carleson measure $μ$ on the unit disc $\mathbb D$ such that the canonical embedding $$ \operatorname{id}:\mathcal Q_p \longrightarrow \mathcal T_{p,2}^2(μ) $$ is not bounded. The main ingredient is a new family of $\mathcal Q_p$ test functions that encodes Cantor-type structures on the unit circle $\mathbb T$ into the analytic behavior of functions in $\mathcal Q_p$. This construction is inspired by ideas developed in a recent work of the first and third authors on composition operators on $\mathcal Q_p$ spaces.

math.CV↗

On the holomorphic differential operator $\frac{d}{dz}: Q_K\to L^q(WdA)$

In this paper, we obtain non-testing characterizations, in terms of dyadic capacity gauges, of the boundedness and compactness of the differentiation operator $$ \frac{d}{dz}:Q_K\longrightarrow L^q(W\,dA), \qquad 0<q<\infty. $$ We also characterize the limiting case as $q\to0^+$, formulated in terms of a logarithmic geometric mean, while the endpoint $q=\infty$ is treated separately using a standard testing argument. These results greatly extend the previous work on ${\mathcal Q}_p$-spaces to the general setting of $Q_K$-spaces. As applications, we characterize composition operators and Volterra-type integral operators between different $Q_K$-spaces. In particular, the off-diagonal characterization established here, together with the previously established diagonal case, completely resolves Zhao's 2009 open question on composition operators between ${\mathcal Q}_p$-spaces.

math.CV↗

On the growth of Bloch functions

We prove that there exist two Bloch functions $f_1$ and $f_2$ on $\mathbb D$ such that $$ |f_1(z)|+|f_2(z)| \geq \left(\log\frac{1}{1-|z|}\right)^{1/2}, \qquad z\in\mathbb D, $$ thereby resolving an open problem posed in 2008 by Girela, Peláez, Pérez-González and Rättyä. Our proof is based on a new Szegő-type recursion involving $\mathbb C^2$-valued polynomials and their reciprocal polynomials.

math.CV↗

Staleness-Learning Rate Scaling Laws for Asynchronous RLHF

High-throughput RLHF systems often decouple rollout generation from policy optimization, leading to the use of stale rollouts during learner updates. In this work, we study the effect of such staleness in asynchronous GRPO. We make the behavior policy explicit in the GRPO surrogate objective and distinguish between the surrogate-gradient mapping used by the learner and the true total derivative of a distribution-dependent population objective. Under assumptions of local boundedness, distributional smoothness, and behavior-policy smoothness, we show that stale rollouts introduce a per-step surrogate-gradient bias of order O(S * eta), where S denotes the maximum rollout lag and eta denotes the learning rate. We further derive a conditional collapse-time scaling law: when within-cycle drift remains below a batch-level clipping radius, collapse is governed primarily by cumulative learner drift T * eta; when the stale-rollout constraint is active, stability instead depends explicitly on S * eta. This yields a two-constraint stability condition eta << min{R_batch / (S * G_upd), R_crit / (T * G_upd)}, explaining why the maximum stable learning rate may appear weakly dependent on staleness in the horizon-limited regime.

cs.LG↗

Proximity-Induced Skyrmion Stabilization at the Cu2OSeO3/Bi2Se3 Interface

We investigate proximity-induced magnetic interactions at the interface between the topological insulator Bi2Se3 and the chiral magnetic insulator Cu2OSeO3, with particular focus on the low temperature skyrmion phase. Broadband ferromagnetic resonance spectroscopy reveals enhanced stability of noncollinear spin textures in the Cu2OSeO3/Bi2Se3 heterostructure compared with bare Cu2OSeO3. In addition to an extra resonance mode in the tilted conical phase that is absent in bare Cu2OSeO3, field cycling resolves two counterclockwise skyrmion resonance branches separated by approximately 238 MHz, consistent with the coexistence of a bulk skyrmion lattice and an interfacial skyrmion phase stabilized by proximity-induced exchange coupling and enhanced interfacial Dzyaloshinskii-Moriya interactions. The finite frequency separation indicates that the two skyrmion phases occupy distinct magnetic energy landscapes while retaining similar resonance character. Resonant elastic x-ray scattering measurements further confirm that the interfacial skyrmion phase spans a broader magnetic-field range than the bulk phase, demonstrating enhanced stability and ordering of topological spin textures at the interface. These findings establish interface engineering as a promising route for extending the stability regime of skyrmion and tilted-conical phases in topological-magnetic heterostructures.

cond-mat.mtrl-sci↗

Lusztig sheaves and integrable highest weight modules in the symmetrizable case

This paper continues the work of \cite{fang2023lusztigsheavesintegrablehighest} and \cite{fang2023lusztigsheavestensorproducts}. For a symmetrizable generalized Cartan matrix $C$ and the corresponding quantum group $\mathbf{U}$, we consider an associated quiver $Q$ equipped with an admissible automorphism $a$. We construct a category $\widetilde{\mathcal{Q}/\mathcal{N}}$ obtained from localizations of Lusztig sheaves for the corresponding framed and $2$-framed quivers with automorphism. The Grothendieck groups of these categories realize the integrable highest weight module $L(λ)$ and the tensor product $L(λ_1)\otimes L(λ_2)$ of integrable highest weight $\mathbf{U}$-modules. After quotienting by traceless objects, Lusztig sheaves yield the signed canonical bases of $L(λ)$ and $L(λ_1)\otimes L(λ_2)$. As applications, we recover symmetrizable crystal structures on Nakajima quiver varieties, Nakajima tensor product varieties, and Lusztig nilpotent varieties of preprojective algebras.

math.RT↗

Geometric realization of affine bases: the Kronecker quiver case

In this paper, we study the transition matrix between the PBW basis and the canonical basis for the negative part of the quantized enveloping algebra of the Kronecker quiver from a geometric viewpoint. Building on Lusztig's geometric construction of the canonical basis, we construct sheaf-complex realizations of PBW basis elements by means of flag sheaf complexes over the strata $X(α,m)$ of representation varieties. Our first goal is to give a geometric description of the simple constituents appearing in the restrictions of these flag sheaf complexes to the strata $X(α,m)$. This allows us to compare the PBW-type sheaf complexes with the simple perverse sheaves $IC(X(α),L_χ)$ arising in Lusztig's construction. Using this description together with a purity result for the relevant $\mathbb{F}_q$-structures, we obtain another proof that the elements defined by Lusztig's perverse sheaves indeed form a basis of the composition algebra.Our second goal is to make the transition coefficients between the PBW basis and the canonical basis geometrically explicit. More precisely, we show that these coefficients are governed by the multiplicities of local systems in the restrictions of intersection cohomology complexes to smaller strata. As a consequence, the transition matrix from the canonical basis to the PBW basis is upper triangular with diagonal entries equal to $1$, and its coefficients admit a direct geometric interpretation. In particular, in the Kronecker quiver case we recover the triangularity of the transition matrix and obtain positivity properties of the corresponding coefficient polynomials.

math.QA↗

The iterated geometric Green's formula

Fang, Lan, and Xiao established the geometric Green's formula as a categorical isomorphism for arbitrary semisimple complexes. In this short note, we generalize their work to multi-step compositions. Specifically, we establish the iterated geometric Green's formulas for the composition of an $(n-1)$-fold restriction and an induction, as well as its dual.

math.RT↗