Search arXiv⌕ Search

arXiv subjects

Tianle Liu

Publications and source records attributed to Tianle Liu.

At least 19 recordsLinked to original sources

RoLA: Rotary-Positioned Low-Rank Linear Attention for Efficient Diffusion Transformers

Diffusion Transformers (DiTs) achieve strong video generation quality, but their dense spatiotemporal self-attention scales quadratically with sequence length and quickly becomes the dominant inference bottleneck. Sparse low-rank hybrids alleviate this cost by combining a local sparse branch with a global compressed branch. In video DiTs equipped with 3D Rotary Position Embeddings (RoPE), the global branch faces a structural compatibility issue: when RoPE is applied before a nonlinear feature map, the rotation and nonlinearity generally do not commute, making it difficult to keep a query-independent linear summary while preserving relative rotary geometry. Existing work often sidesteps this issue by replacing genuine cross-token global aggregation with coordinate-conditioned surrogates or learnable absolute positional modules. These compromises can be effective, but they approximate relative decay from absolute coordinates and introduce extra positional parameters. We propose \textbf{RoLA}, a rotary-positioned low-rank linear-attention branch that keeps genuine cross-token aggregation while remaining compatible with a reusable linear summary. The design applies RoPE \emph{outside} the nonlinear low-rank feature map and reuses a truncated subset of the pre-trained rotary schedule matched to the low-rank bottleneck. This yields a linear-time low-rank global branch with relative positional behavior by design and no additional positional parameters; the full sparse--low-rank module still includes the fixed-sparsity sparse branch. Experiments on open-source video DiTs show that the resulting method remains competitive in generation quality at 90\% sparsity while achieving 2.63$\times$ end-to-end inference speedup on Wan2.1-14B (720p, 81 frames, measured on an NVIDIA H100 GPU).

cs.CV↗

SparkDiffusion: Mitigating the High-Sparsity Trap --- A Unified Framework for up to $265\times$ Single-GPU Acceleration of Visual Generation

Video diffusion transformers are expensive because attention dominates long spatiotemporal token sequences. We identify the \emph{high-sparsity trap}: at extreme attention sparsity, step-local training losses keep decreasing while terminal generation quality stagnates or degrades. The trap is one of supervision: the dominant terminal errors originate in the high-noise structure-generation stage, and terminal-aligned training corrects terminal errors that substantially extended step-local training cannot. This yields a simple staging principle: \emph{first adapt the sparse architecture into a coarse prior, then correct the terminal distribution}. We instantiate the principle as \method, a unified acceleration framework for visual generation that combines a short sparse warm-up, few-step trajectory-mixed distillation, and FP8 quantization with fused kernels. \method sustains $97\%$ attention sparsity with strong visual quality on long-sequence 720P generation across Wan2.1/Wan2.2 backbones and T2V/I2V tasks, and $90\%$ sparsity on Wan2.1-T2V-1.3B-480P. With 3-step CFG-free inference, \method achieves a $265\times$ end-to-end speedup over the 50-step CFG dense baseline for Wan2.1-T2V-14B-720P on a single RTX~5090 ($220\times$ on H100), and denoises a Wan2.1-T2V-1.3B-480P video in $1.3$s.

cs.CV↗

A Heavily Right Strategy for Statistical Inference with Dependent Studies in Arbitrary Dimensions

We leverage recent advances in heavy-tail approximations for global hypothesis testing with dependent studies to construct approximate confidence regions without modeling or estimating their dependence structures. A non-rejection region is a confidence region but it may not be convex. Convexity is appealing because it ensures any one-dimensional linear projection of the region is a confidence interval, easy to compute and interpret. We show why convexity fails for nearly all heavy-tail combination tests proposed in recent years, including the influential Cauchy combination test. These insights motivate a heavily right strategy: truncating the left half of the Cauchy distribution to obtain the Half-Cauchy combination test. The harmonic mean test also corresponds to a heavily right distribution with a Cauchy-like tail, namely a Pareto distribution with unit power. We prove that both approaches guarantee convexity when individual studies are summarized by Hotelling $T^2$ or $χ^{2}$ statistics (regardless of the validity of this summary) and provide efficient, exact algorithms for implementation. Applying these methods, we develop a divide-and-combine strategy for mean estimation in any dimension and construct simultaneous confidence intervals in a network meta-analysis for treatment effect comparisons across multiple clinical trials. We also present many open problems and conclude with epistemic reflections.

stat.ME↗

Wasserstein-p Bounds via Cumulant-Based Edgeworth Expansion for $α$-Mixing Random Fields

We establish normal-approximation bounds in the Wasserstein-$p$ distance for $α$-mixing random fields, extending results previously available mainly for independent or locally dependent variables. For standardized sums indexed by $T\subset\mathbb{Z}^d$, we obtain rates $O(|T|^{-β})$, where $β\in(0,1/2]$ depends on $p$, the moment order, the dimension, and the mixing decay; the optimal square-root rate holds under sufficiently fast polynomial mixing. For integer $p\geq2$, a refined bound retains the stronger mixing power produced by the expansion, covers the endpoint of finite $(p+2)$-th moments, and identifies exact logarithmic corrections at critical decay. Under finite-range dependence, the endpoint moment condition gives the optimal rate for every real $p\geq1$. The proof combines a cumulant-based Edgeworth expansion with Stein's method. Its main technical tool is a constructive graph calculus that organizes high-order dependence and preserves cancellations before mixing bounds are applied. The resulting Wasserstein estimates also imply two-regime non-uniform normal-approximation bounds with sharper sample-size decay in the far tail. A Lean 4/Mathlib companion machine-checks the principal theorem interfaces and the complete paper-internal genogram argument for the expansions, relative to seven explicitly identified literature and standard inputs.

math.PR↗

Intrinsic Restriction Traces and Toric Dynamics

We prove Morita invariance of the Campbell--Lind--Malkiewich--Ponto--Zakharevich restriction-system trace after passage to perfect modules. It therefore defines an intrinsic integral restriction trace for an exact endofunctor of a small idempotent-complete stable $\infty$-category. On $π_0$, the $m$-th ghost is the laced trace of the $m$-fold iterate, compatibly with Frobenius. For lattice-graded algebras, the ghost targets have twisted cocenters in degree zero; over the open parameter torus, this applies to the cyclic bimodule of Dinkins--Karpov--Krylov. For finite monomial endomorphisms of toric varieties, we construct motivic restriction classes whose ghosts are sums over cones fixed by the iterates. In one example, two classes have the same first ghost, while their second ghosts differ after rational Betti realization.

math.AG↗

Wasserstein-p Bounds in the Central Limit Theorem Under Weak Dependence

The central limit theorem is one of the most fundamental results in probability and has been successfully extended to locally dependent data and strongly-mixing random fields. In this paper, we establish its rate of convergence for transport distances, namely for arbitrary $p\ge1$ we obtain an upper bound for the Wasserstein-$p$ distance for locally dependent random variables and strongly mixing stationary random fields. Our proofs adapt the Stein dependency neighborhood method to the Wasserstein-$p$ distance and as a by-product we establish high-order local expansions of the Stein equation for dependent random variables. Finally, we demonstrate how our results can be used to obtain tail bounds that are asymptotically tight, and decrease polynomially fast, for the empirical average of weakly dependent random variables.

math.PR↗

Stein Kernels and Normal Approximation for Log-Concave Bilinear Forms

Jiang, Lee, and Vempala conjectured that if $X,Y\in\mathbb{R}^n$ are independent isotropic log-concave random vectors, then $W_2(L(\langle X,Y\rangle),N(0,n))$ is bounded by a universal constant. Subject to Theorems 1.2 and 2.5 of arXiv:2607.24164v1, we prove this conjecture and a rectangular bilinear-form extension. For independent isotropic log-concave $X\in\mathbb{R}^m$, $Y\in\mathbb{R}^n$, and nonzero $B\in\mathbb{R}^{m\times n}$, put \[ r_4(B)=\frac{(\operatorname{Tr}(B^\top B))^2} {\operatorname{Tr}((B^\top B)^2)}. \] We construct a nonnegative scalar Stein kernel for $X^\top B Y/\|B\|_F$ whose squared $L^2$ discrepancy is at most $20/r_4(B)$, and consequently obtain the same bound for squared $2$-Wasserstein distance to $N(0,1)$. The proof develops an exact covariance identity and deficit decomposition for trace observables of moment-map Stein kernels, together with a stability theorem for positive Stein kernels under log-concave approximation. Taking $B=I_n$ yields \[ W_2^2\left(L\left(\frac{\langle X,Y\rangle}{\sqrt{n}}\right),N(0,1)\right)\leq\frac{20}{n}, \] which is the Jiang--Lee--Vempala conjecture.

math.PR↗

Kac's Walk on Rotation Matrices Mixes in $\boldsymbol{Θ(n^2)}$ Steps: A Proof Discovered with AI

Let $N=\binom n2=\dim\mathrm{SO}(n)$. We prove that the coordinate-plane Kac walk on $\mathrm{SO}(n)$ has total-variation mixing time of order $N$: for every fixed $0<\varepsilon<1$, \[ t_{\mathrm{mix}}^{(n)}(\varepsilon)=Θ_\varepsilon(n^2). \] The lower bound is the dimensional singularity obstruction before $N$ steps. The upper bound removes the final logarithm from the previously known $O(n^2\log n)$ estimate. The proof combines the discrete Malliavin coupling and low-degree pseudo-mixing inputs with a new log-free analysis of the derivative shells. Its static core is a circuit-anchored, arbitrary-spectrum root/pass identity for the physical five-box prime. Keeping one normalization base per original circuit permits simultaneous scalar regluing without paying for artificial cuts. Its temporal core is an exact chronological calculus: passive singleton runs acquire a coboundary/Riesz gain, while root-interrupted components are allocated by vertex-labelled packets before absolute values are taken. The curvature split into pure-Weyl and Ricci parts is kept at its physical tensor type. All-Weyl packets retain a full $N^{-1}$ resource; mixed packets contain a typed $O(n^{-1/2})$ Ricci debit; and the final packetless Ricci cell is closed by a joint invariant-column estimate on its two root-hit circuits and an exact causal restoration of the marked root time. These estimates yield an $O(n)$ squared first-derivative shell and a summable all-order marked-shell expansion through logarithmic degree. The resulting score energy is $O(n/c^2)$ after $cN$ steps. A weighted submersion integration-by-parts argument and the Haar log-Sobolev inequality then give the uniform total-variation upper bound. No cutoff profile or cutoff window is asserted.

math.PR↗

Universal Beta Incidence Angles: Cauchy Rigidity and Infinite Arrangements

Let $U$ be Haar-uniform on $\mathbb S^{p-1}$, let $a_1,\ldots,a_k$ be arbitrary nonzero vectors, and let $w_1,\ldots,w_k$ be simplex weights. Define \[ g(U)=\sum_{j=1}^k w_j\frac{a_j}{a_j^\top U}, \qquad N(U)=\frac{g(U)}{\|g(U)\|}. \] We prove the universal incidence law \[ \{U^\top N(U)\}^2\sim\operatorname{Beta}\!\left(\frac12,\frac{p-1}{2}\right), \] independently of the number, arrangement, rank, or overcompleteness of the directions and of the weights. Thus a deterministic, generally non-Haar function of $U$ has the same squared-cosine law as an independent Haar direction. One proof combines a Herglotz--Cauchy boundary principle, a Haar-random two-plane with one common phase, and an exact Beta--Cauchy tangent-projection equivalence. A second proof specializes the positive-semidefinite Pillai--Meng identity. The planar structure leads to converses: plane-conditional Cauchy laws recover positivity, while for signed measures an exact phase-cancellation deficit equals twice the hidden negative mass. This yields local-to-global rigidity under a phase-norming condition strictly weaker than injectivity and an unconditional exclusion of negative atoms. The law extends to probability measures under almost-sure reciprocal integrability. We characterize this condition by an exact Wiener--Dini belt series, prove finite Shannon entropy to be the sharp universal criterion for countable weights, and give an entropy--geometry extension for clustered measures. Every compact carrier of zero one-dimensional Hausdorff measure is admissible, whereas a nonzero rectifiable arc component forces divergence on a set of positive Haar measure. In orthogonal coordinates, the theorem also gives a weight-free scaled $F$ law for Pearson divergence from a fixed simplex vector to a $\operatorname{Dirichlet}(1/2,\ldots,1/2)$ vector.

math.PR↗

Copula Geometry and Second-Order Calibration of Heavily Right Aggregation

Heavy-tailed $p$-value combination tests are attractive under unknown dependence because their null tails can be first-order robust even when the exact dependent null distribution is unavailable. That robustness does not resolve calibration: $\Pr\{T>q(α)\}=α+o(α)$ neither quantifies the remaining size error nor determines its sign. We develop a second-order calibration theory for positive Half-Cauchy and reciprocal, or harmonic-mean, aggregation. After exact marginal standardization, extremal dependence is represented by an index-one exponent measure on the coordinate-face lattice. Its support separates axial, full-interior, proper-face, and mixed geometries. A Möbius decomposition shows how the noncompact weighted half-space probes every face, while quantitative face limits determine whether the correction is integrable, critically amplified, or nonintegrable. With common-factor inversion and a tail-to-calibration map, this yields machinery applicable beyond individual parametric copula families. The results include an integrable hidden-face transfer theorem, a sharp fixed-dimensional local-corner theorem with explicit higher-face control, a common-heavy-factor theorem, and size and critical-value expansions relative to independence. Gaussian, standard multivariate-$t$, positive Clayton, and max-linear pair-shock models exhibit distinct power, logarithmic, radial--angular, and proper-face mechanisms; the Gaussian weighted-half-space transfer is conditional on explicit face-boundary hypotheses. Thus dependence geometry determines the rate, coefficient, and direction of the calibration error left unresolved by first-order validity. Asymptotics use $t\to\infty$, equivalently vanishing significance levels.

math.ST↗

Two-flag degenerations and real circles tangent to three conics

Three general plane conics admit $184$ complex tangent circles, and it had been conjectured that at most $136$ of them could be real. We construct an explicit strongly general triple of smooth conics over $\mathbb{Q}$ with exactly $160$ real tangent circles; the same count therefore occurs on a nonempty Euclidean chamber. The construction combines a fourfold splitting theorem for two flagged double-line degenerations with exact Sturm--Tarski, elimination, and interval certificates. We also show that the Grothendieck--Witt-valued count is $92\mathbb{H}$ and express its real local signs, up to a fixed orientation convention, in terms of curvature differences, residual intersection divisors, and contact normals. The exact-arithmetic code and certificate data are archived in the accompanying repository.

math.AG↗

GS-Playground: A High-Throughput Photorealistic Simulator for Vision-Informed Robot Learning

Embodied AI research is undergoing a shift toward vision-centric perceptual paradigms. While massively parallel simulators have catalyzed breakthroughs in proprioception-based locomotion, their potential remains largely untapped for vision-informed tasks due to the prohibitive computational overhead of large-scale photorealistic rendering. Furthermore, the creation of simulation-ready 3D assets heavily relies on labor-intensive manual modeling, while the significant sim-to-real physical gap hinders the transfer of contact-rich manipulation policies. To address these bottlenecks, we propose GS-Playground, a multi-modal simulation framework designed to accelerate end-to-end perceptual learning. We develop a novel high-performance parallel physics engine, specifically designed to integrate with a batch 3D Gaussian Splatting (3DGS) rendering pipeline to ensure high-fidelity synchronization. Our system achieves a breakthrough throughput of 10^4 FPS at 640x480 resolution, significantly lowering the barrier for large-scale visual RL. Additionally, we introduce an automated Real2Sim workflow that reconstructs photorealistic, physically consistent, and memory-efficient environments, streamlining the generation of complex simulation-ready scenes. Extensive experiments on locomotion, navigation, and manipulation demonstrate that GS-Playground effectively bridges the perceptual and physical gaps across diverse embodied tasks. Project homepage: https://gsplayground.github.io.

cs.RO↗

A Forward-Inverse Dynamic Game Framework for Enhanced Multi-Agent Trajectory Planning

This paper studies feedback Nash equilibrium (FBNE) seeking for multi-agent trajectory planning in nonlinear dynamical systems with unknown agents' objectives and state-dependent inter-agent coupling. While dynamic game theory provides a principled framework for such problems, existing approaches typically assume fully rational agents with known objectives or rely on fixed regularization, limiting their ability to capture bounded rationality and spatially varying interaction intensity in safety-critical settings. To this end, we propose a KL-regularized dynamic game with a state-dependent weight that adaptively balances optimality and behavioral priors. To infer unknown cost parameters from demonstrated behaviors, we develop a context-aware inverse game module based on maximum-entropy inverse reinforcement learning with physics-informed regularization, ensuring structural consistency with the forward game. We establish per-iteration well-posedness of the regularized local game and show that the adaptive weighting function remains Lipschitz continuous under bounded nominal-trajectory updates. Numerical simulations and multi-robot experiments on cooperative navigation and merging scenarios validate the effectiveness of the proposed framework.

cs.RO↗

A Fan Algorithm for Chow-Witt Rings of Smooth Projective Toric Varieties

Let $R$ be a real closed field and let $X_Σ$ be a smooth projective split toric variety. For an explicitly chosen basis rigidification $\mathfrak r$ of its Picard grading, we prove that the pair $(Σ,\mathfrak r)$ determines the resulting total Chow--Witt ring and give a terminating finite algorithm for all groups, products, twists, and forgetful maps. The ring is identified with an explicit fibre product of the Picard-graded diagonal $\mathbf{I}$-cohomology ring and the integral inverse images of twisted Bockstein kernels over the mod-$2$ Chow ring. Using real cycle classes, we identify the first factor with the cohomology of the real toric variety with all sign local systems. We construct the mod-$2$ cycle map on invariant divisors as explicit deck-transition cocycles and obtain finite signed boundary and product matrices directly from the fan.

math.AG↗

LADY: Linear Attention for Autonomous Driving Efficiency without Transformers

End-to-end autonomous driving has emerged as a promising paradigm. However, state-of-the-art methods rely heavily on Transformer architectures. The inherent quadratic complexity of Transformers restricts their ability to model long-range spatial and temporal dependencies, particularly on resource-constrained edge platforms. Given the inherent demand for efficient temporal modeling in autonomous driving, this computational bottleneck severely constrains real-time deployment. While linear attention mechanisms offer a computationally efficient alternative, existing architectures are predominantly limited to self-attention, lacking the cross-modal capabilities essential for autonomous driving. In this work, we propose LADY, the first fully linear attention-based generative model for end-to-end autonomous driving. LADY incorporates a novel, lightweight linear cross-attention (LICA) mechanism to enable effective cross-modal interaction while preserving linearity. A key advantage of our framework is its ability to fuse long-range temporal contexts during inference with constant computational and memory costs ($O(1)$), regardless of the historical sequence length. Experiments on the NAVSIM and Bench2Drive benchmarks demonstrate that LADY achieves performance comparable to state-of-the-art methods, delivering competitive planning accuracy with significantly reduced latency. Furthermore, efficiency benchmarking on edge devices validates the model's feasibility for resource-limited scenarios.

cs.AI↗

Torus-enriched Motivic Bruhat Complexes and Maximal Compact Groups

Bruhat decompositions give cellular models for split algebraic groups, flag varieties, and maximal compact groups, but motivic boundaries retain orientation and torus-translation data lost in the flag quotient. Over a perfect field of characteristic zero, let the group be connected, split, semisimple, and simply connected. Fixing a Borel subgroup with split maximal torus and unipotent radical, we construct a torus-enriched motivic cellular complex for the basic affine space and compute its boundary in every degree. Each cover in Bruhat order contributes a two-face operator determined by a transported coroot, a tail determinant weight, and an explicit Milnor--Witt frame degree. Bott--Samelson purity proves the formula, while the unipotent torsor identifies the complex with that of the group. Over the real numbers, realization identifies it at chain level with the extended-Weyl complex of a maximal compact subgroup, while torus augmentation gives the flag complex. A single motivic complex therefore interpolates between the two incidence theories. A finite torus-support filtration makes this explicit; after inversion of two it splits by the characters of the component group of the real split torus, and the support spectral sequence degenerates. Calculations in the rank-three special linear and exceptional rank-two cases exhibit the first higher differentials beyond the previously known range.

math.AG↗

Cellular $\mathbb{A}^1$-Homology from Bruhat Boundary Matrices of Split Semisimple Flag Varieties

Let $k$ be a perfect field of characteristic different from 2, we compute the cellular $\mathbb{A}^1$-homology of the flag varieties $G/P_Θ$ attached to split semisimple simply connected groups over $k$ and describe the differentials in the cellular $\mathbb{A}^1$-chain complex concretely. The construction applies uniformly to the type $A$ coefficient formula, to the type $B_n,C_n,D_n$ for $n\leq 7$, and to the exceptional types for $F_4,E_6,E_7$. Under real realization over $k=\mathbb{R}$, this computation recovers the corresponding results of real flag manifolds. We also provide a detailed computation for $SL_3/B$ and an application to the full split flag variety of type $F_4$.

math.AG↗

Autonomous End-to-End SOH Prediction Services for Battery Systems via Temporal-Contrastive Representation Learning

Accurate state of health (SOH) estimation is a critical diagnostic service for lithium-ion battery management. However, reliance on labor-intensive manual feature engineering and opaque black-box models hinders scalable industrial deployment. To address this, we introduce TC-SOH: a modular, plug-and-play service architecture for autonomous, end-to-end SOH prediction. TC-SOH employs a temporal-contrastive mechanism and a cross-window prediction pretext task to extract degradation-relevant representations directly from raw operational data. To improve transparency, we connect model efficacy with representation diagnostics: visualization, sensitivity analysis, redundancy analysis, bidirectional probing, future-SOH probing, and temporal shuffling show that learned features overlap with selected expert descriptors while retaining additional SOH-relevant variation, and that ordered temporal context improves subsequent-SOH prediction. Across four public datasets, TC-SOH outperforms the considered physics-informed and data-driven baselines, reducing MAPE by 1.91 times and RMSE by 2.13 times.

cs.LG↗