Search arXivSearch

arXiv subjects

Kai Jiang

Publications and source records attributed to Kai Jiang.

At least 19 recordsLinked to original sources

High-Performance Scaled P-Type SnOx Transistor by Atomic Layer Deposition with CFET Integration

The development of high-performance p-type oxide semiconductors is essential for realizing complementary logic for monolithic 3D integration, yet p-type oxide semiconductors still exhibit substantially inferior performance compared with their n-type counterparts. In this work, we demonstrate high-performance p-type SnOx transistors by atomic layer deposition (ALD), as back-end-of-line compatible devices for monolithic 3D integration. The SnOx transistors exhibit high field-effect mobility of 6.9 cm2/Vs, low subthreshold swing (SS) of 185 mV/dec, decent on/off ratio (ION/IOFF) of 1.8*104 and high bias stability. By scaling the channel length down to 80 nm, a high on-current of 38.7 mA/mm at VDS of -1 V is achieved. It is understood that precursor and reaction engineering to suppress Sn4+ component in SnOx film are the key for performance enhancement. Furthermore, a complementary field-effect transistor with ALD SnOx p-FET vertically stacking on ALD In2O3 n-FET is also demonstrated, achieving maximum voltage gain of 21 V/V at VDD of 4 V. These findings suggest ALD SnOx as a promising candidate for scaled high-performance BEOL p-type transistors.

cond-mat.mtrl-sci

Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation

We present Vidu S2, which comprises Vidu S2-Avatar, a real-time interactive digital-character model, and Vidu S2-Editing, a real-time video editing model. Moreover, we explore the feasibility of real-time spatial video generation for both Vidu S2-Avatar and Vidu S2-Editing. Compared with Vidu S1, Vidu S2-Avatar supports real-time 720p video generation, generation with dynamic references that can be updated at any moment, and stronger instruction following, such as dancing. Vidu S2-Editing supports editing a video stream in real time, including style rendering, clothing replacement, character replacement, and background replacement. Experiments show that Vidu S2 outperforms all baselines. A playable online demo is available at https://vidu.com/vidu-stream.

cs.CV

Regional Frequency Constrained Dispatch Method Considering Spatial-joint Stochastic Disturbances and Contingencies

The increasing penetration of renewable energy challenges frequency stability due to high variability and declining inertia. Traditional frequency security constrained dispatch methods fail to capture regional frequency heterogeneity and spatially correlated stochastic disturbances, resulting in inaccurate frequency security enforcement. To address this, a regional frequency constrained dispatch method is proposed, considering the spatially joint stochastic disturbances and contingencies. Firstly, a multi-regional frequency response model is constructed, incorporating the Vine-Copula based characterization of regional stochastic disturbances and regional frequency support. Then, a progressive latent-bottleneck physics-informed neural network is applied to characterize differential frequency nadir terms and regional frequency support integral terms via an encoder-decoder architecture, which enables optimization compatible expressions of system frequency dynamics. Finally, a day ahead dispatch model is developed in which regional stochastic frequency constraints are embedded using a CVaR-based formulation. A case study on the IEEE 118-bus system shows that the proposed method outperforms the unified center of inertia embedded approach in mitigating regional frequency violations, with the regional frequency nadir and RoCoF improved by 32.76% and 2.06%, respectively.

eess.SY

Superintegrability of stratified symplectic spaces

We define the notion of superintegrability of a Hamiltonian system on a stratified symplectic space. We focus on spin Calogero-Moser-Sutherland (sCMS) systems, where the phase space is a stratified symplectic space obtained by the Hamiltonian reduction of the cotangent bundle over a compact Lie group, and demonstrate that the sCMS systems for $SU(3)$ are superintegrable.

math-ph

A High-Accuracy Numerical Homogenization Framework for Quasiperiodic Hamilton--Jacobi Equations

In this work, we develop an accurate numerical homogenization framework for computing effective Hamiltonians of quasiperiodic Hamilton--Jacobi equations (QHJEs) with convex Hamiltonians of the form $H(x,p) = |p|^k/{k}-f(x), ~k>1$, where $f$ is quasiperiodic. Computing effective Hamiltonians in the quasiperiodic setting requires solving QHJEs posed on the whole space. Their solutions generally possess neither translational symmetry nor decay and may exhibit low regularity. These features pose substantial challenges for numerical computation. To address these difficulties, we introduce a quasiperiodic boundary condition, which allows the original whole-space problem to be treated on a bounded domain while preserving quasiperiodicity at the boundary. We then propose an SL--FPR scheme that combines a semi-Lagrangian approximation with the finite points recovery method and establish stability and error estimates for the resulting scheme. We also extend the quasiperiodic homogenization result from the quadratic case to general $k>1$ and apply the proposed method to accurately approximate the corresponding effective Hamiltonians. Numerical experiments illustrate the convergence and applicability of the method and validate the extended homogenization results.

math.NA

TurboT2VA: Fast Large-Scale Text-to-Video-Audio Generation via Score-Regularized Consistency Distillation

Joint text-to-video-audio generation produces synchronized visual and acoustic content, but the long sampling trajectories and heterogeneous multimodal computation of large models make inference prohibitively expensive. We present TurboT2VA, a distillation and inference framework for accelerating a 19B-parameter joint video-audio model. Large-scale T2VA distillation is challenged by modality-imbalanced optimization, the difficulty of continuous-time consistency training at scale, and the quality--diversity trade-off. TurboT2VA addresses these issues with per-modality normalization and a progressive curriculum comprising discrete consistency warm-up, continuous consistency refinement, and joint consistency--distribution matching. The curriculum first establishes a stable, diverse generation trajectory and only then introduces distribution-level refinement. On LTX-2, four-step distillation reduces generator latency from 50.52s to 2.51s at the standard evaluation resolution of 512$\times$768, achieving a 20.1$\times$ speedup while maintaining strong visual quality, audio fidelity, diversity, and video-audio synchronization. We further develop an architecture-aware inference stack that combines guarded W8A8 and fused operators, padded-text compaction, and modality-aware sparse attention while preserving dense cross-modal and text-conditioning paths. Under the high-resolution deployment setting at 1024$\times$1792, the complete stack reduces generator latency from 318.74s to 5.83s on one NVIDIA H20, achieving a 54.67$\times$ generator-only speedup. Inference code and generation demos are available at https://github.com/thu-ml/TurboDiffusion/tree/main/turbot2va.

cs.CV

Accurately computing quasiperiodic parabolic equations within finite-size domains via modeling quasiperiodic boundary conditions

Quasiperiodic systems exhibit long-range order without decay and are naturally posed on the whole space. However, in practical applications, computations are performed on finite domains, making the choice of boundary conditions that preserve the global quasiperiodic structure a key modeling challenge. In particular, conventional boundary conditions contain no information about the quasiperiodic field beyond the computational domain. Traditional periodic boundary conditions (PBCs) suffer from Diophantine errors due to the rational approximation of irrational numbers, limiting their accuracy. Motivated by this, we propose a class of quasiperiodic boundary conditions (QBCs) for quasiperiodic problems, which avoid the limitations caused by traditional Diophantine errors. By exploiting a homomorphism between a low-dimensional physical domain and a high-dimensional torus, QBCs effectively capture the long-range structure at the boundaries. To validate the proposed approach, we apply QBCs to solve quasiperiodic parabolic equations (QPEs) within finite-size domains and establish rigorous convergence results. Numerical experiments demonstrate that QBCs substantially reduce the influence of Diophantine errors. When employed to model finite-size QPEs and combined with suitable numerical discretizations, they enable accurate and efficient computations for both high- and low-regularity cases, while exhibiting improved convergence compared with PBCs.

math.NA

StageWAM: Joint-Embedding Stage Prediction for World-Action Models in Robot Manipulation

Generalist robot policies aim to map multimodal observations and linguistic task instructions to actions across diverse tasks. However, existing methods typically represent the future as a fixed, short video-action chunk. This short-term future captures local scene evolution for action execution, but it does not explicitly describe the stage-level future that specifies how a task should progress from its current stage to the next. We therefore distinguish two complementary futures for robot manipulation: a short-term physical future to capture local scene evolution and a stage-level semantic future to represent task progress. We introduce StageWAM, which augments a Motus-based World Action Model (WAM) with Stage-JEPA, a goal-conditioned Joint-Embedding Predictive Architecture (JEPA) predictor. Given the current observation and task instruction, Stage-JEPA uses a frozen V-JEPA2 encoder to extract the current-state representation and predicts the latent target of the next inferred stage. Across 50 RoboTwin 2.0 tasks in clean and randomized environments, StageWAM achieves 90.25% overall success and reduces the mean number of execution steps in successful rollouts by 5.97% relative to the strongest baseline.

cs.RO

Miles: Metric Learning with Expandable Subspace for Pre-Trained Model-Based Class-Incremental Learning

Class Incremental Learning (CIL) aims to learn new concepts consistently from a data stream without forgetting. Unlike typical CIL methods which need to learn a model from scratch, pre-trained model (PTM) can easily adapt to a new task with fine-tuning. However, existing PTM-based CIL methods fail to achieve a trade-off between performance and computational expenditure, i.e., they either adopt the same parameter space so that leading catastrophic forgetting, or expand a new branch for each task but adding more computational cost. To this end, we propose MetrIc Learning with Expandable Subspace (Miles) to harness the prior information within pre-trained knowledge, thereby orchestrating an efficient expansion of the parameter space through guided optimization. Specifically, it decouples the learnable modules with the pre-trained model, exploiting prior information from intermediate features of the backbone network to enable more flexible parameter expansion. Then, a central loss is adopted to guide the new category to cluster towards the corresponding prototype in the new task subspace while incorporating an auxiliary distance regularization term to maintain metric equilibrium across tasks. Extensive experiments on six benchmark datasets demonstrate that Miles achieves state-of-the-art performance in various CIL settings.

cs.CV

MobileSAM2: Lightweight Segment Anything for Spatial Intelligence

The recent large video foundation model, SAM2, enables segment anything in both images and videos, serving as a powerful base model for various applications. However, many of such use cases require to operate on resource-constrained devices like mobile phones and laptops. In this work, we aim to make SAM2 more mobile-friendly by distilling the heavyweight SAM2 into a lightweight model, facilitating segment anything in both images and videos on mobile devices. To this end, we propose Hypergraphical Knowledge Distill (HyperKD), which introduces the idea of hypergraph into knowledge distillation, aiming to effectively model and transfer SAM2's generalizable and comprehensive knowledge. HyperKD consists of Temporal HyperKD and Granularity HyperKD that construct hypergraphs to explicitly model and extract the generalizable temporal knowledge and the comprehensive multi-granularity knowledge from SAM2 respectively, which are then distilled into the lightweight student model by aligning it with the constructed hypergraphs. Besides, we present MobileSAM2, a new family of lightweight SAM2 that balances efficiency and effectiveness via searching the best model architectures with HyperKD during model size reduction. Extensive experiments validate MobileSAM2 across multiple benchmarks and show promising generalization performance on embodied AI tasks.

cs.CV

Vidu S1: A Real-Time Interactive Video Generation Model

We introduce Vidu S1, a real-time interactive video generation model supporting voice control of digital characters. Users can control video generation content at any moment through voice instructions. Vidu S1 supports infinite-length real-time video generation without blurring, drift, or visual distortion. Built with TurboDiffusion and TurboServe, Vidu S1 outputs 540p real-time videos at up to 42 FPS on regular consumer GPUs. Users can upload custom images of real people, anime, and pets, and choose different voice tones for personalized experiences. Experiments show that Vidu S1 achieves the best performance across all test metrics while fully meeting real-time inference requirements. A playable online demo is available at https://vidu.com/vidu-stream.

cs.CV

High-Mobility and High-Reliability Top-Gate Oxide Semiconductor Transistors by Oxygen Engineering

In this work, we investigate the role of oxygen (O) on the performance of top-gate (TG) atomic-layer-deposited (ALD) oxide semiconductor transistors. The results reveal distinct defect characteristics and positive bias temperature instability (PBTI) degradation mechanisms between oxygen-rich (O-rich) and oxygen-deficient (O-poor) devices. It is found that an O-rich device fabrication process followed by O-free annealing can effectively achieve TG indium-rich (In-rich) oxide semiconductor transistors with high mobility, high reliability and high stability in hydrogen environment because O-rich process can suppress oxygen vacancies and their interaction with hydrogen, while O-free annealing plays a critical role in minimizing the formation of O-rich defects such as oxygen dimers (O-O bonds). Consequently, TG In-rich transistors with high mobility, steep subthreshold slope, and high PBTI reliability at high temperature are demonstrated. The understanding of O-rich defects provides a new insight to overcome the mobility-stability trade-off.

cond-mat.mtrl-sci

Latent Noise Mask for Reducing Visual Redundancy in Multimodal Large Language Models

Multimodal large language models (MLLMs) often fail in fine-grained visual reasoning, as question-relevant visual cues are diluted by dense and redundant image tokens. Recent multimodal reasoning methods usually extend chain-of-thought from language models into visual or latent spaces, seeking to add intermediate reasoning states while overlooking the negative impact of redundant visual tokens. We propose LatEnt Noise maSk (Lens), a question-conditioned visual evidence purification framework that empowers MLLMs to reason with cleaner visual cues in latent space. Lens introduces a lightweight Lens Evidence Token (LET) to score which visual tokens support the current question and preserve them during decoding. Guided by the LET scores, it injects adaptive latent noise into low-relevance tokens, softly suppressing distractors without changing the model backbone or token sequence. With only one temporary learnable control token and a lightweight noise generator, Lens adds minimal overhead while improving the base MLLM by 2.4-6.4 points on most VQA datasets and by 4.1-6.4 points on grounding tasks. These results show that multimodal reasoning can benefit more directly from cleaner question-relevant visual evidence than from simply extending the reasoning trace.

cs.CV

TurboServe: Serving Streaming Video Generation Efficiently and Economically

Streaming video generation is emerging as a new serving workload in which users interact with long-lived sessions that generate video progressively, chunk by chunk. Unlike offline video generation or typical LLM serving, streaming video generation must preserve session state across active and idle periods, repeatedly schedule ongoing sessions, and deliver each chunk under a tight latency target. This creates two key serving challenges in multi-user, multi-GPU environments: session duration heterogeneity, where long-running sessions make placement decisions suboptimal over time, and temporal user-demand heterogeneity, where the number of active sessions fluctuates sharply across bursts and idle periods. We present TurboServe, the first serving system designed specifically for streaming video generation workloads. TurboServe formulates serving as an online scheduling problem that jointly coordinates session placement and GPU provisioning. Its closed-loop scheduling algorithm combines a migration-aware placement controller, which rebalances sessions across GPUs to reduce the maximum per-chunk latency, with a load-driven autoscaling controller, which adapts the GPU budget to workload variation for improved cost efficiency. To support these decisions at runtime, TurboServe implements coalesced chunk processing for batching concurrent active sessions on the same GPU, GPU-CPU offloading for session suspension and resumption, and NCCL-based GPU-GPU migration for online rebalancing. We evaluate TurboServe on real-world production traces from Shengshu Technology across multiple model sizes and GPU clusters with up to 64 NVIDIA B300 GPUs. Compared with baseline serving configurations, TurboServe reduces worst-case per-chunk latency by 37.5% and total GPU operating cost by 37.2% on average. Our code is publicly available at https://github.com/shengshu-ai/TurboServe.

cs.DC

Structure and energetics of grain boundaries in self-assembled double-gyroid block copolymer networks

Grain boundaries (GBs) are ubiquitous defects in crystalline materials. However, they remain less explored in block copolymer ordered phases. Here, we develop a self-consistent field theory framework to investigate GB structure and energetics in double-gyroid (DG) diblock copolymer networks. The GB energy landscape is obtained as a function of GB orientation, which reveals multiple local minima representing distinct network-switching GBs. Remarkably, the global minimum is a previously unidentified asymmetric-tilt network-switching GB (ATNS), exhibiting a lower energy than the experimentally observed $(422)$ twin boundary (TB). Comparative analyses of representative low- (ATNS, $(422)$ TB) and high-energy twist ($(0\bar{1}\bar{1})$, $(100)$ TNSs) GBs reveal that, unlike enthalpy-dominated hard matter, GB stability in DG networks is predominantly entropy-driven. Twist-type GBs generate new nodes and disrupt nodal coplanarity, causing chain packing frustration and large entropy penalties. Conversely, the ATNS preserves favorable network connectivity and minimizes conformational constraints on polymer chains, making it the energetically preferred GB.

cond-mat.soft

Bayesian Sparse Regression for Microbiome-Metabolite Data Integration

Numerous studies have shown that microbial metabolites, which represent the products of bacteria in the human gut, play a key role in shaping cancer risk and response to treatment. However, metabolite data typically contain a large proportion of missing values, which may result from either low abundance or technical challenges in data processing. Moreover, given the compositionality of microbiome data, where the observed abundances can only be interpreted on a relative scale, standard variable selection methods are not applicable. In this project, we propose a novel Bayesian regression method to address these challenges in the integration of metabolite and microbiome data. Key features of our proposed model include modeling the two different mechanisms of missingness for the metabolite data and adopting a Bayesian prior designed to address the compositional characteristics of microbiome data. We demonstrate on simulated data that our proposed model can accurately impute the unobserved true metabolite values and correctly select the relevant microbiome predictors. We further illustrate our method using real data from a study focused on understanding the interplay between the microbiome and metabolome in colorectal cancer.

stat.AP

Geometric construction of superintegrable Poisson projection chains via Poisson centralizers

We introduce a geometric framework for constructing superintegrable systems from Poisson centralizers (commutants) in the Lie-Poisson algebra $S(\mathfrak{g})$ of a complex semisimple Lie algebra. Starting from a chain of reductive subgroups, we study the corresponding invariant Poisson subalgebras and their Poisson centers, and formulate superintegrability in terms of a \emph{Poisson projection chain} of affine Poisson varieties. For a maximal torus $T\subset G$, we prove that the inclusions $S(\mathfrak{g})^G\subset S(\mathfrak{g})^T\subset S(\mathfrak{g})$ determine a superintegrable chain and identify the associated quotient maps $\mathfrak{g}\xrightarrow{\chi_T}\mathfrak{g}//T\xrightarrow{\rho}\mathfrak{g}//G$. The rank (transcendence degree) computations yield the expected dimension split between commuting Hamiltonians and first integrals, and we describe the corresponding symplectic leaves in the intermediate space. Several examples illustrate how the centralizer generators organize into explicit superintegrable Poisson chains.

math-ph

Defect annihilation mechanism in the formation of dodecagonal quasicrystals

Understanding defect evolution is essential to the structural stability of quasicrystals, yet the kinetics of defect repair remain poorly understood. Here, by combining the string method and the spring pair method, we determine the minimum energy path from defective to defect-free dodecagonal quasicrystals using a particle model with the Lennard-Jones-Gauss potential. We find that defect annihilation proceeds via three stages: phason flip, aggregation and decomposition of shield-like defects. These sequential transformations are driven by potential energy gradients and accompanied by an increase in structural symmetry. The three stages act synergistically in promoting defect annihilation, offering new insights into the microscopic repair mechanisms of quasicrystals.

cond-mat.mtrl-sci