Search arXiv⌕ Search

arXiv subjects

Zijian Liu

Publications and source records attributed to Zijian Liu.

At least 19 recordsLinked to original sources

Design of a combined polarimetric and velocimetric measurement for viscoelastic stress, and its constitutive resolving power

A polarimeter does not report stress. It reports a retardance and an azimuth, and converting those to a stress pair fixes both the calibration that is required and the covariance that the subsequent inference must carry. We set out that observation chain for planar viscoelastic flow, combine it with velocimetry, and ask what the combined measurement can resolve. The calibration constant is fixed by three separately measurable quantities (path length, stress-optic coefficient and wavelength) rather than being fitted. Propagating the polarimetric errors to first order gives a stress covariance that is anisotropic and site dependent even for independent homoscedastic inputs, and whose conditioning degrades as the retardance approaches zero, where the linearization itself stops describing the measurement. Wrapping imposes a separate design bound. Two numerical studies follow, both in ideal calibrated stress coordinates under a prescribed covariance rather than the propagated polarimetric one. Holding the total scalar count fixed and varying the split between velocimetric and optical sites, an unequal allocation favoring optical sites outperforms either pure configuration. We then ask what the measurement resolves between constitutive models compatible with the same velocity data. In self-consistent pressure-driven flow of the finitely extensible nonlinear elastic Peterlin model, at extensibility L^2=50 and 3% noise, the optical channel rejects the velocity-compatible Oldroyd-B family in 78.1% of realizations at the Deborah number De=2, the ratio of the relaxation time to the flow timescale, and in 100% at De=4, while rejecting only about 5% below De=0.3. The resolving power of a design is therefore a strong function of Deborah number and must be quoted with it.

physics.flu-dyn↗

A closed-form solution for streaming and Lagrangian transport in a deforming circular cavity

Streaming from a deforming cavity wall serves micromixing, pumping and particle handling. We solve it in closed form in a two-dimensional circular cavity, for any azimuthal wall mode $m$, as $\mathrm{Wo}^2 \to 0$. A biharmonic inversion against the Reynolds stress, corrected by the second-order slip a moving wall imposes, gives the Lagrangian mean a tracer follows for a deforming no-slip wall, $ψ_L = -[m(5m+4)a_m^2/(128(m+2)(2m+1))]\,r^{2m}(r^2-1)^2\sin 2mθ$, with a companion form for a shear-free interface. For a single mode the factor $(5m+4)/(m+2)$ relating it to the auxiliary solution $ψ_2$ is the same for every member of the co-phased prescribed-velocity family; at $m=2$ the physical Eulerian mean peaks an order of magnitude above $ψ_2$, with opposite sign. The no-slip cell centers lie at $r^2 = m/(m+2)$, and at large $m$ the peak streamfunction falls as $m^{-2}$, the peak speed as $m^{-1}$. At fixed radial wall-velocity amplitude the ranking over $m$ follows the wall kinematics: an externally driven wall peaks at $m=1$, a wall with zero first-order surface strain at $m=3$. Mode superpositions invert without degenerating, each harmonic carrying its own correction. At finite $\mathrm{Wo}$ the first-order field stays closed form in Bessel functions and the mean flow reduces to quadrature; the construction approaches the $m=2$ Rayleigh limit on a separate tangentially driven boundary problem. An independent finite-element solver, with the closed form withheld, reproduces $ψ_2$ with quadratic mesh convergence.

physics.flu-dyn↗

Structural identifiability and stress reconstruction from incomplete optical maps with velocimetry

Reconstructing the stress field of a planar viscoelastic flow from optical measurements loses its direct evidence wherever optical coverage is interrupted, and no improvement in optical precision restores an observation that was never made. We characterize what a second, velocity channel adds, and what neither channel can supply. Two calibrated optical components determine the local deviatoric stress pointwise, while velocity constrains spatial stress variation through momentum balance, so the two channels are complementary rather than redundant. The isotropic part of the stress is unobservable to both: the divergence of an isotropic field is a pure gradient, which the Leray projection annihilates, so every representable isotropic mode lies in the joint null space. That accounts for the null space exactly when the optical field is complete, and bounds it from below otherwise, since finite incomplete sampling and aperture zeros can remove further directions. We verify the count directly on three discretizations. In paired synthetic tests with finite measurement apertures, spatially correlated noise and optical stripe dropout, adding velocity reduces the mean whole-domain deviatoric error from 50.40% to 27.82% at 3% reference noise, and the improvement survives shared gaps, inverse-grid refinement at fixed physical sampling, and a constitutively generated stress field. The improvement does not rest on how the regularization parameter is chosen: it holds under both the expected-norm discrepancy rule and generalized cross-validation, and we report each selection with its position in the search interval, which is where the two rules differ.

physics.flu-dyn↗

Characterizing variation bounding in discrete-time Hankel operators

We investigate the $k$-variation bounding property of the discrete-time Hankel operator, i.e., its invariance under the set of signals with a variation (number of sign changes) of at most $k$. Building on existing sign-consistency criteria, it is shown that this property is equivalent to the external positivity of $k+1$ explicitly realized linear discrete-time systems. Thus, making the property tractable via numerical and analytic certificates. We also derive dominant-pole restrictions for the case of $k=1$. A three-node thermal example illustrates the results and distinguishes variation bounding from variation diminishing.

math.DS↗

Random Reshuffling Dominates Stochastic Gradient Descent

Stochastic Gradient Descent ($\textsf{SGD}$) is one of the most classical optimization algorithms with favorable theoretical guarantees, yet the practical implementation of $\textsf{SGD}$ differs subtly from its well-known form and is often referred to as Shuffling Stochastic Gradient Descent ($\textsf{Shuffling SGD}$). A particularly popular strategy in $\textsf{Shuffling SGD}$ is Random Reshuffling ($\textsf{RR}$), which has achieved great empirical success across numerous experiments. Despite its strong performance, $\textsf{RR}$ has long been considered a heuristic due to a lack of theoretical support. Over the last decade, people have finally established provable convergence rates for $\textsf{RR}$, thus justifying its observed superiority. However, for smooth convex optimization, two clouds over the convergence theory of $\textsf{RR}$ remain to this day. More precisely, according to the current theory, $\textsf{Shuffling SGD}$ under $\textsf{RR}$ converges only when the stepsize is smaller than a threshold proportional to $1/n$, where $n$ is the number of summands in the objective (or the number of data points). Consequently, the optimally tuned theoretical rate of $\textsf{Shuffling SGD}$ under $\textsf{RR}$ is strictly worse than that of $\textsf{SGD}$ when the number of epochs is smaller than another threshold proportional to $n$. These two restrictions heavily limit the applicability of existing theories and leave a critical mismatch with practice. In this work, for the first time, we prove that $\textsf{RR}$ dominates $\textsf{SGD}$ in smooth convex optimization under any reasonable stepsize after any finite number of epochs, thereby addressing a longstanding open question.

math.OC↗

ShareLock: A Stealthy Multi-Tool Threshold Poisoning Attack Against MCP

With the rapid evolution of LLM-driven agents, Model Context Protocol (MCP), an open protocol bridging LLMs with external tools, has quickly become foundational to modern agent ecosystems. However, the expanding adoption of MCP has also introduced novel security concerns such as Tool Poisoning Attack (TPA), which exploit LLM-server interactions to inject malicious prompts. Existing poisoning schemes typically adopt a monolithic plaintext embedding paradigm, which fails to withstand manual inspection or automated detectors. Current research still lacks a systematic analysis on multi-tool poisoning, where multiple tools can be exploited cooperatively to disperse detection risk. In this paper, we introduce ShareLock, a multi-tool threshold poisoning framework that utilizes Shamir's threshold scheme to ensure exceptional stealth and fault tolerance. ShareLock distributes the malicious instruction as benign-looking secret shares across multiple tool descriptions, achieving both information-theoretic secrecy and attack robustness against moderate auditing. After a covert reconstruction trigger is planted during server update, the aggregated shares reconstruct the hidden instruction, resulting in critical breaches of system assets or private data. To evaluate the realistic threat of ShareLock, we constructed a comprehensive benchmark encompassing four multi-tool scenarios and conducted extensive experiments across mainstream LLMs on two distinct MCP clients. Our results demonstrate that ShareLock significantly outperforms existing single-tool poisoning strategies in tool description-based detection while maintaining an average attack success rate exceeding 90%.

cs.CR↗

Adam Converges in Nonsmooth Nonconvex Optimization

Adam is one of the most widely implemented and influential modern optimizers. Why is it effective across different optimization problems in practice? This question arguably lies at the center of the optimization community over the last decade and has motivated a substantial body of work aimed at understanding its convergence behavior. However, existing studies have mainly focused on the convergence rate of Adam in smooth nonconvex optimization, which unfortunately does not adequately capture practical settings, since many real-world problems are nonsmooth, such as those arising in training neural networks. Thus, these studies cannot fully explain the popularity and empirical success of Adam. Recently, an insightful and powerful framework called Online-to-Nonconvex Conversion has opened a new way to analyze Adam for nonsmooth nonconvex optimization. Unfortunately, prior works along this line share two common limitations. First, all of them ignore the important bias-correction term in the original Adam algorithm. Second and more importantly, many of them require extra operations that are not used in Adam, such as a clipping step. Therefore, the convergence guarantee for the original Adam method still remains unclear. In this work, we present the first finite-time analysis for the classical form of Adam, i.e., with the bias-correction step and without further algorithmic modifications, and prove that a randomly scaled learning rate ensures a convergence rate of $1/T^{\frac{2}{13}}$ for nonsmooth nonconvex optimization. Moreover, our result provably applies to the modern heavy-tailed noise regime, which is closer to practice. Interestingly, our theory is established under the parameter choice $β_1=β_2$, aligning with the recent empirical studies.

math.OC↗

Co-policy: Responsive Human-Robot Co-Creation for Musical Performances

Art has long stood as a pivotal expression of human creativity. Embodied artificial intelligence offers a route for generative models to participate in that creativity through physical action rather than disembodied digital content. In robotic music co-creation, it is challenging to connect semantic musical understanding with real-time and physically executable performance. We present Co-policy, a framework for human-robot musical co-creation that separates semantic intent grounding, constrained musical variation, and visuomotor execution. To ground musical semantics, Co-policy uses pre-inference semantic anchors and a fine-tuned Qwen-vl planner (F-Qwen) to transform speech, live musical seeds, and visual observations into structured co-creation plans. To support low-latency execution, Co-policy introduces a Gaussian-Mixture Visuomotor Policy (GMP), implemented as a conditional mixture-density policy that maps target notes and visual context to multimodal robot actions in a single forward pass. Unlike robotic playback systems that merely reproduce user-specified notes, Co-policy generates complementary musical responses under both musical and physical constraints. Real-robot chime experiments, ablations, and expert evaluation show improved intent alignment, execution accuracy, and response frequency over diffusion-policy and ablated baselines, supporting physically grounded action generation as a key requirement for embodied human-AI co-creation.

cs.RO↗

Shadow Completion in Celestial OPEs

We argue that celestial OPEs must be supplemented by shadow-basis operators. Although the shadow transform does not introduce new bulk degrees of freedom, it provides a distinct primary state in the boundary celestial theory. From OPE consistency, we show that the ordinary celestial OPE does not close on Mellin-basis exchanges alone. Rather, the same exchanged bulk particle must also appear through its shadow-basis representative. This leads to a shadow-completed OPE, with the shadow OPE coefficient fixed by the ordinary collinear coefficient through the universal shadow factor. We discuss the corresponding boundary Hilbert-space interpretation, extend this argument to gluons and gravitons, and verify the shadow exchange directly in tree-level regular celestial amplitudes, including a scalar $2\rightarrow n$ analysis and an explicit five-point example.

hep-th↗

HiGR: Industrial-Scale Hierarchical Generative Slate Recommendation Framework in Tencent

Slate recommendation, which presents users with a ranked item list in a single display, is ubiquitous across mainstream online platforms. While recent generative recommendation methods have shown strong potential in modeling item sequences with semantic IDs, directly applying them to industrial-scale slate recommendation faces a fundamental disconnect: entangled SID spaces confound high-level list planning, fine-grained autoregressive decoding over long sequences limits semantic planning efficiency, and token-level objectives misalign with holistic slate quality. In this paper, we propose HiGR, an industrial-scale hierarchical generative framework for slate recommendation that bridges this disconnect through a co-designed pipeline. First, HiGR learns structured SIDs via a Prefix-Contrastive Residual Quantized VAE (PCRQ-VAE). By enforcing high-level prefixes to capture shared semantics, PCRQ-VAE creates a controllable discrete space that acts as a prerequisite for efficient planning. Leveraging this structured space, our Hierarchical Slate Decoder (HSD) shifts autoregressive modeling from entangled token-level decoding to coarse-grained preference embeddings. This design significantly reduces inference latency while allowing explicit global slate structure planning. Finally, this stable planning space enables an ORPO-based listwise alignment mechanism to optimize triple-objective implicit feedback-ranking fidelity, genuine user interest, and diversity. Extensive offline experiments show that HiGR outperforms state-of-the-art baselines by over 10% in offline recommendation quality while achieving a $5\times$ inference speedup. Online A/B tests on Tencent platforms further improve watch time by 1.22% and video plays by 1.73%. HiGR has been deployed on multiple Tencent platform surfaces, serving hundreds of millions of users and proving its industrial-scale applicability.

cs.IR↗

AIR-VLA+: Decoupling Movement and Manipulation via Cascaded Dual-Action Decoders with Asymmetric MoE for Aerial Robots

Aerial manipulation systems have long suffered from representation coupling in end-to-end control, as platform-level Unmanned Aerial Vehicle (UAV) movement and end-effector-level arm manipulation differ substantially in action scale, dynamics, and control objectives. In this paper, we propose AIR-VLA+, a flow matching action generation architecture specifically designed for aerial manipulation, featuring cascaded dual-action decoders and an asymmetric feature-level Mixture of Experts (MoE). We construct cascaded manipulation and movement decoders, allowing the UAV to unidirectionally observe the manipulator's intent during movement to achieve workflow coordination, while isolating the impact of UAV movement information backpropagation on arm manipulation stability. Addressing the characteristic that UAV movement is highly dependent on high-level semantics and responsible for task state transitions in aerial manipulation, we design an input feature enhancement module for the UAV movement decoder. This module introduces an implicit visual grasp projector to perceive the interaction state between the gripper and the object, and injects compressed global semantic features. Within the UAV movement decoder, we deploy an implicit MoE architecture, enabling different movement experts to spontaneously exhibit capacity inclinations for various task stages during training. Through dense soft blending computation on the feature manifold, the UAV movement is endowed with stronger task-stage adaptability. Experiments on the standardized AIR-VLA benchmark demonstrate that our method comprehensively surpasses all baselines with an overall average score of 48.0. The overall task completion score improves by 80.2\% compared to the single-head $π_{0.5}$ policy, effectively mitigating the heterogeneous coordinated control conflicts of composite robots.

cs.RO↗

Can Adaptive Gradient Methods Converge under Heavy-Tailed Noise? A Case Study of AdaGrad

Many tasks in modern machine learning are observed to involve heavy-tailed gradient noise during the optimization process. To manage this realistic and challenging setting, new mechanisms, such as gradient clipping and gradient normalization, have been introduced to ensure the convergence of first-order algorithms. However, adaptive gradient methods, a famous class of modern optimizers that includes popular $\mathtt{Adam}$ and $\mathtt{AdamW}$, often perform well even without any extra operations mentioned above. It is therefore natural to ask whether adaptive gradient methods can converge under heavy-tailed noise without any algorithmic changes. In this work, we take the first step toward answering this question by investigating a special case, $\mathtt{AdaGrad}$, the origin of adaptive gradient methods. We provide the first provable convergence rate for $\mathtt{AdaGrad}$ in non-convex optimization when the tail index $p$ satisfies $4/3<p\leq2$. Notably, this result is achieved without requiring any prior knowledge of $p$ and is hence adaptive to the tail index. In addition, we develop an algorithm-dependent lower bound, suggesting that the existing minimax rate for heavy-tailed optimization is not attainable by $\mathtt{AdaGrad}$. Lastly, we consider $\mathtt{AdaGrad}\text{-}\mathtt{Norm}$, a popular variant of $\mathtt{AdaGrad}$ in theoretical studies, and show an improved rate that holds for any $1<p\leq2$ under an extra mild assumption.

math.OC↗

In-Expectation Convergence of Stochastic Gradient Methods under Heavy-Tailed Noise

Many stochastic gradient methods are believed not to converge when the noise in stochastic gradients has only a finite $p$-th moment for $p\in\left(1,2\right)$, a setting known as the heavy-tailed noise assumption. However, some recent studies have found that Stochastic Gradient Descent ($\textsf{SGD}$), without any modification to its update rule, can surprisingly converge in expectation for convex problems with bounded domains, highlighting the potential of classical stochastic gradient methods. Inspired by this recent progress, we provide a comprehensive study of stochastic optimization under heavy-tailed noise and establish new in-expectation convergence results for Stochastic Mirror Descent ($\textsf{SMD}$) and Accelerated Stochastic Mirror Descent ($\textsf{ASMD}$) in convex optimization, and for $\textsf{SGD}$ and Stochastic Gradient Descent with Momentum ($\textsf{SGDM}$) in nonconvex optimization. Notably, our results not only hold without algorithmic changes but also avoid restrictive assumptions, such as bounded domains, imposed in prior work. More importantly, our analysis provides a new, elegant, and powerful framework for studying heavy-tailed stochastic optimization, opening a new route to understanding first-order stochastic gradient methods.

math.OC↗

Clipped Gradient Methods for Nonsmooth Convex Optimization under Heavy-Tailed Noise: A Refined Analysis

Optimization under heavy-tailed noise has become popular recently, since it better fits many modern machine learning tasks, as captured by empirical observations. Concretely, instead of a finite second moment on gradient noise, a bounded ${\frak p}$-th moment where ${\frak p}\in(1,2]$ has been recognized to be more realistic (say being upper bounded by $σ_{\frak l}^{\frak p}$ for some $σ_{\frak l}\ge0$). A simple yet effective operation, gradient clipping, is known to handle this new challenge successfully. Specifically, Clipped Stochastic Gradient Descent (Clipped SGD) guarantees a high-probability rate ${\cal O}(σ_{\frak l}\ln(1/δ)T^{1/{\frak p}-1})$ (resp. ${\cal O}(σ_{\frak l}^2\ln^2(1/δ)T^{2/{\frak p}-2})$) for nonsmooth convex (resp. strongly convex) problems, where $δ\in(0,1]$ is the failure probability and $T\in\mathbb{N}$ is the time horizon. In this work, we provide a refined analysis for Clipped SGD and offer two rates, ${\cal O}(σ_{\frak l}d_{\rm eff}^{-1/2{\frak p}}\ln^{1-1/{\frak p}}(1/δ)T^{1/{\frak p}-1})$ and ${\cal O}(σ_{\frak l}^2d_{\rm eff}^{-1/{\frak p}}\ln^{2-2/{\frak p}}(1/δ)T^{2/{\frak p}-2})$, faster than the aforementioned best results, where $d_{\rm eff}\ge1$ is a quantity we call the $\textit{generalized effective dimension}$. Our analysis improves upon the existing approach on two sides: better utilization of Freedman's inequality and finer bounds for clipping error under heavy-tailed noise. In addition, we extend the refined analysis to convergence in expectation and obtain new rates that break the known lower bounds. Lastly, to complement the study, we establish new lower bounds for both high-probability and in-expectation convergence. Notably, the in-expectation lower bounds match our new upper bounds, indicating the optimality of our refined analysis for convergence in expectation.

math.OC↗

Chaotic dynamics of charged particles near weakly magnetized black holes in Einstein-ModMax Theory

This paper presents a systematic study of the chaotic dynamics of charged test particles around purely magnetically charged black holes immersed in a uniform external magnetic field within the framework of Einstein-ModMax theory. By constructing an explicit symplectic integrator, we obtain high-precision numerical solutions of the equations of motion. Combining the observational constraints from the Event Horizon Telescope (EHT) shadow images, we further restrict the parameter ranges of the model. We apply Shannon entropy and MIPP (mutual information for particle pairs) as effective indicators to identify the chaotic behavior of charged test particles in the spacetime of this black hole. Numerical results indicate that these indicators can clearly distinguish between regular and chaotic motion of orbits in strong gravitational field systems. Further analysis reveals that, compared to the key conserved quantities that determine the global dynamical behavior of the system -- energy $E$ and angular momentum $L$, the sensitivity of the system parameters $e^{-ν}$ and $Q_{m}$ to transitions in the orbital dynamical states is significantly reduced. This study provides a new perspective for a deeper understanding of the characterization and evolution mechanisms of chaotic dynamics in strong gravitational fields.

gr-qc↗

Relaxing Anchor-Frame Dominance for Mitigating Hallucinations in Video Large Language Models

Recent Video Large Language Models (Video-LLMs) have demonstrated strong capability in video understanding, yet they still suffer from hallucinations. Existing mitigation methods typically rely on training, input modification, auxiliary guidance, or additional decoding procedures, while largely overlooking a more fundamental challenge. During generation, Video-LLMs tend to over-rely on a limited portion of temporal evidence, leading to temporally imbalanced evidence aggregation across the video. To address this issue, we investigate a decoder-side phenomenon in which the model exhibits a temporally imbalanced concentration pattern. We term the frame with the highest aggregated frame-level attention mass the anchor frame. We find that this bias is largely independent of the input video and instead appears to reflect a persistent, model-specific structural or positional bias, whose over-dominance is closely associated with hallucination-prone generation. Motivated by this insight, we propose Decoder-side Temporal Rebalancing (DTR), a training-free, layer-selective inference method that rebalances temporal evidence allocation in middle-to-late decoder layers without altering visual encoding or requiring auxiliary models. DTR adaptively calibrates decoder-side visual attention to alleviate temporally imbalanced concentration and encourage under-attended frames to contribute more effectively to response generation. In this way, DTR guides the decoder to ground its outputs in temporally broader and more balanced video evidence. Extensive experiments on hallucination and video understanding benchmarks show that DTR consistently improves hallucination robustness across diverse Video-LLM families, while preserving competitive video understanding performance and high inference efficiency.

cs.CV↗

HTNav: A Hybrid Navigation Framework with Tiered Structure for Urban Aerial Vision-and-Language Navigation

Inspired by the general Vision-and-Language Navigation (VLN) task, aerial VLN has attracted widespread attention, owing to its significant practical value in applications such as logistics delivery and urban inspection. However, existing methods face several challenges in complex urban environments, including insufficient generalization to unseen scenes, suboptimal performance in long-range path planning, and inadequate understanding of spatial continuity. To address these challenges, we propose HTNav, a new collaborative navigation framework that integrates Imitation Learning (IL) and Reinforcement Learning (RL) within a hybrid IL-RL framework. This framework adopts a staged training mechanism to ensure the stability of the basic navigation strategy while enhancing its environmental exploration capability. By integrating a tiered decision-making mechanism, it achieves collaborative interaction between macro-level path planning and fine-grained action control. Furthermore, a map representation learning module is introduced to deepen its understanding of spatial continuity in open domains. On the CityNav benchmark, our method achieves state-of-the-art performance across all scene levels and task difficulties. Experimental results demonstrate that this framework significantly improves navigation precision and robustness in complex urban environments.

cs.RO↗

Revisiting the Last-Iterate Convergence of Stochastic Gradient Methods

In the past several years, the last-iterate convergence of the Stochastic Gradient Descent (SGD) algorithm has triggered people's interest due to its good performance in practice but lack of theoretical understanding. For Lipschitz convex functions, different works have established the optimal $O(\log(1/δ)\log T/\sqrt{T})$ or $O(\sqrt{\log(1/δ)/T})$ high-probability convergence rates for the final iterate, where T is the time horizon and δis the failure probability. However, to prove these bounds, all the existing works are either limited to compact domains or require almost surely bounded noise. It is natural to ask whether the last iterate of SGD can still guarantee the optimal convergence rate but without these two restrictive assumptions. Besides this important question, there are still lots of theoretical problems lacking an answer. For example, compared with the last-iterate convergence of SGD for non-smooth problems, only few results for smooth optimization have yet been developed. Additionally, the existing results are all limited to a non-composite objective and the standard Euclidean norm. It still remains unclear whether the last-iterate convergence can be provably extended to wider composite optimization and non-Euclidean norms. In this work, to address the issues mentioned above, we revisit the last-iterate convergence of stochastic gradient methods and provide the first unified way to prove the convergence rates both in expectation and in high probability to accommodate general domains, composite objectives, non-Euclidean norms, Lipschitz conditions, smoothness, and (strong) convexity simultaneously. Additionally, we extend our analysis to obtain the last-iterate convergence under heavy-tailed and sub-Weibull noise.

cs.LG↗