Search arXivSearch

arXiv subjects

Keyu Zhang

Publications and source records attributed to Keyu Zhang.

14 recordsLinked to original sources

Stemma: Induced Decision Regions Reveal LLM Provenance

LLM provenance testing asks whether a suspect LLM belongs to the same lineage as a source. Existing black-box methods largely infer this relationship from response-level characteristics, but these characteristics may shift under adaptation or deployment even when the underlying meaning remains unchanged, weakening the reliability of provenance evidence. To address this limitation, we introduce induced decision regions by mapping open-ended outputs into a finite decision space, thereby abstracting away surface-form variation and reframing provenance testing as measuring the inheritance of decision regions. Empirical analysis shows that the source's induced regions are preserved more strongly in related models than in unrelated models. Building on this signal, we propose Stemma, a practical black-box LLM fingerprinting method that operationalises stability, robustness, and specificity as complementary probe-selection principles for reliably estimating induced decision region inheritance. Across 770 source-suspect pairs drawn from 56 public checkpoints and spanning diverse model-weight transformations, Stemma achieves 0.967 AUC and 87.8% TPR at 1% FPR, substantially outperforming four representative baselines. It further achieves 0.995 AUC and 93.5% TPR at 1% FPR on 1,260 pairs covering 91 deployment instances, demonstrating robustness to diverse inference-time deployment settings.

cs.CR

Equilibrium for Time-inconsistent Mean Field Games: A Systematic Analysis by Entropy Regularization

This paper studies the existence and approximation of equilibria for general time-inconsistent mean field game (MFG) problems in continuous time. To handle the intricate nonlocal equilibrium Hamilton-Jacobi-Bellman (EHJB) system arising from initial-time dependence, such as non-exponential discounting, we develop a vanishing entropy regularization approach. Using entropy regularization, we first characterize the regularized equilibrium through a coupled exploratory equilibrium HJB (EEHJB) equation and a law-dependent stochastic differential equation. By exploiting Schauder fixed-point arguments and tailored parabolic regularity estimates in a suitable functional space involving both value functions and measure flows, we establish the global existence of regularized equilibria under mild assumptions. We then establish convergence as the entropy regularization vanishes. By employing compactness arguments, Young measure techniques, and a duality tool for divergence-form Fokker-Planck equations, we prove that the regularized equilibria converge, up to subsequences, to a mean-field equilibrium of the original MFG. Furthermore, under entropy regularization, we propose a policy iteration algorithm and establish its convergence under short-time-horizon and weak-terminal-interaction conditions.

math.OC

Tailored Prompts, Targeted Protection: Vulnerability-Specific LLM Analysis for Smart Contracts

Smart contracts on blockchains are prone to diverse security vulnerabilities that can lead to significant financial losses due to their immutable nature. Existing detection approaches often lack flexibility across vulnerability types and rely heavily on manually crafted expert rules. In this paper, we present an LLM-based framework for practical smart contract vulnerability detection. We construct and release a large-scale dataset comprising 31,165 professionally annotated vulnerability instances collected from over 3,200 real-world projects across 15 major blockchain platforms. Our approach leverages precise AST-based context extraction and vulnerability-specific prompt design to instantiate customized detectors for 13 prevalent vulnerability categories. Experimental results demonstrate strong effectiveness, achieving an average positive recall of 0.92 and an average negative recall of 0.85, highlighting the potential of carefully engineered contextual prompting for scalable and high-precision smart contract security analysis.

cs.CR

Policy Iteration Achieves Regularized Equilibrium under Time Inconsistency

For a general entropy-regularized time-inconsistent stochastic control problem, we propose a policy iteration algorithm (PIA) and establish its convergence to an equilibrium policy with an exponential convergence rate. The design of the PIA is based on a coupled system of non-local partial differential equations, called the exploratory equilibrium Hamilton--Jacobi--Bellman (EEHJB) equation. As opposed to the standard time-consistent case, policy improvement fails in general and the target value function (now an equilibrium value function) is not even known to exist a priori. To overcome these, we prove that the value functions generated by the PIA form a Cauchy sequence in a specialized Banach space, hence admit a limit, and the rate of convergence is exponential, on the strength of the Bismut--Elworthy--Li formula of stochastic representation. The limiting value function is shown to fulfill the EEHJB equation, which induces an equilibrium policy in a Gibbs form. Such convergence in value additionally implies uniform convergence of the generated policies to the equilibrium policy, again with an exponential rate. As a byproduct, the PIA gives a constructive proof of the global existence and uniqueness of a classical solution to our general EEHJB equation, whose well-posedness has not been explored in the literature.

math.OC

Scalable Dexterous Robot Learning with AR-based Remote Human-Robot Interactions

This paper focuses on the scalable robot learning for manipulation in the dexterous robot arm-hand systems, where the remote human-robot interactions via augmented reality (AR) are established to collect the expert demonstration data for improving efficiency. In such a system, we present a novel method to address the general manipulation task problem. Specifically, the proposed method consists of two phases: i) In the first phase for pretraining, the policy is created in a behavior cloning (BC) manner, through leveraging the learning data from our AR-based remote human-robot interaction system; ii) In the second phase, a contrastive learning empowered reinforcement learning (RL) method is developed to obtain more efficient and robust policy than the BC, and thus a projection head is designed to accelerate the learning progress. An event-driven augmented reward is adopted for enhancing the safety. To validate the proposed method, both the physics simulations via PyBullet and real-world experiments are carried out. The results demonstrate that compared to the baselines, our method not only significantly speeds up the training process, but also achieves much better performance in terms of the success rate for fulfilling the manipulation tasks. By conducting the ablation study, it is confirmed that the proposed RL with contrastive learning overcomes policy collapse. Supplementary demonstrations are available at https://cyberyyc.github.io/.

cs.LG

MARVEL Analysis of the Measured High-resolution Spectra of CO Isotopologues

Carbon monoxide is thought to be the second most abundant molecule in the Universe. This makes observation of both its parent isotopologue ($^{12}$C$^{16}$O) and its stable isotopologues, $^{13}$C$^{16}$O, $^{12}$C$^{18}$O, $^{12}$C$^{17}$O, $^{13}$C$^{18}$O and $^{13}$C$^{17}$O, important in variety of objects. Here the MARVEL (Measured Active Rotational-Vibrational Energy Levels) algorithm is used to determine precise rotational vibrational energy levels for the five minor isotopologues of carbon monoxide in their electronic ground state. A review of 27 literature sources yields 3716, 1454, 89, 728 and 57 validated transitions for $^{13}$C$^{16}$O, $^{12}$C$^{18}$O, $^{12}$C$^{17}$O, $^{13}$C$^{18}$O and $^{13}$C$^{17}$O, respectively, giving 863, 499, 33, 345 and 45 empirically determined, rotation vibration energy levels, respectively.

astro-ph.GA

Mean Field Game with Reflected Jump Diffusion Dynamics: A Linear Programming Approach

This paper develops a linear programming approach for mean field games with reflected jump-diffusion dynamics. We first prove the equivalence between the mean field equilibria in the linear programming formulation and those in the weak relaxed control formulation under some measurability and growth conditions on model coefficients. Building upon the characterization of the occupation measure in the equivalence result, we further establish the existence of linear programming mean field equilibria under fairly general conditions on model coefficients. Finally, a numerical example is presented to illustrate the computation of a mean field equilibrium using the linear programming formulation.

math.OC

RaceTEE: Enabling Interoperability of Confidential Smart Contracts

Decentralized smart contracts enable trustless collaboration but suffer from limited privacy and scalability, which hinders broader adoption. Trusted Execution Environment (TEE) based off-chain execution frameworks offer a promising solution to both issues. Although TEE-based frameworks have made significant progress, prior work has yet to fully explore contract interoperability, a critical foundation for building complex real-world decentralized applications. This paper identifies the key challenges impeding such interoperability and presents practical solutions. Based on these insights, we introduce RaceTEE, a novel framework that leverages off-chain TEE-enabled nodes to efficiently execute confidential, long-lived smart contracts with interactions of arbitrary complexity among contracts. We implement a RaceTEE prototype using Intel SGX, integrate it with Ethereum, and release it as open source. Evaluation across diverse use cases demonstrates its practicality and effectiveness.

cs.CR

Major-Minor Mean Field Game of Stopping: An Entropy Regularization Approach

This paper studies a discrete-time major-minor mean field game of stopping where the major player can choose either an optimal control or stopping time. We look for the relaxed equilibrium as a randomized stopping policy, which is formulated as a fixed point of a set-valued mapping, whose existence is challenging by direct arguments. To overcome the difficulties caused by the presence of a major player, we propose to study an auxiliary problem by considering entropy regularization in the major player's problem while formulating the minor players' optimal stopping problems as linear programming over occupation measures. We first show the existence of regularized equilibria as fixed points of some simplified set-valued operator using the Kakutani-Fan-Glicksberg fixed-point theorem. Next, we prove that the regularized equilibrium converges as the regularization parameter $\lambda$ tends to 0, and the limit corresponds to a fixed point of the original operator, thereby confirming the existence of a relaxed equilibrium in the original mean field game problem. We also extend this entropy regularization method to the mean-field game problem where the minor players choose optimal controls.

math.OC

Constrained portfolio game with heterogeneous agents

We investigate stochastic utility maximization games under relative performance concerns in both finite-agent and infinite-agent (graphon) settings. An incomplete market model is considered where agents with power (CRRA) utility functions trade in a common risk-free bond and individual stocks driven by both common and idiosyncratic noise. The Nash equilibrium for both settings is characterized by forward-backward stochastic differential equations (FBSDEs) with a quadratic growth generator, where the solution of the graphon game leads to a novel form of infinite-dimensional McKean-Vlasov FBSDEs. Under mild conditions, we prove the existence of Nash equilibrium for both the graphon game and the $n$-agent game without common noise. Furthermore, we establish a convergence result showing that, with modest assumptions on the sensitivity matrix, as the number of agents increases, the Nash equilibrium and associated equilibrium value of the finite-agent game converge to those of the graphon game.

math.OC

On time-inconsistent extended mean-field control problems with common noise

This paper studies a class of time-inconsistent mean field control (MFC) problems in the presence of common noise under non-exponential discount and joint law dependence of both state and control. We investigate the closed-loop time-consistent equilibrium strategies for these extended MFC problems and characterize them through an equilibrium Hamilton-Jacobi-Bellman (HJB) equation defined on the Wasserstein space. We first apply the results to the linear-quadratic (LQ) time-inconsistent MFC problems and obtain the existence of time-consistent equilibria via a comprehensive study of a nonlocal Riccati system. To illustrate the theoretical findings, two financial applications are presented. We then examine a class of non-LQ time-inconsistent MFC problems, for which we contribute the existence of time-consistent equilibria by analyzing a nonlocal nonlinear partial differential equation.

math.OC

A Mean Field Game Approach to Relative Investment-Consumption Games with Habit Formation

This paper studies an optimal investment-consumption problem for competitive agents with exponential or power utilities and a common finite time horizon. Each agent regards the average of habit formation and wealth from all peers as benchmarks to evaluate the performance of her decision. We formulate the n-agent game problems and the corresponding mean field game problems under the two utilities. One mean field equilibrium is derived in a closed form in each problem. In each problem with n agents, an approximate Nash equilibrium is then constructed using the obtained mean field equilibrium when n is sufficiently large. The explicit convergence order in each problem can also be obtained. In addition, we provide some numerical illustrations of our results.

math.OC

Equilibrium stochastic control with implicitly defined objective functions

This paper considers a class of stochastic control problems with implicitly defined objective functions, which are the sources of time-inconsistency. We study the closed-loop equilibrium solutions in a general controlled diffusion framework. First, we provide a sufficient and necessary condition for a strategy to be an equilibrium. Then, we apply the result to discuss two problems of dynamic portfolio selection for a class of betweenness preferences, allowing for closed convex constraints on portfolio weights and borrowing cost, respectively. The equilibrium portfolio strategies are explicitly characterized in terms of the solutions of some first-order ordinary differential equations for the case of deterministic market coefficients.

math.OC

Time-inconsistent mean field and n-agent games under relative performance criteria

In this paper we study a time-inconsistent portfolio optimization problem for competitive agents with CARA utilities and non-exponential discounting. The utility of each agent depends on her own wealth and consumption as well as the relative wealth and consumption to her competitors. Due to the presence of a non-exponential discount factor, each agent's optimal strategy becomes time-inconsistent. In order to resolve time-inconsistency, each agent makes a decision in a sophisticated way, choosing open-loop equilibrium strategy in response to the strategies of all the other agents. We construct explicit solutions for the $n$-agent games and the corresponding mean field games (MFGs) where the limit of former yields the latter. This solution is unique in a special class of equilibria.

q-fin.MF