Search arXivSearch

arXiv subjects

Bohan Zhang

Publications and source records attributed to Bohan Zhang.

At least 19 recordsLinked to original sources

Tikhonov-Regularized Second-Order Primal-Dual Dynamics for Convex-Concave Bilinear Saddle Point Problems

This paper develops a Lyapunov-based coefficient-compatible framework for a class of second-order primal-dual dynamical systems with Tikhonov regularization. Unlike existing accelerated inertial primal-dual dynamics, whose analyses are typically carried out under prescribed damping profiles, the proposed framework characterizes the interplay among the viscous damping coefficient, the time-scaling coefficient, the linear extrapolation parameter, and the Tikhonov regularization parameter through differential compatibility conditions, thereby accommodating time-varying coefficients beyond prescribed damping profiles. Within this framework, we investigate two cases in which the Tikhonov regularization parameter decays to zero at different rates. When the Tikhonov regularization parameter decays to zero relatively rapidly, the fast convergence rate of the primal-dual gap is preserved, whereas a slower decay rate yields refined asymptotic convergence estimates. We further show that a sufficiently slow decay ensures strong convergence of the trajectories to the minimum-norm saddle point. To establish the latter property, we introduce a moving-saddle Lyapunov framework centered at the instantaneous regularized saddle point and derive verifiable coefficient conditions ensuring the required tracking behavior. Finally, three numerical examples are presented to illustrate the theoretical results.

math.OC

ORBITALIF: An Efficient Spiking Federated Learning Framework for Onboard Cloud Removal

Low-earth-orbit (LEO) satellites enable high-resolution, large-scale Earth observation for applications such as disaster monitoring and environmental surveillance. However, cloud coverage often obscures the Earth's surface, and conventional cloud-removal pipelines that download cloudy images to ground stations for processing suffer from limited contact windows, constrained satellite-to-ground bandwidth, and high latency. In this work, we propose a novel satellite federated learning framework for cloud removal across LEO constellations, named orbital attention leaky integrate-and-fire (OrbitALIF). OrbitALIF performs both onboard training and inference using a compact 2.30,M-parameter spiking neural network (SNN) backbone with an adaptive gated fusion module (AGFM) and a spectral-spatial hybrid attention module (SHAM), combined with a decentralized federated learning strategy that shares model weights via inter-satellite links. Our experiments show that OrbitALIF achieves competitive cloud removal quality while consuming only 0.287,mJ per inference on neuromorphic hardware, a 72.3 times (98.6%) energy reduction versus an equivalent artificial neural network (ANN).

cs.NE

Weightless Fine-Tuning: Personalizing LLMs via Logit-Space Transport

Supervised fine-tuning (SFT) is a standard approach for adapting LLMs to a target distribution, but in settings such as personalization, where each author requires separate weight access, optimization, storage, and retraining, its costs become prohibitive. We propose Weightless Fine-Tuning (WFT), a training-free decoding-time method that approximates the distributional effect of SFT without weight updates. WFT computes supervised residuals on an author's training sequence and transports them to the current prompt through a cross-prefix transport operator estimated from dropout-induced cross-covariance. The operator captures how a perturbation at one context propagates to predictions at another, replacing gradient-based parameter updates with logit-space corrections. On three LaMP personalization benchmarks, WFT achieves the best average performance across datasets, matches or exceeds SFT on individual tasks, and outperforms other lightweight baselines on average. In a budget-controlled comparison, WFT approaches SFT performance using less than 7% of the effective computation. Logit-level analysis shows a cosine similarity of 0.875 between the logit shifts induced by WFT and SFT over 95% of the next-token probability mass, suggesting that WFT captures the distributional effect of supervised adaptation without modifying model weights.

cs.LG

A Wearable Stiffness-Rendering Haptic Device with a Honeycomb Jamming Mechanism for Bilateral Teleoperation

This paper addresses the challenge of providing kinesthetic feedback in bilateral teleoperation by designing a wearable, lightweight (20 g), and compact haptic device, the HJ-Haptic, utilizing a honeycomb jamming mechanism for object stiffness rendering. The HJ-Haptic device can vary its stiffness, from 1.15 N/mm to 2.64 N/mm, using a 30 kPa vacuum pressure. We demonstrate its implementation in a teleoperation framework, enabling operators to adjust grip force based on a reliable haptic feedback on object stiffness. A three-point flexural test on the honeycomb jamming mechanism and teleoperated object-grasping tasks were conducted to evaluate the device's functionality. Our experiments demonstrated a small RMSE and strong correlations in teleoperated motion, stiffness rendering, and interaction force feedback. The HJ-Haptic effectively adjusts its stiffness in response to real-time gripper feedback, mimicking the sensation of direct object grasping with hands. The device's use of vacuum pressure ensures operator safety by preventing dangerous outcomes in case of gas leakage or material failure. Incorporating the HJ-Haptic into the teleoperation framework provided the reliable perception of object stiffness and stable teleoperation. This study highlights the potential of the honeycomb jamming mechanism for enhancing haptic feedback in various applications, including teleoperation scenarios, as well as interactions with extended-reality environments.

cs.RO

Consistent Feature Transport for Image Relighting

Image relighting modifies illumination while preserving non-lighting content such as identity and geometry. Existing diffusion-based methods often suffer from unstable illumination changes or inconsistent content preservation under complex lighting, as they lack an explicit mechanism to learn feature transformations between images. We reformulate relighting as an illumination feature transport problem and introduce Consistent Feature Transport (CFT), a training principle that explicitly enforces illumination-consistent transport between source and target image distributions. Built upon rectified flow, CFT jointly models noise-to-image generation and illumination-consistent source-to-target transport through trajectory-level supervision. This dual-transport formulation encourages isolation of illumination-specific variations while preserving content-aligned features. To support complex lighting scenarios, we construct a large-scale portrait relighting dataset with diverse relighting effects. Experiments show consistent improvements over existing state-of-the-art relighting approaches and demonstrate that CFT can generalize to other editing tasks, including style transfer. Code is available at https://github.com/Dixin-Lab/CFT.

cs.CV

Bifurcation reordering programs snap-through symmetry in folded elastic ribbons

Snap-through in slender elastic structures is often viewed as a sudden transition between stable configurations, yet the pathway taken during this transition can differ fundamentally. A structure may snap while preserving symmetry, or first lose symmetry and pass through an asymmetric state before reaching its final configuration. What selects between these pathways remains less well understood, especially when the geometry and loading are themselves symmetric. Here, we show that snap-through symmetry can be programmed by reordering competing bifurcations in folded elastic ribbons. We study an elastic ribbon with two localized folds placed symmetrically about the midpoint and show that varying the fold position changes the relative order of two instabilities: a symmetry-breaking pitchfork bifurcation and a saddle-node bifurcation on the symmetry-preserving branch. When the pitchfork bifurcation occurs first, the ribbon loses symmetry before snapping and follows an asymmetric pathway. Conversely, when the saddle-node bifurcation occurs first, the ribbon loses stability while remaining on the symmetric branch, resulting in a symmetry-preserving transition. Combining experiments, discrete differential geometry simulations and numerical continuation, we map this exchange in bifurcation ordering and construct a phase diagram that predicts the switch between asymmetric and symmetric snap-through regimes. A reduced-order double-mass von Mises truss model captures the same mechanism as a generic competition between symmetry-breaking and symmetry-preserving instabilities. These results establish bifurcation reordering as a geometric mechanism for programming snap-through pathways in slender elastic structures, offering a design principle for multistable systems, morphing structures and instability-based mechanical devices.

cond-mat.soft

Agentic Economic Modeling

We introduce Agentic Economic Modeling (AEM), a framework that aligns synthetic LLM choices with small-sample human evidence for econometric inference. AEM first generates task-conditioned synthetic choices via LLMs, then learns a bias-correction mapping from task features and raw LLM choices to human-aligned choices, upon which standard econometric estimators perform inference to recover demand elasticities and treatment effects. We validate AEM in two experiments. In a large scale conjoint study, using only 10% of the original data to fit the correction model lowers the error of the demand-parameter estimates, while uncorrected LLM choices increase the errors. In a regional field experiment, a mixture model calibrated on 10% of geographic regions estimates a treatment effect of -65$\pm$10 bps on the hold-out regions, closely matching the full human experiment (-60$\pm$8 bps). These results demonstrate AEM's potential to improve RCT efficiency and represent a step toward LLM-based counterfactual generation.

econ.EM

NEP89: Universal neuroevolution potential for inorganic and organic materials across 89 elements

While machine-learned interatomic potentials offer near-quantum-mechanical accuracy for atomistic simulations, many are material-specific or computationally intensive, limiting their broader use. Here, we introduce NEP89, a foundation model based on neuroevolution potential architecture, delivering near-empirical-potential speed and high accuracy across 89 elements. A compact yet comprehensive training dataset covering inorganic and organic materials was curated through descriptor-space subsampling and iterative refinement across multiple datasets. NEP89 achieves competitive accuracy compared to representative foundation models while being three to four orders of magnitude more computationally efficient, enabling previously impractical large-scale atomistic simulations of inorganic and organic systems. In addition to its out-of-the-box applicability to diverse scenarios, including million-atom-scale compression of compositionally complex alloys, ion diffusion in solid-state electrolytes and water, rocksalt dissolution, methane combustion, and protein-ligand dynamics, NEP89 also supports fine-tuning for rapid adaptation to user-specific applications, such as mechanical, thermal, structural, and spectral properties of two-dimensional materials, metallic glasses, and organic crystals.

cond-mat.mtrl-sci

SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows

Spreadsheets are widely used for business analysis, financial modeling, reporting, and decision-making. However, most existing spreadsheet benchmarks evaluate isolated operations such as single-formula generation or local cell edits, and therefore fail to capture end-to-end workflows in realistic business settings. We introduce \textsc{SpreadsheetBench 2}, a workflow-level benchmark for spreadsheet agents that covers three task categories: generation, debugging, and visualization. The benchmark is constructed from authentic business data, including financial reports and corporate filings, and is annotated and validated by domain experts. The benchmark contains 321 tasks; each instance averages 11.8 worksheets and requires 593.5 cell modifications, reflecting large multi-sheet workbooks with cross-sheet dependencies. We evaluate eight frontier large language models under a unified multi-turn agent scaffold, and additionally include several LLM-based spreadsheet products as complementary baselines. Results show that current systems remain far from reliable on real-world workflows: the best model achieves 34.89\% overall task accuracy, and debugging accuracy is as low as 12.00\%. Trajectory analysis and a failure taxonomy further indicate that insufficient spreadsheet inspection and incorrect target-cell selection are the dominant bottlenecks. Together, these findings position \textsc{SpreadsheetBench 2} as a challenging testbed for advancing reliable spreadsheet automation. Project page: https://spreadsheetbench.github.io/

cs.SE

Spin-adapted neural network backflow for symmetry-preserving simulations of strongly correlated electrons

Strongly correlated molecules often contain dense manifolds of low-lying spin states, making total-spin symmetry essential for predictive electronic-structure theory. Neural-network quantum states provide flexible variational wavefunctions, but commonly used fermionic architectures do not enforce this symmetry and can therefore converge to spin-contaminated states with misleading energies and properties. Here we introduce a spin-adapted neural-network backflow (SA-NNBF) ansatz in second quantization, which combines configuration-dependent spatial orbitals with a compressed spin eigenfunction. A projected tensor compression scheme for spin eigenfunctions and a particle-hole representation make variational Monte Carlo calculations with SA-NNBF practical for active spaces containing more than one hundred electrons. Across hydrogen chains and iron-sulfur clusters, SA-NNBF eliminates spin contamination and consistently achieves lower variational energies than standard NNBF with a comparable number of parameters. For the CAS(113e,76o) active-space model of FeMoco, SA-NNBF yields a highly compact spin-adapted variational state, achieving an energy competitive with recent spin-adapted DMRG calculations at bond dimension $D=10000$ while using orders of magnitude fewer parameters. Our work establishes a general framework for developing spin-symmetry-preserving neural-network quantum states for chemically realistic strongly correlated electrons.

physics.chem-ph

LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation

We present LongLive-2.0, an NVFP4-based parallel infrastructure throughout the full training and inference workflow of long video generation, addressing speed and memory bottlenecks. For training, we introduce sequence-parallel autoregressive (AR) training, instantiated as Balanced SP, which co-designs the efficient teacher-forcing layout with SP execution by pairing clean-history and noisy-target temporal chunks on each rank, enabling a natural teacher-forcing mask with SP-aware chunked VAE encoding. Combined with NVFP4 precision, it reduces GPU memory cost and accelerates GEMM computation during training, the proportion of which increases as video length grows. Moreover, we show that a high-quality infrastructure and dataset enable a remarkably clean training pipeline. Unlike existing Self-Forcing series methods that rely on ODE initialization and subsequent distribution matching distillation (DMD), LongLive-2.0 directly tunes a diffusion model into a long, multi-shot, interactive auto-regressive (AR) diffusion model. It can be further converted to real-time generation (4 to 2 denoising steps) with standalone LoRA weights. For inference on Blackwell GPUs, we enable W4A4 NVFP4 inference, quantize KV cache into NVFP4 for memory savings, and boost end-to-end throughput with asynchronous streaming VAE decoding. On non-Blackwell GPU architectures, we deploy SP inference to match the speed on Blackwell GPUs, while the quantized KV cache can lower inter-GPU communication of SP. Experiments show up to 2.15x speedup in training, and 1.84x in inference. LongLive-2.0-5B achieves 45.7 FPS inference while attaining strong performance on benchmarks. To our knowledge, LongLive-2.0 is the first NVFP4 training and inference system for long video generation.

cs.CV

Platform Competition with User-Generated Content

This paper develops a theoretical model of platform competition where user-generated content (UGC) quality arises endogenously from the composition of the user base. Users differ in their relative preferences for content quality and network size, and platforms compete by choosing advertising intensity, which affects user utility through perceived quality. We characterize equilibrium platform choice, identifying conditions under which equilibria are stable. The model captures how platforms' strategic decisions shape user allocation and market outcomes, including coexistence and dominance scenarios. We consider two types of equilibria in advertising levels: Nash equilibria and Stackelberg equilibria, and discuss the industry and policy implications of our results.

econ.TH

Convergence Analysis of Hessian-Damped Tikhonov Regularized Dynamics with Oscillation Control for Convex-Concave Bilinear Saddle Point Problems

In this paper, we propose a class of general second-order primal-dual dynamical systems with Tikhonov regularization and Hessian-driven damping for solving convex-concave bilinear saddle point problems. The proposed dynamical system incorporates five general time-varying terms: viscous damping, time scaling, extrapolation, Tikhonov regularization, and Hessian-driven damping parameters. Under suitable parametric conditions, we analyze the asymptotic convergence properties of the dynamical system by constructing appropriate Lyapunov functions. Specifically, we obtain the convergence rate of the primal-dual gap and the boundedness of trajectories in the proposed dynamical system, and provide some integral estimates. Furthermore, we theoretically prove that the trajectories generated by the dynamical system converge strongly to the minimum-norm solution of the saddle point problem, and fully demonstrate that Hessian-driven damping can effectively alleviate oscillations. Finally, numerical experiments are conducted to verify the validity of the above theoretical results.

math.OC

OpenAI o1 System Card

The o1 model series is trained with large-scale reinforcement learning to reason using chain of thought. These advanced reasoning capabilities provide new avenues for improving the safety and robustness of our models. In particular, our models can reason about our safety policies in context when responding to potentially unsafe prompts, through deliberative alignment. This leads to state-of-the-art performance on certain benchmarks for risks such as generating illicit advice, choosing stereotyped responses, and succumbing to known jailbreaks. Training models to incorporate a chain of thought before answering has the potential to unlock substantial benefits, while also increasing potential risks that stem from heightened intelligence. Our results underscore the need for building robust alignment methods, extensively stress-testing their efficacy, and maintaining meticulous risk management protocols. This report outlines the safety work carried out for the OpenAI o1 and OpenAI o1-mini models, including safety evaluations, external red teaming, and Preparedness Framework evaluations.

cs.AI

PACO: Proxy-Task Alignment and Online Calibration for On-the-Fly Category Discovery

On-the-Fly Category Discovery (OCD) requires a model, trained on an offline support set, to recognize known classes while discovering new ones from an online streaming sequence. Existing methods focus heavily on offline training. They aim to learn discriminative representations on the support set so that novel classes can be separated at test time. However, their discovery mechanism at inference is typically reduced to a single threshold. We argue that this paradigm is fundamentally flawed as OCD is not a static classification problem, but a dynamic process. The model must continuously decide 1) whether a sample belongs to a known class, 2) matches an existing novel category, or 3) should initiate a new one. Moreover, prior methods treat the support set as fixed knowledge. They do not update their decision boundaries as new evidence arrives during inference. This leads to unstable and inconsistent category formation. Our experiments confirm these issues. With properly calibrated and adaptive thresholds, substantial improvements can be achieved, even without changing the representation. Motivated by this, we propose PACO, a support-set-calibrated, tree-structured online decision framework. The framework models inference as a sequence of hierarchical decisions, including known-class routing, birth-aware novel assignment, and attach-versus-create operations over a dynamic prototype memory. Furthermore, we simulate the proxy discovery process to initialize the thresholds during offline training to align with inference. Thresholds are continuously updated during inference using mature novel prototypes. Importantly, PACO requires no heavy training and no dataset-specific tuning. It can be directly integrated into existing OCD pipelines as an inference-time module. Extensive experiments show significant improvements over SOTA baselines across seven benchmarks.

cs.CV

AutonomyLens: A Self-Evolving Simulation-Based Testing Loop for Autonomous Systems

Software engineering practices for validating autonomous cyber-physical systems (e.g., Uncrewed Aerial Vehicles) remain fragmented across scenario design, simulation execution, and telemetry analysis, limiting traceability between requirements, tests, and evidence. This fragmentation reduces reproducibility, slows debugging and iteration, and hinders systematic assurance under complex and evolving environmental conditions. We present AutonomyLens, an LLM-driven framework that integrates scenario specification, simulation execution, and telemetry analysis into a unified validation workflow. AutonomyLens enables developers to translate high-level validation intent into executable, temporally evolving scenarios, automatically run simulations, and perform context-aware analysis of resulting system behavior. The framework introduces (i) a structured representation for mission-level scenarios, (ii) an automated execution pipeline, (iii) analysis mechanisms that align telemetry with scenario context to produce actionable insights, and (iv) counterfactual scenario generation that closes the loop by refining and synthesizing new test cases from observed failures. We describe the early-stage design of AutonomyLens, discuss key challenges in building integrated validation workflows for autonomous systems, and outline how such an approach can improve traceability, reproducibility, and scalability in autonomy validation.

cs.SE

Sparse-RL: Breaking the Memory Wall in LLM Reinforcement Learning via Stable Sparse Rollouts

Reinforcement Learning (RL) has become essential for eliciting complex reasoning capabilities in Large Language Models (LLMs). However, the substantial memory overhead of storing Key-Value (KV) caches during long-horizon rollouts acts as a critical bottleneck, often prohibiting efficient training on limited hardware. While existing KV compression techniques offer a remedy for inference, directly applying them to RL training induces a severe policy mismatch, leading to catastrophic performance collapse. To address this, we introduce Sparse-RL empowers stable RL training under sparse rollouts. We show that instability arises from a fundamental policy mismatch among the dense old policy, the sparse sampler policy, and the learner policy. To mitigate this issue, Sparse-RL incorporates Sparsity-Aware Rejection Sampling and Importance-based Reweighting to correct the off-policy bias introduced by compression-induced information loss. Experimental results show that Sparse-RL reduces rollout overhead compared to dense baselines while preserving the performance. Furthermore, Sparse-RL inherently implements sparsity-aware training, significantly enhancing model robustness during sparse inference deployment. The corresponding training data and code are publicly available on the repository.

cs.LG

Learning through Creation: A Hash-Free Framework for On-the-Fly Category Discovery

On-the-Fly Category Discovery (OCD) aims to recognize known classes while simultaneously discovering emerging novel categories during inference, using supervision only from known classes during offline training. Existing approaches rely either on fixed label supervision or on diffusion-based augmentations to enhance the backbone, yet none of them explicitly train the model to perform the discovery task required at test time. It is fundamentally unreasonable to expect a model optimized on limited labeled data to carry out a qualitatively different discovery objective during inference. This mismatch creates a clear optimization misalignment between the offline learning stage and the online discovery stage. In addition, prior methods often depend on hash-based encodings or severe feature compression, which further limits representational capacity. To address these issues, we propose Learning through Creation (LTC), a fully feature-based and hash-free framework that injects novel-category awareness directly into offline learning. At its core is a lightweight, online pseudo-unknown generator driven by kernel-energy minimization and entropy maximization (MKEE). Unlike previous methods that generate synthetic samples once before training, our generator evolves jointly with the model dynamics and synthesizes pseudo-novel instances on the fly at negligible cost. These samples are incorporated through a dual max-margin objective with adaptive thresholding, strengthening the model's ability to delineate and detect unknown regions through explicit creation. Extensive experiments across seven benchmarks show that LTC consistently outperforms prior work, achieving improvements ranging from 1.5 percent to 13.1 percent in all-class accuracy. The code is available at https://github.com/brandinzhang/LTC

cs.CV