Search arXiv⌕ Search

arXiv · 2609.30718

NEMSim: Learning Control-Conditioned Multi-Event Physical Dynamics via Executable Event-Mechanism Priors

Abstract

High-fidelity simulation of control-conditioned multi-event physical systems is computationally expensive, especially across broad control spaces and long trajectories. In these systems, macroscopic evolution emerges from localized discrete events whose intensities and effects depend on process controls and evolving local states, while the available system knowledge is typically expressed as event-attribute descriptions. Purely data-driven surrogates must infer these event effects from limited trajectory coverage, which can hinder generalization to unseen control regimes. Physics-guided methods instead primarily build on equation-level constraints or differentiable solvers rather than discrete event-rule priors. We therefore propose NEMSim (Neural Event-Mechanism Simulator), which compiles predefined event-attribute descriptions into an executable transition structure linking control-dependent event intensities, prior-guided mechanism attribution, and state-dependent responses. To enable evaluation of control-conditioned multi-event dynamics with explicit system knowledge, we construct a 3D KMC-based benchmark pairing high-fidelity trajectories with explicit event rules, standardized splits, and evaluation protocols. Across three settings, NEMSim reduces Avg. RMSE by 58.9%-81.3% relative to the strongest baseline in each setting. It also remains best in the data-efficiency study with training-data fractions down to 10%. Mechanism analyses further show that these gains arise from executable rule integration rather than prior access or architecture alone.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Junsong Yu, Junjie Xie, Pengwei Liu, Dong Ni. 2026-09-25. NEMSim: Learning Control-Conditioned Multi-Event Physical Dynamics via Executable Event-Mechanism Priors. https://arxiv.org/abs/2609.30718

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Sample-Efficient Optimization over Generative Priors via Coarse Learnability

We study zeroth-order optimization where solutions must minimize a cost $d(s)$ while maintaining high probability under a complex generative prior $L(s)$ (e.g., a parameterized model). This reduces to sampling from a target distribution proportional to $L(s) e^{-T \cdot d(s)}$. Since classical model-based optimization (MBO) lacks finite-sample guarantees for expressive approximate learners, we introduce "coarse learnability", a flexible statistical assumption requiring only that a learned model covers the target's probability mass within a polynomial factor. Leveraging this assumption, we design an iterative MBO algorithm called \alift with a sample correction step that provably approximates the target using only a polynomial number of samples. We apply this framework to globally optimizing non-convex objectives bounded by a quadratic envelope in $R^n$, where we show this assumption is naturally satisfied for a family of "optimistic" posterior distributions. To reach global $\varepsilon$-optimality, this implies a sample complexity of $\widetilde{O}(\log 1/\varepsilon)$, a rate characteristic of optimistic space-partitioning methods. We further justify coarse learnability as an assumption for generative priors theoretically, proving that in simple settings, parametric maximum likelihood estimation and over-smoothed kernel density estimators naturally satisfy it. Finally, one motivation for our framework comes from inference-time alignment. Though our primary contribution pertains to the theoretical foundations of MBO, we provide qualitative evidence that, in simple settings, even primitive LLMs can shift their distributions toward lower-cost regions when fine-tuned with zeroth-order feedback.

cs.LG↗

The Devil in the Details: Emergent Misalignment, Format and Coherence in Open-Weights LLMs

Prior work has shown that fine-tuning models on a narrow domain with misaligned data can lead to broad misalignment - a phenomenon termed "emergent misalignment" (Betley et al. 2025). While all tested models were susceptible to emergent misalignment, some models showed more resistance than others. Specifically the Qwen-2.5 family proved to be relatively resistant, while GPT-4o exhibited the strongest misalignment. In this paper we evaluate if current-generation open-weights models exhibit similar resistance to the Qwen-2.5 family and measure misalignment robustness over a range of model architectures and scales. We replicate the effect across nine modern open-weights models (Gemma 3 and Qwen 3 families, 1B-32B parameters). Models fine-tuned on insecure code generation show a 0.68% misalignment rate (compared to 0.07% for base models), matching the lower end of prior open-model results but dramatically lower than GPT-4o's 20%. We identify a critical format-dependent vulnerability: requiring JSON output doubles misalignment rates compared to natural language prompts (0.96% vs 0.42%). This suggests that structural constraints may bypass safety training by reducing the model's 'degrees of freedom' to refuse. These findings confirm emergent misalignment as a reproducible phenomenon in modern open-weights models, with rates substantially lower than observed in proprietary systems.

cs.LG↗

Shapley-based Data Valuation for LLM Alignment via Sequential Preference Optimization

Data valuation is a natural framework for understanding which data sources matter most when aligning a Large Language Model (LLM) from multiple sources. The standard game-theoretic approach treats each source, or equivalently each preference dataset, as a player in a cooperative game and assigns it a contribution score through the Shapley value. In practice, however, Shapley-based valuation is computationally prohibitive because it requires aligning a separate model for every possible coalition of sources, i.e., an exponential number of alignments. We address this challenge for Direct Alignment Algorithms (DAAs), including IPO, which learn through log-policy ratios with respect to a reference policy. We show that, when a model is aligned sequentially source by source, exact optimization makes each stage contribute additively to the log-probability of a full response, up to a prompt-dependent normalization constant. This allows the log-probability assigned by any coalition to a fixed response to be reconstructed from the base policy and the policies trained on each source individually. This reduces the alignment cost of Shapley-based valuation from exponential to linear, since only one model per source needs to be trained to evaluate coalition scores. We test whether this theoretical property remains approximately valid under finite training across several base models and real-world data sources. We finally compute the Shapley values of these sources under multiple reward models, showing how their estimated contributions vary across evaluation criteria.

cs.LG↗