Search arXiv⌕ Search

arXiv · 2610.06877

When Can World Models Recover Physical Laws?

Abstract

Accurate prediction does not establish that a world model has recovered a physical law: distinct dynamics can generate identical records under the same observation protocol. We formulate law recovery on a fixed physical domain under an explicit catalog of experiments, sensor uncertainty, and an acquisition budget. A rate--distortion converse separates the information needed to describe a law from the information the apparatus can reveal. Its constructive counterpart gives a finite response codebook and an explicit decoding budget. On compact world classes, uniform recovery is possible exactly when every pair of different laws is experimentally distinguishable; equivalently, the apparatus can recover all the entropy of every finite law source. An inverse response modulus quantifies stability. For Lipschitz fields on a $d$-dimensional state--action domain, noisy full-state readouts after resets require minimax budget $Θ(\varepsilon^{-(d+4)/2})$ for squared law error $\varepsilon$, compared with $Θ(\varepsilon^{-(d+2)/2})$ for direct field observations. Exact crossing-time symmetries establish the lower bound under adaptive experiment selection and arbitrary durations with constant inputs. Reproducible synthetic cases illustrate the separate roles of intervention, calibration, and repeated measurement. Together, the results identify which evidence supports a claim of physical-law recovery and the cost of acquiring it.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Ye Yuan, Jun Liu. 2026-09-15. When Can World Models Recover Physical Laws?. https://arxiv.org/abs/2610.06877

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

Examples of statistical practice in elementary particle physics for comparison with Mayo's severity concept

The field of elementary particle physics (also known as high energy physics) has strong traditions in frequentist inference, heavily influenced by Fisher and by Neyman and Pearson. While I am not aware of direct use and citation of Deborah Mayo's severity concept in the field, one might expect that ideas similar to severity were independently developed. Here, I describe typical practice that addresses the same issues that severity was developed to address, with some tentative comparisons with severity. I invite experts on severity to explain how they would apply severity in these cases and to discuss any differences.

stat.OT↗

Generative AI in Publishing: An Editors' Panel on Ethics and Policies

Generative artificial intelligence (GenAI) is reshaping scholarly research faster than journals have developed stable norms for its use. This article presents an edited thematic account of a 2026 Joint Statistical Meetings panel that brought together editorial perspectives from mathematical statistics, data science, biomedical statistics, and general statistical scholarship. The discussion examines journal policies, disclosure, authorship and research integrity, peer-review confidentiality, researcher training, editorial workload, access, and possible future models of scholarly publishing. Panelists shared commitments to human accountability, the protection of confidential submissions, and disclosure of consequential assistance, while offering different recommendations on assistance with research ideas and proofs, disclosure requirements, automated review, and policy enforcement. By distinguishing shared principles from unresolved implementation questions, the article clarifies the choices facing statistical publishing and outlines an agenda for evaluating policies and practices as GenAI evolves. The account seeks to foster continued discussion of GenAI in scientific communication and encourage statistical organizations to develop more robust operational standards.

stat.OT↗

Data Quality Assessments: A Theoretically Structured Overview of Approaches and Methods

The quality of data is crucial for both practice and academia, and this holds for descriptive statistics, AI and advanced analytics alike. This study addresses an apparent gap in the literature and as such presents a full and theoretically grounded overview of data quality assessment approaches and methods, as well as their defining characteristics and inter-relationships. For this purpose a broad typology is introduced that employs two theoretical dimensions, namely the evaluation logic (formal versus informal) and the assessment driver (norms versus data). This yields four high-level approaches and 15 methods for assessing data quality. The research results are relevant for academia, as they provide a theory-based overview and definition of the ways that data quality can be evaluated. The study is also relevant for practice because it allows professionals to make informed decisions on using these methods, e.g. as part of an audit or broad data quality assessment strategy.

stat.OT↗