What Can Latent World Models Know? Physical Information in Multimodal Predictive Representations
A central premise of latent world models is that predicting the future encourages representations to internalize the physics of their environment. We ask which physical quantities are accessible in learned latent states, how this depends on training, and how those quantities relate to the model's predictions. We present PokeWorld, a simulated environment in which a robot finger pushes objects whose mass, drag, and contact stiffness vary across episodes while remaining visually identical. We first measure parameter recovery from raw observation sequences, then train matched action-conditioned world models. Prediction targets can strongly shape latent content. For example, contact stiffness becomes decodable when touch is predicted, while providing touch only as an input does not. Longer-horizon prediction improves estimation of position and velocity from learned representations relative to single-step prediction. Drag is recoverable from raw observations, but has weak linear readout from the learned latent states. Yet the models' glide forecasts depend systematically on drag. Experiments on RH20T reproduce the same input-target dependence on real multimodal robot data. Together, these results show that observations, prediction targets, and prediction horizons shape both which physical quantities are accessible in latent states and how they affect future predictions.