Learning Spatiotemporal Sensitivity in Video LLMs via Counterfactual Reinforcement Learning
Video large language models (Video LLMs) can achieve strong video-QA accuracy without reliably tracking spatiotemporal dynamics. A model may answer a motion question from static cues, for example, and give the same prediction even after the underlying motion is reversed. Correctness-based reinforcement learning does not directly address this problem because it rewards the final answer without requiring the policy to respond to task-relevant video dynamics. We propose Counterfactual Relational Policy Optimization (CRPO), which explicitly trains Video LLMs to respond to controlled changes in the visual input. For each training example, CRPO constructs a counterfactual video, such as a horizontally flipped or temporally reversed version of the original, and jointly optimizes rollouts from both videos under a shared policy. Factual supervision anchors what the model should answer, while counterfactual supervision constrains when that answer should change: predictions should change when an intervention alters task-relevant dynamics and remain stable when the queried property is preserved. This coupling provides a direct behavioral learning signal without requiring ground-truth labels for transformed videos or annotated spatiotemporal reasoning traces, while discouraging indiscriminate answer changes. To evaluate this property, we introduce DyBench, a paired counterfactual benchmark with 3{,}014 videos and a strict pair-accuracy metric. Across paired spatiotemporal evaluations and standard video benchmarks, CRPO improves sensitivity to motion and temporal changes while improving performance on general video understanding. The gains also extend to segment reordering, a transformation never used during training, suggesting that CRPO learns sensitivity to video dynamics beyond the training interventions. The project website can be found at https://ddz16.github.io/crpo.github.io/ .