Axiomatic World Modeling for Physics Reasoning
Xinye Yang ⋅ Zhenyang Liu ⋅ Yuxuan Wang ⋅ Yuanyuan Lei
Abstract
Recent research-level physics-reasoning benchmarks for LLMs reveal a failure mode that lies not in knowledge or execution, but in modeling the world being reasoned about. When faced with unseen problem setups, models often silently alter given quantities, fabricate laws, or inject unstated assumptions to force the problem into familiar patterns, and subsequent reasoning then proceeds within this corrupted world. We define this premise-level divergence as modeling drift: the gap between a model's internal axiomatic world model (AWM) and the true physical world specified in the problem. Although RLVR and PRM have yielded substantial gains in mathematics and code, their reward shapes do not directly supervise world modeling in axiomatic domains such as physics. We propose Reinforcement Learning with Axiomatic World Modeling (RLAWM), which optimizes the policy by penalizing divergence between its inferred axiomatic world model and the true physics world. RLAWM probes physics reasoning at two hierarchical levels (modeling and solving), each scored by a physically grounded reward. The modeling reward targets abstract knowledge formulation, evaluating whether the model captures the true axiomatic world before attempting calculation. The solving reward targets instantiated knowledge reasoning, checking whether the model's concrete derivation remains consistent with the axiomatic world. This two-level structure separates what world the model believes it is reasoning in from how it reasons within that world. The framework operates exclusively during training, with no extra cost during inference. RLAWM improves performance by up to $+10.8$ points on PhysReason and $+5.2$ on PHYSICS over the baseline. RLAWM also exhibits strong zero-shot generalization, yielding a $2.45{\times}$ gain over the strongest baseline on unseen physics problems.
Chat is not available.
Successful Page Load