What Post-Training Must Preserve: Outcome-Stable, Force-Variable Interpretation
Manodyna K H ⋅ Samarth Ramesh
Abstract
We study a target for post-training under changing language contexts: preserve what form fixes while allowing discourse to change what discourse should change. Outcome-Role-Attitude separates these two requirements. In context-controlled minimal pairs, the same utterance keeps a stable outcome space while role assignment and interactional force vary. Layerwise analysis shows corresponding trajectory divergence. We do not introduce a new post-training algorithm; instead, the paper provides an empirical and formal diagnostic for adaptation objectives that must distinguish legitimate context sensitivity from unwanted instability.
Chat is not available.
Successful Page Load