Leveraging Psychophysical Attentional Distribution for Gaze-Augmented Reward Modeling
Abstract
Integrating human feedback into language models has gained increasing attention as a means to align model outputs with human preferences, with Reinforcement Learning from Human Feedback (RLHF) serving as a prominent framework for preference-based alignment. Recent studies have incorporated eye-tracking (ET) data as an additional supervisory signal for reward model training, grounding reward learning on human reading behavior. Building on this approach, we extend prior fixation-centric models by leveraging established findings from cognitive psychology on human letter recognition. Specifically, we model visual attention as a spatially graded distribution centered on the current gaze location, and incorporate this representation into reward model training. Our results show that the reward model trained with the proposed gaze distribution achieves higher preference prediction accuracy over a baseline model. Moreover, the fitted attentional distribution reflects key properties of human attention during reading, suggesting that it captures cognitively meaningful aspects of attention allocation.