Enhancing VLM Reward Models Through Structure-Aware Fine-Tuning
Abstract
Designing effective reward functions remains a major bottleneck in Reinforcement Learning (RL). Recent work uses Vision-Language Models (VLMs) as reward models, replacing manual reward engineering with text-observation similarity. However, these rewards are often noisy and brittle, limiting their direct utility for policy learning. We present Structure-Aware Fine-Tuning (SAFT), a simple self-supervised method that refines VLM reward signals online without ground- truth supervision during fine-tuning. SAFT uses LoRA adapters to impose task- inherent structural priors on the VLM reward landscape. Across varying base model capabilities, SAFT denoises rewards, accelerates policy convergence, improves alignment with ground-truth rewards as measured by EPIC distance, and reduces the need for human preference annotation. These results suggest that VLM reward failures are not always purely semantic, but can also arise from structural brittleness, supporting task structure as a scalable supervision signal for text-conditioned RL and underscoring the broader value of incorporating task structure as a general inductive bias.