WATERFALL: Workflow for Adaptive Training with Evolutionary Reward Formulation and Automated Learning Loops
Abstract
Reward design remains a central bottleneck in reinforcement learning (RL), particularly in sparse-reward, long-horizon, and partially observable settings requiring sequential dependency chaining, memory, and deceptive affordance disambiguation. Recent foundation-model approaches reduce manual reward engineering; however, most assume explicit goal-conditioning, privileged task descriptions, expert demonstrations, or access to ground-truth metrics. These assumptions are particularly problematic in partially observable environments, where task-relevant objectives, affordances, and dynamics must be discovered through interaction rather than disclosed a priori. We formalise this problem setting by introducing WATERFALL, an iterative population-based workflow for automated reward discovery operating under the strict observability constraints inherent in POMDP environments. WATERFALL does so without access to semantic task descriptors, pre-disclosed environment dynamics, goal conditioning, or ground-truth metrics. WATERFALL achieves this by (i) synthesising diverse programmatic reward candidates via persona-conditioned generation, (ii) evaluating candidate behaviour from raw visual rollouts utilising a Swiss-system evolutionary tournament workflow judged by vision-language models to identify elites, and (iii) leveraging longitudinal assessment histories to iteratively mutate, escalate, or prune candidates. We evaluate WATERFALL across fully observable and reveal-gated partially observable environments against sparse-reward RL, intrinsic-motivation baselines, as well as state-of-the-art foundation-model reward-synthesis baselines. Our results indicate that WATERFALL can synthesise, evaluate, and iteratively refine reward programs without access to privileged semantic information or ground-truth scalar supervision. Enforcing this strict information contract is especially important in the intended reveal-gated POMDP regime.