RECIPE: Procedural Planning via Grounding in Instructional Video
Abstract
Visual planning asks a model to generate the remaining steps of a procedure in natural language given a partial video context and a goal. Progress on this task is bottlenecked by annotation: clean labeled datasets are small, domain-narrow, and encode a single execution trajectory per example, even though many valid orderings often exist. Large-scale instructional video corpora such as HowTo100M offer orders of magnitude more procedural content, but supervised fine-tuning on pseudo-labels extracted from their noisy ASR narrations fails: segmentation and alignment errors propagate into training, and the resulting supervision is still single-trajectory. We identify a key asymmetry. Extracting clean step labels from noisy video is hard, but verifying whether a generated step sequence is temporally grounded in ASR transcripts is comparatively cheap and scales to millions of videos via precomputed text embeddings. We exploit this asymmetry in RECIPE, which uses grounding quality as a reward signal for Group Relative Policy Optimization (GRPO), turning the noisy corpus into a verification signal rather than a labeling source. The framework applies uniformly to two input configurations of the planner: a Socratic pipeline in which a frozen vision-language model rewrites the video into a textual history fed to the planner, and a Video configuration in which the planner consumes video tokens directly. It applies equally to annotated and weakly supervised training regimes. We evaluate on seven procedural benchmarks using a reference-based LLM-as-judge protocol that scores generated plans across six procedural-quality criteria. RECIPE-RL improves over the base checkpoint at every scale we test (0.5B, 3B, 7B) and on every benchmark, with macro-accuracy gains of +7 to +8 points in-domain at every scale and up to +16 points zero-shot. It substantially outperforms supervised fine-tuning on both annotated and pseudo-labeled continuations (the latter actually degrades the base checkpoint), and is robust to fully replacing human annotations with VLM-derived pseudo-traces. Plugging RECIPE-RL into the proposal stage of VidAssist [14] improves over the strongest zero-shot baseline in our comparison at every horizon on the Visual Planning for Assistance benchmark, and a diversity analysis on COIN shows that RECIPE-RL preserves the generation variety that supervised fine-tuning collapses.