Visually Grounded Reinforcement Learning for LLM Programmatic Animation Generation in Interactive Execution Environments
Abstract
Generating programmatic animations with libraries such as Manim poses unique challenges for large language models (LLMs), requiring spatial reasoning, temporal sequencing, and adaptation to interactive execution environments. However, how different post-training objectives influence an LLM's capacity to leverage dynamic feedback during generation remains under-researched. This study introduces ManimTrainer, a framework that combines Supervised Fine-Tuning (SFT) with visually grounded Reinforcement Learning via Group Relative Policy Optimisation (GRPO) using a unified code-and-visual reward signal, alongside ManimAgent, an inference pipeline featuring Renderer-in-the-loop (RITL) and documentation-augmented execution feedback. Evaluating 17 open-source sub-30B LLMs across nine combinations of training and inference strategies on ManimBench, this study presents a unified study investigating how training types interact with test-time execution feedback. The results show that SFT and GRPO improve visual quality to a comparable degree under vanilla inference, while GRPO also preserves the policy's responsiveness to extrinsic error signals during self-correction. Under interactive RITL feedback, the ordering is consistent across inference types, with GRPO ahead of the base and SFT models by a median of 2.7 pp and 1.3 pp across the 17 models. Additionally, the analysis shows that the correlation between code and visual metrics strengthens with SFT and GRPO but weakens under inference-time execution loops, highlighting the complementary roles of training objectives and interactive environments in programmatic generation.