HyperSkill: Training-Free Omnimodal GRPO via Hypergraph-Indexed Skill-Library Evolution
Abstract
Reinforcement learning sharpens LLM reasoning, but parameter updates are expensive, forgetful, and opaque; training-free libraries swap gradients for a textual store yet inherit trajectory-level admission and flat structure. We close these gaps with HyperSkill, a Training-Free Omnimodal GRPO framework that keeps the backbone frozen and replaces gradient updates with auditable edits to a Hypergraph-Indexed Skill-Library, where skills are advantage-bearing nodes and compositions are hyperedges. Three mechanisms drive its evolution: a counterfactually verified, step-level distillation admits new skills; moving-average updates with hyperedge growth and pruning follow retrieval feedback; and an advantage-weighted retriever with similarity and capacity gates falls back to zero-skill. Across 11 text/VL/omnimodal benchmarks, HyperSkill beats every baseline by 6.9 F1 (15.6%) without touching a parameter. Our software and data are publicly available.