Advancing Affordance-Grounded Creative Tool Use in Large Multimodal Models
Abstract
Large multimodal models (LMMs) have advanced rapidly in perception and reasoning, yet it remains unclear whether they can discover visually grounded, physically feasible solutions in open-ended environments. We first introduce MM-CreativityBench, a benchmark for affordance-grounded creative tool use, where models must inspect scenes, entities, and parts to identify non-obvious object uses grounded in visual evidence. Our evaluation shows that current LMMs often fail to sustain grounded exploration: they overlook relevant entities, under-examine critical parts, or hallucinate unsupported attributes. To address these failures, we propose affordance-grounded alignment, framing creative tool use as a preference learning problem. Using Direct Preference Optimization and supervision from an affordance knowledge base, we train models to favor visually grounded attribute–affordance reasoning over hallucinated alternatives while improving exploration efficiency. Our results yield consistent gains in correct entity and part selection, and reduce hallucination and grounding errors. These findings position grounded creativity as a core capability for future multimodal agents: the ability to adapt to unfamiliar environments and solve problems beyond memorized patterns or surface-level plausibility, moving closer to real human-like intelligence.