HierRR: Enhancing Instruction Alignment in Open-Vocabulary Indoor Scene Synthesis via Agentic Hierarchical Reasoning and Reflection
Abstract
Open-vocabulary indoor scene synthesis aims to generate plausible layouts from arbitrary user instructions while ensuring physical feasibility and semantic consistency. Existing methods directly infer spatial relations between objects from instructions and solve for layouts based on predefined rules, achieving progress in physical feasibility. However, they often struggle to precisely align with complex instructions due to insufficient comprehension of spatial relations and limited capability to translate them into layouts. In this paper, we propose a Hierarchical Reasoning and Reflection agent framework (HierRR), which leverages a relational hierarchy to enable multi-granularity spatial relation comprehension and attributable layout reasoning, improving semantic consistency while maintaining physical feasibility. Specifically, HierRR introduces the spatial transformation chain‑of‑thought reasoning module that explicitly models the mapping from spatial relations to executable geometric placements, mitigating semantic drift during reasoning. Furthermore, the semantic- and vision-guided reflection module is employed to iteratively refine spatial relations and layout reasoning from local to global along the hierarchy, effectively bridging the gap between language instructions and geometric layouts. Extensive comparative experiments and ablation studies demonstrate that HierRR generates more plausible and semantically consistent indoor scene layouts than existing methods, while generalizing effectively to complex scenes with numerous objects and diverse spatial constraints.