PhysGraphNet: Physical-State Scene Graphs via Latent Graph Reasoning and Counterfactual Supervision
Zhengtao Yao ⋅ Runhao Li ⋅ Yan Wen ⋅ Guang Yang ⋅ Siheng Wang ⋅ Chenhao Wei ⋅ Rongchao Zhang ⋅ Guoqing Ma ⋅ Haoyan Xu ⋅ Junhao Dong
Abstract
Manipulation planners depend on accurate predicates about physical state---is an object occluded? is a container open? is a path blocked?---but such predicates are hard to read off an image. Contrastively-trained vision--language models capture semantic content but collapse to near-zero binary F1 on these predicates, and even GPT-4o few-shot falls below random on multiple-choice physical-state questions. We present \ours, which predicts physical-state scene graphs from a single image and a natural-language goal. A frozen CLIP ViT, adapted with LoRA, feeds a heterogeneous graph of object, relation, and learnable memory tokens into a Latent Graph Reasoning Transformer (\lgrt); training combines supervised, counterfactual-margin, contrastive, and calibration losses with offline counterfactual-pair augmentation. On \pacbench, \ours reaches $\mu$-F1$\,{=}\,0.651$ and $99.3\%$ MCQ accuracy, a $+0.170$ absolute $\mu$-F1 gain over the strongest no-graph supervised baseline, with $99.6\%$ counterfactual directional consistency. Component ablations attribute $+0.426$ $\mu$-F1 to the \lgrt and $+0.392$ to the counterfactual loss. A \manipbench-trained variant scores $74.4\%$ MCQ against $30.0\%$ for the best zero-shot VLM on that benchmark's 4-choice format.
Chat is not available.
Successful Page Load