Timezone: »
Vision-and-Language Navigation (VLN) requires an agent to navigate in a real-world environment following natural language instructions. From both the textual and visual perspectives, we find that the relationships among the scene, its objects, and directional cues are essential for the agent to interpret complex instructions and correctly perceive the environment. To capture and utilize the relationships, we propose a novel Language and Visual Entity Relationship Graph for modelling the inter-modal relationships between text and vision, and the intra-modal relationships among visual entities. We propose a message passing algorithm for propagating information between language elements and visual entities in the graph, which we then combine to determine the next action to take. Experiments show that by taking advantage of the relationships we are able to improve over state-of-the-art. On the Room-to-Room (R2R) benchmark, our method achieves the new best performance on the test unseen split with success rate weighted by path length of 52%. On the Room-for-Room (R4R) dataset, our method significantly improves the previous best from 13% to 34% on the success weighted by normalized dynamic time warping.
Author Information
Yicong Hong (Australian National University)
Cristian Rodriguez (Australian National University)
Yuankai Qi (University of Adelaide )
Qi Wu (University of Adelaide)
Stephen Gould (ANU)
More from the Same Authors
-
2023 Poster: Revisiting Implicit Differentiation for Learning Problems in Optimal Control »
Ming Xu · Timothy Molloy · Stephen Gould -
2023 Poster: LoRA: A Logical Reasoning Augmented Dataset for Visual Question Answering »
Jingying Gao · Qi Wu · Alan Blair · Maurice Pagnucco -
2022 Spotlight: Lightning Talks 6B-2 »
Alexander Korotin · Jinyuan Jia · Weijian Deng · Shi Feng · Maying Shen · Denizalp Goktas · Fang-Yi Yu · Alexander Kolesov · Sadie Zhao · Stephen Gould · Hongxu Yin · Wenjie Qu · Liang Zheng · Evgeny Burnaev · Amy Greenwald · Neil Gong · Pavlo Molchanov · Yiling Chen · Lei Mao · Jianna Liu · Jose M. Alvarez -
2022 Spotlight: On the Strong Correlation Between Model Invariance and Generalization »
Weijian Deng · Stephen Gould · Liang Zheng -
2022 Poster: Learning Distinct and Representative Modes for Image Captioning »
Qi Chen · Chaorui Deng · Qi Wu -
2022 Poster: On the Strong Correlation Between Model Invariance and Generalization »
Weijian Deng · Stephen Gould · Liang Zheng -
2021 Poster: Landmark-RxR: Solving Vision-and-Language Navigation with Fine-Grained Alignment Supervision »
Keji He · Yan Huang · Qi Wu · Jianhua Yang · Dong An · Shuanglin Sima · Liang Wang -
2021 Poster: Debiased Visual Question Answering from Feature and Sample Perspectives »
Zhiquan Wen · Guanghui Xu · Mingkui Tan · Qingyao Wu · Qi Wu -
2021 Poster: Rethinking conditional GAN training: An approach using geometrically structured latent manifolds »
Sameera Ramasinghe · Moshiur Farazi · Salman H Khan · Nick Barnes · Stephen Gould -
2018 Poster: Partially-Supervised Image Captioning »
Peter Anderson · Stephen Gould · Mark Johnson -
2009 Poster: Region-based Segmentation and Object Detection »
Stephen Gould · Tianshi Gao · Daphne Koller -
2009 Spotlight: Region-based Segmentation and Object Detection »
Stephen Gould · Tianshi Gao · Daphne Koller -
2008 Oral: Cascaded Classification Models: Combining Models for Holistic Scene Understanding »
Geremy Heitz · Stephen Gould · Ashutosh Saxena · Daphne Koller -
2008 Poster: Cascaded Classification Models: Combining Models for Holistic Scene Understanding »
Geremy Heitz · Stephen Gould · Ashutosh Saxena · Daphne Koller -
2008 Poster: Learning Bounded Treewidth Bayesian Networks »
Gal Elidan · Stephen Gould -
2008 Demonstration: High-Accuracy 3D Sensing for Mobile Manipulators »
Stephen Gould · Morgan Quigley · Siddarth Batra · Ellen Klingbiel · Quoc V Le · Andrew Y Ng -
2008 Spotlight: Learning Bounded Treewidth Bayesian Networks »
Gal Elidan · Stephen Gould -
2007 Demonstration: Holistic Scene Understanding from Visual and Range Data »
Stephen Gould · Morgan Quigley · Andrew Y Ng · Daphne Koller -
2006 Demonstration: Peripheral-Foveal Vision for Real-time Object Recognition »
Benjamin Sapp · Stephen Gould · Adrian Kaehler · Gary R Bradski · Andrew Y Ng