Language-Grounded into Model-Free 6D Pose Estimation for unseen Object Grasping
Abstract
Determining the precise position and orientation of an object in 3D space, known as 6D pose estimation, is a fundamental capability for robotic manipulation and augmented reality. Existing approaches typically require pre-built, object-specific models or pre-annotated segmentation masks, limiting their applicability to unseen objects. We propose an integrated pipeline that removes those dependencies by combining a locally hosted large language model, served via Ollama, that grounds free-form user instructions into structured object queries, with an open-vocabulary segmentation model capable of detecting and masking arbitrary objects from those queries without prior training on the target categories, and a model-free pose estimation framework that reconstructs object geometry from a single reference image and aligns it to recover the full 6D pose. Given only a reference image and a natural-language instruction, our system automatically interprets the user's intent, detects, segments, reconstructs, and localises any object in 6D without requiring any 3D scan, CAD model, or manual annotation. Experimental evaluation demonstrated that the pipeline achieved performance on the benchmark Hand-Object 3D dataset.