HANSEL: From Web Agent Trajectories to Interactive Evidence
Abstract
AI web agents can perform complex, multi-step tasks such as searching for products, comparing options, and making purchases on behalf of users. However, verifying the correctness of an agent’s output remains difficult. Existing transparency mechanisms, such as full trajectory logs and screenshots, treat verification as a passive reading task, leaving users to sift through overwhelming logs or trust potentially unfaithful explanations. We present HANSEL (Highlighting Agent Navigation Steps as Evidence Links), a system that extracts interactive, verifiable evidence from web-agent trajectories. Given an agent trajectory, HANSEL extracts evidence pages and snippets that support the agent’s final answer. HANSEL further provides an interactive interface that reconstructs these evidence pages with relevant page states preserved (e.g., applied filters, search queries, and scroll positions), allowing users to directly inspect and interact with the supporting evidence. When the agent’s answer cannot be traced to any visited page, HANSEL explicitly flags this gap. A technical evaluation on 96 tasks from AssistantBench and Online-Mind2Web shows that HANSEL achieves 80.2% precision and 86.3% recall in identifying evidence pages, while reducing trajectory volume by 54.3%.