Synthetic Web: Benchmarking Language Agents under Adversarial Search Ranking
Abstract
Language agents increasingly act as web-enabled systems that search, browse, and synthesize information from diverse sources. However, these sources can include unreliable or adversarial content, and the robustness of agents to adversarial ranking remains poorly understood. Existing benchmarks evaluate functional navigation or static factuality but cannot causally isolate this vulnerability and often confound retrieval-time reasoning with memorized knowledge. We introduce Synthetic Web Benchmark, a controlled environment of procedurally generated web ecosystems designed to evaluate retrieval-time reasoning and source criticism. The benchmark comprises thousands of hyperlinked articles with ground-truth labels, process-level interaction traces, and contamination filtering to ensure that answers cannot be recovered from pretraining alone. By injecting a single high-plausibility misinformation article at a specified rank, we measure the causal effect of adversarial exposure under minimal intervention. Across six frontier models and thousands of evaluation instances, we observe catastrophic failures: accuracy collapses despite unrestricted access to truthful evidence, accompanied by limited search escalation, weak cross-source synthesis, and severe miscalibration. These results show that current agents struggle to arbitrate conflicting sources even when sufficient evidence is available, revealing fundamental limitations in retrieval-based reasoning. The benchmark provides a reproducible testbed for studying epistemic robustness and developing more reliable web agents.