GeoFidelity-Bench: Evaluating Block-Conditioned Geographic Fidelity of Street-View Generation
Abstract
Text-to-image models can generate plausible street scenes from prompts such as “a street in New York City,” but it remains unclear whether these images resemble the requested place or merely reflect a generic city stereotype. Existing evaluation metrics mainly measure image quality or text alignment, and do not directly assess geographic fidelity. We introduce GeoFidelity-Bench, a benchmark for evaluating whether generated street-view images match a target location at the level of named street blocks. The benchmark contains 112 street blocks from 25 cities across six continents, paired with 7563 curated Mapillary reference images and geographic metadata. Its block-level design allows us to compare city-only prompts, prompts with street and neighborhood names, and prompts that additionally include raw GPS coordinates. Across six open-source text-to-image models, adding named local text improves similarity to the target location, while raw GPS coordinates add only a small and uneven gain. Same-city control prompts further show that the improvement is not explained by prompt length alone: corrupting local tokens reduces retrieval performance, with stronger evidence for neighborhood-level or combined local information than for the street token alone. The benchmark also shows that current generators remain far from matching real local street-view variation. GeoFidelity-Bench provides a controlled testbed for studying location-conditioned generation and for measuring progress beyond generic city-level plausibility.