Figure 1: KilometerVision — the first benchmark probing city-scale spatial understanding from real-world videos spanning 1km distances. Left: Real-world hour-long video of a walking tour in Istanbul and relevant questions. Centre: Grounding the video on the map instantly reveals loop closures, distances between landmarks, or sequences of streets visited. Right: Using the landmark-route-map paradigm from cognitive science, we define tasks to comprehensively evaluate visual city-scale spatial intelligence in video-language models.
We push the frontier of large-scale spatial intelligence in Vision-Language Models (VLMs) and introduce the first benchmark that probes geographical layout understanding from real-world videos, spanning up to 1km distances. Inspired by the cognitive science literature, we evaluate models against the hierarchical stages of human spatial awareness: anchoring via landmarks, connecting them through routes, and integrating these into global mental maps. Extensive experiments reveal a fundamental divergence in how current AI models process spatial information. Instead of utilising true path integration or forming geometric survey knowledge, we find that VLMs rely almost entirely on 2D visual recognition and text-matching to bypass complex spatial reasoning. The benchmark will be made publicly available.
Our benchmark evaluates VLMs across three cognitive stages of spatial awareness based on 288 hours of real-world YouTube Walking Tour videos grounded via Google's Visual Positioning System (VPS):
We evaluated state-of-the-art models including Gemini 2.5 Flash, PLM-8B, Claude Opus, and GPT-5 against a human baseline. Our experiments reveal:
@inproceedings{mahendran2026kilometer,
title={Kilometer Vision: A New Frontier for Large-Scale Spatial Intelligence in VLMS},
author={Mahendran, Aravindh and King, Michael and Grimes, Matthew Koichi and Yang, Antoine and Zhu, Tyler and Heyward, Joseph and Han, Tengda and Ginosar, Shiry and Sun, Chen and Damen, Dima and Osindero, Simon and Snavely, Noah and Lynen, Simon and Carreira, Jo{\~a}o and P{\v{a}}tr{\v{a}}ucean, Viorica},
booktitle={European Conference on Computer Vision (ECCV)},
year={2026}
}