KilometerVision: A New Frontier for Large-Scale Spatial Intelligence in VLMs

Accepted at ECCV 2026
Aravindh Mahendran1*, Michael King1*, Matthew Koichi Grimes1*, Antoine Yang1, Tyler Zhu2,
Joseph Heyward1, Tengda Han1, Shiry Ginosar3, Chen Sun1, Dima Damen1,
Simon Osindero1, Noah Snavely1, Simon Lynen4, João Carreira1, Viorica Pătrăucean1✉
1Google DeepMind     2Princeton University     3Toyota Technological Institute at Chicago     4Google
*core contributors     corresponding author    
📄 Paper (PDF) 💻 Benchmark (eval.ai challenge)
KilometerVision Teaser Figure

Figure 1: KilometerVision — the first benchmark probing city-scale spatial understanding from real-world videos spanning 1km distances. Left: Real-world hour-long video of a walking tour in Istanbul and relevant questions. Centre: Grounding the video on the map instantly reveals loop closures, distances between landmarks, or sequences of streets visited. Right: Using the landmark-route-map paradigm from cognitive science, we define tasks to comprehensively evaluate visual city-scale spatial intelligence in video-language models.

Abstract

We push the frontier of large-scale spatial intelligence in Vision-Language Models (VLMs) and introduce the first benchmark that probes geographical layout understanding from real-world videos, spanning up to 1km distances. Inspired by the cognitive science literature, we evaluate models against the hierarchical stages of human spatial awareness: anchoring via landmarks, connecting them through routes, and integrating these into global mental maps. Extensive experiments reveal a fundamental divergence in how current AI models process spatial information. Instead of utilising true path integration or forming geometric survey knowledge, we find that VLMs rely almost entirely on 2D visual recognition and text-matching to bypass complex spatial reasoning. The benchmark will be made publicly available.

The Landmark-Route-Map Paradigm

Our benchmark evaluates VLMs across three cognitive stages of spatial awareness based on 288 hours of real-world YouTube Walking Tour videos grounded via Google's Visual Positioning System (VPS):

Key Findings

We evaluated state-of-the-art models including Gemini 2.5 Flash, PLM-8B, Claude Opus, and GPT-5 against a human baseline. Our experiments reveal:

Citation

@inproceedings{mahendran2026kilometer,
  title={Kilometer Vision: A New Frontier for Large-Scale Spatial Intelligence in VLMS},
  author={Mahendran, Aravindh and King, Michael and Grimes, Matthew Koichi and Yang, Antoine and Zhu, Tyler and Heyward, Joseph and Han, Tengda and Ginosar, Shiry and Sun, Chen and Damen, Dima and Osindero, Simon and Snavely, Noah and Lynen, Simon and Carreira, Jo{\~a}o and P{\v{a}}tr{\v{a}}ucean, Viorica},
  booktitle={European Conference on Computer Vision (ECCV)},
  year={2026}
}