The 4th Perception Test Challenge at ECCV 2026
Do large multimodal models truly understand the spatial structure of the world they perceive? Given that they are largely trained on passive, internet-scale data, how do they represent environments across scales, from the tabletop in a tea preparation video to the city-scale layout of an hour-long walking tour?
Put your model to test and win prizes totalling 20K EUR!
NEW this year: KilometerVision track probing spatial intelligence in walking tour videos, including questions on distance estimation, landmark recognition, compass, map understanding.
NEW this year: KilometerAudio track probing multimodal audio-video understanding in hour-long walking tour videos.
Humans are remarkably good at navigation, but most studies of human navigation use sparse environments and simple tasks that do not fully assess naturalistic navigation. Here we sought to recover and characterize the human navigation network more completely. Participants first learned to navigate in a large virtual reality city, and then performed a taxi driver task in the MRI scanner. Brain activity was modeled in terms of 38 separate feature spaces spanning many different aspects of navigation. Results show that naturalistic navigation is supported by a network of 11 distinct regions distributed across visual, parietal, and prefrontal cortices, and organized into functional gradients that transform perception to motor responses. To understand how the human network differs from deep neural network (DNN) models for navigation, we modeled the brain data using features drawn from two navigational DNNs, Learning from All Vehicles and InterFuser. In both cases the perception and action modules explain substantial variance in the brain data, but the internal modules that transform perceptual inputs to action outputs explain little brain variance and they map inconsistently across participants. These results suggest that DNNs may not navigate as efficiently as humans because they do not use human-like algorithms or representations.
Abstract to be announced soon.
Abstract to be announced soon.