FROM OBSERVATION TO PREDICTION

SpatialMindFrom Sight to Foresight:
Predictive Spatial Reasoning
in Vision-Language Models

1 University of Illinois at Chicago2 Bosch

† This work was done during an internship at Bosch.‡ Project lead.

Paper, code & data coming soon

TL;DR Ground spatial reasoning in metric geometry, understand observed dynamics, and anticipate what comes next.

SpatialMind reasons from a current distance estimate through observed motion to future distance prediction. A driving example and an eight-task benchmark comparison illustrate the approach.
From sight to foresight. SpatialMind connects current spatial states, observed dynamics, and future prediction in a shared reasoning framework.

Abstract

Predicting future spatial states supports collision avoidance and timely decision-making in dynamic environments. However, existing vision-language models (VLMs) and benchmarks for spatial reasoning primarily focus on observed scenes, leaving predictive spatial reasoning beyond the observed interval underexplored. To this end, we introduce SpatialMind, a metric-scale VLM for spatial reasoning and future prediction. Its metric depth adapter anchors spatial reasoning to real-world scale, while its progressive state chain establishes current spatial states and observed dynamics as the foundation for future prediction.

Given a video prefix, SpatialMind predicts distances, motion directions, and spatial relations in both observed and unseen future frames. For training and evaluation, we build a scalable data engine that grounds entity descriptions in metric geometry to generate question-answer pairs and state supervision. Using this engine, we construct the SpatialMind-30K dataset and the SpatialMind-2K benchmark, both covering driving and everyday egocentric scenes. The benchmark spans eight tasks across three levels: current-state understanding, observed-dynamics understanding, and future prediction. Experiments show that SpatialMind substantially outperforms both general and spatially specialized models on our benchmark while achieving competitive zero-shot performance on VSI-Bench, OSI-Bench, and VLM4D.

  1. L1

    Current state

    Where are things now?

    Motion & spatial state
  2. L2

    Observed dynamics

    How are things changing?

    Motion distance & relative motion
  3. L3

    Future prediction

    What happens next?

    Future direction, distance & relation

Grounded in geometry.
Reasoning through time.

An observed RGB video and a spatial question are the starting point. Metric geometry provides scale; a progressive state chain connects the evidence to the answer.

SpatialMind architecture with metric depth refinement, separate geometry and motion encoders, anchor time alignment, visual-geometry fusion, and progressive state-chain reasoning.
Model architecture. Global and local depth corrections, temporal geometry fusion, and structured reasoning connect video observations to metric spatial predictions.
01 Metric depth adaptation
Refine depth estimates using multiview geometry, with global alignment and local corrections that anchor spatial features to real-world scale.
02 Temporal geometry fusion
Encode geometry and camera motion separately, align their features in time, and fuse them into the RGB visual tokens through cross-attention.
03 Progressive state chains
Ground the target, establish its current state, and reason through observed dynamics to a future prediction. Future frames are withheld from model inputs.

Eight tasks. Two worlds.

Driving scenes and everyday egocentric activities, organized into a common hierarchy of spatial reasoning.

30,677Training QA pairs
2,000Benchmark questions
8Tasks across 3 levels
The SpatialMind data engine links target selection and verified entity descriptions to metric geometry, temporal spatial metadata, and question-answer generation across eight tasks.
SpatialMind data engine. A Select–Describe–Verify process grounds natural-language descriptions in metric geometry. Semantic propagation and temporal alignment support question generation and state-chain supervision from Waymo and Ego-Exo4D videos.

Results

SpatialMind improves predictive spatial reasoning on SpatialMind-2K and transfers to three external benchmarks without additional fine-tuning.

52.65% overall score

+6.95 pp over Qwen3-VL-8B + SFT

Table 1. Results on SpatialMind-Benchmark (%)
ModelOverallL1 · Current stateL2 · Observed dynamicsL3 · Future prediction
MotionSpatialEgo
distance
Object
distance
Relative
motion
DirectionDistanceRelation
Proprietary Models (API)
GPT-5.6 Luna22.6020.2814.3412.5925.8744.7627.2714.0425.61
Claude Sonnet 527.8016.4323.089.7926.5749.3030.0723.5144.21
Gemini 3.8 Flash30.1921.6218.6619.3949.6552.8520.2841.3422.91
GPT-6-Astra (medium)39.0038.4613.9919.5851.7573.7839.8629.1252.28
General-purpose VLMs
Qwen2.5-VL-3B14.405.955.2510.852.8032.2028.7017.8513.00
Qwen3.5-9B10.3010.859.455.6013.309.454.2020.358.05
LLaVA-NeXT-7B14.405.256.655.957.0030.1031.5015.0518.60
Qwen2.5-VL-7B17.603.8517.152.4530.8033.6025.9010.1527.70
Cosmos3-8B18.955.602.409.8032.9035.3032.203.9043.20
Qwen3-VL-8B20.203.1513.307.3523.1038.5024.5021.0034.70
Spatially specialized VLMs
SR-3D-8B9.602.456.3012.602.1024.8514.708.753.50
SpatialRGPT-8B9.706.655.255.252.8017.8534.303.8510.20
SpatialBot-3B11.802.807.701.0534.3023.4528.702.8013.00
VLM-3R-7B14.007.901.054.207.7031.5029.402.4532.65
SpaceR-7B14.705.6014.705.2514.7033.2521.703.5022.45
ViLaSR-7B14.807.709.107.007.0032.2016.806.0030.15
SR-ReaL-8B14.905.9511.5511.908.4031.5032.207.0516.15
4D-RGPT-8B15.403.5011.9012.957.0033.9527.3012.2516.15
SpatialReasoner-7B15.807.207.705.6012.6033.2530.108.4027.00
SFT on SpatialMind training data
Qwen3-VL-8B + SFT45.7036.0257.3423.7737.0656.9957.3440.6457.89
SpatialMind (Ours)52.6538.3265.4630.1552.3468.4265.1342.8964.48

Within each model group, bold marks the best result and underline marks the second best. Numerical tasks use mean relative accuracy (MRA); multiple-choice tasks use accuracy (ACC).

Scroll horizontally to view all eight tasks.

A CLOSER LOOK · SECTION 5.1

Three observations

Different model families. A shared challenge: reasoning at real-world scale.

  1. 01Proprietary Models

    Astra sets the pace.

    39.00%

    GPT-6-Astra (medium) · Overall

    GPT-6-Astra leads proprietary models on six of eight tasks. The two exceptions are current-state spatial estimation and future-distance prediction, both requiring absolute metric distances.

  2. 02General-Purpose Models

    Categories come easier.

    20.20%

    Qwen3-VL-8B · Best overall in this group

    Direction and relation questions are easier than metric estimation. Current-state measurements and ego-motion distance remain challenging: categorical future accuracy does not establish metric understanding.

  3. 03Spatial Expert Models

    Geometry alone falls short.

    9.60–15.80%

    Overall range across spatial experts

    Depth and 3D priors still leave weaknesses in current-state measurements, ego-motion distance, and future-distance prediction. Geometric priors alone do not ensure metric accuracy across driving and egocentric scenes.

BibTeX

SpatialMind
@misc{wang2026spatialmind,
  title={From Sight to Foresight: Predictive Spatial Reasoning
         in Vision-Language Models},
  author={Wang, Feiran and Wang, Xiaoqi and Li, Ziwei and
          He, Wenbin and Yan, Yan and Ren, Liu},
  year={2026},
  url={https://brack-wang.github.io/spatialmind/}
}

Figure preview