FROM OBSERVATION TO PREDICTION
SpatialMindFrom Sight to Foresight:
Predictive Spatial Reasoning
in Vision-Language Models
Paper, code & data coming soon
TL;DR Ground spatial reasoning in metric geometry, understand observed dynamics, and anticipate what comes next.
Abstract
Predicting future spatial states supports collision avoidance and timely decision-making in dynamic environments. However, existing vision-language models (VLMs) and benchmarks for spatial reasoning primarily focus on observed scenes, leaving predictive spatial reasoning beyond the observed interval underexplored. To this end, we introduce SpatialMind, a metric-scale VLM for spatial reasoning and future prediction. Its metric depth adapter anchors spatial reasoning to real-world scale, while its progressive state chain establishes current spatial states and observed dynamics as the foundation for future prediction.
Given a video prefix, SpatialMind predicts distances, motion directions, and spatial relations in both observed and unseen future frames. For training and evaluation, we build a scalable data engine that grounds entity descriptions in metric geometry to generate question-answer pairs and state supervision. Using this engine, we construct the SpatialMind-30K dataset and the SpatialMind-2K benchmark, both covering driving and everyday egocentric scenes. The benchmark spans eight tasks across three levels: current-state understanding, observed-dynamics understanding, and future prediction. Experiments show that SpatialMind substantially outperforms both general and spatially specialized models on our benchmark while achieving competitive zero-shot performance on VSI-Bench, OSI-Bench, and VLM4D.
- L1
Current state
Where are things now?
Motion & spatial state - L2
Observed dynamics
How are things changing?
Motion distance & relative motion - L3
Future prediction
What happens next?
Future direction, distance & relation
Grounded in geometry.
Reasoning through time.
An observed RGB video and a spatial question are the starting point. Metric geometry provides scale; a progressive state chain connects the evidence to the answer.
- 01 Metric depth adaptation
- Refine depth estimates using multiview geometry, with global alignment and local corrections that anchor spatial features to real-world scale.
- 02 Temporal geometry fusion
- Encode geometry and camera motion separately, align their features in time, and fuse them into the RGB visual tokens through cross-attention.
- 03 Progressive state chains
- Ground the target, establish its current state, and reason through observed dynamics to a future prediction. Future frames are withheld from model inputs.
Eight tasks. Two worlds.
Driving scenes and everyday egocentric activities, organized into a common hierarchy of spatial reasoning.
Results
SpatialMind improves predictive spatial reasoning on SpatialMind-2K and transfers to three external benchmarks without additional fine-tuning.
52.65% overall score
+6.95 pp over Qwen3-VL-8B + SFT
| Model | Overall | L1 · Current state | L2 · Observed dynamics | L3 · Future prediction | |||||
|---|---|---|---|---|---|---|---|---|---|
| Motion | Spatial | Ego distance | Object distance | Relative motion | Direction | Distance | Relation | ||
| Proprietary Models (API) | |||||||||
| GPT-5.6 Luna | 22.60 | 20.28 | 14.34 | 12.59 | 25.87 | 44.76 | 27.27 | 14.04 | 25.61 |
| Claude Sonnet 5 | 27.80 | 16.43 | 23.08 | 9.79 | 26.57 | 49.30 | 30.07 | 23.51 | 44.21 |
| Gemini 3.8 Flash | 30.19 | 21.62 | 18.66 | 19.39 | 49.65 | 52.85 | 20.28 | 41.34 | 22.91 |
| GPT-6-Astra (medium) | 39.00 | 38.46 | 13.99 | 19.58 | 51.75 | 73.78 | 39.86 | 29.12 | 52.28 |
| General-purpose VLMs | |||||||||
| Qwen2.5-VL-3B | 14.40 | 5.95 | 5.25 | 10.85 | 2.80 | 32.20 | 28.70 | 17.85 | 13.00 |
| Qwen3.5-9B | 10.30 | 10.85 | 9.45 | 5.60 | 13.30 | 9.45 | 4.20 | 20.35 | 8.05 |
| LLaVA-NeXT-7B | 14.40 | 5.25 | 6.65 | 5.95 | 7.00 | 30.10 | 31.50 | 15.05 | 18.60 |
| Qwen2.5-VL-7B | 17.60 | 3.85 | 17.15 | 2.45 | 30.80 | 33.60 | 25.90 | 10.15 | 27.70 |
| Cosmos3-8B | 18.95 | 5.60 | 2.40 | 9.80 | 32.90 | 35.30 | 32.20 | 3.90 | 43.20 |
| Qwen3-VL-8B | 20.20 | 3.15 | 13.30 | 7.35 | 23.10 | 38.50 | 24.50 | 21.00 | 34.70 |
| Spatially specialized VLMs | |||||||||
| SR-3D-8B | 9.60 | 2.45 | 6.30 | 12.60 | 2.10 | 24.85 | 14.70 | 8.75 | 3.50 |
| SpatialRGPT-8B | 9.70 | 6.65 | 5.25 | 5.25 | 2.80 | 17.85 | 34.30 | 3.85 | 10.20 |
| SpatialBot-3B | 11.80 | 2.80 | 7.70 | 1.05 | 34.30 | 23.45 | 28.70 | 2.80 | 13.00 |
| VLM-3R-7B | 14.00 | 7.90 | 1.05 | 4.20 | 7.70 | 31.50 | 29.40 | 2.45 | 32.65 |
| SpaceR-7B | 14.70 | 5.60 | 14.70 | 5.25 | 14.70 | 33.25 | 21.70 | 3.50 | 22.45 |
| ViLaSR-7B | 14.80 | 7.70 | 9.10 | 7.00 | 7.00 | 32.20 | 16.80 | 6.00 | 30.15 |
| SR-ReaL-8B | 14.90 | 5.95 | 11.55 | 11.90 | 8.40 | 31.50 | 32.20 | 7.05 | 16.15 |
| 4D-RGPT-8B | 15.40 | 3.50 | 11.90 | 12.95 | 7.00 | 33.95 | 27.30 | 12.25 | 16.15 |
| SpatialReasoner-7B | 15.80 | 7.20 | 7.70 | 5.60 | 12.60 | 33.25 | 30.10 | 8.40 | 27.00 |
| SFT on SpatialMind training data | |||||||||
| Qwen3-VL-8B + SFT | 45.70 | 36.02 | 57.34 | 23.77 | 37.06 | 56.99 | 57.34 | 40.64 | 57.89 |
| SpatialMind (Ours) | 52.65 | 38.32 | 65.46 | 30.15 | 52.34 | 68.42 | 65.13 | 42.89 | 64.48 |
Within each model group, bold marks the best result and underline marks the second best. Numerical tasks use mean relative accuracy (MRA); multiple-choice tasks use accuracy (ACC).
Scroll horizontally to view all eight tasks.
A CLOSER LOOK · SECTION 5.1
Three observations
Different model families. A shared challenge: reasoning at real-world scale.
-
01Proprietary Models
Astra sets the pace.
39.00%
GPT-6-Astra (medium) · Overall
GPT-6-Astra leads proprietary models on six of eight tasks. The two exceptions are current-state spatial estimation and future-distance prediction, both requiring absolute metric distances.
-
02General-Purpose Models
Categories come easier.
20.20%
Qwen3-VL-8B · Best overall in this group
Direction and relation questions are easier than metric estimation. Current-state measurements and ego-motion distance remain challenging: categorical future accuracy does not establish metric understanding.
-
03Spatial Expert Models
Geometry alone falls short.
9.60–15.80%
Overall range across spatial experts
Depth and 3D priors still leave weaknesses in current-state measurements, ego-motion distance, and future-distance prediction. Geometric priors alone do not ensure metric accuracy across driving and egocentric scenes.
| Benchmark | Average / Overall | Breakdown |
|---|---|---|
| VSI-Bench | 46.4 | Measurement: 52.1 · Spatiotemporal: 52.4 |
| OSI-Bench | 35.0 | Static: 30.2 · Dynamic: 33.4 |
| VLM4D | 53.4 | Real: 53.5 · Synthetic: 52.9 |
Each benchmark retains its original scoring convention. Scores across different benchmarks are not directly comparable.
BEYOND OUR BENCHMARK · SECTION 5.2
What transfers to new scenes?
Three external benchmarks, evaluated without additional fine-tuning.
-
01VSI-Bench
Metric understanding transfers indoors.
Benchmark focusVideo-based spatial understanding, with an emphasis on metric measurement and spatiotemporal reasoning.
SpatialMind leads measurement and spatiotemporal reasoning among the compared models, while remaining competitive overall. These results support the transfer of metric reasoning to indoor scenes.
-
02OSI-Bench
Stronger open-world spatial inference.
Benchmark focusOpen-world spatial inference in both static and dynamic metric settings.
SpatialMind leads overall and in static metric reasoning among the compared models, with competitive dynamic performance. This supports its ability to reason about spatial relationships in open-world environments.
-
03VLM4D
Motion reasoning across visual domains.
Benchmark focusSpatial and temporal understanding across real and synthetic videos, including motion and perspective changes.
SpatialMind leads overall and on both real and synthetic videos among the compared models, supporting the transfer of motion and perspective understanding across visual environments.
BibTeX
@misc{wang2026spatialmind,
title={From Sight to Foresight: Predictive Spatial Reasoning
in Vision-Language Models},
author={Wang, Feiran and Wang, Xiaoqi and Li, Ziwei and
He, Wenbin and Yan, Yan and Ren, Liu},
year={2026},
url={https://brack-wang.github.io/spatialmind/}
}