Construct the scene state
Organize static geometry, observed 3D object trajectories, coarse spatial supports, and semantic context in a shared world frame.
† Corresponding author
TL;DR From an observed video to editable 3D trajectories: evolve an explicit 4D scene state to guide controllable future video generation.
Video world models aim to preserve scene structure and predict how dynamic objects evolve beyond visual observations. We present Kepler4D, a framework for future video generation through explicit 4D scene state evolution. Given a monocular video, Kepler4D constructs a shared 3D representation of background geometry, object motion histories, coarse spatial supports, and semantic context. Chain-of-Motion summarizes observed motion and uses a vision-language model to select structured speed and heading decisions and decide whether to bound object-center height from below. A deterministic rollout converts these decisions into future object trajectories for inspection and editing before synthesis. We render the evolving proxies into geometric controls for a pretrained video generator, separating coarse object motion from the synthesis of appearance and articulation. Experiments on real-world videos demonstrate that Kepler4D enables controllable object motion and plausible future rollout while preserving scene consistency.
Make future object motion explicit, inspectable, and editable before synthesis.
Organize static geometry, observed 3D object trajectories, coarse spatial supports, and semantic context in a shared world frame.
Chain-of-Motion infers structured motion decisions. Deterministic rollout expands them into future trajectories for inspection and editing.
Render depth, object control signals, and masks under a target camera path to guide a pretrained video generator.
Explore automatic continuation, language-guided motion, and novel viewpoints.
Predict future object motion from the observed video and scene context.
Use language to specify how an object should move before generating the future frames.
Comparison methods receive the same language command. Kepler4D additionally uses its inferred trajectories and rendered geometric controls.
Render the reconstructed scene from a target camera path.
If you find this work useful, please consider citing it.
@misc{wang2026kepler4d,
title={{Kepler4D}: Controllable Future Video Generation via 4D Scene State Evolution},
author={Wang, Feiran and Duan, Bin and Wu, Junyi and Liu, Gaowen and Yan, Yan},
year={2026},
url={https://github.com/Brack-Wang/kepler4d}
}