Kepler4D: Controllable Future Video
Generation via 4D Scene State Evolution

1 University of Illinois at Chicago 2 University of Michigan Ann Arbor 3 Cisco

† Corresponding author

TL;DR From an observed video to editable 3D trajectories: evolve an explicit 4D scene state to guide controllable future video generation.

Observe motion. Evolve scene state. Control the future. A shared scene representation supports automatic future evolution, language-guided trajectory editing, and generation under different camera viewpoints.

Abstract

Video world models aim to preserve scene structure and predict how dynamic objects evolve beyond visual observations. We present Kepler4D, a framework for future video generation through explicit 4D scene state evolution. Given a monocular video, Kepler4D constructs a shared 3D representation of background geometry, object motion histories, coarse spatial supports, and semantic context. Chain-of-Motion summarizes observed motion and uses a vision-language model to select structured speed and heading decisions and decide whether to bound object-center height from below. A deterministic rollout converts these decisions into future object trajectories for inspection and editing before synthesis. We render the evolving proxies into geometric controls for a pretrained video generator, separating coarse object motion from the synthesis of appearance and articulation. Experiments on real-world videos demonstrate that Kepler4D enables controllable object motion and plausible future rollout while preserving scene consistency.

From scene state to future video

Make future object motion explicit, inspectable, and editable before synthesis.

Pipeline overview. Background geometry anchors the scene, while evolving object proxies provide foreground motion guidance. The video generator synthesizes appearance and local articulation.
01

Construct the scene state

Organize static geometry, observed 3D object trajectories, coarse spatial supports, and semantic context in a shared world frame.

02

Evolve object motion

Chain-of-Motion infers structured motion decisions. Deterministic rollout expands them into future trajectories for inspection and editing.

03

Render and generate

Render depth, object control signals, and masks under a target camera path to guide a pretrained video generator.

One scene state, different futures

Explore automatic continuation, language-guided motion, and novel viewpoints.

Predict future object motion from the observed video and scene context.

Automatic future evolution. All methods use the observed video without externally supplied object trajectories or motion commands; methods with camera control receive the same target camera path. In this example, Kepler4D depicts the runner moving forward and vaulting over the railing, as in the reference.

Use language to specify how an object should move before generating the future frames.

Different instructions, different trajectories. “Go Straight,” “Turn Left 60°,” and “Turn Left 90°” specify alternative futures for the same observed video. The second column shows the inferred object trajectories.

Comparison methods receive the same language command. Kepler4D additionally uses its inferred trajectories and rendered geometric controls.

Render the reconstructed scene from a target camera path.

Static novel view synthesis on DL3DV. We compare Kepler4D with GEN3C, TrajectoryCrafter, and SPMem under target camera motion. Each row shows the input view, target camera path, held-out reference view, and generated results.

BibTeX

If you find this work useful, please consider citing it.

Kepler4D
@misc{wang2026kepler4d,
  title={{Kepler4D}: Controllable Future Video Generation via 4D Scene State Evolution},
  author={Wang, Feiran and Duan, Bin and Wu, Junyi and Liu, Gaowen and Yan, Yan},
  year={2026},
  url={https://github.com/Brack-Wang/kepler4d}
}

Figure