Same law, different geometry

How Does Geometry Enter Generated Motion?

Weihan Li1, Junhao Wu2, Yuhan Song1, Xiaofeng Lin1, Xinlei Chen3

1The University of Tokyo2RWTH Aachen University3Harbin Institute of Technology, Shenzhen

Under a fixed physical law, the visible geometry of a scene determines how motion must change. We ask how video generators realize this relationship. We fix the law and the initial state and change only the geometry drawn in the first frame, within matched families of tracks and deflectors, and compare each generated trajectory with the simulator prediction for that geometry. Paired interventions change one thing at a time: a local bump, the height of a barrier, the words of the prompt, the length of the clip.

Across nine image-to-video models, geometry is preserved and shapes the motion: the speed of the ball follows the drawn undulation of a track. A physical state would carry this response forward, and here the generated motion parts from the law. The mean slope barely accelerates the ball, successive contacts fail to compose through a consistent state, an edit ahead of the ball alters its motion before it arrives, and the ball climbs over barriers higher than its release point. Two global conditions organize the global trajectory: text strongly controls the destination, while clip length strongly controls timing in the open-weight models tested. The pattern persists with photographed first frames. Current video generation thus behaves as geometry-conditioned motion synthesis whose evolution of state differs systematically from that of a fixed physical law.

  1. response
  2. state
  3. goal
  4. time

One frame fixes a world

A world is a triple W=(Ω,L,x0) of visible geometry, local law and initial state. The first frame shows the geometry and the initial state: a white track with a barrier in its middle, and a ball at rest near the top. The law, gravity with rolling contact, is the same for every geometry of the family.

The simulator gives the motion the law requires

A simulator integrates the law on the drawn geometry, and a 64-member ensemble over material parameters and the starting point gives a tolerance band. Here the barrier is 0.70 of the release height, and the ball rolls over it to the end of the track.

Change only the geometry

Keep the law and the initial state and raise the barrier to 1.20 of the release height. The same law now turns the ball back before the top. Each piece of geometry acts when the ball reaches it, and what it does depends on the speed and direction the ball brings. Use the switch to go back and forth.

Give the same frame to a video model

The first frame goes to an image-to-video model with a prompt that is identical for every geometry of the family, so the barrier reaches the model only through the image. The generated clip plays in step with the simulator.

Read the motion against physics

We track the ball into the coordinates of the board. The blue dashed line marks the release height, the white dashed line is the simulator's path and the coloured trace the generated ball; each trace runs from faint to full as time passes. MiniMax-H3 crosses both barriers, as it does with every seed. Across the models, 43 of 49 physically impassable crossings are made.

Four readings, each a switch

Response: geometry is preserved and shapes the motion. State: a physical state would carry this response forward, and here the generated motion parts from the law. Goal: the outcome named in the prompt controls the destination. Time: clip length controls the timing in the open-weight models tested. Each reading below is a switch like this one.

1 / 6

Each panel keeps the law and the initial state and changes one thing between generations that share everything else, including the seed where the model allows it. The simulator and a generated video play side by side and switch together.

response

The speed follows the drawn shape

In the track family, 83–100% of each model's videos are intact. In them the drawn undulation keeps its amplitude to within 1% and the ball stays on the track. On undulating tracks the speed of the generated ball rises and falls with the drawn profile. Relative to each model's own tempo, the gain at 3 cycles per metre is 0.54–1.24 of the physical one. Eight models lead the physical phase by 15°–41°, at least twice the floor of 6°–7.5°, and MiniMax-H3 responds in phase. The response repeats across seeds. The mean slope contributes little: in the generated motion the acceleration along the straight track is at most 0.65 of what the ball's own transit time implies, and below 0.3 in seven models.

Drawn track
state

Does the modulation form a local dynamics?

A state-mediated physical dynamics makes three testable predictions here: successive events compose through the state, only geometry the ball has reached can affect it, and its evolution obeys the energy bound.

state · composition

Successive contacts fail to compose

With three deflectors the simulator's ensemble keeps a single order of contacts. The closed models reproduce it in 12–24% of their videos and the open-weight models in none. In MiniMax-H3, Wan 2.2 and Cosmos 3 Nano the ball seldom reaches a second contact: it slides off the first deflector and drops. In the two Seedance models the exit direction is off by 19°–31° at the first, the second and the third contact alike, so the chain is lost through independent local deviations, and in Wan 3.0 the deviation grows from 10° to 28°. A contact law calibrated on each model's single-deflector videos improves the prediction of the three-deflector scenes for one model only, so the deviations do not reduce to a consistent wrong material.

Deflectors
state · locality

Edits act ahead of the ball

We add a 15 mm bump or dip to one section of a track and generate with the same seed as for the unedited track. Physics changes the speed at the edit by −11% and +8% and leaves the motion before it unchanged. MiniMax-H3 slows at the bump by 2.4%, a fifth of the physical change, and reaches the edit at the unchanged time. In LTX-2.5 the ball reaches either edit 3–6% sooner and slows at the dip, where physics speeds up. In three open-weight models the path before the ball reaches the edit changes 1.7–4.0 times as much as under an edit of equal size placed off the track, beyond the seed-matched null. The shape ahead of the ball takes part in the motion from the start.

Edit
state · energy

The ball crosses barriers higher than its release point

Physics passes the barriers below the release height and none above it. Six models cross every barrier they were given, and overall 43 of 49 physically impassable crossings are made. In physics the speed kept over the top falls from 0.53 of the speed at the foot to zero as the barrier rises. In the models it stays at the same level at every height, for example 0.42–0.46 in Seedance 2.0 mini.

Barrier height
state · energy

The ball climbs past its release height

On a ramp pair the physical ball turns back at 0.93–0.98 of its release height. In 63 of 72 readable videos the generated ball rises above the release height, and in 37 it runs to the top of the ramp. The mean height ratio per model is 1.15–2.36. The ball slows on a rise, as the shape suggests, by an amount unrelated to the height gained.

Up-ramp

The same 20° descent and drop to the valley; only the up-ramp angle differs. The closed models were run at 15° and 60°.

goal

The named outcome controls the destination

The prompt of the deflector family says that the ball bounces on the deflectors until it comes to rest in the tray. Under this prompt the ball reaches the tray in 98–100% of the trackable videos of seven models, whether or not the contacts were right. The generated motion itself leads there less often: continued by the simulator from the state in which the ball leaves its last correct contact, it reaches the tray in 11–84% of the videos of six of these models, which exceed the simulator by 16–87 percentage points. We then generate the same scenes, first frames and seeds with a prompt that describes the scene and leaves the outcome unnamed. Tray arrival falls in all eight models, by 33 to 96 points. The excess over the simulator is then at most 27 points in every model, and in MiniMax-H3 and the two Seedance models it is gone. Wan 3.0 retains the most, with 67% arrival under the neutral prompt. The drawn tray by itself barely attracts the ball: when the tray is shifted, the landing point follows it by at most 0.42 of the shift, and with the tray removed no model brings the ball to rest where it was.

Prompt

Prompt given to the model

time

Clip length controls the timing

For four open-weight models we change only the length of the clip, from 5 s to 3 or 8 s, with the scene, prompt and seed fixed. A physical event keeps its duration, so its tempo is independent of the clip length. The generated tempo grows with the clip length with an exponent γ of 0.85–1.04 in three models, where 1 means that the event is scaled to the clip. MiniMax-H3 reaches the end of the track at 90–93% of the clip at every length, LTX-2.3 at 55–59% and the distilled LTX-2.5 at about half. Plotted against normalized time, the three motions of one seed fall on one curve. Wan 2.2 doubles its speed at 3 s and keeps its 5 s pace at 8 s.

Clip length
photographs and regeneration

Photographed first frames and regeneration

photographed first frame

The pattern holds on photographs

We built eight of the scenes, photographed them, measured their geometry from the photographs and ran the simulator on the measured geometry. With the photographs as first frames, the readings of six of the nine models stay within their seed-to-seed spread of the readings on renders of the same geometry. The three changes are specific: LTX-2.5 moves the camera on photographs, MiniMax-H3 shows larger exit residuals, and Seedance 2.0 a higher gain.

First frame
Scene
regeneration

Regeneration returns the model's own motion

Wan 2.2 (A14B), one source and seed at its three noise levels.

We give two models a source video whose structure is intact and whose dynamics are wrong, add noise and let the model regenerate it. A motion slowed twofold keeps its tempo. A violated contact is changed: a ball that passes through a drawn deflector is made to hit it. What replaces the source is the model's own trajectory for that scene, with the order of contacts still wrong and a restitution of 0.16–0.45 in place of the prescribed 0.85. For a source with a wrong restitution, the regenerated motion of Wan 2.2 moves 0.66 of the way toward the model's own image-to-video output of the same scene and seed, and 0.10 toward physics.

Noise level
  • source
  • physics
  • its own image-to-video output
  • regenerated
Taken together, current video generators use the visible geometry to shape motion locally, the text to choose its outcome, and the clip length to pace it. A fixed law prescribes a different organization: a state that is updated where the ball is, composed across events and bounded by energy. These results motivate evaluating physics through how controlled changes propagate through generated motion, beyond whether the resulting clip appears physically plausible.

Citation

@article{li2026geometry,
  title   = {How Does Geometry Enter Generated Motion?},
  author  = {Li, Weihan and Wu, Junhao and Song, Yuhan and Lin, Xiaofeng and Chen, Xinlei},
  journal = {arXiv preprint arXiv:2610.05135},
  year    = {2026}
}