Use the chapter buttons to jump directly to each part of the overview video.
Generated task videos give a robot a useful picture of what should happen next. Yet the robot can progress faster or slower than the clip. The future may show the correct task and still give guidance for the wrong moment. On eight CALVIN tasks, a five-frame early shift reduces generated-future success from 69.8% to 34.2%, below the 54.0% future-free policy.
RAFC treats temporal compatibility as a control problem. It estimates which nearby phase of a generated future to use and how strongly to trust it, without receiving a shift label.
Future-Experience Conditioning first builds a 16-frame future from language grounding, a robot-free digital-twin rollout, and video diffusion. This video is generated once at task initialization, then reused during execution.
At each control step, RAFC compares a static first-frame fallback with three dynamic views at offsets −2, 0, +2. All four candidates pass through the same frozen behavior-cloning policy. A learned gate sets the future-trust coefficient and weights the three nearby phases; a bounded residual actor corrects the blended action. The gate and actor learn from task reward without temporal-alignment supervision.
The model can reduce reliance on a misleading dynamic future while preserving useful task information from the same generated clip.
The six nonzero test offsets are off-grid relative to RAFC's local {−2, 0, +2} candidate bank. The applied shift is hidden from the gate. Across the shifted conditions, RAFC raises task-balanced success from 51.8% to 73.7%. Uniform averaging of the identical four-branch bank reaches 66.7%, so learned weighting adds 7.0 points.
| Method | −5 | −3 | −1 | 0 | +1 | +3 | +5 | Avg. shifted |
|---|---|---|---|---|---|---|---|---|
| Generated future | 34.2 | 52.9 | 64.6 | 69.8 | 65.3 | 51.9 | 41.6 | 51.8 |
| Uniform averaging | 49.6 | 68.9 | 78.4 | 79.4 | 79.0 | 68.8 | 55.5 | 66.7 |
| RAFC | 61.7 | 76.8 | 81.2 | 82.3 | 81.5 | 76.0 | 65.2 | 73.7 |
Success (%). Full paper reports mean ± standard deviation across three training seeds; the table shows means for readability. Shifted average excludes zero.
With distinct windows from independently generated longer videos, shifted success rises from 69.7% to 75.6%. Under 0.75× and 1.25× rate warps, the average rises from 58.6% to 73.2%. An unannounced step-40 timing change yields 75.4% with RAFC, compared with 65.7% uniform averaging and 68.9% with a constant learned gate.
The same generated-future interface is tested on a Franka Panda under natural timing variation. Each task has 20 trials per method.
| Task | Plain FEC | FEC + RAFC |
|---|---|---|
| Open kettle | 3/20 (15%) | 9/20 (45%) |
| Close kettle | 5/20 (25%) | 11/20 (55%) |
| Close microwave door | 8/20 (40%) | 14/20 (70%) |
| Overall | 16/60 (26.7%) | 34/60 (56.7%) |
Timing offsets are not imposed or measured in these physical trials; controlled phase interventions are evaluated in simulation.