RAFC

Reliability-Aware Future Conditioning for Temporally Robust Robot Manipulation
University of BremenUniversity of North TexasToyota Motor North America

Overview

Use the chapter buttons to jump directly to each part of the overview video.

Why reliability-aware future conditioning?

Generated task videos give a robot a useful picture of what should happen next. Yet the robot can progress faster or slower than the clip. The future may show the correct task and still give guidance for the wrong moment. On eight CALVIN tasks, a five-frame early shift reduces generated-future success from 69.8% to 34.2%, below the 54.0% future-free policy.

RAFC treats temporal compatibility as a control problem. It estimates which nearby phase of a generated future to use and how strongly to trust it, without receiving a shift label.

RAFC paper overview showing generated futures, a static fallback, three phase candidates, learned trust and phase weights, and residual control.
RAFC system overview from the paper. The controller weighs nearby temporal hypotheses and a static fallback before acting.

How does RAFC work?

Future-Experience Conditioning first builds a 16-frame future from language grounding, a robot-free digital-twin rollout, and video diffusion. This video is generated once at task initialization, then reused during execution.

At each control step, RAFC compares a static first-frame fallback with three dynamic views at offsets −2, 0, +2. All four candidates pass through the same frozen behavior-cloning policy. A learned gate sets the future-trust coefficient and weights the three nearby phases; a bounded residual actor corrects the blended action. The gate and actor learn from task reward without temporal-alignment supervision.

at = blended base action + bounded residual correction

The model can reduce reliance on a misleading dynamic future while preserving useful task information from the same generated clip.

Temporal robustness on CALVIN

The six nonzero test offsets are off-grid relative to RAFC's local {−2, 0, +2} candidate bank. The applied shift is hidden from the gate. Across the shifted conditions, RAFC raises task-balanced success from 51.8% to 73.7%. Uniform averaging of the identical four-branch bank reaches 66.7%, so learned weighting adds 7.0 points.

Method−5−3−10+1+3+5Avg. shifted
Generated future34.252.964.669.865.351.941.651.8
Uniform averaging49.668.978.479.479.068.855.566.7
RAFC61.776.881.282.381.576.065.273.7

Success (%). Full paper reports mean ± standard deviation across three training seeds; the table shows means for readability. Shifted average excludes zero.

Beyond clipped phase shifts

With distinct windows from independently generated longer videos, shifted success rises from 69.7% to 75.6%. Under 0.75× and 1.25× rate warps, the average rises from 58.6% to 73.2%. An unannounced step-40 timing change yields 75.4% with RAFC, compared with 65.7% uniform averaging and 68.9% with a constant learned gate.

Real-robot evaluation

The same generated-future interface is tested on a Franka Panda under natural timing variation. Each task has 20 trials per method.

TaskPlain FECFEC + RAFC
Open kettle3/20 (15%)9/20 (45%)
Close kettle5/20 (25%)11/20 (55%)
Close microwave door8/20 (40%)14/20 (70%)
Overall16/60 (26.7%)34/60 (56.7%)

Timing offsets are not imposed or measured in these physical trials; controlled phase interventions are evaluated in simulation.