[Paper Notes] ForceDelta-VLA: Distilling Force-Conditioned Action Corrections for Contact-Rich Manipulation

16 minute read

Published:

This post supports English / 中文 switching via the site language toggle in the top navigation.

TL;DR

Contact-rich manipulation often needs two different time scales. A slow VLA should plan the task motion, while a fast controller should react to force changes between complete action-chunk updates. ForceDelta-VLA makes this split explicit and learns the fast part from an existing demonstration dataset.

A frozen force-conditioned teacher is paired with a learned force-agnostic reference mode. Their matched predictions define a force-correction target. A second delay-correction target accounts for the mismatch between an old reference action and the current reference prediction, including the change in reference state. A lightweight student then reads recent wrench history, proprioception, cached task context, and the delayed reference action to predict both corrections.

Across five single-arm and four bimanual real-robot tasks, ForceDelta-VLA reaches 82.2% mean success, compared with 54.4% for ForceVLA and 70.6% for direct execution of the temporal teacher. On successful trials, mean peak contact force drops by roughly 26% relative to ForceVLA on both platforms. The correction pathway runs in 2.43 ms, while the reference-action pathway takes about 189.7 ms.

Paper and source version

ForceDelta-VLA: Distilling Force-Conditioned Action Corrections for Contact-Rich Manipulation is by Ju Dong, Yu Fu, Jian Chen, Yimeng Liu, Haocheng Zhao, Lei Zhang, Kaixin Bai, Liding Zhang, Diwen Zheng, Alois Christian Knoll, Angela P. Schoellig, and Jianwei Zhang, from the University of Hamburg, University of Science and Technology of China, and Technical University of Munich. These notes follow arXiv:2609.18242v1, submitted September 16, 2026. See the paper PDF. Results below are author-reported.

1. Why split reference motion from contact correction?

During USB insertion, the broad approach direction can remain useful for several control cycles while the local alignment must react to new contact. A single force-aware VLA that regenerates a complete action chunk at every query may be too slow for this local response.

ForceDelta-VLA decomposes the commanded pose into three parts:

\[P(A^{cmd})=P(A^{ref})+\Delta\hat A^{force}+\Delta\hat A^{delay}.\]

$A^{ref}$ is a reusable motion predicted by the slow pathway. $\Delta\hat A^{force}$ responds to current wrench feedback. $\Delta\hat A^{delay}$ corrects the fact that the reference was generated earlier, from another state, and may be stale by the time it is executed. Gripper commands remain in the reference action; the corrections modify the pose coordinates.

The split addresses an under-supervision problem. Demonstrations record the final action, yet do not label which component represents task motion and which component responds to contact. The paper constructs those labels from paired teacher predictions instead of collecting new correction demonstrations or using on-policy reinforcement learning.

2. Three-stage training pipeline

The method starts from a ForceVLA-style force-aware VLA and replaces its instantaneous force embedding with a causal temporal convolutional network. The teacher receives visual observations, language, robot state, and a 100 ms wrench history, then predicts a 50-step action chunk.

Stage 1: temporal force-conditioned teacher

The frozen teacher is trained on teleoperated manipulation trajectories:

\[A^{cond}_{T,t,1:H}=T_\theta(V_t,L,S_t,z^F_t;\epsilon),\]

where $z^F_t=E_F(F_t)$ encodes the recent wrench history and $\epsilon$ is the initial flow noise.

Stage 2: force-agnostic reference mode

The reference pathway must represent unavailable wrench input, not a measured zero wrench. It replaces the force history with a learned missing-force token $z^\emptyset_F$ and uses a low-rank adapter on the pose channels of the action expert. With the teacher backbone frozen, the adapter learns the original flow-matching objective:

\[A^{ref}_{T,t,1:H}=T^{ref}_{\theta,\phi}(V_t,L,S_t,z^\emptyset_F;\epsilon).\]

This produces the slow reference action that can be cached and reused while the fast student reacts to contact.

Stage 3: correction distillation

At a correction time $t$, the system compares two current predictions that share the cached visual-language prefix, robot state, and flow noise: one force-conditioned and one force-agnostic. The difference defines the force target. A separate target compares the current force-agnostic prediction with the previously stored reference and aligns the change in reference state.

3. Force and delay targets

Let $k$ be the latest reference query whose result is available at correction time $t$. The stored reference was generated from state $S_k$ and prefix $E_k$. Reusing this prefix while updating the state and force history gives:

\[A^{cond}_{T,t|k,1:K}=T_\theta(E_k,S_t,z^F_t;\epsilon_k)_{1:K},\] \[A^{ref}_{T,t|k,1:K}=T^{ref}_{\theta,\phi}(E_k,S_t,z^\emptyset_F;\epsilon_k)_{1:K}.\]

The force-correction target is the matched pose difference:

\[\Delta A^{force}_{T,t|k,j}=P(A^{cond}_{T,t|k,j})-P(A^{ref}_{T,t|k,j}).\]

The delay target compares the current reference prediction with the old reference segment and adds a state-alignment term $\Gamma(S_t,S_k)$:

\[\Delta A^{delay}_{T,k\rightarrow t,j}= P(A^{ref}_{T,t|k,j})-P(A^{ref}_{k}(t_j))+\Gamma(S_t,S_k).\]

The second target is more than a timestamp offset. It learns the action correction needed because the robot has moved away from the state in which the stored reference was generated.

4. One lightweight student with two correction heads

Each cached reference stores its action chunk, query state, pooled visual-language context, and a validity mask. The student also receives an interpolated reference-action segment, recent wrench history, the current robot state, and timing features describing cache age and interpolation phase.

A shared attention module processes projected context tokens and a learned correction query. It is evaluated twice:

  • the force branch sees all inputs and predicts $\Delta\hat A^{force}$;
  • the delay branch masks the force token and predicts $\Delta\hat A^{delay}$.

The mask prevents the delay output from directly using force history. Separate heads and separate targets preserve the meaning of the two corrections, while the shared module keeps the student compact.

The normalized distillation loss is

\[\mathcal L_{distill}=\sum_{j=1}^{K}w_j\left[ \ell_{pose}(\Delta\hat A^{force}_{t,j},\Delta A^{force}_{T,t|k,j}) +\lambda_{delay}\ell_{pose}(\Delta\hat A^{delay}_{t,j},\Delta A^{delay}_{T,k\rightarrow t,j}) \right],\]

with temporally decaying weights $w_j$.

flowchart LR
    A[Vision + language + state + wrench history] --> B[Slow force-conditioned teacher]
    A --> C[Force-agnostic reference mode]
    B --> D[Force target]
    C --> D
    C --> E[Cached reference action]
    E --> F[Reference-state and delay target]
    D --> G[Fast correction student]
    F --> G
    H[Recent wrench + proprioception + cached context + timing] --> G
    G --> I[Force correction + delay correction]
    E --> J[Reference action plus corrections]
    I --> J
    J --> K[Robot executor]

5. Asynchronous schedule replay

At deployment, the teacher and student run independently. The teacher periodically refreshes a cached reference action. The student reads the latest completed reference, predicts corrections, and updates the executor while the teacher computes its next chunk.

Training must reproduce this timing. ForceDelta-VLA samples reference-query periods and inference latencies before constructing targets, then replays the same fixed schedule during optimization. At each correction time, the student sees the most recent reference that would actually be available under that schedule.

This matters because a student trained only with fresh reference actions would encounter a distribution shift at deployment. The reference can be old, its state can be misaligned, and some correction steps can expire before a result arrives. The executor therefore selects the most recent valid student query, interpolates the reference at the corresponding timestamps, clips the two corrections separately, and skips expired steps.

6. Real-robot evaluation

The evaluation contains five single-arm tasks on a 7-DoF Franka Panda and four bimanual tasks on an X Square Robot. The tasks cover object flipping, USB insertion, cabinet opening, button pressing, whiteboard wiping, plug removal, plug insertion, drawer opening, and bimanual wiping.

Each task has 200 demonstrations. Images, robot states, and commanded actions are recorded at 30 Hz; wrench estimates are timestamped at 100 Hz. Both platforms use joint-torque-based wrench estimates supplied by the robot, without additional wrist force/torque sensors. Each method is evaluated with 20 trials per task.

MethodMean success
$\pi_{0.5}$47.2%
ForceVLA54.4%
TA-VLA57.2%
ImplicitRDP45.0%
Temporal Teacher70.6%
ForceDelta-VLA82.2%

ForceDelta-VLA improves over ForceVLA on all nine tasks. The largest gains are reported on USB Insertion (40% → 80%), Button Pressing (50% → 85%), and Plug Insertion (35% → 70%).

TaskForceVLAForceDelta-VLA
Object Flipping65%90%
USB Insertion40%80%
Cabinet Opening60%80%
Button Pressing50%85%
Whiteboard Wiping, single70%90%
Plug Removal50%80%
Plug Insertion35%70%
Drawer Opening60%80%
Whiteboard Wiping, bimanual60%85%

7. Force, latency, and delay results

Relative to ForceVLA, mean peak contact force over successful trials drops by 4.2 N on the single-arm platform and 4.3 N on the bimanual platform, approximately 26% in both cases. Completion time decreases by 5.2 s and 5.9 s, respectively.

PathwayForward latency
ImplicitRDP12.86 ms
Temporal Teacher176.4 ms
ForceDelta reference action189.7 ms
ForceDelta correction2.43 ms

The student is fast enough to run between 100 Hz robot command transmissions, while the reference generator operates at roughly 5 Hz under serial inference. On USB Insertion and Plug Insertion, adding 200 ms of reference-action delay reduces ForceDelta-VLA success by 15 percentage points, compared with a 25-point drop for the temporal teacher. Mean peak force rises by 3.7 N for ForceDelta-VLA and 9.6 N for the teacher.

This is the clearest evidence for the two-rate design: a stale reference hurts, yet a fast correction layer absorbs part of the delay instead of regenerating a complete action chunk at every update.

8. Ablations and unseen objects

The ablations support both the target construction and the decomposition. Removing force correction lowers success by 15 points on the single-arm platform and 17.5 points on the bimanual platform, while increasing peak force. Removing learned delay correction lowers success by 7.5 points on both platforms. Zeroing the force input performs worse than simply disabling the force output, showing that the active branch uses the wrench history.

A single combined-correction head loses 10 points on both platforms. Fast full-action distillation also underperforms correction distillation. The result favors reusing the reference action and supervising two separate correction meanings.

Replacing the paired teacher-derived force target with a demonstration-derived target reduces success by 12.5 and 17.5 points. Removing asynchronous schedule replay reduces success by 7.5 points on single-arm tasks and 12.5 points on bimanual tasks.

On unseen objects across USB Insertion, Object Flipping, and bimanual Whiteboard Wiping, ForceDelta-VLA reaches 66.7% success, compared with 40.0% for the Temporal Teacher and 16.7% for ImplicitRDP. Its mean peak force on successful trials is 13.6 N, 8.2 N below the teacher.

9. Limits and takeaway

ForceDelta-VLA is a fast local adaptation layer. It depends on the slow reference for major strategy changes, so a jam that requires retreating substantially, changing approach direction, or re-establishing contact can still fail. Insertion failures often occur when the reference keeps advancing after the connector jams; cabinet and drawer failures can start before stable handle contact is established. The method also uses torque-derived wrench estimates rather than dedicated force/torque sensors, and the unseen-object evaluation is small.

My main takeaway is that force feedback becomes easier to distill when the model is asked to correct a useful reference action. The student does not regenerate the entire task trajectory; it learns how recent contact and state mismatch should bend the next few pose commands. The most promising next step is a stronger recovery policy that lets corrections request retreat or a new approach, together with high-precision force/torque sensing for delicate insertion and assembly.