[Paper Notes] Residual Fault Adaptation for Dexterous In-Hand Manipulation Under Runtime Joint Faults

17 minute read

Published:

This post supports English / 中文 switching via the site language toggle in the top navigation.

TL;DR

A dexterous hand can execute a healthy manipulation policy perfectly until one joint starts receiving altered commands. The object may slip before a conventional fault-diagnosis module has identified the problem. Residual Fault Adaptation (RFA) keeps a frozen healthy teacher for nominal behavior and adds a recurrent residual policy that infers corrections from proprioception and command-response history.

Training randomizes six hidden command-channel faults: joint locking, range restriction, intermittent dropout, reduced command gain, command delay, and command bias. An adaptive sampler emphasizes fault modes with lower recent success. A separate Direct FIDR policy supplies a training-only distributional reference through a small KL term; it is absent during deployment.

On a 16-DoF LEAP Hand rotating a cube, RFA reaches 91.26% post-onset success under a fixed mixed-fault simulation protocol, compared with 87.73% for the healthy policy and 87.98% for Direct FIDR. Under healthy actuation it reaches 99.11%. A targeted real-robot study with software-injected faults demonstrates zero-shot deployment and measures command response and cube rotation; it does not report hardware task-success or contact-force metrics.

Paper and source version

Residual Fault Adaptation for Dexterous In-Hand Manipulation Under Runtime Joint Faults is by Linan Deng, Xing Liu, Lin Hong, Feng Hua, Guijun Ma, Zuogong Yue, and Fumin Zhang from the Hong Kong University of Science and Technology and Huazhong University of Science and Technology. These notes follow arXiv:2609.17404v1, submitted September 15, 2026. See the paper PDF. Results below are author-reported.

1. The fault is hidden in the command channel

The paper studies a 16-DoF LEAP Hand rotating a cube around the world $z$-axis. The policy produces a normalized relative action $a_t\in[-1,1]^{16}$, which becomes a candidate joint-position target $\hat q_{t+1}$. A runtime fault changes the target delivered to the low-level controller, $\tilde q_{t+1}$, before the measured joint position $q_{t+1}$ responds:

\[a_t\rightarrow \hat q_{t+1}\rightarrow \tilde q_{t+1}\rightarrow q_{t+1}.\]

The deployed policy receives no fault label, affected-joint mask, severity, onset time, or controller-switch signal. It must infer that a command channel has changed from the mismatch between what was requested and what the joints actually did.

This is especially difficult for in-hand manipulation. A single joint can alter several fingertip contacts, redistribute forces, change the available manipulation wrench, and cause a held object to slip. The disturbance also begins mid-episode, after the policy has already established a contact strategy.

2. Six runtime fault modes

Each fault instance is defined by the affected joint set $J$, fault mode $\kappa$, severity parameters $\xi$, and onset step $t_{onset}$. The reported training and evaluation distribution uses one affected joint at a time.

ModeCommand-channel effect
Joint lockingHolds the delivered target at the measured position at onset for a sampled duration
Range restrictionProjects the target into a reduced joint interval
Intermittent dropoutDrops updates and holds the previous delivered target
Reduced gainDelivers only a fraction of the target change
Command delayDelivers an older nominal target from the command history
Command biasAdds a signed offset to the nominal target

For example, a reduced-gain fault uses

\[\tilde q_{t+1,j}=\tilde q_{t,j}+g_j(\hat q_{t+1,j}-\tilde q_{t,j}),\]

while an intermittent dropout uses a Bernoulli update mask $m_{t,j}$:

\[\tilde q_{t+1,j}=m_{t,j}\hat q_{t+1,j}+(1-m_{t,j})\tilde q_{t,j}.\]

The fault is sampled at reset but activates later, so every faulted episode begins with healthy actuation and transitions into degraded execution.

3. Teacher-anchored residual control

RFA has three training stages.

Stage I: healthy policy

A recurrent healthy policy $\pi_H$ is trained without runtime faults and then frozen. Its native observation contains three consecutive frames of joint positions and nominal candidate targets:

\[o^H_t=[b_{t-2},b_{t-1},b_t],\qquad b_t=[q_t,\hat q_t].\]

For a 16-DoF hand, this is a 96-dimensional history. The healthy policy supplies the nominal action pathway throughout deployment.

Stage II: Direct FIDR reference

A second recurrent policy $\pi_F$ is initialized from the healthy checkpoint and trained directly under fault-injection domain randomization (FIDR). It receives a 144-dimensional observation and predicts a complete action under faults. Once trained, it is frozen. Direct FIDR is a distributional reference for Stage III, not a deployed component.

Stage III: residual policy

The residual policy augments the healthy observation with command-response features:

\[e_t=\hat q_t-q_t,\qquad \Delta e_t=e_t-e_{t-1},\qquad \Delta q_t=q_t-q_{t-1}.\]

The resulting observation is

\[o^R_t=[(o^H_t)^\top,e_t^\top,(\Delta e_t)^\top,(\Delta q_t)^\top]\in\mathbb R^{144}.\]

These features expose the consequence of a faulty command channel without revealing which fault occurred. A recurrent policy can integrate the deviations over time, so explicit diagnosis is unnecessary.

4. Bounded composition and training objective

At each step, the healthy teacher produces a nominal action mean $\mu^H_t$ and the residual policy produces a correction mean $\mu^R_t$. The teacher is clipped to $[-1,1]$ and the residual is passed through a scaled hyperbolic tangent. The correction is then limited by the remaining action headroom:

\[\mu^C_{t,j}=\tilde\mu^H_{t,j}+\delta_{t,j},\]

where

\[\delta_{t,j}=\begin{cases} \bar\mu^R_{t,j}(1-\tilde\mu^H_{t,j}), & \bar\mu^R_{t,j}\ge 0,\\ \bar\mu^R_{t,j}(1+\tilde\mu^H_{t,j}), & \bar\mu^R_{t,j}<0. \end{cases}\]

The composed action remains in $[-1,1]^{16}$. The residual mean head starts at zero, so the initial RFA controller is equivalent to the clipped healthy teacher. This gives training a stable nominal starting point.

RFA is trained with recurrent PPO plus a low-weight KL regularizer on valid fault-active samples:

\[\mathcal L=\mathcal L_{PPO}+ \frac{\lambda_{ref}}{\max(1,|B_{act}|)} \sum_{t\in B_{act}}D_{KL}(P^C_t\Vert P^F_t),\]

with $\lambda_{ref}=0.005$. PPO optimizes task return; the Direct FIDR distribution acts as a fault-conditioned action prior. The KL term is masked out on healthy samples and does not force RFA to imitate Direct FIDR everywhere.

flowchart LR
    A[Healthy policy training] --> B[Frozen healthy teacher]
    A --> C[Initialize Direct FIDR]
    C --> D[FIDR training + freeze]
    B --> E[Nominal action]
    F[Command-response history] --> G[Recurrent residual policy]
    E --> H[Bounded teacher + residual composition]
    G --> H
    D -. training-only KL reference .-> H
    H --> I[Faulty command channel]
    I --> J[Measured joint response]
    J --> F

5. Adaptive fault sampling

FIDR first randomizes whether an episode is faulted, which joint is affected, the fault mode, its severity, and the onset time. During RFA training, adaptive sampling changes only the relative probabilities of the six fault modes.

For each mode and affected finger group, the sampler tracks an exponential moving average of recent fault-active success. Every 500 environment control steps, difficulty is defined as $d_k=1-\frac{1}{4}\sum_g\hat s_{k,g}$ and the mode probability is normalized as

\[w_k=\frac{d_k}{\sum_{k'=1}^{6}d_{k'}}.\]

A mode with lower recent success receives more training samples. The sampler does not change the probability of an episode being faulted, the severity range, the onset distribution, or the single-joint constraint. The fixed mixed-fault test benchmark uses balanced cases and does not inherit the learned training weights.

This distinction matters. Adaptive sampling is a curriculum mechanism, not an evaluation-time fault detector.

6. Simulation protocol and metrics

The simulation uses Isaac Lab, a 30 Hz control loop, 4,096 parallel environments, 20–120 second episodes, and recurrent PPO. Faulted and healthy checkpoints are evaluated without online learning, auxiliary fault diagnosis, ground-truth fault information, or controller switching.

The fixed mixed-fault benchmark balances six modes, all 16 joints, and three normalized severity levels $\alpha\in{0.25,0.50,0.75}$. Post-onset evaluation uses a 10-second window. The primary labels are mutually exclusive:

  • Success rate (SR): the object is not dropped and its mean world-frame $z$-axis rotation rate reaches at least 5°/s;
  • Drop rate (DR): the object is dropped in the evaluation window;
  • Non-drop failure rate (NDFR): the object is not dropped but the rotation criterion is not met;
  • Fault-retention ratio (FTR): mixed-fault SR divided by healthy-condition SR.

The evaluation therefore distinguishes retaining the object from actually continuing the manipulation.

7. Main simulation results

MethodConditionSR ↑DR ↓NDFR ↓FTR ↑
Healthy policyHealthy97.46%1.86%0.68%—
Healthy policyMixed fault87.73%1.55%10.72%90.02%
Direct FIDRHealthy99.23%0.17%0.60%—
Direct FIDRMixed fault87.98%1.82%10.20%88.65%
RFAHealthy99.11%0.50%0.39%—
RFAMixed fault91.26%0.83%7.91%92.08%

Under mixed faults, RFA improves SR by 3.53 points over the healthy policy and 3.28 points over Direct FIDR. It also reduces drop rate and non-drop failure rate. Under healthy actuation, RFA remains close to the best reference and exceeds the healthy checkpoint’s SR by 1.65 points.

RFA has the highest mean SR in all six fault categories. The largest degradation of the healthy policy occurs for command bias and range restriction on the index finger, with SR reductions of 51.80 and 36.98 points relative to matched healthy operation. RFA improves SR over the healthy policy in 21 of 24 fault-mode–finger combinations, although the gains are not uniform.

The result is a passive fault-tolerance claim: the deployed controller does not identify the fault explicitly. It uses the time history of command-response mismatch to adjust its action distribution.

8. Ablations

The ablations separate the roles of adaptive sampling, command-response features, recurrence, and the reference KL.

VariantMixed-fault SRDRNDFR
Direct FIDR87.98%1.82%10.20%
RFA, stationary sampling, no KL90.01%2.27%7.72%
RFA, adaptive sampling, no KL90.96%1.62%7.42%
RFA, stationary sampling, with KL90.55%1.66%7.80%
RFA, adaptive sampling, with KL91.26%0.83%7.91%
RFA without command-response block89.14%1.35%9.51%
RFA feed-forward residual90.46%1.73%7.81%

Removing the 48-dimensional command-response block reduces SR by 2.12 points and raises NDFR. Replacing the recurrent residual policy with a feed-forward policy reduces SR by 0.80 points and increases DR. Adaptive sampling contributes mainly through improved success and lower drops. The KL reference shifts the balance between drop and non-drop failures; its effect is not a uniform improvement across every metric.

9. Real-robot study and limits

The physical study uses a LEAP Hand with software-injected command-channel faults. The setup records nominal targets, delivered targets, measured joint positions, and cube rotation. The video protocol contains 210 recordings from three controllers, seven conditions, and ten trials per combination. Faulted trials contain pre-fault, fault-active, and recovery intervals.

The pooled rotation rates over faulted trials are:

ControllerPre-faultFault-activeRecovery
Healthy policy20.09 ± 7.58°/s16.71 ± 10.54°/s20.34 ± 8.50°/s
Direct FIDR34.11 ± 8.63°/s29.23 ± 9.76°/s32.89 ± 9.74°/s
RFA29.27 ± 9.18°/s25.37 ± 6.95°/s29.01 ± 8.40°/s

These physical numbers are descriptive because pre-fault rates differ across controllers. The study demonstrates zero-shot deployment and command-response behavior; it does not measure contact force, object drops, or hardware task-success rates.

The scope is also narrow: single-joint software-level command-channel faults. Concurrent multi-joint failures, electrical faults, friction changes, backlash, actuator degradation, and physically induced failures are outside the reported training and evaluation distribution.

My main takeaway is that the teacher–residual composition gives dexterous manipulation a practical passive fault-tolerance interface. The healthy policy preserves nominal behavior, while a recurrent residual learns how a command should change after the measured joints stop following the requested target. The next challenge is recovery that can request a retreat, change the manipulation direction, or re-establish contact after a jam, together with physical fault injection and force sensing.