[Paper Notes] WM-Craftnet: World Synesthesia Model for Generalizable and Robust Dexterous In-Hand Manipulation

18 minute read

Published:

This post supports English / 中文 switching via the site language toggle in the top navigation.

TL;DR

Dexterous in-hand manipulation is a state-estimation problem hidden inside a control problem. A hand observes joint motion, intermittent touch, and noisy partial depth, then must infer object geometry, contact evolution, drift, and slip quickly enough to keep manipulating. WM-Craftnet learns that hidden interaction state with a World Synesthesia Model (WSM): a Dreamer-style recurrent state-space model trained on proprioception, tactile contact, wrist depth, actions, and several predictive targets.

The design uses the world model in a specific way. It does not optimize the actor through imagined rollouts. The WSM’s deterministic recurrent state is detached, compressed from 512 dimensions to 16, and supplied as context to a PPO actor. Noisy depth enters the model while clean simulated depth is the reconstruction target, turning denoising into task-relevant representation learning. A WSM pretrained on nine objects then initializes learning on 49 new objects.

The results support three claims. On the nine-object $z$-axis benchmark, the full pretrained WSM reaches 753.3 return, compared with 386.9 for the strongest raw-sensor baseline in that table. On 20 real objects, WM-Craftnet succeeds in 175/200 trials, while the four baselines score between 33/200 and 53/200. The mechanism ablations also matter: noisy-depth supervision, removing recurrent input, and replacing recurrent context with a per-frame feature all reduce return.

The boundary is equally useful. Severe simulated perturbation recovery reaches only 14.1%. Axis-specific policies use separately trained WSMs, the 49-object result includes downstream adaptation and a newly learned controller, and screwdriver use remains qualitative. I read WM-Craftnet as a strong predictive-state recipe for sustained visuotactile control, with long-horizon recovery and broader skill transfer still open.

Paper and source version

WM-Craftnet: World Synesthesia Model for Generalizable and Robust Dexterous In-Hand Manipulation is by Jie Yin, Zeyuan Zhao, Xiaojing Tan, Yang Liu, Chiyu Wang, and Xinyang Gu of Sharpa Robotics. The paper was accepted to CoRL 2026.

These notes follow the 15-page arXiv:2609.07002v1 PDF, submitted September 7, 2026. The arXiv entry and official project page provide the primary materials. Reported numbers come from the paper; I have not reproduced the simulation training or hardware experiments.

1. Infer the interaction state, then control it

Continuous in-hand rotation exposes the limits of a single observation. Depth reveals global hand-object geometry but is partial and noisy near contact. Binary tactile sensing localizes contact without directly revealing object pose or shape. Proprioception records how the hand moved; recent actions explain which motion produced the sensory change. Their history contains more information than any frame alone.

WM-Craftnet formalizes the deployed observation as

\[o_t=\left(p_t^{\mathrm{stack}},c_t^{\mathrm{stack}},D_t,u_{t-1}\right),\]

where $p_t^{\mathrm{stack}}$ and $c_t^{\mathrm{stack}}$ are short histories of normalized proprioception and binary contact, $D_t$ is the current noisy wrist-depth input, and $u_{t-1}$ is the previous joint-target command. The actor also receives a commanded rotation axis $a\in{\pm x,\pm y,\pm z}$ and a projected WSM feature. Object identity and privileged object state are absent from the actor input.

The target is signed angular progress around the commanded axis:

\[\Delta\theta_t=\operatorname{proj}_a\!\left( \log\!\left(R^o_{t+1}(R^o_t)^{-1}\right) \right).\]

This task requires maintaining the object inside a controllable contact region while accumulating target-axis rotation. A memorized finger gait can work for one geometry and initial pose; offsets, changed mass distribution, or external force break its timing. WSM supplies a recurrent estimate that can change with observed interaction dynamics.

2. The World Synesthesia Model is a predictive state encoder

Each WSM update receives single-step proprioception $p_t$, tactile/contact $c_t$, noisy depth $I_t$, previous action $u_{t-1}$, and an episode-start flag $m_t$. MLP branches encode the low-dimensional signals, a CNN encodes depth, and a multimodal encoder produces $e_t$. A recurrent state-space model then updates stochastic state $z_t$ and deterministic memory $h_t$:

\[e_t=f_{\mathrm{enc}}(p_t,c_t,I_t),\] \[(z_t,h_t)=f_{\mathrm{wm}}(z_{t-1},h_{t-1},u_{t-1},e_t,m_t).\]

The actor-side world-model feature is $w_t=h_t$. A small MLP maps the 512-dimensional recurrent state through $512\rightarrow64\rightarrow32\rightarrow16$. The resulting 16-dimensional vector is detached before entering the policy. Gradients from PPO therefore do not reshape the world model through the actor interface.

flowchart TD
    A["Proprioception p_t"] --> E["Multimodal encoder"]
    B["Binary tactile c_t"] --> E
    C["Noisy wrist depth I_t"] --> E
    D["Previous action u_{t-1}"] --> F["Dreamer-style RSSM"]
    E --> F
    G["Previous recurrent state"] --> F
    F --> H["Predict clean depth, proprioception, touch, reward, pose, value and shape"]
    F --> I["Detached h_t: 512 → 16"]
    J["Current deployable observation o_t"] --> K["Asymmetric PPO actor"]
    L["Commanded axis a"] --> K
    I --> K
    K --> M["Relative 22-DoF joint targets"]

This separation gives the two learning systems distinct jobs. WSM learns which latent state makes multimodal interaction predictable. PPO learns which action to take given current sensors, the command, and that state. The critic can use simulator-only information during training; the actor and WSM inputs remain deployable.

3. Noisy input and clean target make denoising part of dynamics learning

The rollout buffer stores

\[\mathcal D=\{(p_t,c_t,I_t,I_t^{\mathrm{clean}},u_{t-1},r_t,m_t)\}_{t=1}^{T}.\]

Input depth is cropped around the hand-object workspace and corrupted with temporally correlated dropout, Gaussian noise, and small image rotations. The image decoder must reconstruct the clean crop-only simulated target. The core objective is

\[\mathcal L_{\mathrm{wm}}= \mathcal L_{\mathrm{img}}+ \mathcal L_{\mathrm{prop}}+ \lambda_c\mathcal L_{\mathrm{tac}}+ \lambda_r\mathcal L_{\mathrm{reward}}+ \lambda_{\mathrm{dyn}}\mathcal L_{\mathrm{dyn\text{-}KL}}+ \lambda_{\mathrm{rep}}\mathcal L_{\mathrm{rep\text{-}KL}}+ \mathcal L_{\mathrm{aux}}.\]

The reconstruction and reward heads ask the recurrent state to preserve geometry, body configuration, contact, and task progress. The full model adds simulator-supervised object pose, value, and object-shape heads. Shape uses a basis-point-set-to-mesh displacement vector. These labels are training objectives; decoded clean depth, pose, and shape do not become deployment-time actor inputs.

That distinction prevents an easy misreading of “world model.” WM-Craftnet performs an observation update at every control step and uses the current deterministic state as policy context. The paper does not train PPO on latent imagined trajectories. Its world model behaves like a multimodal, action-conditioned state estimator whose internal state has been shaped by future-relevant prediction questions.

Online training alternates the two paths. PPO rollouts enter a replay ring buffer; after each PPO epoch, contiguous chunks update the RSSM. For prior transfer, the encoder, RSSM, depth/proprioceptive decoders, and reward predictor initialize the downstream model and continue adapting on new rollouts.

4. Control and sim-to-real share one deployable interface

The actor computes

\[u_t=\pi_\theta\!\left(o_t,a,\psi(h_t)\right),\]

and applies the output as a relative target update,

\[q_t^{\mathrm{target}}=q_{t-1}^{\mathrm{target}}+\alpha u_t.\]

Signal-level smoothing and joint-limit clamping precede execution. The Sharpa Wave hand has 22 DoF. Hand control, wrist depth, tactile input, simulation observations, and the policy all run at 10 Hz; simulation physics runs at 60 Hz. The policy MLP has widths $[512,256,256]$.

The reward combines signed spin with terms for object velocity, useful fingertip contact, finger distance, torque, work, action size, tracking, hand pose, object drift, off-axis motion, and prolonged lack of spin. Reset conditions terminate drops, excessive axis deviation, sustained low spin, and horizon completion. All baselines use the same reward, randomization, initialization, and reset strategy.

For transfer, the authors identify joint dynamics from real 1 Hz sinusoidal commands, then tune simulated PD gains to match amplitude and phase. The reported average sim-real joint-tracking error after calibration is below $0.2^\circ$. Domain randomization covers object mass, friction, PD gains, observation and action noise, reset pose, external force, and gravity direction. Tactile thresholds, dropout, and latency are also varied. This combination is important: the learned state can denoise only the variation represented by data and predictive supervision.

5. The ablations separate memory, prediction, and geometry supervision

The main nine-object $z$-axis study averages each simulation entry over 128 evaluation episodes and reports 95% confidence intervals. The selected rows below expose the mechanism more clearly than a single best score.

VariantReturn ↑Episode length ↑Rotation rate ↑Off-axis ↓Angular variation ↓
Touch Dexterity386.9362.51.0181.6011.764
In-Hand Rotation, depth + touch236.1272.50.8821.9862.487
WM-Craftnet from scratch414.3264.20.7421.2851.735
LSTM, no predictive objective621.0424.31.1861.2971.392
Noisy-depth supervision708.0431.11.2651.2991.456
No $h_{t-1}$ recurrent input705.4434.31.2821.1251.185
Per-frame encoder-decoder feature667.6417.41.2541.3231.343
Full pretrained WSM753.3435.21.2931.2251.324

Several conclusions survive the metric trade-offs. Recurrence alone helps: the LSTM exceeds raw-sensor baselines. Predictive multimodal training adds another large gain. Clean-depth targets improve return by 45.3 over noisy-depth reconstruction. Supplying the recurrent state to the policy is stronger than a per-frame encoder-decoder feature.

The full model does not dominate every stability column. Removing $h_{t-1}$ produces the lowest off-axis motion and angular variation in this controlled table, while return drops by 47.9. The evidence supports a task-performance advantage for the full recurrent model, plus a stability–progress trade-off that deserves separate reporting.

Test-time tactile masking degrades performance gradually: full WSM scores 753.3 return, 25% dropout scores 743.4, and all-zero tactile scores 724.9. The small gap does not make touch irrelevant. Modality-head controls score 684.5 with proprioception, 698.9 with proprioception plus touch, and 753.3 with the full depth–touch state. Tactile contact is one cue inside a representation jointly trained to infer interaction state.

Auxiliary heads offer a second view. Removing all pose, value, and shape heads gives 688.1 return. The value head alone reaches 737.0; all heads recover 753.3. Simulation-only labels can therefore improve the deployable feature without appearing at test time, a familiar asymmetric-learning pattern applied inside the world model.

6. Pretraining transfers predictive structure, followed by adaptation

The WSM is first trained on nine $z$-axis objects. Its predictive components initialize a downstream experiment with 49 new objects, while the actor–critic and task-specific heads are learned for that distribution. The transferred model continues updating from downstream rollouts.

After 3,000 epochs, each downstream object rotates $9.37\pm0.13$ rad per episode on average, compared with 3.28 rad for the no-prior baseline. The training fall rate for the prior-based run decreases from 6% to 0.3%. In a separate five-block scaling diagnostic, rotation rate rises from 1.10 when trained from scratch to 1.34 with the nine-object prior and 1.46 with the 49-object prior.

This is evidence for reusable initialization, not frozen zero-shot control over 49 objects. The transfer contains three moving parts: predictive weights arrive from the smaller distribution, the WSM adapts to new rollouts, and a downstream controller is trained. A useful follow-up would hold the WSM frozen, vary downstream data size, and report how quickly representation reuse reduces controller sample complexity.

The paper’s t-SNE plots show partially object-dependent regions, shared regions, and locally coherent temporal trajectories. Object labels color the visualization and are absent from policy inputs. These plots are consistent with an interaction representation that carries both geometry and phase; they cannot establish which physical variable each dimension encodes.

7. Rotation generalizes across objects, while recovery remains hard

The $x$-, $y$-, and $z$-axis policies are evaluated directly on four held-out objects per axis. Each axis uses its own training object set and separately trained WSM.

Held-out simulation setBest non-WSM rotation rateWM-Craftnet rotation rateBest non-WSM returnWM-Craftnet return
$x$ axis0.9231.002246.0333.6
$y$ axis0.3780.715193.0203.5
$z$ axis0.6960.842301.7533.1

The strongest baseline can differ between the two metric columns. WM-Craftnet leads both metrics for all three axes. On the seen-object axis benchmarks, it also obtains the lowest fall rate for $x$ and $y$ rotation. These results suggest that recurrent interaction inference matters most when gravity and changing contacts remove the passive support available in palm-up $z$-axis rotation.

Hardware evidence is broad for standard $z$-axis rotation. Across 20 objects and 10 trials per object, WM-Craftnet records 175/200 successes. Blind RL, Touch Dexterity, depth-only RL, and In-Hand Rotation obtain 41, 33, 45, and 53 successes. In the smaller four-object table, WM-Craftnet rotates the seen duck by 16.179 rad in 20 seconds with 10/10 success. On the unseen double-notched block it reaches 4.32 rad and 8/10; the WSM-denoised-depth version of the In-Hand Rotation baseline reaches 0.80 rad and 0/10. Denoising helps perception, while recurrent task context supplies additional control-relevant history.

The perturbation test gives the sober number. Under severe simulated disturbances, WM-Craftnet recovers in $14.1\%\pm6.0\%$ of trials, versus 6.2% for the strongest baseline success rate. Recovered trials take 2.56 seconds on average, the fastest result. The relative improvement is real; most disturbed episodes still fail.

8. What the paper establishes—and what it leaves open

Strong evidence: predictive multimodal state learning improves a fixed PPO-style control pipeline; clean-depth targets matter; the learned components are useful initialization for a larger object distribution; and the resulting controller transfers to a real 22-DoF hand across a substantial object set.

Open questions:

  • The main tasks are short-horizon rotations. Long-duration drift, jamming, and exits from the hand workspace remain frequent failure modes.
  • Severe recovery has a low absolute success rate and a wide confidence interval.
  • Axis transfer is not demonstrated with one shared axis-conditioned model; each axis has a separately trained WSM and policy.
  • The 49-object study measures transfer plus adaptation, so it does not isolate frozen representation reuse or zero-shot control.
  • Screwdriver translation and rotation show interface breadth qualitatively. They do not yet form a quantitative tool-use benchmark.
  • Baselines are reimplemented on the Sharpa hand under shared training settings. This improves control inside the study, while leaving cross-implementation sensitivity and compute-matched comparisons worth checking.

Takeaways for research and practice

The most reusable idea is to define the latent state by questions the controller needs answered. Clean geometry, contact, reward, pose, value, and shape each constrain a different ambiguity in partial observation. Their predictions can be discarded at deployment while the recurrent state carries the useful structure.

For a new dexterous task, I would preserve the three-way separation: deployable noisy sensors feed the recurrent model, privileged simulation labels shape predictive heads, and the policy receives a compact stopped-gradient feature. I would then add two evaluations early. The first freezes the representation to measure genuine reuse. The second creates a graded recovery benchmark—pose offsets, force impulse, contact loss, and jamming—so progress is visible before disturbances reach an almost unrecoverable regime.

WM-Craftnet also suggests a practical criterion for calling a world model useful. Pixel reconstruction quality is secondary. The decisive test is whether its state lets the same controller remain coherent when geometry, contact, and sensory quality shift. Here the ablations and 200-trial hardware study make a credible case; the 14.1% recovery result shows exactly where the next model must improve.