[Paper Notes] DexTouch-WM: Learning Action-Conditioned Tactile World Models from Human Touch for Dexterous Robot Manipulation
Published:
This post supports English / 中文 switching via the site language toggle in the top navigation.
TL;DR
DexTouch-WM tests whether collecting more human touch can improve a robot world model while keeping robot data fixed. With 5 hours of robot interaction and human data increased from 0 to 100 hours, Contact-IoU rises from 0.415 to 0.588, while visual PSNR increases from 23.473 to 27.096. Human and robot pretraining tasks are disjoint.
The transfer depends on a deliberately shared interface: matching pressure-sensor layouts, human motion retargeted to robot hand configurations, and a generative model that jointly predicts video and tactile changes. Its most useful architectural details are anatomical tactile tokenization, residual prediction anchored to the initial touch, and action conditioning at each modality’s temporal resolution. Downstream results need more care: better agreement as a policy evaluator does not guarantee better synthetic training data.
Paper and source version
These notes follow the eight-page arXiv:2609.20649v2 PDF, dated September 18, 2026. The authors are Yan Qin, Yue Chen, Wenwei Lin, Shujia Liu, Chuqiao Lyu, Kailun Su, Weiyang Jin, Chenze Yu, Ping Luo, Wenbo Ding, Tianxing Chen, and Renjing Xu, affiliated with HKUST (Guangzhou), Xspark AI, Peking University, the University of Hong Kong, and Tsinghua University. This version is an arXiv preprint; no conference acceptance is stated.
The versioned paper record and PDF are the sources for the mechanisms and reported numbers below. I have not reproduced training or robot experiments.
1. Make human touch compatible with robot supervision
Images can show a hand holding an object while hiding whether its fingers are loading, slipping, or losing contact. Learning those transitions from robot interaction alone is expensive. DexTouch-WM uses wearable sensing to collect human interaction, then makes its observations and action coordinates compatible with the robot training interface.
The human collection rig combines Moxian piezoresistive gloves, Manus finger tracking, Vive wrist tracking, two wrist cameras, and a head camera. Gloves record 360 taxels per hand at 60 Hz; the model retains 320 per hand, consisting of five $4\times4$ fingertip pads and one $15\times16$ palm pad. The robot hand uses the same sensing layout. The sensor acquisition frequency should not be read as a demonstrated world-model inference rate.
Human wrist poses and 25-keypoint hand skeletons undergo calibrated coordinate alignment, geometry normalization, and inverse-kinematics retargeting to WujiHand2. Both domains then use the paper’s 67-dimensional pose representation:
\[a_t=[e_t^L,e_t^R,q_t^L,q_t^R,e_t^c]\in\mathbb R^{67}.\]Here $q_t^L,q_t^R$ are absolute 20-DoF hand configurations; wrist and camera poses use first-frame-relative representations. Camera motion supplies egocentric viewpoint context. This is a pose-conditioned prediction interface: the future motion sequence is supplied to the model.
Shared sensor topology removes the need to learn a separate tactile-domain mapping in this setup. It still leaves differences in hand mechanics and object interaction. The robot data anchors the model in that target domain.
2. Predict video and contact changes together
The modeled distribution is
\[p_\theta(v_{1:T},x_{1:T}\mid v_0,x_0,a_{1:T},\ell),\]where $v$ is RGB, $x$ is bilateral tactile pressure, and $\ell$ is an instruction. A video expert initialized from Wan2.2-TI2V-5B exchanges information with a lightweight tactile expert through masked joint self-attention in Mixture-of-Transformers blocks. Each stream retains its own feed-forward pathway. The video VAE and pretrained tactile codec are frozen during world-model training; the first visual latent stays clean as context.
flowchart TD
H["Human RGB, pressure and tracked motion"] --> A["Shared tactile layout and retargeted robot poses"]
R["Robot RGB, pressure and recorded motion"] --> A
A --> V["Video VAE latents"]
A --> T["Split-Hands tactile latents"]
V --> M["Video and tactile experts with joint attention"]
T --> M
P["Instruction and actions at two temporal resolutions"] --> M
M --> O["Future RGB and bilateral pressure maps"]
Preserve the sensing anatomy
Placing every taxel on one image canvas creates artificial neighbors between physically disconnected finger and palm regions. Split-Hands uses one shared encoder across the ten fingertips and another across the two palms. Pad identity embeddings retain which region produced each token. Contact-weighted reconstruction pretraining prevents the many inactive taxels from overwhelming the sparse pressure events.
Figure 4 compares this representation with a single-grid encoder: Split-Hands has lower validation reconstruction error and faster Contact-IoU convergence. The useful design principle is to share local pressure features while preserving the physical identity of each sensing region.
Anchor tactile dynamics to the initial observation
With tactile encoder $E_x$ and decoder $D_x$, the model predicts
\[z_t=E_x(x_t),\qquad r_t=z_t-z_0,\qquad \hat x_t=D_x(z_0+\hat r_t).\]The clean $z_0$ remains available throughout prediction. A sustained initial grasp therefore provides a baseline, and the modeled residual describes subsequent changes in pressure. This also clarifies the meaning of “residual”: every step is anchored to the initial encoding, so $r_t$ is not a consecutive-frame difference accumulated through time.
Match action timing to each latent stream
The causal video VAE compresses future frames by a factor of four, while tactile tokens retain the observation rate. Action conditioning follows those two clocks:
\[c_j^v=\phi_v(a_{4j-3:4j}),\qquad c_t^x=\phi_x(a_t), \qquad c_0^v=c_0^x=0.\]Four-frame pose chunks condition video latents; frame-level poses condition tactile latents. AdaLN injects these signals into each expert’s timestep modulation. This preserves finer contact timing without requiring the video stream to operate at the same latent rate.
The training objective combines conditional flow-matching losses:
\[\mathcal L=\mathcal L_{\mathrm{video}}^{\mathrm{FM}} +\lambda_{\mathrm{tac}}\mathcal L_{\mathrm{tactile}}^{\mathrm{FM}}.\]The first visual latent is excluded from prediction loss. Tactile loss supervises residuals on contact or temporally changing tokens, limiting the contribution of static empty regions. Joint attention provides cross-modal interaction without an extra alignment loss. These equations describe the paper’s objective; the PDF does not provide a complete optimizer, batch-size, and training-schedule recipe.
The action-injection ablation is informative. Adding touch with cross-attention conditioning leaves visual PSNR almost unchanged but reduces Trajectory Accuracy from 0.884 to 0.850. AdaLN brings it to 0.891, and raises Contact-IoU from 0.388 to 0.415. Joint prediction needs an effective action interface to preserve motion consistency.
3. What scales when the robot budget stays fixed?
The pretraining corpus contains about 100 human hours across 50 tasks and 5 robot hours across six tasks. Four additional downstream tasks have separate human and robot demonstrations, explaining the paper’s totals of 54 human tasks and 10 robot tasks. The robot platform uses a Tianji arm with a 20-DoF Wuji hand.
Table I evaluates all scaling variants on the same held-out robot episodes from the six pretraining tasks. Thus, the experiment establishes transfer from disjoint human tasks into robot-domain prediction on held-out episodes; it does not establish zero-shot prediction on arbitrary new robot tasks.
| Metric | 0h human | +10h | +50h | +100h |
|---|---|---|---|---|
| Visual PSNR ↑ | 23.473 | 24.032 | 26.874 | 27.096 |
| Trajectory Accuracy ↑ | 0.891 | 0.900 | 0.962 | 0.962 |
| Geometry Error ↓ | 0.131 | 0.112 | 0.102 | 0.100 |
| Tactile MSE ↓ | 0.093 | 0.108 | 0.038 | 0.029 |
| Contact-IoU ↑ | 0.415 | 0.417 | 0.534 | 0.588 |
| Contact-F1 ↑ | 0.551 | 0.554 | 0.659 | 0.706 |
The same Split-Hands + AdaLN architecture and 5h robot set underlie all four columns. Adding 10h barely changes contact overlap and worsens tactile MSE. Most gains arrive at 50–100h. Visual trajectory accuracy has already plateaued by 50h, while contact prediction continues improving. The evidence supports scaling in this measured range, with neither uniform gains at every increment nor an established asymptotic scaling law.
4. A useful evaluator can still overestimate performance
For each downstream task, 300 robot demonstrations are split into $A$, $B$, and $C$, with 100 trajectories each. Starting from the same human-plus-robot pretrained checkpoint, WM-Robot adapts on $B+C$; WM-Mix adapts on $B$ plus 100 task-specific human demonstrations. The name WM-Robot refers to adaptation data: both models inherit human pretraining.
FTP-1, $\pi_{0.5}$, and X-VLA train on the same $A+B$ robot demonstrations. Each policy then runs in the real world and in both imagined environments, with predicted observations feeding subsequent actions. Each policy–task pair receives ten matched rollouts per environment and five independent human raters per rollout. Scores reflect normalized task progress, with a penalty for obvious physical inconsistencies in imagined rollouts. They are not generally binary success rates.
Across four tasks, WM-Mix improves mean Pearson correlation with real scores from 0.646 to 0.844, and lowers mean maximum rank violation (MMRV) from 0.025 to 0.017. This supports replacing half the robot adaptation demonstrations with human demonstrations in the tested evaluator setup. Ranking errors persist on Stand Bottle, and each task’s correlation uses only three policy scores.
Absolute calibration is weaker. On Place Shoes, $\pi_{0.5}$ scores 0.450 in reality and 0.900 in WM-Mix. A model can preserve policy ordering while giving optimistic scores. The separately reported pooled correlation of 0.925 compares WM-Robot with WM-Mix; it is not their correlation with reality. These distinctions matter if imagined evaluations will determine which policy gets hardware time.
5. Synthetic observations have uneven training value
The data-generation experiment compares policies trained on 200 real trajectories with policies trained on 100 real plus 100 synthetic trajectories. Each synthetic episode reuses a held-out $A$ trajectory’s initial observations and recorded action sequence, replacing its future RGB and tactile observations with model predictions. Split $A$ is excluded from world-model training. This tests observation synthesis along recorded actions, so it does not establish collecting the synthetic half without access to real motion sequences.
The four-task mean real-robot scores are:
| Policy | 200 real | 100 real + 100 WM-Robot | 100 real + 100 WM-Mix |
|---|---|---|---|
| FTP-1 | 0.731 | 0.688 | 0.694 |
| $\pi_{0.5}$ | 0.625 | 0.563 | 0.494 |
| X-VLA | 0.500 | 0.506 | 0.400 |
These are the paper’s rounded means from Table IV. FTP-1 retains much of its all-real performance with WM-Mix data, but $\pi_{0.5}$ loses more, and X-VLA reaches zero on Stand Bottle with that mixture. Better prediction metrics and better evaluator agreement do not settle whether synthetic observations contain the cues a particular policy needs to learn.
My reading is that the strongest result is the controlled transfer experiment: aligned human touch improves robot contact prediction at a fixed robot-data budget. For reuse, I would start with the sensing correspondence, anatomical tokenizer, and timing alignment. I would require a policy-specific real-robot check before replacing demonstrations with generated observations. Evidence across different sensor layouts, multiple robot-hand embodiments, and broader policy sets would be needed before treating this as a general substitute for robot data.
