[Paper Notes] Dex-X: Learning Visual-Tactile Dexterous Manipulation From Human Videos with Simulated Interaction
Published:
This post supports English / 中文 switching via the site language toggle in the top navigation.
TL;DR
Human videos show where a hand and object move, but they do not record the contact forces that keep a grasp stable. DEX-X fills that gap by reconstructing a monocular video, retargeting the motion to a Franka FR3 and Sharpa Wave hand, and then replaying the task as a physical interaction in simulation. A privileged RL policy gets fingertip forces there. A second policy learns from the expert and runs on a depth point cloud, proprioception, and real tactile readings.
The state expert averages 65.9% success over six simulated task categories. Without real-world fine-tuning, the student reaches 28/30 on cube picking, 24/30 on cup pouring, 22/30 on cup lifting, and 16/30 on squeegee table cleaning. On cube picking, removing tactile input lowers success from 28/30 to 11/30.
The part I would keep is the use of simulation as a missing-modality generator. The main qualification is that the simulator does not recover touch directly from pixels. It supplies force feedback while a robot policy follows a reconstructed motion prior under a modeled contact system. The deployed policy still receives that retargeted reference. Its tactile input contains five force magnitudes; spatial pressure and shear are absent.
Paper and source version
Ruoqu Chen, Feixiang Ruan, Liu Cao, Zihao Wang, Botian Xu, Shiqin Tong, Jiajun Liu, Mingzhi Pei, Chenyu Zhang, Wanli Xing, Kaifeng Zhang, and Mengdi Xu wrote Dex-X: Learning Visual-Tactile Dexterous Manipulation From Human Videos with Simulated Interaction. The affiliations listed in the paper are Tsinghua University, Shanghai Qizhi Institute, Sharpa, Tongji University, and Renmin University.
These notes follow the 22-page arXiv:2609.07747v2, revised September 9, 2026. The project page provides the system overview and rollout videos. I read the paper and appendix; I did not rerun the training or hardware evaluation.
1. “Tactile completion” happens through interaction
DEX-X begins with monocular RGB demonstrations. For every sequence, FoundationPose estimates the object’s 6D pose, WiLoR initializes the human hand, and a MANO optimization refines the hand trajectory. A joint hand-object pass reduces temporal jitter and penalizes mesh interpenetration. Object meshes are prepared separately through 3D scanning or single-image reconstruction.
This processing yields wrist poses, 3D hand keypoints, and object poses at 30 Hz. None of those channels contains measured touch. DEX-X creates the missing signal later: it transfers the trajectory to a robot in IsaacLab, lets a policy interact with the object, and reads force from simulated fingertip contact sensors. The force label therefore depends on the simulator, robot morphology, controller, and contact path explored by RL. It is physically structured supervision, though it isn’t a recovered measurement of what the demonstrator felt.
That distinction matters. The human video supplies task timing and a motion prior; simulation supplies a plausible robot-side contact experience. Calling this tactile completion is useful because it highlights what the video lacks, as long as “completion” isn’t read as frame-by-frame tactile reconstruction.
flowchart TD
A["Monocular human video"] --> B["Object pose + MANO hand reconstruction"]
B --> C["Spatial augmentation and robot retargeting"]
C --> D["Privileged PPO expert in simulation"]
E["Simulated fingertip forces and contact dynamics"] --> D
D --> F["DAgger action supervision"]
G["Depth scene + real fingertip forces"] --> H["Point-cloud student at 30 Hz"]
F --> H
H --> I["Zero-shot real-robot execution"]
2. Retargeting preserves the interaction, then filters what the robot cannot reach
Before retargeting, each reconstructed trajectory receives one shared planar transform. The yaw perturbation is sampled from $[-10^\circ,10^\circ]$, and the object is moved to a workspace anchor plus a translation sampled within $\pm5$ cm on each planar axis. Applying the same transform to the wrist, hand keypoints, and object keeps their relative geometry intact.
Retargeting then runs in two stages. The first holds the hand in a nominal pose and fits arm joints 2–7 to the human wrist trajectory; the base-yaw joint stays fixed. The second releases all seven arm joints and optimizes the 22 hand joints against MANO keypoints. The thumb and index finger get the largest weights, and distal points matter more than proximal ones. This bias makes sense for grasp acquisition, but it also tells us that the transfer objective already encodes a view of which parts of the demonstration matter.
The hand limits use empirically measured hardware ranges, which can be tighter than the URDF. After optimization, any augmented trajectory with a mean end-effector position error above 8 cm is discarded. Kinematic retargeting is thus a proposal mechanism plus a reachability filter. It does not need to produce a dynamically successful motion; the RL stage gets that job.
3. The expert tracks the video while learning contact
The state expert uses PPO with an asymmetric actor-critic across 4,096 parallel environments. Its 557-dimensional actor observation contains proprioception, a 311-dimensional wrist-and-hand reference, the target object pose, fingertip-to-object distances, a 128-dimensional BPS object encoding, a noisy current object pose, and tactile channels. The critic adds 148 privileged dimensions, including ground-truth object dynamics and five future target states.
The active tactile observation is deliberately small: one scalar contact-force magnitude per fingertip. Each value averages two recent samples. Training perturbs those readings with 20% multiplicative Gaussian noise, drops each finger with 5% probability, and occasionally holds the previous value. Fifteen slots for contact positions remain zero in the reported model.
The policy outputs 29 commands: seven arm joint deltas and 22 absolute hand targets. It runs at 30 Hz. The reward keeps the demonstrated motion visible throughout training:
\[r_t=r_t^{\mathrm{wrist}}+2r_t^{\mathrm{hand,abs}}+r_t^{\mathrm{hand,rel}} +r_t^{\mathrm{object}}+r_t^{\mathrm{contact}}+r_t^{\mathrm{action}} +r_t^{\mathrm{success}}+r_t^{\mathrm{collision}}.\]The contact group rewards fingertip force, approaching the object, and low slip. Object mass spans 0.01–0.15 kg; the training also randomizes center of mass, friction, PD gains, action delay, pose noise, bias, latency, dropout, and tactile sensing. Motion tracking keeps the task recognizable while randomized physics forces the policy to find contacts that survive perturbation.
This is the paper’s central compromise. Pure retargeting stays close to the human motion and often fails physically. Unconstrained RL could abandon the demonstration. DEX-X uses the reference as a dense scaffold, then lets contact-aware RL correct the executable details.
4. Distillation replaces privileged geometry with a visual-tactile point cloud
The expert sees object geometry and state that are inconvenient or unavailable at deployment. DEX-X trains one multi-task student with DAgger. The student drops the BPS vector, reference fingertip distances, and noisy current object pose, leaving 417 scalar features. It keeps proprioception, fingertip forces, and the retargeted motion reference, including the target object pose.
The replacement visual branch contains 1,055 points:
| Point set | Count | What it carries |
|---|---|---|
| Depth-derived scene | 1,024 | XYZ, point type, zero force |
| Wrist and fingertips | 6 | XYZ, point type, zero force |
| Fingertip surface samples | 25 | XYZ, point type, corresponding fingertip force |
A shared PointNet converts the $1055\times5$ array into a 64-dimensional feature, so the final student input has 481 dimensions. The 25 tactile points do not report 25 independent contact measurements. Five surface samples per finger repeat that finger’s scalar force, placing the reading near the relevant geometry. The original five force values also remain in the vector observation.
During DAgger iteration $k$, execution mixes expert and student actions:
\[a_t^{\mathrm{exec}}=\beta_k a_t^{\mathrm{teacher}}+(1-\beta_k)a_t^{\mathrm{student}}, \qquad \beta_k=(0.85)^k.\]The paper uses 30 iterations, 4,096 rollout steps per iteration, and a replay buffer capped at 200,000 transitions. Mean-squared action regression trains the student for eight epochs after each rollout. This schedule moves data collection gradually onto the states caused by the student itself.
5. The simulation average hides two revealing reversals
The benchmark totals 14 human demonstrations across pick-up, peg insertion, tool use, in-hand rotation, in-hand translation, and bimanual handover. The expert is compared with DAPG, a ManipTrans-style residual imitation policy, and direct kinematic retargeting.
| Category | DEX-X overall | DAPG | ManipTrans | Kinematic retargeting |
|---|---|---|---|---|
| Pick up | 89.6% | 66.7% | 1.5% | 1.8% |
| Peg insertion | 43.1% | 52.8% | — | 0.0% |
| Tool use | 74.3% | 70.1% | 35.0% | 24.1% |
| In-hand rotation | 36.8% | 0.7% | 56.9% | 0.0% |
| In-hand translation | 64.8% | 12.5% | 11.5% | 0.0% |
| Bimanual handover | 86.6% | — | 4.7% | — |
| Reported average | 65.9% | 40.6%* | 21.9% | 5.2% |
* DAPG’s average covers the five single-hand categories.
The overall result is strong, especially for sustained tool contact and handover. Still, DEX-X does not win every row. DAPG is 9.7 points better on peg insertion, and ManipTrans is 20.1 points better on in-hand rotation. Those reversals are useful. Force-aware RL helps most where a policy must acquire or preserve contact through a long interaction; a well-matched imitation objective can still be better for a narrower trajectory-tracking task.
The table also distinguishes entered-stage success for reach, grasp, and manipulation from overall success. DEX-X reaches 85.5% on average, grasps 78.0%, enters successful manipulation at 66.2%, and finishes at 65.9%. Most of its loss occurs before the final manipulation segment.
6. The modality ablations are more convincing than the representation comparison
In simulation, the point-cloud policy with force reaches 44% strict success and 58% relaxed success. The strict criterion requires final position within 3 cm; the relaxed criterion uses 5 cm, with a $30^\circ$ threshold for rotation tasks. A depth-based force policy gets 32% strict success. Removing force from the point-cloud model lowers relaxed success from 58% to 45%. Scalar force, binary contact, and 3D force encodings perform similarly in the reported comparison, so the deployed model uses scalar force to match the hardware.
The cleaner test happens on the real cube-picking task. All four variants come from the same teacher:
| Deployment observation | Success |
|---|---|
| Full visual-tactile | 28/30 (93.3%) |
| No vision | 14/30 (46.7%) |
| No tactile | 11/30 (36.7%) |
| Proprioception only | 8/30 (26.7%) |
Vision helps the hand find the cube; tactile feedback helps it keep the cube after contact. The 30-trial sample is too small to rank the two modalities precisely, but either ablation loses roughly half or more of the successful trials. That is direct evidence that the student uses both inputs on hardware. Proprioception and the reference alone do not carry the teacher’s behavior.
7. Real transfer works best near the training geometry
DEX-X deploys on the 29-DoF arm-hand system with no task-specific real-world fine-tuning. Cube picking succeeds in 28/30 trials, cup pouring in 24/30, cup lifting in 22/30, and squeegee manipulation in 16/30. The squeegee failure analysis locates the main losses early: 93% of trials reach the grasp without collision, 67% grasp the tool, and 53% retain it without a later collision or drop. All trials that pass that third gate complete the final manipulation. For the cube, final success stays at 93%. Here, zero-shot means that the policy receives no further training on real task data; hardware preparation still includes measuring the hand’s joint ranges and calibrating simulated damping with sinusoidal trajectories from the real FR3.
The appendix compares simulated and real cube-picking rollouts. Mean wrist-trajectory error is 2.83 cm, with 0.46–0.82 cm cross-episode standard deviation. Error grows from 1.1 cm during approach to 2.7 cm during grasp and 3.8 cm during lift, mainly along the load-bearing $x$ axis. The real wrist also lifts about 5.7 mm higher. The authors attribute the pattern to arm compliance under payload, object-placement variation, and camera calibration. Fingertip-force profiles show similar timing and allocation across domains, although that part of the analysis is qualitative.
Geometry generalization is much less uniform:
| Cube-picking object | Training status | Success |
|---|---|---|
| Training cube | Seen | 28/30 (93.3%) |
| Thin cube | Unseen | 23/30 (76.7%) |
| Square cube | Unseen | 8/30 (26.7%) |
| Big duck | Unseen | 8/30 (26.7%) |
| Small duck | Unseen | 7/30 (23.3%) |
The thin cube result supports local geometric transfer. The larger shape and category changes expose the boundary: success falls to 23–27%, only a little above the proprioception-only result on the training cube. “Zero-shot generalization” is accurate here, but broad object generality would overstate the evidence.
8. What I would carry forward
DEX-X makes simulation do more than provide cheap rollouts. It turns visual motion into a robot interaction and generates the sensory channel the source data never contained. That is a useful pattern for other missing modalities: reconstruct what can be observed, execute under a model, and train a deployable policy on the model’s extra signals.
The first experiment I would run next concerns the reference. The deployed student still consumes 311 dimensions of retargeted wrist-and-hand motion plus the target object pose. Replacing that stream with a sparse task goal, a short video embedding, or an object-centric phase variable would show how much skill has entered the policy and how much remains in the trajectory supplied at test time.
I would also test richer tactile data without changing the rest of the pipeline. The current representation repeats five contact-force magnitudes across 25 surface points; it cannot state where contact occurs within a fingertip, which way shear acts, or whether the surface is beginning to slip. Those signals should matter most on the squeegee task, where grasp acquisition and post-grasp retention dominate failure.
Finally, scale is still a research question. Fourteen demonstrations establish that the method can span several task families, not that it can already absorb Internet-scale video. The difficult part will be deciding which reconstructions and retargeted variants deserve expensive physical interaction in simulation, especially when object meshes, tracking quality, and task phase are uncertain.
