[Paper Notes] Morphometric Imitation: From Morphology and Contact Aware Hand Retargeting to Sim-to-Real Visuomotor Policy

14 minute read

Published:

This post supports English / 中文 switching via the site language toggle in the top navigation.

TL;DR

Morphometric Imitation transfers a reconstructed human hand–object interaction through three representations: a robot-compatible geometric reference, a physically executable teacher trajectory, and a policy driven by depth observations. Morphometric optimization (MMO) first reshapes MANO to match the target hand, recovers the original contact locations, and solves robot inverse kinematics. Residual reinforcement learning then corrects that reference using object-motion and contact supervision. A point-cloud visuomotor student learns from the resulting simulated demonstrations.

The central result is that reference quality changes what downstream RL can learn. With the same residual-RL formulation, MMO references improve success over the strongest baseline from 53.2% to 82.5% on Dex3, 68.2% to 69.9% on Allegro, and 56.5% to 91.8% on Sharpa. On hardware, category-specific Sharpa policies succeed in 268 of 300 trials (89.33%), using no real-world training data. The experiment covers ten categories with three real instances each; it does not establish a single policy spanning categories or hardware transfer to all three hands.

Paper and source version

Tara Sadjadpour, Siming He, C.K. Wolfe, Haozhi Qi, Lea Wilken, S. Shankar Sastry, Claire Tomlin, and Jitendra Malik, University of California, Berkeley. These notes follow the 24-page arXiv:2609.28660v1, submitted September 23, 2026, including its evaluation appendices. The source is an arXiv preprint; no conference acceptance is assumed. See the paper PDF and official project page. Results below are reported by the authors and have not been independently reproduced here.

1. Preserve the interaction through three changes of representation

Human and robot hands differ in palm size, finger lengths, finger count, and joint structure. Matching fingertip vectors can reproduce a recognizable gesture while missing the object region that makes the grasp work. Even a geometrically plausible grasp can penetrate an object, slip, or collide with the table when executed under physics. Finally, a teacher that observes exact object state cannot be deployed directly from a depth camera.

The method assigns these problems to successive stages:

flowchart TD
    A["Human MANO motion + object mesh and trajectory"] --> B["Match MANO morphology to the robot"]
    B --> C["Recover demonstrated contact locations"]
    C --> D["Linear blend retargeting + arm-hand IK"]
    D --> E["Residual PPO under randomized physics"]
    E --> F["Simulated point clouds + commanded joint targets"]
    F --> G["ManiFlow visuomotor student"]
    G --> H["Depth and proprioception on real Sharpa hardware"]

The input is already reconstructed 3D interaction data. In the experiments, it comes from ten GRAB motion-capture trajectories, covering reaching, grasping, lifting, and selected reorientation motions. Recovering this input from unconstrained monocular video remains future work.

2. MMO builds a hand model that can carry contact correspondences

Match morphology once per robot hand

The intermediate representation is Scaled MANO, with six scale parameters: one for the palm and one for each finger. Palm scaling is applied globally about the wrist; each finger receives an additional adjustment about its MCP joint. Skinning weights softly identify the vertices belonging to each finger, and the template and corrective blend shapes are scaled consistently.

URDF-derived length ratios initialize the scales. The optimization then adjusts scales, global rotation, translation, and finger pose to align corresponding robot and MANO joints and fingertips:

\[\mathcal L_{\mathrm{align}}=w_j\mathcal L_j+w_f\mathcal L_f.\]

MANO shape coefficients stay fixed at zero. A correspondence map handles different skeletons; missing joints can receive interpolated phantom positions, and multiple human fingers can map to one robot finger. Levenberg–Marquardt solves this alignment once per robot hand. Its output is a robot-shaped MANO mesh that retains the original vertex topology.

That topology is useful: a vertex associated with a demonstrated contact still has a corresponding vertex after the hand proportions change.

Recover contact after changing the hand’s proportions

For each reference frame, human mesh vertices within 5 mm of an object vertex define the contact set. Their original positions become targets for the morphology-aligned hand. The method optimizes global and local pose while keeping the learned scales fixed:

\[\mathcal L_{\mathrm{contact}} =w_c\mathcal L_c+w_d\mathcal L_d +w_{\mathrm{table}}\mathcal L_{\mathrm{table}}+w_p\mathcal L_p, \qquad \mathcal L_c=\frac{1}{|H_c|}\sum_{v\in H_c} \|\hat p_v-p_v\|_2^2.\]

The other terms preserve distances between coupled fingers, penalize table penetration, and regularize rotations toward the human pose. The coupled-finger term is omitted for five-fingered hands. Sequential initialization from the previous solution helps temporal consistency.

A useful detail concerns reaching and retreating, when the contact set is empty. MMO takes the vertex indices from the most-contacted frame, then uses those vertices’ positions in the current human frame as targets. This gives the future grasp region a continuous pre-grasp target without treating the hand as already touching the object.

Recover robot poses from a dense skeleton target

Linear blend retargeting transfers the aligned MANO motion to robot joints and fingertips. Blend weights are computed once by heat diffusion on a skeleton graph. Each robot point is transformed by a weighted combination of MANO joint transformations. Parent–child directions and orthogonalization then construct link-pose targets, including the wrist.

PyRoki jointly solves arm and hand IK with position, orientation, joint-limit, velocity, self-collision, and hand–table costs. Additional surface sampling strengthens table-collision checking. The reported 180 frames/s for single-hand and 150 frames/s for bimanual retargeting on an RTX 4090 exclude IK; these numbers are not end-to-end policy-training throughput.

3. Residual RL turns the reference into executable demonstrations

In ManiSkill, a PPO teacher predicts a joint-space correction:

\[q_t^{\mathrm{cmd}}=q_t^{\mathrm{ref}}+\Delta q_t.\]

Its privileged observations include robot state, observed and reference object poses, current and future reference joint configurations and contact states, physical properties, table clearance, and previous commands. The reference supplies a useful grasp and motion; the learned residual adapts them to contact dynamics.

The reward multiplies object tracking and contact-count agreement:

\[D_t=\frac{1}{N_p}\sum_{k=1}^{N_p} \|T_{\mathrm{ref},t}p_k-T_{\mathrm{obs},t}p_k\|_2, \qquad r_t=e^{-\alpha D_t} \left(1-\frac{|N_{\mathrm{goal},t}-N_{\mathrm{obs},t}|}{N_{\max}}\right),\] \[N_{\mathrm{goal},t}=\min(N_{H,t},N_{\max}).\]

Here ADD measures object-pose error through transformed mesh points; $N_{H,t}$ is the number of contacting human fingers and $N_{\max}$ is the robot’s finger count. The contact reward matches the number of contacting fingers. It does not directly optimize the dense contact-patch locations used in MMO. Fine contact geometry enters through the reference, while RL combines coarser contact supervision with object motion and physical constraints.

Early termination uses smoothed tracking and contact errors, plus immediate termination for table collision or excessive hand–object force. Randomization covers controller gains, friction, mass, viable center-of-mass locations, anisotropic object scale, initial pose, reference starting time, and external object wrenches. Initial positions span 10 × 10 cm, with yaw within ±15°. Each category receives its own teacher; reported training takes 60–90 minutes per policy on an RTX 4090.

4. Distillation replaces privileged state with depth and proprioception

Teacher rollouts provide commanded joint targets, proprioception, and point clouds rendered from one depth camera. RANSAC removes the table plane. Demonstration collection retains domain randomization but starts at the beginning of the reference trajectory instead of a random timestep.

The student uses ManiFlow’s point-cloud encoder, DiT-X action generator, and joint flow-matching and consistency training. Its inputs are the current and previous point clouds, current joint state, and previous command. It predicts joint-target chunks and executes three actions at 10 Hz before replanning.

Point-cloud randomization exposes the policy to incomplete geometry and synthetic table residue and cable returns. The hardware input contains 512 points. RGB images shown in the paper are illustrative; the policy receives depth-derived point clouds and proprioception.

This stage learns the teacher’s commanded joint targets, including the residual correction. It does not require the deployed student to estimate privileged physical parameters explicitly or execute the residual teacher online.

5. What the experiments establish

Better contact geometry helps, with different gains across hands

Location-aware F1 requires agreement in frame, object-surface location, and mapped hand part. Surface-area weights prevent densely tessellated regions from dominating. Failed retargeting attempts receive zero F1; continuous geometry comparisons use sequences successfully retargeted by every method. F1 and patch-distance means therefore require attention to their respective aggregation rules.

The following values summarize Tables II and III; the best baseline is selected separately for each metric:

HandBest baseline F1MMO F1Best baseline RL successMMO RL success
Dex3, 3 fingers10.2%38.0%53.2%82.5%
Allegro, 4 fingers22.0%30.3%68.2%69.9%
Sharpa, 5 fingers26.7%37.2%56.5%91.8%

All methods feed the same residual-RL formulation, making this a useful test of the reference’s contribution. The improvement is substantial on Dex3 and Sharpa and modest on Allegro. The authors attribute Allegro’s difficulties to its larger hand and thicker fingers, including reduced table clearance. Finger count alone does not predict success.

Contact fidelity also remains imperfect. At the 5 mm threshold, MMO F1 is only 30.3–38.0%. Appendix Table VII reports substantial kinematic penetration, confirming the need for physical refinement. Appendix Table VIII’s paired 95% interval for Sharpa’s F1 advantage over Contact PyRoki crosses zero; the positive mean is not a uniformly conclusive separation across the ten selected demonstrations. These intervals describe demonstration variation, not independent policy-training seeds.

Object pose and contact are complementary

Table IV removes each information source from observations, rewards, and termination conditions together. It is broader than a reward-only ablation:

Teacher informationDex3 successAllegro successSharpa success
Object pose only51.2%47.6%76.6%
Contact only42.5%41.6%54.3%
Both82.5%69.9%91.8%

Object motion specifies task progress; contact information helps retain a workable interaction. Their joint use produces the best success and contact-aware ADD on every hand. This does not mean every contact metric wins: on Allegro, Table III gives Position a lower final patch distance than MMO despite its slightly lower task success.

Hardware transfer has a clear scope

The real system is a KUKA iiwa14 with a Sharpa Wave hand and an Intel RealSense L515. Ten category-specific policies are each tested on three physical instances at ten initial poses. Success requires avoiding hard table collisions, lifting at least 5 cm, completing the expected motion, and maintaining a stable grasp for at least 5 seconds.

Sharpa evaluationSuccessTrials per category
Privileged teacher, simulation91.78%2,048
Visuomotor student, simulation93.75%64
Visuomotor student, hardware89.33%30

The student–teacher comparison uses different rollout counts and should not be read as proof that distillation improves control. Hardware success is 4.42 percentage points below the simulated student. Every category reaches at least 80%; lightbulbs achieve 30/30. No failures from hard table collisions were observed in these trials.

The transparent wineglasses require ping-pong balls placed in their bowls because the glass itself returns no usable depth. Their 28/30 successes demonstrate operation from partial observations under this setup. They do not establish unassisted transparent-object perception.

6. What I would carry into another manipulation system

My main takeaway is to treat the kinematic reference as part of the learning system. A morphology-aware contact reference can reduce the burden on RL before any policy architecture changes. The intermediate hand mesh is valuable because it preserves correspondences while accommodating different robot proportions.

Evaluation should retain separate measures for contact geometry, physical feasibility, task completion, and visual deployment. Improved F1 can coexist with penetration; a physically successful grasp can depart from the human patch. Measuring all stages makes those tradeoffs visible.

The remaining generalization gap is concrete. Thin tape rolls differ from the GRAB torus in ways anisotropic scaling cannot reproduce, and an unusually thick alarm clock can trigger an insufficiently open pre-grasp. Appendix D also places 22 of 29 on-grid failures in the near row of the tested workspace. That suggests targeted expansion of geometry and pose coverage as a practical next experiment, although the paper does not test that remedy.

The demonstrated result is a complete simulation-to-hardware pipeline from ten reconstructed interactions. Unified multi-category learning, broader spatial coverage, monocular-video reconstruction, and real deployment on other hands remain open extensions.