[Paper Notes] WEAVE: Learning Whole-Body Dexterous Loco-Manipulation from Human–Object Interactions

19 minute read

Published:

This post supports English / 中文 switching via the site language toggle in the top navigation.

TL;DR

Whole-body loco-manipulation is unforgiving. A humanoid must walk to an object, place several underactuated fingers well enough to hold it, move the object, and keep balancing while the contact forces change. A human demonstration shows the desired interaction, but its hand pose is not automatically a stable robot grasp. The missing step is to repair the demonstration at the level of contact.

WEAVE builds that repair into the data pipeline. It first retargets captured SMPL-X human–object motion to a Unitree G1, refines the arm and finger configuration with a differentiable force-closure objective, and generates a locomotion prefix that approaches the interaction from different directions. A single reference-conditioned policy then learns all nine object families with explicit object geometry and contact observations. The policy commands 29 body DoFs and 12 finger DoFs at 50 Hz.

The headline results are strong inside the simulator. WEAVE reaches 92.5% success and 96.3% progress on its training interactions. On held-out interaction sequences involving the same nine objects, it reaches 65.0% success and 84.5% progress without additional training. When trained directly on the evaluation collection, one shared nine-object policy scores 95.3%, versus 91.5% for nine object-specific specialists given the same aggregate number of PPO iterations.

The most revealing result is not the four-point success gap. The specialists usually track reference poses more accurately, yet finish fewer interactions. WEAVE suggests that a larger shared policy can trade a little imitation fidelity for a better recovery margin. That distinction matters for contact-rich robotics: a chair carried successfully along a nearby motion is better than a perfectly imitated pose that loses the chair.

The scope is also precise. All evaluation is in simulation. The actor receives ground-truth object pose, contact state, and a geometry descriptor; the held-out set contains new sequences and approach variations for known objects, not unseen object categories. The released 9,474 simulator rollouts (23.23 hours) are valuable physical interaction data, but they are not real-robot trials. WEAVE establishes a promising contact-aware skill-acquisition layer, not yet an autonomous vision-language humanoid system.

Paper and source version

WEAVE: Learning Whole-Body Dexterous Loco-Manipulation from Human–Object Interactions is by Liu Cao, Xingze Wu, Jingzhi Cui, Botian Xu, Mingzhi Pei, Ruoqu Chen, and Mengdi Xu, with affiliations at Tsinghua University, Dalian University of Technology, and The Chinese University of Hong Kong.

These notes follow arXiv:2609.16683v1, submitted September 15, 2026. The paper PDF, project page, official code, and released dataset are the primary sources. Numerical results below come from the paper; I have not reproduced its Isaac Sim training runs.

1. The hard part is preserving the interaction across embodiments

A whole-body human motion can be retargeted link by link and still fail as manipulation data. Human and robot limbs have different lengths, joint limits, palm shapes, finger coupling, and reachable contact patterns. Small errors at the torso may look harmless. A centimeter of fingertip error can turn a force-bearing grasp into surface contact that immediately slips.

WEAVE formulates the task as a closed-loop reference-tracking problem. Each captured sequence provides paired SMPL-X body motion and an object trajectory. After retargeting, the policy must approach the object, establish a multi-finger grasp, transport it toward the reference configuration, and maintain balance throughout the interaction.

The embodiment is a Unitree G1 with 29 actuated body joints and two Inspire dexterous hands with 12 actuated finger DoFs. The hands remain underactuated: the controller commands proximal finger joints, while intermediate and distal joints follow fixed mimic couplings. That detail limits how literally human finger articulation can be transferred.

The system has two main stages:

flowchart TD
    A["Captured SMPL-X motion + object trajectory"] --> B["Whole-body inverse kinematics"]
    B --> C["Contact-aware arm and hand refinement"]
    C --> D["Kimodo locomotion-prefix generation"]
    D --> E["Robot–object references + contact labels"]
    E --> F["Unified PPO reference tracker"]
    G["Robot proprioception"] --> F
    H["Object pose + fingertip geometry + BPS-SDF"] --> F
    F --> I["29 body + 12 finger position targets at 50 Hz"]
    I --> J["Physics rollouts across nine objects"]

This separation is sensible. Reference construction solves a geometric question—what robot configuration could preserve the observed interaction? Reinforcement learning solves a dynamical one—how can the robot execute that reference under balance and contact constraints?

2. Retarget the body, then repair the grasp

The first pass uses GMR-style whole-body inverse kinematics to align selected robot links with their SMPL-X counterparts. The pelvis is constrained only in the horizontal plane; its height is optimized from ground contact. Joint-limit, velocity, and acceleration penalties keep the motion feasible and temporally smooth.

Ordinary whole-body IK preserves the arm and wrist trajectory but cannot guarantee a stable fingertip arrangement. WEAVE therefore freezes the retargeted pelvis and legs and jointly refines both arms and fingers. Its hand objective can be summarized as

\[\mathcal L_{\text{hand}} =\lambda_a\sum_{t,k}c_{t,k}\lVert p_{t,k}-\bar p_{t,k}\rVert -\lambda_q\sum_t\widehat Q_{\mathrm{FC}}(q_t^h) +\lambda_p\sum_{t,k}\left[-(p_{t,k}-\bar p_{t,k})^\top n_{t,k}\right]_+ +\mathcal L_{\mathrm{reg}}.\]

The attraction term moves robot fingertips toward corresponding human contact points. The penetration term keeps them from solving that objective by entering the mesh. Regularization preserves smooth motion and a reasonable distance from the initial retargeting.

The force-closure term is the crucial addition. In simplified form,

\[Q_{\mathrm{FC}}(\xi_t) =\min_{\lVert w\rVert_2=1} \max_{f\in\mathcal F(\xi_t),\,\lVert f\rVert_1\le 1} w^\top G(\xi_t)f.\]

It asks how well the available contact forces can oppose the weakest disturbance wrench. Positive force closure means the grasp can resist arbitrary wrench directions under the model. WEAVE differentiates an approximation of this score through the hand configuration, so refinement rewards contact arrangements that can hold an object, not merely touch its surface.

This objective encodes an important hierarchy. Pose correspondence is useful for initializing a human-like interaction. Contact attraction recovers the intended fingertip region. Force closure decides whether that arrangement has a chance of working under load.

3. A locomotion prefix turns a local interaction into loco-manipulation

Captured human–object clips often begin close to the object. Training only on the interaction segment would teach grasping and transport while skipping the approach. WEAVE uses Kimodo to synthesize a locomotion prefix from a sampled approach direction to the first interaction pose:

\[\tau_{\mathrm{ref}}=\tau_{\mathrm{pre}}\oplus\tau_{\mathrm{int}}.\]

The resulting reference spans walking, reaching, grasping, and transporting. Each robot link also receives a ternary contact label—separated, neutral, or in contact—derived from link-to-object distance thresholds. Neutral labels are ignored by the contact-matching reward, which avoids supervising ambiguous near-contact frames.

The authors generate three approach variations per captured training interaction and five per evaluation interaction. This makes the evaluation collection broader in both direction and generated prefix. It is useful stress testing, but it also introduces an intentional distribution mismatch. Part of the 92.5% to 65.0% success drop reflects wider approach coverage, not just failure to generalize the original human interaction.

4. The policy observes contact and geometry directly

The tracker is a contact- and geometry-aware asymmetric actor–critic trained with PPO. At control step $t$,

\[a_t\sim\pi_{\mathrm{track}}\!\left( \cdot\mid o_t^{\mathrm{prop}},o_t^{\mathrm{obj}}, \hat x_{t:t+H}^{\mathrm{robot}}, \hat x_{t:t+H}^{\mathrm{obj}} \right).\]

Robot proprioception includes base angular velocity, projected gravity, joint positions and velocities, and the previous action. The reference command provides a short horizon of robot joint configuration, pelvis pose, object pose, and contact labels.

Object observation is richer than a pose vector. It contains ground-truth relative object pose, fingertip-to-surface vectors, binary contact flags, and a Basis Point Set signed-distance-field descriptor (BPS-SDF) of object geometry. This interface lets one policy distinguish a thin lamp stand from a broad table and select contact accordingly. It also places object perception outside the problem: pose, geometry, and contact are already available to the actor in simulation.

The reward combines three groups:

  • robot and object tracking for pelvis, body links, joints, and object pose;
  • hand opposition and contact-label matching for the grasp;
  • penalties for foot sliding, abrupt actions, and non-finger joint-limit violations.

Hand opposition rewards the thumb and opposing fingers for approaching different sides of the object when a grasp is expected. Contact matching compares measured hand contact with the non-neutral reference labels:

\[r_{\mathrm{contact}} =\frac{\sum_h m_{t,h}\left(1-\lvert y_{t,h}-\tilde c_{t,h}\rvert\right)} {\sum_h m_{t,h}+\epsilon}.\]

The critic additionally receives base linear velocity and current tracked-body poses. Domain randomization covers robot and object friction and restitution, torso center of mass, and finger actuator properties. Simulation runs at 200 Hz and the policy emits position targets at 50 Hz for low-level PD control.

Episodes terminate when pelvis or object position error grows too large, projected gravity violates balance thresholds, ankle or wrist vertical error exceeds 0.25 m, or an expected hand contact is absent for ten consecutive control steps. These guards keep on-policy samples close to the reference. They also mean the benchmark says little about recovery after a large contact loss: such states are deliberately cut off.

The network uses a SimBaV2 backbone. Two-dimensional weight matrices are optimized with Muon, while biases and other parameters use AdamW. This optimizer choice returns in the paper’s final ablation.

5. The release is large, but it is simulator data

The reference collection contains nine everyday object families and 9,474 trajectories in total. The project reports 7,869 training references covering 19.56 hours and 1,605 evaluation references covering 3.67 hours.

ObjectTraining trajectoriesEvaluation trajectories
Tripod1,131150
White chair1,066175
Wood chair987230
Clothes stand861145
Small table824210
Floor lamp805165
Large box791230
Large table767180
Small box637120
Total7,8691,605

The paper describes the release as physically executed rollouts because these are policy executions inside a physics simulator, with robot–object trajectories and contact annotations. The conclusion states that evaluation is carried out entirely in simulation. I would therefore call this a physics-grounded simulator dataset, not hardware interaction data.

That distinction does not make the release unimportant. Contact-rich humanoid trajectories are expensive to collect, and the dataset can support policy learning or physically consistent human–object motion generation. Its clean state, object geometry, and contact labels are precisely the supervision that real video usually lacks.

6. Read the 65% result with the split definition attached

WEAVE first trains one multi-object policy on the training references and evaluates it on both splits. A second policy is trained directly on the evaluation collection as an oracle for how executable those references are.

Policy and evaluationSuccess rateProgress rateInterpretation
Train policy → training interactions92.45%96.30%Fit to the collected training references
Train policy → held-out interactions64.98%84.52%Transfer to new sequences and wider approach variants
Evaluation oracle → evaluation interactions95.26%98.23%Direct fit to the evaluation references

The held-out result is meaningful: a single policy can execute many human–object sequences it did not train on. It is narrower than object-category generalization. Both splits use the same nine named objects, and the actor receives their geometry descriptor and ground-truth state. The experiment tests unseen interaction sequences for those same objects, plus a broader distribution of approach prefixes.

The oracle is especially informative. Its 95.26% success shows that most evaluation references can be learned when included in training. The large train-policy gap therefore points to coverage and distribution shift more strongly than to impossible retargeting. A natural next experiment is to vary approach diversity while holding the underlying human interactions fixed, then vary the interaction split while holding approach sampling fixed.

Tracking error alone understates this gap. On held-out sequences, body and object tracking errors remain fairly close to training values while completion falls by more than 27 percentage points. Contact-rich execution has thresholds: a modest local error can cross from stable support to a dropped object.

7. One shared policy beats nine specialists—and imitates less exactly

The paper’s cleanest comparison trains on the evaluation collection itself. Nine object-specific specialists each receive 3,000 PPO iterations. The unified policy receives 27,000 iterations, matching the specialists’ aggregate budget. Observations, actions, reward, network, and PPO configuration are otherwise held constant; neither side receives dedicated hyperparameter tuning.

ObjectUnified policySingle-object specialist
Floor lamp100.0%100.0%
Small table98.6%96.2%
Large table98.3%96.1%
Clothes stand95.9%95.9%
Small box95.0%90.0%
Large box92.2%89.1%
White chair96.6%90.9%
Wood chair91.7%79.2%
Tripod90.0%90.0%
Pooled95.3%91.5%

The pooled rate is weighted by the number of trajectories, so objects with more evaluation clips contribute more. The unified policy ties on three objects, improves on six, and regresses on none. Shared walking, balance, transport, and contact structure apparently outweigh interference at this scale.

Yet specialists produce lower tracking error on seven of the eight reported tracking metrics. They fit each object’s reference more closely while succeeding less often. The authors argue that the unified policy has learned broader recovery margins: related shapes and approach phases act as mutual augmentation, discouraging brittle memorization of one exact grasp.

I find this result more important than the pooled score by itself. It says the objective should be judged by interaction completion, not pose imitation alone. A useful follow-up would perturb the robot or object during transport and measure recovery probability directly. That would test the recovery-margin explanation instead of inferring it from success and tracking-error disagreement.

8. Muon helps early learning; the evidence is deliberately narrow

The final experiment crosses two backbones, MLP and SimBaV2, with two optimizer choices, AdamW and Muon. All four configurations train for 3,000 iterations on the small-table task.

Both Muon configurations improve episode length and reward earlier than their AdamW counterparts. SimBaV2 adds a smaller gain. The paper’s explanation is plausible: Muon’s orthogonalized momentum may prevent one reward direction from dominating gradients when body tracking, object tracking, and finger contact differ in scale. SimBaV2’s normalization may help when the state distribution changes abruptly at contact onset.

This is a useful engineering clue, not a broad optimizer verdict. The ablation covers one object task, reports training curves, and does not present repeated-seed uncertainty. Success-rate comparisons across all nine objects would be needed to separate faster reward optimization from more reliable skill completion.

9. What WEAVE establishes—and what remains open

WEAVE makes three convincing contributions.

First, contact-aware retargeting repairs the information that ordinary pose correspondence loses. The force-closure objective gives hand refinement a physical target before reinforcement learning begins.

Second, a single geometry-conditioned tracker can absorb a diverse reference collection without decomposing locomotion, balance, and dexterous grasping into object-specific controllers. The unified-versus-specialist result is evidence for positive transfer across contact-rich skills.

Third, the project releases code and a substantial simulator rollout dataset with contact annotations. That makes the work more useful than a benchmark result alone.

The current boundary is equally clear:

  • No real-robot evaluation. Contact dynamics, calibration error, actuator delay, and hand wear remain untested.
  • Privileged object input. The actor consumes ground-truth object pose, contact flags, and geometry; onboard perception and state estimation are absent.
  • Known objects at test time. The 65% result concerns unseen sequences of the same nine objects, not new categories.
  • Reference-conditioned behavior. A higher-level system must still select, generate, or compose an interaction reference.
  • Limited hand embodiment. Fixed mimic coupling bounds finger-level fidelity, and heavy or highly articulated objects lie outside the reference and randomization range.
  • Restricted recovery evidence. Early termination keeps training stable but removes severe failure states from the learned distribution.

This is why I see WEAVE as a motor-skill substrate for a future vision-language-action system. A planner could choose an interaction and a perception module could estimate the object state; WEAVE addresses the difficult middle layer that converts a contact-rich reference into whole-body motor execution.

What I would test next

The first test should preserve the policy and replace privileged object input in stages. Add measured pose noise and latency, then train a perception student from RGB-D or proprioceptive contact history, and finally run on hardware. Report failure causes separately for perception, grasp establishment, transport, and balance.

The second should isolate generalization. Hold the approach sampler fixed while testing new human interaction clips; then hold the interaction fixed while testing new approaches; finally introduce unseen object geometry. The current aggregate test score mixes the first two effects and never reaches the third.

The third should test the paper’s most interesting hypothesis: shared training creates recovery margin. Apply controlled pushes, pose offsets, friction changes, and brief contact loss to unified and specialist policies. Plot success as a function of perturbation magnitude, alongside tracking error. If the unified policy keeps its lead as references become less reachable, the fidelity-versus-completion story becomes a measured mechanism.