[Paper Notes] SUGAR: A Scalable Human-Video-Driven Generalizable Humanoid Loco-Manipulation Learning Framework

20 minute read

Published:

This post supports English / 中文 switching via the site language toggle in the top navigation.

TL;DR

Human videos contain task logic: where a person walks, when contact begins, how the object moves, and what a successful interaction looks like. They also contain bad robot supervision. Occlusion corrupts pose estimates, human-to-robot retargeting creates penetrations, and reconstructed contact can violate physics. Training a humanoid directly on those trajectories gives it detailed targets that cannot actually be executed.

SUGAR treats the extracted motion as a coarse prior and repairs it in simulation. A privileged reinforcement-learning Refiner converts each human-object trajectory into a physically valid robot-object execution. A Command Tracker absorbs the motor skill, while a diffusion-based Command Generator learns to produce short command chunks from the current object state and an optional goal. The reference video disappears at deployment.

The strongest evidence comes from tasks where reference replay breaks down. On the held-out simulation set, SUGAR reaches 69.6% on Carry Box, 99.2% on Pick Bottle, and 86.3% on Stand Bottle; both reference-tracking baselines score zero on all three. The real Unitree G1 completes 46 of 60 trials across six tasks using MoCap state observations. More human videos help: averaging the six test-task success rates in Table 2 gives 58.6% with 20 videos per task, 74.3% with 50, and 83.5% with 100.

I read SUGAR as a strong recipe for turning imperfect demonstrations into task-specific closed-loop skills. The paper does not yet establish a vision-language generalist. Hardware deployment uses state input from motion capture, training appears to be separate for each task, and the evidence for novel objects, disturbance recovery, and long-horizon behavior is mostly qualitative. Those limits define the next useful experiment.

Paper and source version

SUGAR: A Scalable Human-Video-Driven Generalizable Humanoid Loco-Manipulation Learning Framework is by Tianshu Wu, Xiangqi Kong, Yue Chen, Qize Yu, Hang Ye, Jia Li, Yizhou Wang, and Hao Dong, from Peking University and Beihang University.

These notes follow the 18-page arXiv:2605.20373v1, submitted May 19, 2026. The paper PDF, project page, and official repository provide the primary materials. Numerical results below come from the paper; I have not reproduced the Isaac Sim training or hardware trials.

1. The usable part of a human video is the task structure

Task-specific reinforcement learning can produce impressive humanoid behavior, but every new task brings reward design and environment work. Teleoperation supplies embodiment-consistent demonstrations at the cost of operators and specialized hardware. Reference tracking offers another shortcut: reconstruct a human motion, retarget it, and ask the robot to replay it. That shortcut inherits the reconstruction errors and binds inference to one recorded trajectory.

SUGAR keeps the information that survives noisy reconstruction. A carrying video still reveals the rough body path, the object’s motion, and the interval during which the hands should support the box. Those signals define the interaction even when individual poses are inaccurate. Simulation then supplies the missing physical test: can a robot execute a nearby motion while balancing, respecting contact, and moving the object?

The paper evaluates six coarse whole-body tasks: Carry Box, Push Box, Kick Box, Pick Bottle, Stand Bottle, and Sit Chair. For each task, the authors collect 100 human videos for training and 30 for testing. The resulting scale is 600 training videos and 180 held-out videos across the study. This is large enough to test data scaling from 20 to 100 videos per task, though still far from internet-scale learning.

flowchart TD
    A["Raw human videos"] --> B["Kinematic priors P: human motion, object pose, contact"]
    B --> C["Privileged RL Refiner"]
    C --> D["Refined skills R: feasible robot-object executions"]
    D --> E["Train 50 Hz Command Tracker"]
    E --> F["Closed-loop rollout dataset D"]
    F --> G["Train 10 Hz diffusion Command Generator"]
    G --> H["Generator + Tracker + PD control on Unitree G1"]

2. Stage one: build a kinematic interaction prior

The extraction pipeline estimates both sides of the interaction. SAMBody recovers the human motion, which is aligned to depth and refined with ICP. SAMObj generates an object mesh; the mesh scale is fitted to the captured point cloud, and FoundationPose estimates its 6D trajectory.

Contact is the extra signal that makes object interaction different from ordinary motion imitation. The pipeline asks a vision-language model whether a task-specific body part is in direct physical contact with the named object. The prompt explicitly rejects intention and near-contact. Severe occlusion makes that judgment unreliable for kicking, so the authors infer contact when object velocity crosses a threshold. Temporal filtering smooths the reconstructed trajectories.

The output is

\[\mathcal P=\{\hat\tau^i\}_{i=1}^{N}, \qquad \hat\tau=\{(\hat p_t^R,\hat p_t^O,\hat l_t)\}_{t=1}^{T},\]

where the hat marks quantities recovered from video. Each prior contains human motion, object motion, and a contact label. It is structured enough to express the task and too noisy to serve as a robot demonstration.

“Fully automated” should be read at the clip-processing level. The VLM prompt still receives a task definition through [BODY_PART] and [OBJECT], and the kicking case uses a task-dependent velocity heuristic. SUGAR removes frame-by-frame manual labels; it has not removed all task specification.

3. Stage two: let physics edit the demonstration

The Refiner is a privileged reference-tracking policy,

\[\pi_r\!\left(a_t^r\mid o_t^R,o_t^O,o_t^{\mathrm{priv}},\hat\tau^i\right),\]

trained with PPO in Isaac Sim. It sees simulator state and future reference information unavailable to the deployable actor. Its job is to stay close to the recovered task while producing a trajectory that obeys dynamics. Successful rollouts form

\[\mathcal R=\{\tau^i\}_{i=1}^{N}, \qquad \tau=\{(p_t^R,p_t^O,l_t,c_t)\}_{t=1}^{T}.\]

The recorded command is

\[c_t=[q_t^{\mathrm{cmd}},v_t^{\mathrm{cmd}},\omega_t^{\mathrm{cmd}},l_t],\]

combining joint positions, root linear and angular velocities, and contact state. This command becomes the interface between high-level intent and low-level execution in stage three.

The reward has three groups:

\[r=r_{\mathrm{track}}+r_{\mathrm{int}}+r_{\mathrm{reg}}.\]

The tracking terms cover robot pose and velocity plus object pose and velocity. Interaction terms preserve object-to-body geometry and reward agreement between measured contact force and the video-derived contact label. Regularizers penalize foot slip, unwanted contacts, joint acceleration, torque, action changes, and limit violations. The design is shared across the six tasks; task success criteria and the extraction prompt still vary.

The contact reward matters because pose similarity can hide a failed manipulation. In the paper’s Carry Box sequence, the policy without interaction reward bends like the demonstrator but never lifts the box. It has matched the visible body motion while missing the event that gives the motion its purpose.

Progressive State Pool Initialization

Ordinary Reference State Initialization samples a point on the recovered trajectory. A bad reconstructed frame may place a hand inside the box or start the robot from a dynamically impossible configuration. Starting every episode at the first frame avoids that problem and makes late stages hard to reach.

The Progressive State Pool stores intermediate states that the Refiner has already visited successfully. Training can restart from these physically checked milestones. The pool moves the curriculum forward without trusting arbitrary states from the raw video prior.

The authors also randomize friction, restitution, joint offsets, base center of mass, and object mass. Object mass ranges from 0.5 to 2 times nominal. Random pushes perturb the robot and, during active contact, the object. This part of training teaches compensation that the clean reference trajectory never demonstrates.

4. Stage three: distill reference tracking into autonomous control

The Refiner can repair a clip, but it still consumes a reference trajectory and privileged state. SUGAR removes both dependencies through two policies with different jobs.

The Command Tracker maps deployable robot history, object pose, and command $c_t$ to joint targets. Its actor receives five-step histories of root angular velocity, joint positions and velocities, actions, and projected gravity, plus the current object pose and command. A PD controller converts targets to torque. Training begins with behavior cloning from the Refiner, warms up actor and critic, then switches to PPO. Initialization gradually moves from refined reference states to the progressive pool.

Both Refiner and Tracker use three-layer MLPs with widths [512, 256, 128]. They train with 4,096 environments for 30,000 PPO iterations. The asymmetric actor-critic setup gives privileged state to the Refiner and to the Tracker’s critic; the deployed Tracker actor keeps the smaller observation set.

The Command Generator solves the task-level problem:

\[\pi_g(c_{t:t+7}\mid o_t^O,c_{t-1},g).\]

It is a state-based Diffusion Policy with a 12-block Diffusion Transformer. Given the current object state, previous command, and optional target object state $g$, it predicts eight future commands. The system executes four, replans at 10 Hz, and linearly interpolates the commands for the 50 Hz Tracker. Chunking smooths motion; frequent replanning keeps the policy responsive.

This hierarchy is practical. The generator chooses how the interaction should progress. The tracker handles balance, contacts, and motor execution. A single monolithic policy would have to learn both time scales from the same data and observation interface.

5. Train the generator on states the tracker actually reaches

There is still a distribution gap inside the hierarchy. If the generator trains on ideal refined states, the deployed Tracker may land a few centimeters away. The next command then comes from a state absent from the generator’s training set, and the error can accumulate.

SUGAR rolls out the frozen Tracker using commands from each refined skill and records the object state it actually reaches:

\[\mathcal D=\{\tau_i^*\}_{i=1}^{M}, \qquad \tau_i^*=\{(\tilde o_t^O,c_t,g)\}_{t=1}^{T}.\]

The generator learns from $\tilde o_t^O$, not the ideal reference object state. This is a compact form of execution-aware imitation. It gives the generator examples of the drift produced by its own downstream controller and makes replanning useful instead of cosmetic.

This is the part of SUGAR I would reuse first. It applies beyond humanoids: whenever a planner emits commands to an imperfect learned tracker, train the planner on tracker rollouts. The condition changes if rollout errors enter unsafe regions or cover only a narrow part of deployment; then additional intervention data or online aggregation is needed.

6. Read the simulation results task by task

Table 1 compares SUGAR with ResMimic and HDMI, two reference-trajectory tracking baselines. The baselines receive a demonstration trajectory at inference. SUGAR receives only an optional goal object state after training. The interfaces differ, so the result supports the complete autonomous pipeline more directly than a controlled comparison of identical policy classes.

Held-out simulation taskBetter reference baseline SRSUGAR SRSUGAR final position errorReal G1
Kick Box18.5%76.0%0.265 m7/10
Push Box54.6%70.0%0.325 m6/10
Carry Box0.0%69.6%0.326 m7/10
Sit Chair20.6%99.6%-9/10
Pick Bottle0.0%99.2%-9/10
Stand Bottle0.0%86.3%-8/10

“Better reference baseline” selects the larger test success rate between ResMimic and HDMI for each task. Final position error is reported only for the three target-placement tasks. All simulation methods use the same train/test video split, according to the paper.

Carry Box is the clearest stress test. Both baselines obtain 0% on training and held-out trajectories, while SUGAR reaches 84.5% and 69.6%. Reference tracking cannot compensate for a coarse hand-object reconstruction. The physics refiner and interaction reward can discover an executable nearby behavior.

Sit Chair and Pick Bottle are already relatively easy for some ablations, with many values above 94%. Kick, Push, and Carry expose larger differences in object displacement and contact stability. A single average would hide that structure.

7. More videos improve coverage, with one irregular point

The scaling experiment trains on 20, 50, and 100 videos per task. The mean below is my calculation from the six held-out success rates in Table 2.

Training videos per taskMean held-out SR across six tasksKickPushCarrySitPickStand
2058.6%32.735.033.590.094.266.4
5074.3%63.152.361.098.995.375.3
10083.5%76.070.069.699.699.286.3

Every held-out task improves from 20 to 50 to 100 videos. Training-set Push Box briefly falls from 37.1% at 20 videos to 30.1% at 50 before reaching 83.6% at 100, so the experiment is not perfectly monotonic at every individual entry. The held-out trend is clean.

The experiment demonstrates useful scaling over a fivefold data range. It does not yet identify a scaling law, and the six task families are fixed. The next question is whether a shared model can absorb additional tasks and objects without training a new three-policy stack for each one.

8. The ablations reveal four different failure modes

Removing the Refiner and learning directly from kinematic priors lowers the mean held-out success from 83.5% to 70.8%. Direct learning still works on Sit Chair and Pick Bottle, while Kick and Push fall to 46.3% and 41.7%. The value of physics repair grows with precise object displacement and unstable contact.

Removing the interaction reward is more specific. Pick Bottle collapses from 99.2% to 0%, Carry drops from 69.6% to 60.3%, and the mean across tasks becomes 61.8%. Motion tracking alone does not tell the policy that the object must remain supported.

Removing interaction robustness enhancement leaves Pick Bottle near 98% but hurts Push Box and Stand Bottle by roughly 23 percentage points each. Domain variation and perturbations contribute most where changes in friction, mass, or impact alter the outcome.

Replacing the Progressive State Pool with start-only or raw-reference initialization gives mean held-out rates of 78.9% and 79.8%. The gap is smaller in the aggregate, though task-level effects vary. Progressive initialization is a training stabilizer; it is not the sole source of capability.

9. What the real-robot result establishes

The policies are trained entirely in Isaac Sim and transferred to a Unitree G1. Across ten trials per task, the robot records 7/10 Kick, 6/10 Push, 7/10 Carry, 9/10 Sit, 9/10 Pick, and 8/10 Stand. Summed across tasks, that is 46/60, or 76.7%. Sixty trials give a useful hardware check, while per-task estimates remain coarse.

The paper shows sequences in which the robot resumes after a failed bottle pickup, continues under human disturbance, and handles boxes, chairs, and bottles with changed appearance and geometry. These examples support the claimed behaviors. They are not accompanied by a separate recovery rate, perturbation benchmark, or quantified object-generalization table.

Real-world perception comes from MoCap. The deployed generator observes object pose relative to the robot root, while the tracker also consumes proprioceptive history. Appearance is absent from the policy input, and object geometry is not represented explicitly. Zero-shot transfer to a visually different object is therefore expected; transfer to substantially different geometry depends on the physical tolerance learned through randomization and feedback.

10. Generalizable does not yet mean one generalist policy

Appendix D reports the cost for each individual task: about 20 GPU hours for the Refiner, 20 for the Tracker, and 5 for the Command Generator on a single RTX 5090. Combined, the reported recipe is roughly 45 GPU hours per task. The dataset and tables also organize training by task.

I therefore interpret generalization as transfer across initial states, goals, disturbances, and some object changes within a learned task. The paper does not show a single policy switching among all six skills from a task instruction, or executing a seventh unseen interaction. The framework scales data collection better than teleoperation; model and compute scaling across task count remain open.

The authors state three limitations directly. Extracted priors support coarse interactions and do not yet capture fine manipulation. Data efficiency is low. A state-based policy makes deployment less convenient than a policy that consumes vision and language.

There is another systems issue. Stage one already uses strong perception models and depth to produce object trajectories offline, yet stage-three deployment bypasses that uncertainty with MoCap. Replacing MoCap with onboard perception will introduce delay, occlusion, object identity errors, and pose jumps precisely during contact. Those errors should enter both the Tracker’s state distribution and the Generator’s rollout dataset.

What I would test next

The cleanest next experiment would keep the trained controller fixed and replace MoCap gradually: first inject measured pose noise and delay, then run an external RGB-D tracker, and finally move perception onboard. Report task success, recovery time, object-pose error during contact, and the fraction of failures caused by perception versus control.

For learning, I would train one shared Tracker across the six refined-skill datasets and condition a shared Generator on a task or language embedding. If the shared model retains the reported per-task success, SUGAR starts to look like a route from scalable human video to a reusable loco-manipulation policy. If it suffers interference, the refined dataset still gives a valuable place to study routing, adapters, or continual learning.

Fine-grained contact should be the third test. Binary contact is enough to say “hold the box” or “kick now.” Assembly, tool use, and dexterous handover need contact location, force direction, and phase. Adding those signals would show whether physics refinement can move from repairing coarse trajectories to producing supervision for precision interaction.