Embodied Model Training: Interactive Map
Select a model to explore shared training components and data sources. Use the language switch in the navigation for English / 中文.
Four stages of embodied-model training
A shared map of capabilities and data—not four mandatory training runs.
Select a model to highlight shared components. Gray means not mapped here, not proven absent.
STAGE 1 · PRE-TRAINING
Semantic foundations
Language, objects, scenes, common sense, and task intent.
LLM / VLM Web text and image priors
Visual / video representations Semantic encoders and tokenizers
STAGE 2 · PRE-TRAINING
Physical-world priors
Interaction, temporal change, affordances, and transferable motion priors.
Action-free world learning Predict observations or latent transitions without robot action labels
Ego / human video Unlabeled video or video with human motion labels
High-fidelity physical interaction Wearable human sensorimotor streams
Simulation / synthetic Synthetic trajectories and diverse environments
STAGE 3 · PRE-TRAINING / MID-TRAINING
Action foundations / grounding
Learn a transferable action representation; align it with executable robot controls.
Broad robot trajectories Multi-task, multi-scene, or multi-robot co-training
Aligned human–robot data Retargeted action spaces or paired human–robot play
High-fidelity human interface Tracked motion, gloves, handheld grippers, or exoskeletons
STAGE 4 · FINE-TUNING / POST-TRAINING
Deployment adaptation
Adapt to the target robot, task, and environment; improve reliability and recovery.
Target robot demonstrations Task demonstrations; supervised adaptation / behavior cloning
Deployment / controller adaptation Latency, control frequency, and hardware integration
On-robot RL + corrections Reward feedback, advantage estimation, and expert intervention
World / value refinement Rollout-based prediction and value updates for planning
The data continuum
Toward more explicit action alignment—not a quality ranking Web text / image Semantic context; little direct action supervision
Action-free video Observed interactions and world changes
Human motion signals Ego + reconstructed pose or sensorimotor capture
Robot-aligned human Retargeting or co-designed action interfaces
Robot-native Robot commands and states; contact when recorded
π0.6 / π*0.6: the RL highlight refers to π*0.6 / RECAP, not supervised π0.6. Training a controller, running feedback control, and adapting model weights are distinct operations. See the article for version scope and evidence.
Read the bilingual article: model recipes, evidence, and the limits of this map →