Embodied Model Training: Interactive Map

← All apps

Select a model to explore shared training components and data sources. Use the language switch in the navigation for English / 中文.

Four stages of embodied-model training

A shared map of capabilities and data—not four mandatory training runs.

Select a model to highlight shared components. Gray means not mapped here, not proven absent.

STAGE 1 · PRE-TRAINING

Semantic foundations

Language, objects, scenes, common sense, and task intent.
LLM / VLM Web text and image priors
Visual / video representations Semantic encoders and tokenizers
STAGE 2 · PRE-TRAINING

Physical-world priors

Interaction, temporal change, affordances, and transferable motion priors.
Action-free world learning Predict observations or latent transitions without robot action labels
Ego / human video Unlabeled video or video with human motion labels
High-fidelity physical interaction Wearable human sensorimotor streams
Simulation / synthetic Synthetic trajectories and diverse environments
STAGE 3 · PRE-TRAINING / MID-TRAINING

Action foundations / grounding

Learn a transferable action representation; align it with executable robot controls.
Broad robot trajectories Multi-task, multi-scene, or multi-robot co-training
Aligned human–robot data Retargeted action spaces or paired human–robot play
High-fidelity human interface Tracked motion, gloves, handheld grippers, or exoskeletons
STAGE 4 · FINE-TUNING / POST-TRAINING

Deployment adaptation

Adapt to the target robot, task, and environment; improve reliability and recovery.
Target robot demonstrations Task demonstrations; supervised adaptation / behavior cloning
Deployment / controller adaptation Latency, control frequency, and hardware integration
On-robot RL + corrections Reward feedback, advantage estimation, and expert intervention
World / value refinement Rollout-based prediction and value updates for planning

The data continuum

Toward more explicit action alignment—not a quality ranking
Web text / image Semantic context; little direct action supervision
Action-free video Observed interactions and world changes
Human motion signals Ego + reconstructed pose or sensorimotor capture
Robot-aligned human Retargeting or co-designed action interfaces
Robot-native Robot commands and states; contact when recorded

π0.6 / π*0.6: the RL highlight refers to π*0.6 / RECAP, not supervised π0.6. Training a controller, running feedback control, and adapting model weights are distinct operations. See the article for version scope and evidence.

Read the bilingual article: model recipes, evidence, and the limits of this map →