[Paper Notes] One Policy to Run Them All: URMA for Multi-Embodiment Locomotion
Published:
This post supports English / 中文 switching via the site language toggle in the top navigation.
TL;DR
One Policy to Run Them All introduces URMA (Unified Robot Morphology Architecture), an end-to-end multi-task reinforcement-learning architecture that uses one policy for legged robots with different numbers of joints, feet, and morphology types. Instead of padding every robot into one maximum-size vector or maintaining a separate network head for each platform, URMA describes each joint and foot with morphology metadata, aggregates variable-length observations with attention, and decodes one action for every joint.
The paper trains a single PPO policy on 16 simulated robots: nine quadrupeds, five humanoids, one biped, and one hexapod. The policy transfers zero-shot to held-out simulated robots and to three real quadrupeds. Its strongest practical result is that the same policy can control the Unitree A1, MAB Honey Badger, and MAB Silver Badger after simulation training with domain randomization, with no real-world fine-tuning. The main limitation is coverage: transfer works best near the training distribution, and the experiments do not yet include exteroceptive sensing or humanoid hardware deployment.
Paper Info
The paper is “One Policy to Run Them All: an End-to-end Learning Approach to Multi-Embodiment Locomotion” by Nico Bohlinger, Grzegorz Czechmanowski, Maciej Piotr Krupka, Piotr Kicki, Krzysztof Walas, Jan Peters, and Davide Tateo. It appeared at CoRL 2024 and was published in PMLR 270.
1. Why Multi-Embodiment Locomotion Is Difficult
Deep reinforcement learning has produced strong locomotion controllers for individual quadrupeds, bipeds, and humanoids. Reusing those policies across robots is difficult because the embodiment changes the dimensionality and meaning of both observations and actions. A quadruped may have a different number of joints from another quadruped; a humanoid adds a different topology and foot set; a hexapod changes the correspondence problem again.
Two common workarounds have clear costs. A padding-based policy puts every robot into a maximum-length vector, but one coordinate can represent different joints on different robots. A multi-head policy gives each platform its own encoder and decoder, but a new morphology requires a new head and a new mapping into the shared representation. Both approaches make zero-shot transfer to an unseen morphology awkward.
URMA treats embodiment as part of the input structure. It learns a shared locomotion representation while retaining the per-joint information needed to produce the correct low-level action.
2. The URMA Architecture
URMA splits a robot observation into fixed-size general observations and variable-length robot-specific observations. For locomotion, the latter are divided into joint observations and foot observations.
Each joint has an observation vector (o_j) and a description vector (d_j). The description can include the joint’s rotation axis, relative position, torque and velocity limits, control range, and other characteristic properties. The description tells the network what a joint is; the observation tells it what that joint is currently doing.
The joint encoder maps descriptions into attention keys and joint observations into values. A learnable-temperature attention operation produces a fixed-size latent representation:
[ \bar z_{\text{joints}} = \sum_{j\in J} \frac{\exp(f_\phi(d_j)/(\tau+\epsilon))} {\sum_{k\in J}\exp(f_\phi(d_k)/(\tau+\epsilon))} f_\psi(o_j). ]
The same mechanism encodes feet. The two variable-length latents are concatenated with the general observations and sent through a shared core network:
[ \bar z_{\text{action}} =h_\theta(o_g,\bar z_{\text{joints}},\bar z_{\text{feet}}). ]
The universal morphology decoder then combines this action latent with each joint’s description and its individual latent. It outputs a mean and standard deviation for every joint, so the policy can produce an action vector whose size follows the current robot:
[ a_j\sim \mathcal N\bigl(\mu_\nu(d^a_j,\bar z_{\text{action}},z_j),\; \sigma_\nu(d^a_j)\bigr). ]
The architecture therefore has one shared encoder, one shared core, and one shared decoder. The variable-length part is handled through attention and per-joint decoding instead of padding or platform-specific heads.
3. Training Objective and Practical Recipe
The paper formulates each robot embodiment as a task in multi-task reinforcement learning. If there are (M) embodiments, the objective is the average expected discounted return:
[ J(\theta)=\frac{1}{M}\sum_{m=1}^{M}J_m(\theta), \qquad J_m(\theta)=\mathbb E_{\tau\sim\pi_\theta} \left[\sum_{t=0}^{T}\gamma^t r_m(s_t,a_t)\right]. ]
The policy is trained with PPO in MuJoCo. The authors use 48 parallel environments, three for each of 16 robots, and train for 100 million simulation steps per robot. Adding a robot requires adjusting reward coefficients, controller gains, and domain-randomization ranges; the neural architecture itself remains unchanged. A time-dependent curriculum gradually increases penalty terms, allowing the shared policy to learn basic locomotion before handling stricter gait shaping.
Domain randomization covers embodiment and environment dynamics. This is essential for sim-to-real transfer: the policy should learn motion patterns that survive changes in friction, mass, actuator behavior, and other physical parameters.
4. Experiments
The training set contains nine quadrupeds with three joint configurations, five humanoids with five configurations, one biped, and one hexapod. The baselines are a multi-head architecture and a padding architecture with a one-hot task ID.
URMA learns faster and reaches a higher final return than training separate policies for each robot. It eventually outperforms the multi-head baseline, although the attention encoder learns more slowly at the beginning because it must discover how to route joint information. The padding baseline performs substantially worse because the task ID alone does not resolve the structural mismatch between different observation and action spaces.
For zero-shot evaluation, a policy trained without Unitree A1 transfers well to A1. It also transfers to MAB Silver Badger, whose embodiment includes an additional spine joint and lacks the foot observations used by the training robots. After zero-shot evaluation, URMA remains ahead during fine-tuning on Silver Badger because it starts from a better representation.
The authors also remove all foot observations at test time. URMA retains stronger performance than the baselines, suggesting that the morphology-aware representation can degrade gracefully when part of the sensor input disappears.
In the real world, the same simulation-trained policy controls Unitree A1, MAB Honey Badger, and MAB Silver Badger on pavement, grass, and plastic turf with small inclines. Honey Badger is unseen during training. Its gait is weaker than the gaits of the two training-set robots, but it still locomotes robustly without further fine-tuning.
5. What the Paper Establishes
The paper’s central result is architectural: variable-size low-level control spaces can be handled by a single end-to-end policy when each joint is represented together with a meaningful description. This creates a reusable correspondence between joints across embodiments. A knee on one robot and a differently placed knee on another do not need to occupy the same fixed input coordinate; their descriptions let the network route their information into a shared latent space.
The experiments also show a useful training-efficiency effect. Sharing representations across embodiments lets the policy benefit from data collected for other robots. A single multi-embodiment run can reach useful performance faster than training every robot from scratch, while the same representation provides a starting point for zero-shot and few-shot transfer.
6. Limitations
Generalization still depends on training coverage. A robot that is far outside the training distribution can remain difficult, even though its joint descriptions fit the architecture. The method also relies on reward and controller settings being adapted when a new robot is added; the network is shared, but the training configuration is not completely automatic.
The experiments omit exteroceptive sensors such as cameras and depth inputs, so the demonstrated policy focuses on proprioceptive locomotion rather than navigation over complex terrain. Real-world deployment is limited to quadrupeds, and the paper does not yet establish transfer to a physical humanoid. Finally, the theoretical analysis supports the sample-efficiency intuition for shared encoders and decoders, but it does not remove the empirical need for a sufficiently broad embodiment distribution.
Takeaway
URMA is a morphology-aware interface between a robot’s variable joint structure and a fixed neural policy. Its contribution is a practical recipe: describe each joint, aggregate variable-length observations with attention, decode actions per joint, and train across many embodiments with PPO and domain randomization.
This makes URMA an important predecessor to later multi-embodiment locomotion systems. Compared with approaches centered on long-context online adaptation, URMA’s main focus is representing and transferring across robot structures. Its longer-term value is the idea that a low-level locomotion policy can be shared at the level of joint descriptions, providing a foundation on which larger embodiment-scaling and adaptation systems can build.
