SysID: Solving the Hidden Dynamics Between a Policy and a Dexterous Hand
Published:
This post supports English / 中文 switching via the site language toggle in the top navigation.
TL;DR
I used to think that once a reinforcement-learning policy produced an action, the hard part was over. For a dexterous hand, that action is only a request sent into another system: a position controller, a motor driver, a transmission, a tendon or linkage, a contact surface, and a collection of delays and saturations. The hardware decides what the request actually becomes.
That hidden system is why identification matters so much for high-dynamic sim-to-real. The policy can be excellent in simulation and still fail on hardware because the same command produces a different motion, contact force, or recovery response. For a computer-science student, system identification is the missing bridge between “the model outputs an action” and “the robot moves as expected.”
This post starts with two optimizers that are often used for black-box identification—CEM and CMA-ES—and then follows the PACE workflow: excite the real robot, replay the same trajectory in simulation, fit a compact physical parameter set, and use the aligned simulator for policy learning. PACE was developed for sim-to-real across robotic systems, and its core ideas transfer naturally to dexterous hands once the hand-specific dynamics are made explicit.
1. The Action Is Not the Motion
In a software environment, an action often looks complete. The policy emits a vector, the environment consumes it, and the next state is returned. That interface makes the action feel like a cause with a predictable effect.
A real hand breaks that illusion. A command such as “move the fingertip three millimeters” may pass through a chain like this:
policy output
-> action scaling and safety limits
-> joint or tendon controller
-> motor current and torque loop
-> transmission, elasticity, backlash, friction
-> finger links and contact geometry
-> object and environment
Every arrow can change the result. The same target may be reached with a different delay, a different overshoot, a different contact force, or no motion at all when friction and saturation dominate. A dexterous hand adds coupling: one tendon can influence several joints, one contact can redistribute load across the fingers, and a small calibration error can change the entire grasp.
This is why sim-to-real is not just a matter of making the policy more intelligent. The simulator must expose a response that is close enough to the hardware for the policy’s learned assumptions to remain useful.
For a CS student, this is an important shift in mental model. The policy is a function that proposes commands. The robot is a dynamical system that interprets those commands. Identification estimates the hidden part between the two.
2. CEM and CMA-ES: Two Ways to Search Without Gradients
CEM and CMA-ES solve the same outer problem:
\[a^* = \arg\min_a J(a),\]where \(a\) is a vector of simulator parameters and \(J\) is a trajectory-level error measured after running a rollout. The simulator may contain friction, clipping, contacts, delay, and discontinuities, so differentiating through the entire hardware-matched pipeline is often inconvenient or unreliable.
Both methods keep a probability distribution over candidate parameters:
sample parameters -> run simulations -> measure error -> update distribution
The important difference is how the distribution learns from a generation of samples.
CEM: fit the distribution to the best samples
The Cross-Entropy Method samples \(N\) candidates from a current distribution \(q_t(\theta)\). After evaluating them, it keeps an elite fraction—say the best 5–20 percent for a minimization problem—and fits the next distribution to those elite samples.
For a diagonal Gaussian, the update is conceptually simple:
\[\mu_{t+1}=\operatorname{mean}(\theta_{elite}), \qquad \sigma^2_{t+1}=\operatorname{var}(\theta_{elite}).\]Smoothing is usually added so that the variance does not collapse after one lucky generation. The key hyperparameters are the population size, the elite ratio, the smoothing coefficient, the minimum variance, and the boundary rule.
CEM is easy to explain and easy to adapt. A discrete parameter can use a categorical distribution, and a mixed parameter vector can use different distributions for different blocks. This is useful for a hand when some quantities are naturally integer-valued, such as a delay in control cycles, a mode switch, or a discrete tendon routing choice.
The basic diagonal version has a limitation: it treats parameters as independent. If a good simulator requires a larger damping together with a larger effective inertia, the distribution may need many samples to discover that combination. A full-covariance CEM can learn the relationship, but it pays a higher computational and numerical cost.
CEM also has a characteristic failure mode. Elite selection can make the distribution narrow too quickly around a locally good region. Smoothing, variance floors, multiple restarts, or a small amount of injected exploration are practical ways to keep the search alive.
CMA-ES: learn the shape of useful moves
CMA-ES also samples a population, usually from a multivariate Gaussian:
\[\theta_i = m + \sigma \mathcal{N}(0,C),\]where \(m\) is the current mean, \(\sigma\) is a global step size, and \(C\) is a covariance matrix. It ranks the candidates, moves the mean toward the better ones, and adapts both \(C\) and \(\sigma\) over time. Evolution paths remember whether successful steps keep pointing in a consistent direction.
The covariance matrix is the practical distinction. It can rotate and stretch the search distribution, so the optimizer can follow a narrow, slanted valley in parameter space. This matters for identification because parameters are rarely independent. In a dexterous hand, motor inertia, damping, tendon compliance, friction, controller gains, and delay can compensate for one another over a particular excitation range.
CMA-ES therefore tends to be a strong default for continuous, moderately sized, non-separable black-box problems. Its price is memory and computation: a full covariance matrix grows quadratically with the number of parameters. It also needs deliberate handling for bounds and integer parameters.
The two methods can be summarized like this:
| Question | CEM | CMA-ES |
|---|---|---|
| What drives the update? | Elite samples and a fitted distribution | Rank-weighted steps, covariance adaptation, and step-size control |
| Default parameter dependence | Often independent in diagonal CEM | Explicitly modeled through \(C\) |
| Mixed or discrete variables | Natural with categorical or hybrid distributions | Requires rounding, enumeration, or a hybrid design |
| Main risk | Early distribution collapse | Expensive covariance adaptation and poor treatment of discrete variables |
| Best starting point | Low-dimensional or highly batched search | Continuous parameters with strong coupling |
Neither optimizer identifies physics by itself. It only chooses which parameter vector to try next. The data collection, simulator model, loss, and parameter boundaries determine what can actually be identified.
3. Why a Dexterous Hand Is a Difficult Identification Problem
A legged robot and a dexterous hand share the same identification logic, but the hand makes several hidden effects more visible.
First, a hand has many degrees of freedom packed into a small mechanical structure. The parameters are coupled through tendons, gears, linkages, and shared contact forces. Second, the operating regime changes quickly. A finger can move freely for one moment and become constrained by an object in the next. Third, contact is part of the task itself. Removing all contact makes the actuator dynamics easier to observe, but it does not tell us how the hand behaves while grasping, sliding, or making an insertion.
This suggests a staged identification strategy:
- Free-space identification: estimate actuator, joint, tendon, and delay parameters while the hand is away from objects.
- Contact identification: add controlled contact experiments to estimate compliance, contact friction, force offsets, and object-dependent effects.
- Task validation: replay grasping and manipulation trajectories and check the errors that matter to the policy.
The order matters. If we fit everything at once from a single grasp, the optimizer can use actuator friction to explain an unmodeled contact event, or use a joint bias to compensate for a wrong object pose. The resulting parameters may reduce one trajectory’s error while losing physical meaning.
For a fully actuated hand, a compact parameter vector might begin with
\[\theta = [I_a, d, \tau_f, q_b, T_d],\]where \(I_a\) is effective inertia, \(d\) is viscous damping, \(\tau_f\) is Coulomb friction, \(q_b\) is encoder or assembly bias, and \(T_d\) is delay. A tendon-driven or underactuated hand may need additional terms for tendon stiffness, backlash, coupling, transmission efficiency, or joint-limit behavior. The goal is still to keep the parameter set small enough that the experiment can distinguish its effects.
This is where identification becomes a modeling decision. More parameters can make the simulator look more expressive, yet they can also make the problem unidentifiable. If two parameters produce the same change in joint trajectories under the available excitation, no optimizer can recover both reliably.
4. How PACE Turns This into a Repeatable Loop
PACE gives a useful template for the full process. Its central idea is to fit a compact, physically meaningful simulator so that a replayed trajectory looks like the real trajectory. The published method uses fixed-base, contact-free joint excitation and CMA-ES for parameter fitting; the framework itself is presented as agnostic to the particular optimizer.
The loop looks like this:
flowchart TD
A[Real hand: fixed-base excitation] --> B[Record targets and measured joint states]
B --> C[Sample simulator parameters with CEM or CMA-ES]
C --> D[Replay the same targets in parallel simulations]
D --> E[Compute trajectory and contact-aware losses]
E --> F[Update the search distribution]
F --> C
F --> G[Use the fitted simulator for RL]
G --> H[Validate the policy on hardware]
H --> I[Collect informative failure data]
I --> C
PACE’s original excitation stage fixes the base, keeps the limbs away from contacts, and drives the joints with synchronized chirp position targets. This design removes long-horizon locomotion drift and unmeasured external forces. The simulation receives the same target sequence, so a direct time-domain comparison is meaningful.
For each candidate parameter vector \(p_e\), the simulator produces a trajectory \(q^{sim}_{k,e}\). The basic loss is the time-averaged joint-position error:
\[\ell_e = \frac{1}{K}\sum_{k=1}^{K} \left\|q^{real}_k-q^{sim}_{k,e}\right\|^2.\]PACE maps normalized optimizer variables from \([-1,1]\) into physical bounds. In the public ANYmal example, the candidate population is evaluated in thousands of parallel simulation environments. The current implementation initializes CMA-ES at the center of the normalized range, uses a broad initial step size, accumulates the trajectory error, and then calls tell once a full excitation sequence has finished.
A PACE-style hand pipeline would preserve this structure while changing the experiment and the model:
- excite individual fingers first, then coupled finger groups;
- include position, velocity, current or torque when the hardware exposes them;
- use several amplitudes and frequency bands so friction, compliance, and delay become observable;
- add controlled contact trajectories after free-space fitting;
- evaluate both tracking error and task-relevant force or grasp stability error;
- treat delay as a discrete or hybrid parameter instead of silently rounding a continuous sample;
- validate on held-out objects and motion patterns.
The optimizer still sees a black-box loss. The intelligence of the identification process comes from choosing experiments that separate the parameters.
5. What This Changes for a CS Student
A CS background makes it natural to focus on the policy, the neural network, and the training objective. Those are visible in code. The hardware response is harder to see because it is distributed across firmware, motor drives, mechanics, sensors, and contact.
System identification adds a missing layer of literacy. It asks questions that look physical but directly affect learning:
- Does the action represent a position target, a torque target, a tendon displacement, or a command to another policy?
- How much delay exists between issuing a command and observing its result?
- Which parameters are actually observable from the available sensors?
- Is a tracking error caused by the policy, the actuator model, the contact model, or the experiment?
- Does the simulator reproduce the distribution of failure, not only the average trajectory?
These questions also change how I think about “better models.” A larger policy cannot reliably compensate for a simulator that teaches the wrong response distribution. It may memorize a correction for one hand, one object, or one controller setting, then fail when the hardware changes. A smaller policy trained in a better aligned simulator can be more transferable because its action assumptions are closer to reality.
The most useful interface is therefore a loop:
policy -> controller -> hardware -> observation
^ |
| v
simulator <- identified dynamics <- data
The policy sits inside a physical loop. Identification makes the loop visible enough to model, measure, and improve.
6. What I Want to Remember
CEM and CMA-ES are easy to mistake for the main idea because their names appear beside the optimization results. They are only the search machinery. The harder work is deciding what the hand should be excited with, which parameters are meaningful, what the sensors can reveal, and which loss corresponds to the behavior we care about.
PACE makes that lesson concrete. Its CMA-ES implementation is useful because it fits a non-smooth trajectory objective in a massively parallel simulator. Its deeper contribution is the identification loop: controlled excitation, a compact end-to-end model, direct replay, and a simulator that is aligned before reinforcement learning begins.
For dexterous hands, I would treat PACE as a starting architecture rather than a finished recipe. Free-space excitation can identify the hidden actuator response. Contact experiments must then teach the simulator how compliance, friction, and force redistribution shape manipulation. CEM may be easier when the hand has discrete modes or a small parameter set; CMA-ES is attractive when the continuous parameters are coupled. In both cases, the experiment design determines the quality of the answer.
A policy does not send motion into an empty mathematical space. It sends a request into a physical machine with memory, delay, friction, and contact. System identification is the process of learning that machine well enough that simulation and hardware begin to share the same meaning of an action.
References: PACE repository · PACE paper
