[Paper Notes] DexMani: Human-Derived Manipulability Guidance for Dexterous Rotation
Published:
This post supports English / 中文 switching via the site language toggle in the top navigation.
TL;DR
Sustained rotation is a contact-transition problem. A finger must release, move, and make contact again while the remaining contacts support the object; the new hand configuration must also leave useful rotational directions available for the next step. DexMani represents this future-facing property as the short-horizon evolution of contact-conditioned rotational manipulability.
Human demonstrations supervise an energy model over desirable changes in a six-dimensional manipulability descriptor. During robot reinforcement learning, a base PPO policy proposes an action, the target hand uses its own kinematics and active contacts to evaluate eight nearby alternatives, and the lowest-energy alternative becomes a clipped hint for a learned residual policy. The human joint trajectory is never used as a robot action target.
This shared prior improves simulated performance across three rotation tasks and four robot-hand designs. On LEAP Hand, DexMani averages 57.5% success, 5.6 percentage points above the strongest baseline; on cap unscrewing across Shadow, Allegro, and XHand, it averages 43.4%. The real system closes the loop at 20 Hz, though its success rates—6/10 for cap unscrewing, 3/10 for free-object rotation, and 1/10 for faucet turning—show that the sim-to-real gap remains substantial.
Paper and source version
Xiaoyang Chen, Shengcheng Luo, Haoran Guo, Jiaming Jiang, Wanlin Li, Ziyuan Jiao, and Chenxi Xiao wrote DexMani: Human-Derived Manipulability Guidance for Dexterous Rotation. The authors are affiliated with Shanghai Jiao Tong University, ShanghaiTech University, the Beijing Institute for General Artificial Intelligence, and Beihang University.
These notes follow the 16-page arXiv:2608.00554v1 PDF, submitted August 1, 2026. The official project page provides method diagrams and videos for the human demonstrations, simulated tasks, cross-hand experiments, and real deployment. No conference acceptance is stated in this version.
I have read the paper and supplementary material and checked the numerical results against the project page. I have not rerun the training or hardware experiments.
1. Rotation quality depends on the next contact state
Cap unscrewing, object spinning, and faucet turning all require repeated finger gaiting. Immediate angular progress is only half of each decision. The fingers must end in a configuration that can continue generating rotation about the task axis. DexMani calls this configuration- and contact-dependent capability contact-conditioned rotational manipulability.
This framing changes what transfers from a person to a robot. Pose retargeting asks a robot to reproduce a human configuration or fingertip trajectory, so differences in joint layout, range of motion, and hand proportions must be resolved explicitly. DexMani transfers a preference in a shared task-space coordinate system: given the current visual–tactile context, which direction should rotational capability evolve? Each hand can realize that direction through its own joints and contacts.
The proposal has three layers:
- compute an analytic rotational-manipulability label from human hand kinematics and active fingertip contacts;
- learn an energy landscape over short-horizon changes in that label; and
- reuse the frozen landscape as local action guidance during robot RL.
The distinction between state and evolution matters. Maximizing the current manipulability score can create a locally broad rotational workspace and still lead into a poor sequence of contact transitions. DexMani learns how successful human trajectories reshape that workspace over time.
2. A contact-gated object-rotation descriptor
At time $t$, let $J^{\mathrm{hum}}_{c,t}$ stack the positional Jacobians of the active human fingertip contacts. The paper constructs a contact-space capability matrix and maps it through the grasp matrix into object-rotation space:
\[C^{\mathrm{hum}}_{c,t} =W_t^{1/2}J^{\mathrm{hum}}_{c,t}H_h (J^{\mathrm{hum}}_{c,t})^\top W_t^{1/2},\] \[M^{\mathrm{hum}}_{\omega,t} =P_\omega(G_t^+)^\top C^{\mathrm{hum}}_{c,t}G_t^+P_\omega^\top +\epsilon_m I_3.\]$W_t$ activates contacts from the tactile signal, $H_h$ scales joint directions, $G_t^+$ is a damped pseudoinverse of the grasp matrix, and $P_\omega=[0_{3\times3}\ I_3]$ selects rotation from the six-dimensional object twist. In the released experimental formulation, inactive fingertips are removed before the matrices are built, so $W_t=I$ over the remaining contacts; all human joint directions receive equal weight, $H_h=I$.
The grasp matrix uses fingertip positions relative to the active-contact centroid:
\[G_t= \begin{bmatrix} I_3 & \cdots & I_3\\ [r_{1,t}]_\times & \cdots & [r_{N_t,t}]_\times \end{bmatrix}, \qquad G_t^+=G_t^\top(G_tG_t^\top+\lambda_G I_6)^{-1}.\]Centering at the detected contacts removes the need for an object-center estimate. The resulting symmetric positive-definite $3\times3$ matrix describes available object rotation about the wrist-frame axes. DexMani converts it to a six-vector using log-Euclidean coordinates:
\[m_t=\operatorname{vech}(\log M_{\omega,t})\in\mathbb{R}^6.\]Human and robot descriptors use the same wrist-frame convention. Every evaluated target direction is the hand-frame $z$-axis, $d=[0,0,1]$. This alignment makes the six coordinates comparable across embodiments, but it also narrows the evidence: the experiments do not establish one prior that handles arbitrary axes or coordinate conventions.
The contact abstraction is deliberately compact. A 256-taxel human glove is reduced to five binary fingertip states; robot sensors are also grouped by finger and thresholded. Four-finger hands fill the little-finger entry with zero. The transfer therefore needs a hand kinematic model and finger-level contact detection, without requiring taxel correspondence, contact normals, or matching joint spaces.
3. Learn a direction field from human rotation
The collection system records data at 30 Hz using two 640 × 480 RGB cameras, a Manus Quantum MetaGlove for 21 hand-joint positions, a Meta Quest 3 controller for the global wrist pose, and a 256-channel piezoresistive tactile glove. The dataset contains more than 100,000 frames over 27 objects. Recorded motions also drive a MANO hand in simulation to create additional rendered RGB observations.
An eight-frame visual–tactile history enters a transformer encoder. RGB frames become image-patch tokens; the five binarized fingertip contacts become tactile tokens. The fused feature $z_t^{VT}$ represents the visible interaction and the active-contact pattern. Reconstruction and future-manipulability prediction serve as auxiliary pretraining objectives.
For the energy objective, the positive example is the normalized short-horizon change in log-manipulability:
\[\hat v_t^+ =\frac{m_{t+\Delta}-m_t} {\|m_{t+\Delta}-m_t\|_2+\epsilon_{\mathrm{num}}}.\]The implementation uses $\Delta=4$ raw frames. Near-stationary samples are excluded because normalization would amplify measurement noise. The conditioning context is
\[c_t=[z_t^{VT},m_t,d_t],\]and $E_\theta(\hat v\mid c_t)$ assigns low energy to compatible evolution directions. Each positive is contrasted with eight structured negatives, including random and reversed directions and, in the main formulation, directions associated with mismatched coordinate contexts. The contrastive loss is
\[\mathcal L_E =-\log \frac{\exp[-E_\theta(\hat v_t^+\mid c_t)/\tau]} {\sum_{\hat v\in\mathcal V_t} \exp[-E_\theta(\hat v\mid c_t)/\tau]}.\]Direction normalization discards step size, so another head predicts
\[\alpha_t=\log\!\left(1+\|m_{t+\Delta}-m_t\|_2\right).\]The full objective combines visual–tactile reconstruction, future-state prediction, contrastive energy, and magnitude prediction. After 400 pretraining epochs, the visual–tactile encoder and energy model are frozen. Robot learning can query the human-derived field, while downstream rewards cannot rewrite it.
4. Turn the energy prior into a residual action hint
At a robot control step, the base policy proposes $a_t^0$. DexMani adds eight Gaussian perturbations, clips them to the valid action space, and includes the nominal action to form nine candidates:
\[\mathcal A_t=\{a_t^0\}\cup \left\{\operatorname{clip}_{\mathcal A}(a_t^0+\xi_{t,k})\right\}_{k=1}^{8}.\]Each candidate maps to a target joint configuration. The hand’s fingertip Jacobians estimate the induced contact-point motion, after which DexMani recomputes the robot’s rotational manipulability and evaluates
\[\hat v_t^r(a)= \frac{m^r(q_t+u_t(a,q_t))-m_t^r} {\|m^r(q_t+u_t(a,q_t))-m_t^r\|_2+\epsilon}.\]This estimate uses the current active-contact geometry and discards candidates with negligible predicted change. It avoids a dynamics rollout, making it cheap enough to sit inside RL. Release and re-contact enter at the next control step through updated tactile observations, so the prior evaluates a sequence of local approximations instead of directly predicting contact switches.
The lowest-energy action produces a clipped, stop-gradient bias:
\[a_t^\star=\arg\min_{a\in\mathcal A_t} E_\theta(\hat v_t^r(a)\mid c_t^r), \qquad b_t^E=\operatorname{sg} \left[\operatorname{clip}_{b_{\max}}(a_t^\star-a_t^0)\right].\]A residual policy receives the observation, robot state, nominal action, current manipulability, and $b_t^E$. The executed command is
\[a_t=\operatorname{clip}_{\mathcal A} \left(a_t^0+\lambda_R\Delta a_t\right).\]The energy-selected candidate is not executed directly. The residual policy can interpret or ignore the local hint according to long-horizon reward. Both base and residual policies are trained with the original task reward using PPO; the method adds no manipulability reward term.
flowchart TD
A["Human RGB, wrist/finger pose, and tactile history"] --> B["Active-contact rotational manipulability m"]
A --> C["Visual–tactile context z"]
B --> D["Short-horizon evolution direction"]
C --> E["Contrastive energy prior"]
D --> E
E --> F["Freeze encoder and energy model"]
G["Robot base PPO action"] --> H["Nominal + eight nearby candidates"]
I["Robot kinematics and current contacts"] --> H
H --> J["Candidate manipulability changes"]
F --> K["Low-energy local action bias"]
J --> K
K --> L["Learned residual policy"]
G --> L
L --> M["Native joint-position target"]
M --> I
The extra computation is meaningful. On an RTX 4090, one reported training iteration takes 11.497 s for DexMani and 5.773 s for plain PPO. Local manipulability computation and energy scoring account for 0.434 s. The paper does not further decompose the remaining gap.
5. What transfers across tasks, objects, and hands
Success requires at least a $2\pi$ rotation. Policies are evaluated over three independent seeds, with 1,000 episodes per object and randomized initial hand poses for each seed. Objects in the human demonstrations are excluded from downstream robot training and evaluation; every robot task is further divided into policy-training Seen objects and held-out Unseen objects.
The shared prior is pretrained on human cap-unscrewing and free-object rotation. Human faucet demonstrations are excluded, making Turn Faucet the cross-task test.
| Method | Cap seen | Cap unseen | Object seen | Object unseen | Faucet seen | Faucet unseen | Avg. SR |
|---|---|---|---|---|---|---|---|
| PPO | 11.8 | 5.0 | 26.5 | 21.5 | 33.2 | 20.5 | 19.8 |
| VT Pretraining | 29.0 | 12.7 | 42.6 | 34.0 | 62.4 | 53.5 | 39.0 |
| VTM | 47.0 | 27.8 | 49.5 | 37.9 | 79.0 | 70.1 | 51.9 |
| VTA | 40.1 | 25.1 | 37.4 | 20.3 | 77.2 | 63.5 | 43.9 |
| VTA-E | 24.2 | 11.6 | 29.0 | 19.1 | 67.6 | 42.0 | 32.3 |
| DexMani | 65.1 | 39.3 | 50.7 | 39.2 | 80.2 | 70.3 | 57.5 |
The table reports mean success rates; the paper also provides standard deviations. DexMani leads all six cells, though the margin ranges from large on cap unscrewing to only 0.2 points over VTM on unseen faucets. This pattern supports the value of online guidance most strongly on some contact regimes and gives weaker evidence on others.
Cross-hand evaluation reuses the same prior and cap-unscrewing definition, then trains one policy in each native action space:
| Method | Shadow seen | Shadow unseen | Allegro seen | Allegro unseen | XHand seen | XHand unseen | Avg. SR |
|---|---|---|---|---|---|---|---|
| PPO | 20.8 | 6.8 | 11.8 | 5.2 | 32.5 | 17.2 | 15.7 |
| VTM | 66.4 | 40.9 | 24.3 | 12.5 | 58.0 | 22.7 | 37.5 |
| VTA-E | 52.7 | 24.5 | 17.6 | 9.8 | 45.2 | 18.2 | 28.0 |
| DexMani | 70.0 | 48.2 | 34.6 | 14.8 | 63.4 | 29.5 | 43.4 |
The transfer object is the frozen prior, not the control policy. Four different hands still require embodiment-specific policies, which is less ambitious than zero-shot policy transfer but more reusable than a human-action target tied to one joint layout.
6. The ablations isolate evolution, context, and long horizon
The mechanism study holds the residual-policy interface fixed and swaps the guidance signal:
| Guidance | Cap seen | Cap unseen | Faucet seen | Faucet unseen | Avg. SR |
|---|---|---|---|---|---|
| Zero Guidance | 18.6 | 9.4 | 48.6 | 20.2 | 24.2 |
| Context-Shuffled | 14.0 | 5.3 | 22.5 | 7.3 | 12.3 |
| Greedy-M | 54.5 | 36.3 | 67.8 | 44.7 | 50.8 |
| DexMani | 65.1 | 39.3 | 80.2 | 70.3 | 63.7 |
Zero Guidance shows that a residual network alone cannot explain the gain. Context-Shuffled performs even worse, indicating that a plausible direction applied to the wrong interaction state can be actively harmful. Greedy-M selects the candidate with the largest instantaneous rotational manipulability and is already strong. DexMani’s further gain supports the paper’s main claim: the temporal pattern learned from successful contact transitions carries information beyond the size of the current rotational capability.
Motion metrics add a second view. DexMani obtains the best log dimensionless jerk (LDLJ) and spectral arc length (SPARC) on all three LEAP tasks, with higher values indicating smoother motion. It also leads the task compatibility index (TCI) on cap unscrewing and faucet turning. On free-object rotation, its TCI is 0.36 versus PPO’s 0.43, while its motion is smoother and its success rate is much higher. The prior therefore does not uniformly maximize instantaneous axis-aligned capability; it can trade a local measure for a more successful trajectory.
The physical setup combines a 16-DoF LEAP Hand, a 6-DoF xArm, TwinTac fingertip sensors, RGB observations, and proprioception. Domain randomization covers joint observations, appearance, camera pose, image noise, and tactile force before binarization. The policy receives the same input format as in simulation and receives no real-world fine-tuning.
| Real task | Successes |
|---|---|
| Unscrew Cap | 6 / 10 |
| Rotate Object | 3 / 10 |
| Turn Faucet | 1 / 10 |
These trials establish closed-loop feasibility, not a mature hardware benchmark. Faucet turning combines the largest shifts: the human prior has no faucet demonstrations, the physical faucet is absent from robot-policy training, and sensing and dynamics also change. The resulting 1/10 rate makes the limitation visible instead of hiding it behind selected successful videos.
7. Limits and takeaways for research
The prior is shared; the policies are not. Every task–hand pair needs a fresh PPO training run. A unified controller that consumes embodiment information and transfers directly remains future work.
Candidate scoring is local. The robot predicts one-step manipulability changes under the current active contacts. It does not simulate the release/re-contact event caused by each candidate. Updated tactile state repairs that approximation one control step later, but contact events with delayed benefit may still be difficult to evaluate.
The descriptor sees binary fingertip contact. This improves portability across sensor layouts and discards pressure distribution, contact normals, palm contact, and compliance. Those signals may distinguish stable and unstable transitions that share the same active-finger pattern.
Axis generalization is untested. All experiments encode the desired direction as the wrist-frame $z$-axis. Cross-task and cross-hand results are meaningful within that convention; arbitrary 3D rotation directions need a separate study.
The results are simulation-heavy. The paper evaluates thousands of simulated trials per object and only ten real trials per task. Real success drops sharply, especially for free-object and faucet rotation.
My main takeaway is a useful design pattern for cross-embodiment learning: transfer a task-space derivative that each body can realize locally. DexMani does this with $\Delta m$—the change in contact-conditioned rotational capability—while the robot retains control over its native joints. The energy model expresses a preference, candidate search grounds that preference in the current hand, and the residual policy decides how much to trust it over a longer horizon.
For future work, I would test the same idea with continuous tactile features and a contact-transition predictor, condition a single policy on hand morphology, and randomize the rotation axis during both human-prior and robot-policy training. Those extensions would reveal whether manipulability evolution can become a genuinely general interface between human experience and heterogeneous dexterous hands.
