[Paper Notes] DexJoCo-X: Benchmarking Action Representations for Multi-Hand Dexterous Manipulation
Published:
This post supports English / 中文 switching via the site language toggle in the top navigation.
TL;DR
DexJoCo-X asks how to represent actions when one manipulation policy controls several dexterous hands. It builds a matched simulation benchmark with seven hands, six tasks, and 2,100 accepted demonstrations, then compares native coordinates, function-aligned slots (FAAS), and learned cross-hand latents (DexLatent) within Being-H0.5.
The central result is a task-dependent tradeoff: Native reaches 47.0% overall success, FAAS 47.7%, and DexLatent 33.1%. Native leads on single-arm tasks; FAAS leads on bimanual tasks. This makes a useful case for evaluating an action representation together with its pretrained policy and execution decoder. It provides no evidence yet for transfer to an unseen hand.
Paper and source version
DexJoCo-X: Benchmarking Action Representations for Multi-Hand Dexterous Manipulation is by Xiangwei Jiang, Yao Mu, Lixin Duan, and Wen Li, affiliated with the University of Electronic Science and Technology of China and Shanghai Jiao Tong University. These notes follow the October 2, 2026 preprint, arXiv:2610.03278v1, and its eight-page PDF. The authors list a project website. All experimental numbers below are author-reported.
The benchmark builds on DexJoCo, covered in an earlier post. Its focus shifts to comparing representations across heterogeneous hands under shared tasks, demonstrations, and execution conditions. Being-H0.5, Ego-Pi, and the representation principles are adopted from prior work; the contribution here is the benchmark, collection toolkit, dataset, and controlled comparison.
1. What is shared across hands?
The seven embodiments are XHand, Inspire, Wuji, LEAP, Sharpa Wave, LinkerHand, and Allegro. They retain their own kinematics, actuator limits, joint ordering, and mechanical coupling. The single-arm tasks are bucket lifting, nail hammering, and Hanoi; the bimanual tasks are microwave cooking, iPad unlocking, and photography. Every hand uses the same task objects, initial-scene protocol, and physical success criteria.
A common learning interface synchronizes RGB images, instructions, proprioception, and commands at 20 Hz. Arm actions use world-frame translation and rotation-vector increments. Native state and execution-command layouts are padded to $2\times31$ and $2\times28$, with masks for valid coordinates and active arms. These are native interface dimensions; policy representations can use different layouts. The system supports up to five $256\times256$ camera streams, with a fixed view subset within a backbone comparison.
For hand $h$, let $u_h$ denote its valid native hand command. An encoder produces training targets and a decoder converts predictions back to executable commands:
\[z=e_h(u_h),\qquad \hat u_h=d_h(\hat z).\]This separation is the experimental foundation: observations, arm commands, and execution remain shared while the hand-action encoding changes. A common tensor shape gives a policy a consistent interface, but learning still has to account for the meaning of each hand’s coordinates.
2. The three action representations
Native: preserve each hand’s coordinate order
Native inserts the original command vector into a shared padded layout:
\[z=P_hu_h.\]The decoder selects valid entries, and a mask excludes padding. It is a simple baseline that preserves the original control variables. The policy must learn how those variables relate across embodiments because the same position in the vector need not express the same functional role.
FAAS: assign coordinates to functional slots
FAAS uses a 32-slot hand layout, separate from arm commands, following the functional-slot principle of UniDex. A hand-specific adapter assigns native coordinate $i$ to slot $\sigma_h(i)$:
\[z_{\sigma_h(i)}=s_{h,i}u_{h,i}+b_{h,i}.\]The assignment aligns functional roles; $s_{h,i}$ and $b_{h,i}$ account for direction and offset. Execution applies the corresponding inverse mapping to active coordinates, with dependent joints expanded according to the native hand model. The alignment retains hand-specific command values and coupling constraints.
The practical attraction is a transparent, structured correspondence between the policy output and the hardware. The remaining question is whether this correspondence helps the tasks and pretrained backbone being used. The experiments show that its benefit changes between single-arm and bimanual control.
DexLatent: learn a codec for each hand
DexLatent follows the hand-specific encoder/decoder formulation of XL-VLA. Each encoder $E_h$ maps native commands into a shared latent, while $D_h$ maps policy predictions back to that hand. Codec fitting uses
\[\mathcal L_{\mathrm{codec}} =\lambda_r\mathcal L_{\mathrm{rec}} +\lambda_g\mathcal L_{\mathrm{geom}} +\lambda_p\mathcal L_{\mathrm{prior}}.\]The terms measure native-command reconstruction, cross-hand fingertip geometry through differentiable forward kinematics, and latent-distribution regularization. Geometry is compared over corresponding available digits. The codec is frozen during policy training, so control performance depends on both the policy’s latent predictions and the fitted decoder.
My interpretation is that this introduces another place where geometric similarity and task-relevant contact precision can diverge. The paper reports lower success for DexLatent, but supplies no codec-loss ablation or decoder-error analysis that identifies the cause. Its results therefore constrain this implementation and training setting; they do not establish that learned action latents are generally inferior.
3. Data collection preserves the comparison
Rokoko gloves and Vive trackers provide finger and wrist motion. Redesigned mappings for all seven hands produce native demonstrations, respecting digit correspondence, motion direction, actuator limits, and coupling. Representation encoding happens afterward, so Native, FAAS, and DexLatent reuse the same trajectories.
Reviewed source demonstrations are expanded into randomized scenes using task-specific recipes: scene-relative motions, stage transitions, and hand-specific parameters. Development trials refine these recipes before they are frozen for batch generation. Figure 4 describes GPT-6 assistance in task-rule development; the batch pipeline then executes frozen rules with pose feedback and timing- or state-based transitions.
A trajectory enters the accepted dataset only when all four checks pass:
\[A(\tau)=S(\tau)\land Q(\tau)\land R(\tau)\land D(\tau).\]Here $S$ checks physical task success, $Q$ motion quality, $R$ replay of saved actions, and $D$ data validity, including images, masks, timestamps, and metadata. Failed attempts remain in the audit record. The resulting training set has 50 accepted trajectories per hand–task pair, totaling $7\times6\times50=2{,}100$. Balanced trajectory counts do not imply balanced frame counts: sampling is uniform over applicable frames, so longer trajectories contribute more training frames.
4. Preserving the action head helps, but joint training remains harder
The authors first replace the pretrained 32-dimensional action projection of $\pi_{0.5}$ with an 80-dimensional bimanual output, reserving 40 dimensions per side. That adaptation produces near-zero success in their experiment.
Following Ego-Pi, they then preserve the 32-dimensional projection and interleave commands across tokens:
\[L_t,\ R_t,\ L_{t+1},\ R_{t+1},\ldots\]Each arm–hand command contains at most 28 valid values and fits in one token. Fifty tokens represent 25 bimanual control steps, with left/right commands paired by time for execution. This supports seven separately trained policies, each covering one hand’s six tasks. Mixing all seven hands into one Ego-Pi policy performs substantially below the per-hand models, although the paper does not tabulate that joint model’s score.
Being-H0.5 supplies a different starting point: cross-embodiment pretraining involving human MANO motion and 30 robot embodiments, plus a Mixture-of-Flow architecture with embodiment-aware experts. DexJoCo-X fine-tunes one joint seven-hand policy for each representation. The three Being-H0.5 runs share demonstrations, views, arm commands, sampling, and optimization budgets, making this the strongest controlled comparison in the paper.
5. Results and what they support
Table I reports the following success rates. Each single-arm or bimanual average weights 21 hand–task pairs equally.
| Policy and training scope | Representation | Single-arm | Bimanual | Overall |
|---|---|---|---|---|
| $\pi_{0.5}$ + Ego-Pi, seven per-hand policies | Native | 31.5% | 23.2% | — |
| Being-H0.5, one joint policy | Native | 57.2% | 36.7% | 47.0% |
| Being-H0.5, one joint policy | FAAS | 54.3% | 41.0% | 47.7% |
| Being-H0.5, one joint policy | DexLatent | 40.7% | 25.6% | 33.1% |
Evaluation uses three independent sets of 50 resets per hand–task pair: 150 rollouts per cell and 6,300 per full configuration. Reset sets are shared across methods and disjoint from demonstration-generation seeds. Overall success is the macro-average of the 42 cells. These repeated evaluations are not reported as independent training-seed replications.
FAAS gains 4.3 percentage points over Native on bimanual tasks and loses 2.9 points on single-arm tasks. The overall advantage is only 0.7 points, and the paper provides no significance test or uncertainty interval for that difference. The useful finding is the task-group tradeoff. Native remains a strong baseline inside a model capable of joint multi-hand learning.
Cross-backbone comparisons need a separate reading. Table II uses 300 demonstrations, 5,000 updates, and batch size 128 for each Ego-Pi policy; Being-H0.5 uses 2,100 demonstrations, 120,000 updates, and batch size 8. Learning rates, warmup, prediction horizons, pretraining, and architecture also differ. These are comparisons between complete training systems. They do not isolate the causal contribution of embodiment-aware experts or pretraining. Update counts alone also do not measure relative compute because batch sizes and model costs differ.
6. Research takeaways and limits
The benchmark’s strength is a reusable route from native demonstrations to representation-specific training and a common physical execution interface. It makes it practical to ask whether a proposed alignment improves closed-loop task completion while holding the downstream backbone and data fixed.
The reported policies train on all seven evaluated hands. Physical-robot evaluation and held-out-hand zero-shot/one-shot protocols are future work. Six simulated tasks also leave substantial room for broader contact patterns and task distributions. A shared policy with 47.7% mean success still has substantial failure rates and sharply uneven hand–task performance.
For my own experiments, I would start with Native and explicit masks, compare a functional-slot adapter under the same backbone and training budget, and inspect single-arm and bimanual results separately. For a learned codec, I would additionally measure reconstruction and contact-sensitive decoding errors alongside rollout success. Those are proposed follow-up diagnostics. DexJoCo-X establishes the comparison framework and the observed tradeoff; explaining exactly why a representation succeeds or fails requires further ablations.
