[Paper Notes] OpenDexGrasp: Open-vocabulary Task-Oriented Dexterous Grasping

10 minute read

Published:

This post supports English / 中文 switching via the site language toggle in the top navigation.

TL;DR

A stable grasp is not necessarily a useful grasp. OpenDexGrasp generates a dexterous hand pose from a free-form instruction, multi-view images, and an object point cloud, with the contact region and hand configuration aligned to the intended use. Its central design is to learn affordance grounding and grasp generation in one shared perception-action latent space, so affordance prediction helps training and interpretation without becoming a test-time cascade.

The data follow a Coverage-to-Alignment (C2A) Recipe. OpenDex-Scale contributes broad semantic and geometric coverage through automatic grasp synthesis and vision-language annotation; OpenDex-Align contributes smaller but higher-quality functional demonstrations from teleoperation and category-level transfer. The two subsets contain 1.24M and 27.42K grasps respectively, with functional ratios of 44.93% and 68.02%.

On the paper’s simulation benchmark, OpenDexGrasp reaches 71.65% success, compared with 49.15% for an adapted DexGraspNet 2.0 baseline. In the real-robot evaluation it reaches 72.0% versus 59.0%. These are task-oriented grasping results under the paper’s object and instruction splits; they should not be read as long-horizon manipulation success.

Paper and source version

Jiyao Zhang, Junhan Wang, Tianyu Wang, Zeyuan Chen, Anthony Bolten, Yitong Peng, and Hao Dong, from Peking University, PrimeBot, and BIGAI. These notes follow arXiv:2609.18117v2, dated September 17, 2026. See the paper PDF and official project page. The manuscript identifies itself as a CoRL 2026 paper. All numbers below are author-reported.

1. The target is functional contact

Task-agnostic grasping asks whether a hand can hold an object. Functional grasping asks whether the hold preserves the action implied by the instruction: grasp a kettle by its handle for pouring, keep a spray-bottle trigger accessible, or avoid a knife blade when handing it over. The relevant signal is therefore distributed across language, object parts, viewpoint, geometry, and a continuous high-DoF hand pose.

OpenDexGrasp represents one example as $(u, I, P, a, m)$: a free-form instruction $u$, multi-view RGB observations $I$, an RGB point cloud $P$, a dexterous grasp pose $a$, and an optional point-level affordance label $m$. The pose is parameterized by global translation, a continuous 6D wrist rotation, and hand joint configuration:

\[a=[p,r_{6D},q].\]

The model learns the conditional action distribution $p_\theta(a\mid u,I,P)$. Affordance is auxiliary supervision over the same latent state. Conceptually,

\[p_\theta(m,a\mid u,I,P)=\int p_\theta(m\mid z,u,P)\,p_\theta(a\mid z,u,P)\,p_\theta(z\mid u,I,P)\,dz.\]

This factorization describes shared supervision, not an inference pipeline that first predicts a map and then optimizes a hand pose.

2. OpenDexVerse: coverage first, alignment second

The dataset is designed around a practical tension. Automatic synthesis scales across shapes and grasp modes but can produce noisy functional labels and unnatural contact choices. Human teleoperation provides reliable task contact and natural articulation but is expensive. C2A assigns these sources different jobs.

OpenDex-Scale starts from category-aligned real scanned objects. Following the BODex synthesis pipeline, it samples and optimizes physically plausible dexterous candidates. Each candidate is rendered from three object-centered views and annotated by a vision-language model with object, part, and task descriptions. The result covers 105 categories, 1,110 instances, and 1.24M grasps, of which 557.18K are labeled functional.

OpenDex-Align collects high-quality task-oriented grasps with human teleoperation for three size-aware templates per category. Dense correspondences in category coordinates transfer those demonstrations to nearby compatible instances. It contains 95 categories, 2,770 instances, and 27.42K grasps, including 18.65K functional poses. The smaller set supplies embodied alignment after the broad synthetic prior has been learned.

This ordering matters. The model first sees a wide support of language, geometry, and hand configurations, then its distribution is pulled toward reliable functional behavior. The paper’s recipe is therefore a data curriculum as much as a dataset composition.

3. One latent space for language, geometry, affordance, and action

The architecture uses a pretrained vision-language encoder for the multi-view images and instruction. Selected hidden states retain relationships among task words, object appearance, and view-dependent part evidence. A hierarchical point-cloud encoder supplies a global geometry token while retaining point features for affordance decoding.

A transformer action expert then combines geometry, learned queries, and noisy action tokens. It conditions on the vision-language hidden states through cross-attention and generates the grasp with flow matching. Given target action $a$, Gaussian noise $\epsilon$, and time $t$,

\[x_t=(1-t)\epsilon+ta,\qquad v^\star=a-\epsilon.\]

The action loss is

\[\mathcal L_{act}=\mathbb E_{a,\epsilon,t}\left\|F_\theta(x_t,t,z_P,\{H_\ell\})-(a-\epsilon)\right\|_2^2.\]

An affordance head predicts a score for every object point. Its loss combines focal and Dice terms:

\[\mathcal L_{aff}=\mathcal L_{focal}(\hat m,m)+\lambda_{dice}\mathcal L_{dice}(\hat m,m).\]

The full objective is

\[\mathcal L=\mathcal L_{act}+\lambda_{aff}\mathcal L_{aff},\qquad \lambda_{aff}=0.3.\]

At inference time, the model directly samples a hand pose from language, images, and geometry. There is no separate affordance-to-pose optimization stage.

4. Simulation results

The evaluation separates functional grasps, where the pose must support a downstream use, from non-functional grasps, where any stable hold is acceptable. Each is tested on seen and unseen object categories. The adapted baseline, marked DexGraspNet 2.0*, adds CLIP features and uses the same data and splits.

For the main functional setting, OpenDexGrasp improves on the baseline as follows:

SplitSIV (cm³)PD (cm)SD (cm)SuccessStyle diversityGPT-5 / human score
Seen, baseline5.121.041.5150.88%0.956.48 / 6.10
Seen, OpenDexGrasp1.280.391.3768.07%1.397.43 / 7.85
Unseen, baseline6.101.624.2943.22%0.815.87 / 5.37
Unseen, OpenDexGrasp1.390.481.8062.96%1.416.98 / 6.62

SIV measures solid intersection volume, PD local penetration depth, and SD object displacement after simulation. Lower is better for the first three. Style diversity measures variation across stochastic predictions, so the higher value means the model keeps multiple grasp styles instead of collapsing to one pose.

Ablations expose the role of each ingredient. The full model averages 71.65% success across the simulation splits. Removing affordance grounding lowers this to 67.49%; removing OpenDex-Align lowers it to 69.38%; reducing OpenDex-Scale lowers it to 60.52%. Replacing the pretrained vision-language model with CLIP produces the largest drop, to 57.26%, alongside much worse physical metrics. Broad coverage, embodied alignment, point-level grounding, and a strong VLM each contribute a different part of the result.

The paper also compares direct generation with an affordance-then-optimization variant that uses the same predicted affordance. Direct OpenDexGrasp obtains 71.65% success with 0.93 s average inference time, while optimization obtains 49.15% with 2.84 s. This isolates a useful engineering point: a shared latent action generator can make the affordance signal useful without paying for hundreds of test-time optimization steps.

5. Real-robot transfer

The real setup uses a Sharpa Wave Hand mounted on a Franka Emika Panda, with an Intel RealSense D435 for object pose estimation. For each object and instruction, the authors sample ten valid poses, discard those that would collide with or be occluded by the table, apply a fixed $0.05$-rad closing refinement, and execute each pose once.

Across five seen and five unseen test objects, OpenDexGrasp reaches 72.0% success, while DexGraspNet 2.0* reaches 59.0%. GPT-5 perceptual scores are 7.5 versus 6.7, and human scores are 7.7 versus 6.5. The per-category table shows gains on items such as rice paddle, dustpan, bouquet, pitcher, bottle, hammer, and brush, although performance remains imperfect on several objects.

The evaluation is still a grasp-and-lift style test. It demonstrates that functional contact choices transfer to a physical hand, but it does not establish closed-loop completion of actions such as pouring, brushing, or cutting.

6. What the paper leaves open

The model depends on the visual-language representation and the coverage of its views; errors in part grounding can still affect the hand pose. It predicts an open-loop action and uses lightweight execution-time selection, without tactile feedback or closed-loop correction. OpenDex-Align is also modest compared with the diversity of household tools and long-horizon tasks.

My main takeaway is that task-oriented dexterous grasping benefits from separating semantic-geometric coverage from embodied functional alignment, then joining them in the same action representation. The most convincing evidence is the combination of unseen-category simulation results, the direct-generation ablation, and the real-hand transfer. The next test I would want is closed-loop execution with tactile feedback, where a grasp is judged by completing the instructed use rather than by holding the object successfully.