[Paper Notes] AnyViewDex: View-Invariant Dexterous Manipulation from RGB Observations
Published:
This post supports English / 中文 switching via the site language toggle in the top navigation.
TL;DR
Moving a camera changes the apparent relationship between a fingertip and an object, even when their physical positions stay fixed. AnyViewDex trains an RGB policy to tolerate this change by combining two simulation-only signals: paired views of the same state and the object’s absolute 3D position. Contrastive learning aligns the views; coordinate regression makes the resulting representation retain metric information after global average pooling.
On an xArm7 with a 16-DoF LEAP Hand, the distilled policy succeeds in 368/480 grasping trials (76.7%), using monocular RGB and proprioception without real-world fine-tuning. The evaluation covers eight unseen objects and six camera placements inside the training viewpoint range. The result supports camera repositioning within that range; severe extrapolation still produces grasps displaced from the object.
Paper and source version
AnyViewDex: View-Invariant Dexterous Manipulation from RGB Observations is by Soham Patil, Om Sanjay Gunjal, Sourabh Bhosale, Arhan Chavare, Ramandeep Singh Hora, and Spandan Roy. Patil and Gunjal contributed equally. These notes follow the eight-page arXiv:2609.20107v1, submitted September 17, 2026, including its appendix. I treat it as an arXiv preprint; the source does not identify an accepted venue. The paper links an official project page. All experimental numbers below are author-reported; I have not reproduced the training or robot experiments.
1. What gets aligned, and what the policy actually sees
At each simulated timestep, a fixed canonical camera and a randomized camera render the same physical state. A shared ResNet-18 converts each image into a globally pooled vector. A projection MLP supplies the features for contrastive learning, while the control branch takes the unprojected vector from the randomized camera. The canonical image is a training signal and disappears at deployment.
Writing the normalized projected features explicitly as $z=h(v)/\lVert h(v)\rVert_2$, the paper’s InfoNCE objective can be expressed as
\[\mathcal L_{\mathrm{InfoNCE}} =-\log\frac{\exp((z_t^{\mathrm{canon}})^\top z_t^{\mathrm{rand}}/\tau)} {\sum_j\exp((z_t^{\mathrm{canon}})^\top z_j^{\mathrm{rand}}/\tau)}, \qquad \tau=0.1.\]The denominator compares the matched state with other states in the batch, drawn from other timesteps or parallel environments. Camera parameters are resampled per episode. The objective encourages corresponding states to remain close as the camera moves, but it supplies no explicit units of distance or world-coordinate reference.
This matters because global average pooling removes the explicit feature grid. The network can preserve broad scene identity while losing details needed to place a finger. The authors call this failure spatial collapse. Here that term refers to loss of useful metric information; it should not be read as evidence that every input maps to one identical vector.
2. Absolute coordinate regression gives the representation a metric target
A small auxiliary MLP predicts the object’s translation in the simulator’s world frame from a temporally aggregated state $m_t$:
\[\hat p_t=g_{\mathrm{aux}}(m_t),\qquad \mathcal L_{\mathrm{abs}}=\frac{1}{B}\sum_{i=1}^{B} \left\lVert\hat p^{(i)}-p^{(i)}\right\rVert_2^2.\]The joint objective is
\[\mathcal L_{\mathrm{total}} =\mathcal L_{\mathrm{task}} +\lambda_{\mathrm{InfoNCE}}\mathcal L_{\mathrm{InfoNCE}} +\lambda_{\mathrm{abs}}\mathcal L_{\mathrm{abs}}.\]The same physical object position receives the same target across camera placements. Its prediction gradients update the temporal module and visual encoder along with the control objective. The supervision covers three translation coordinates; object orientation remains the responsibility of the task loss. Ground-truth coordinates and the auxiliary prediction branch are unnecessary for deployed control.
flowchart TD
A["Same simulated state, two camera views"] --> B["Shared ResNet-18 and global pooling"]
B --> C["Projection head: cross-view InfoNCE"]
B --> D["Random-view raw embedding"]
D --> E["Temporal state: frame stack or LSTM fusion"]
E --> F["Policy and task loss"]
E --> G["Auxiliary prediction of world XYZ"]
H["Simulator object XYZ"] --> G
I["Deployment: one RGB camera; proprioception for student"] --> J["Encoder, temporal state, policy"]
J --> K["Robot action"]
Absolute position from an uncalibrated image is still ambiguous in general. This loss teaches the network to use regularities available in its training setup, including the visible robot as a scale reference. It does not remove the need for those visual cues. The paper explicitly reports failures when occlusion or unfamiliar viewpoints undermine them.
There is also a possible shortcut: after contact, joint configuration may reveal where the object is. The distillation episodes therefore begin from a fixed neutral arm pose with randomized object locations, making early coordinate prediction depend on vision. A separate visual probe provides stronger evidence that geometric information actually reaches the encoder.
3. Two training routes share the losses, but use different observations
The Maniwhere RL route runs in MuJoCo and uses DrQ-v2 with a deterministic actor and twin-Q critic. Recent visual embeddings are stacked, and this temporal state contains no proprioception. The auxiliary head predicts the latest frame’s object coordinates. This setup tests whether the representation works with an existing RL control architecture.
The DextrAH-based distillation route runs in Isaac Lab. It concatenates the visual vector with joint positions and velocities, then feeds the result to an LSTM. A privileged PPO teacher supplies action supervision through DAgger. The paper describes the task objective as an uncertainty-weighted MSE between student and teacher stochastic action distributions, with stronger penalties along low-variance teacher dimensions. This recurrent student is the route used for hardware deployment.
The appendix reports a student learning rate of $10^{-4}$, contrastive weight 0.5, a 512-unit LSTM, and policy MLP widths of 512, 512, and 256. It leaves the auxiliary-loss weight out of its condensed hyperparameter table, so the PDF alone is insufficient to reconstruct every training setting. On a single RTX 5090, the authors report about 48 hours for teacher training with 4,096 environments, 14 hours for distillation with 128 environments, and 12 hours per MuJoCo RL task with 256 environments.
Camera coverage differs between the routes. Maniwhere generally samples azimuth from $[-60^\circ,60^\circ]$, with Close Dex using $[0^\circ,120^\circ]$. Distillation uses $[-110^\circ,30^\circ]$, a 140-degree span, plus height, origin, and rotation perturbations. That distribution is part of the method’s operating conditions.
4. The ablations show a strong interaction between the two losses
Hardware evaluation uses the xArm7 and LEAP Hand with an uncalibrated RealSense D455 supplying only RGB. Each method receives 8 objects × 6 viewpoints × 10 trials = 480 trials. Success requires establishing a multi-finger grasp and lifting the object clear of the table. The five conditions total 2,400 trials. Paper, Table III and Section IV-E.
| Training condition | Successful trials |
|---|---|
| Fixed camera | 7/480 |
| Domain randomization only | 145/480 |
| Randomization + absolute-coordinate loss, without InfoNCE | 90/480 |
| Randomization + InfoNCE, without coordinate loss | 40/480 |
| Full AnyViewDex | 368/480 |
The full method improves over domain randomization from 30.2% to 76.7%, approximately 46.5 percentage points. More unusually, adding either representation loss alone lowers performance relative to randomization alone. This is evidence for their interaction in the reported configuration. It also argues against treating either loss as an independently reliable upgrade to an arbitrary RGB policy.
Across the six tested viewpoints, the full model succeeds in 60–64 of 80 trials per view. Those placements are all inside the training cone. The physical study compares matched internal variants; the comparison against a depth-based policy appears in simulation.
5. Small-object localization exposes the pooling tradeoff
The simulation benchmark reports mean success over five seeds. Selected results from Table I show why a single average would hide an important weakness:
| Method | Lift Cube Dex | Pick & Place Dex | Close Dex | Button Dex |
|---|---|---|---|---|
| MV-MWM, RGB | 78.0 ± 5.1 | 34.0 ± 28.9 | 69.5 ± 19.7 | 77.6 ± 14.3 |
| Maniwhere, RGB | 72.5 ± 1.5 | 0.0 ± 0.0 | 17.3 ± 2.7 | 82.4 ± 9.6 |
| Maniwhere, RGB-D | 88.8 ± 8.9 | 76.4 ± 9.2 | 81.5 ± 5.6 | 97.6 ± 1.2 |
| AnyViewDex, RGB | 53.0 ± 5.0 | 72.4 ± 3.6 | 92.1 ± 5.9 | 96.0 ± 2.0 |
Values are percentages, mean ± standard deviation. Bold marks the best RGB result on Lift Cube and the best result overall on the other three tasks.
AnyViewDex has the highest reported RGB-only success on three of four tasks and exceeds the RGB-D baseline on closing a laptop lid. On lifting a small cube, however, it trails both MV-MWM and RGB Maniwhere substantially. The authors attribute this to discarding the spatial grid: a globally pooled vector remains a bottleneck when fine localization is critical, even with 3D supervision.
The contrastive-only ablation adds another qualification. It gets near-zero success on three tasks but retains 88.0% on Button Dex. A fixed-location button press can succeed without the same metric demands as object acquisition. The need for the geometric auxiliary depends on the task.
6. The probe supports geometric encoding, with limited precision
A linear probe on the frozen visual embedding, before proprioceptive fusion, recovers object coordinates with $R^2=0.814$ and 4.73 cm mean Euclidean error over roughly 690,000 in-distribution samples. This supports the claim that metric information exists in the visual representation. It does not establish fingertip-level accuracy: the probe sees one frame, while the deployed controller combines recurrent observations with proprioception and feedback.
Outside the training cone, the probe error rises to about 14.3 cm. The authors report a corresponding physical failure: the hand closes in empty space on a plane displaced from the target. A textureless background close to the object also degrades grasping, and increasing that separation restores performance. These are concrete limits on the learned visual reference frame.
The evidence leaves one causal question open. The authors do not compare the coordinate target with a non-geometric dense auxiliary of matched dimensionality. The probe and ablations support the geometric explanation, but they do not isolate it completely from the optimization benefits of additional supervision.
7. What I would carry into a new system
I would consider this recipe for a compact RGB controller whose camera needs to move within a known workspace. The reported deployment cost is practical: on an RTX 4050 laptop GPU in half precision, the student uses 1–2 GB VRAM, with a 2.74 ms network forward pass and a 5.16 ms end-to-end control step. These are measured computation times, not a demonstrated sustained robot control frequency.
For precision grasping of small objects, I would first retain a low-resolution feature grid and test whether coordinate supervision still helps. For clutter, the single-object XYZ target needs a way to identify which object the policy should manipulate. Before expanding the model, I would also run the missing matched auxiliary control: it would clarify how much of the gain depends on metric geometry itself and guide which simulator labels are worth collecting.
