[Paper Notes] Learning In-Hand Object Reaching to General 6D Poses
Published:
This post supports English / 中文 switching via the site language toggle in the top navigation.
TL;DR
A hand that has already grasped a tool may still need to slide it, turn it, or expose a different working surface. POISE treats this adjustment as palm-relative 6D pose reaching: a fixed wrist and 22-DoF hand use finger motion alone to move the held object toward a target position and orientation in $SE(3)$.
The training recipe has three ideas worth keeping. First, a cache of physics-validated grasps gives reinforcement learning much broader state coverage than perturbing one seed grasp. Second, separate translation and rotation curricula expand the goal range only after the policy succeeds near its current frontier. Third, a wrench-based reward preserves balanced resistance to forces and torques, so reaching one pose is less likely to leave the hand unable to reach the next.
The strongest evidence concerns recovery and continued operation. Diverse initialization raises success from 33.8% to 72.9% when the object begins on the palm without an established grasp. The adaptive curriculum raises strict full-range success from 6.2% to 59.5%. On hardware, the grasp reward raises three-target sequence success from 2/10 to 8/10. These results make POISE a useful in-hand reconfiguration primitive, though every trial still begins after grasp acquisition and the real system depends on a known object mesh and visual pose tracking.
Paper info
“Learning In-Hand Object Reaching to General 6D Poses” is by Junxiao Lin, Tianyue Wu, Jie Yin, Jia Pan, Kaifeng Zhang, and Weiming Zhi, with affiliations at the University of Sydney, Sharpa, and the University of Hong Kong. These notes cover the eight-page arXiv:2609.13761v1, submitted on September 12, 2026.
The project page provides the paper PDF, hardware videos, a browser-based MuJoCo demo, and an overview video. As of September 15, 2026, the project page marks the code as coming soon.
1. The target lives in the palm frame
POISE starts from an already-grasped object and holds the wrist fixed. Let $H$ and $O$ denote the palm and object frames. The controller receives a target palm-relative pose ${}^{H}T_{O,g}=({}^{H}p_g,{}^{H}R_g)$ and must move the current pose ${}^{H}T_{O,t}$ toward it using finger motions. The position and orientation errors are
\[e_p=\left\|{}^{H}p_t-{}^{H}p_g\right\|_2, \qquad e_R=\frac{1}{\sqrt{2}} \left\|\operatorname{Log}\!\left({}^{H}R_g({}^{H}R_t)^\top\right)\right\|_F.\]Here $e_R\in[0,\pi]$ is the geodesic rotation angle. Success requires both errors to stay below 10 mm and $10^\circ$ for 20 consecutive policy steps. The simulator then samples another target without resetting the hand–object state. Training on successive goals matters because a pose can satisfy the current command while ending in a contact arrangement that makes the next command unreachable.
This palm-relative formulation isolates finger dexterity. The policy does not use arm motion to compensate for limited in-hand range, and it does not track a reference hand trajectory. Compared with orientation-only or translation-only tasks, a general $SE(3)$ target couples the two motions: the hand may need to translate the object temporarily to create room for a large rotation, change contacts, and rebuild support afterward.
2. A reset cache doubles as recovery training
A policy initialized from one grasp explores states reachable from that contact arrangement. Small perturbations widen the neighborhood, but they rarely expose the controller to a different way of supporting the object. POISE instead samples object poses across a feasible in-hand workspace, optimizes contacts and joint configurations, converts the result into position-control commands, and runs each candidate for two seconds under gravity. Only grasps that retain the object enter the cache.
The cache makes state coverage part of the dataset design. This becomes especially clear in the recovery experiment. The narrow and diverse policies each use 6,842 reset states and differ only in how those states are distributed. On independently generated held-out grasps with paired 30 mm/$180^\circ$ targets, diverse initialization raises first-target success from 40.1% to 51.5% and completed goals per episode from 0.69 to 1.00.
The harder test starts from 720 shared trials in which the object rests on the palm and no grasp has been established. Reaching the target requires the fingers to build contact, lift the object, and resume pose control. Success rises from 33.8% to 72.9%, while mean recovery-and-reach time among successful trials falls from 11.63 s to 9.45 s.
I find this result more interesting than the held-out-grasp gain. It shows that the initialization distribution can teach a recovery behavior without a separate recovery policy or an explicit state machine. There is a boundary to the claim: the object begins on the palm, so this is recovery from contact loss inside the hand, not retrieval after the object falls away.
3. Geometry conditions one policy; the curriculum opens the workspace
The actor is a three-layer ELU MLP. It receives three-frame histories of measured and commanded finger joints, estimated palm-frame object pose and velocity, palm-frame gravity, visual-pose confidence, the target pose, current-to-target translation and rotation, and a geometry descriptor. An asymmetric critic gets clean simulator state, fingertip information, joint torques, and contact forces during training. Actions increment commanded joint positions, and the deployable policy runs at 20 Hz.
Object shape enters through a 64-dimensional basis point set (BPS) descriptor. Fixed query points are shared across canonical object frames; each component records the normalized distance from one query point to the nearest object surface. A multi-object policy receives this vector without a categorical object ID, so the same network can adjust its finger coordination for different contact surfaces and edges.
Large coupled pose changes are initially too sparse for useful exploration. POISE grows the maximum rotation from $5^\circ$ to $180^\circ$ and translation from 10 mm to 30 mm. Goal axes and directions remain random. At each stage, target magnitudes are sampled near the frontier, within the exposed range, or at the current limit in a 0.6/0.3/0.1 ratio. Frontier success of 0.4 advances the corresponding bound by $10^\circ$ or 10 mm, with separate curriculum progress for each object.
Direct full-range training is the failed path that makes the curriculum result convincing. Under the strict tolerance, it reaches only 6.2% success on full-range targets and 3.4% on maximum-change targets. The curriculum reaches 59.5% and 55.3%, respectively.
flowchart TD
A["Physics-validated grasp cache"] --> B["Random stable reset"]
B --> C["Actor observation history"]
D["BPS object geometry"] --> C
E["Adaptive palm-relative 6D goal"] --> C
C --> F["PPO finger policy at 20 Hz"]
F --> G["Incremental joint-position command"]
G --> H["Object motion and contact changes"]
H --> C
I["Pose, completion, grasp, and effort rewards"] --> F
J["FoundationPose on real RGB-D"] --> C
4. Reward rotation first, then tighten translation
The dense pose reward uses separate rotation and position terms:
\[r_t^R=\exp\!\left(-\frac{e_{R,t}}{30^\circ}\right), \qquad r_t^p=\alpha(e_{R,t}) \exp\!\left(-\frac{e_{p,t}}{\sigma(e_{R,t})}\right),\] \[r_t^{\mathrm{pose}}=\frac{1}{6}r_t^R+\frac{1}{4}r_t^p, \qquad (\alpha,\sigma)= \begin{cases} (0.15,30\text{ mm}), & e_R>45^\circ,\\ (0.40,20\text{ mm}), & 25^\circ<e_R\le45^\circ,\\ (1.00,15\text{ mm}), & e_R\le25^\circ. \end{cases}\]When rotation error is large, the position term is broad and lightly weighted. The object can translate while the fingers rearrange contacts. As orientation aligns, the position target becomes narrower and stronger. A completion reward adds 0.5 per step inside the tolerance and a one-time bonus of 45 after the required 20-step dwell.
The unusual part is the grasp term. Each active contact becomes four friction-cone wrench rays, and each ray’s moment is normalized by object size. For each of three force axes and three torque axes, $m_{t,k}$ records the smaller available projection in the positive and negative directions. POISE aggregates the six margins with a generalized mean:
\[Q_t=\left[\frac{1}{6}\sum_{k=1}^{6}(m_{t,k}+\epsilon)^{-8}\right]^{-1/8}-\epsilon.\]The negative exponent makes the weakest bidirectional wrench margin dominate. Contact contributions saturate with force, so squeezing harder cannot raise the score indefinitely. Because achievable quality depends on the object and initial grasp, the reference $Q^\star$ comes from the first four control steps and is clipped to $[0.08,0.35]$:
\[r_t^{\mathrm{grasp}} =1-\operatorname{clip}\!\left( \frac{(Q^\star-Q_t)_+}{Q^\star},0,1 \right)^2.\]The complete reward adds pose, verified completion, and $0.05r_t^{\mathrm{grasp}}$, then penalizes drops, joint torque, and instantaneous mechanical power. In simulation, the grasp term raises $Q$ by about 41%, the weakest force margin by 39%, and the weakest torque margin by 31%. Episode success also moves from 53.7% to 59.1%, so the extra contact objective does not merely trade reaching for static grasp quality.
5. Sim-to-real depends on explicit pose tracking
Training randomizes object scale and mass, inertia, center of mass, friction, PD gains, wrist orientation, and external disturbances. The observation path also receives joint noise, 10 mm-per-axis object-position noise, clipped rotation noise, 0–150 ms pose delay, and 3% per-step pose dropout with hold-last behavior. This separates two transfer problems: the physics has to tolerate contact mismatch, while the policy has to keep working with delayed or missing visual state.
Hardware uses a 22-DoF Sharpa Wave hand and one RealSense D435 RGB-D camera. FoundationPose tracks a known object mesh, and a calibrated camera-to-hand transform expresses the estimate in the palm frame. The observation and action interfaces stay the same as in simulation.
Calling the system vision-based needs that context. The actor does not infer object geometry or pose end to end from pixels. It receives an explicit BPS descriptor and a tracked 6D pose from a model-based visual frontend. This is a reasonable engineering split for studying finger control, but tracking errors remain one of the paper’s stated sources of the real-to-sim performance gap.
6. What the experiments establish
The three design choices survive controlled ablations
| Component | Comparison | Main result |
|---|---|---|
| Diverse resets | Narrow vs. diverse, held-out grasps | 40.1% → 51.5% |
| Diverse resets | Narrow vs. diverse, post-contact-loss recovery | 33.8% → 72.9% |
| Goal curriculum | Full-range sampling vs. curriculum | 6.2% → 59.5% |
| Grasp reward | Without vs. with, simulation episode success | 53.7% → 59.1% |
| Grasp reward | Without vs. with, real three-target sequences | 2/10 → 8/10 |
The hardware ablation deserves a closer read. Ten trials per policy use the same initial grasp and three-target sequence. The grasp reward raises completed targets from 8/30 to 27/30 and cuts median reach time over reached targets from 6.8 s to 4.9 s. Median steady-state errors are slightly larger with the reward: position changes from 4.8 mm to 6.4 mm and orientation from $4.7^\circ$ to $5.9^\circ$. The useful gain is continuity. Distributed contacts prevent progressive contact loss across the sequence; the ablation does not claim better final pose accuracy.
Real demonstrations cover continued reaching, gravity changes, and disturbances
One geometry-conditioned policy is trained across nine shape–size combinations from Cube, Hexagonal Prism, and Square Bifrustum families. On hardware, each object completes at least five random goals in an uninterrupted rollout. The Hammer uses a separate setup with translation curriculum extended to 100 mm because its 215 × 50 × 33.6 mm body has a longer moment arm and workspace.
For user-specified sequences, the Hexagonal Prism reaches four targets with changes up to 35 mm and $100^\circ$ at a mean 3.0 s per target. The Hammer reaches seven targets with changes up to 93 mm and $180^\circ$ at 2.7 s per target. The same wrist-randomized policy is also shown under three fixed wrist orientations, completing 7, 5, and 6 successive goals. A separate demonstration applies two external disturbances after reaching; closed-loop visual feedback reorganizes the contacts and returns the object to the unchanged target.
These are useful capability demonstrations, with limits on what they quantify. The paper reports representative continuous rollouts instead of a large per-object hardware success table. The first three geometries come from the policy’s training families and sizes, while the Hammer uses its own expanded curriculum. BPS conditioning supports one policy across several known geometries here; unseen-shape generalization remains open.
7. Limits and research takeaways
POISE commands the object pose without specifying a target hand posture or contact arrangement. The learned solution can work while looking unnatural. The authors suggest conditioning on both object and grasp targets, which could make the final contact state better suited to a downstream task. They also identify contact-model error and visual 6D tracking error as current transfer bottlenecks, with point-cloud observation distillation, tactile feedback, and adaptation from real interaction as next steps.
The task boundary matters just as much. Every rollout begins with the object already in the hand. Grasp acquisition, recovery from a floor drop, and coordination with arm motion lie outside this formulation. The fixed-wrist setup is useful because it measures what the fingers can do, though a complete tool-use system would eventually decide when to reposition the arm and when to manipulate in hand.
My main takeaway is that initial-state design is part of the controller. The reset cache determines which contact arrangements the policy learns to escape, exploit, or rebuild. In POISE, that choice produces a larger recovery gain than the ordinary held-out-grasp gain. The adaptive curriculum then makes large $SE(3)$ goals learnable, and the grasp reward makes a reached state worth continuing from. Together, the three choices turn single-goal pose matching into repeated in-hand reconfiguration.
