[Paper Notes] ArtManip: Category-Level Articulated In-Hand Manipulation
Published:
This post supports English / 中文 switching via the site language toggle in the top navigation.
TL;DR
Articulated in-hand manipulation couples two objectives that naturally interfere. The hand must move an internal joint while keeping the object’s free-floating base stable. The initial grasp must therefore be both stable and functional: fingers need access to the moving link, useful force directions, and enough support to absorb the reaction force.
ArtManip turns that initialization problem into part of the training distribution. It constructs two-link articulated objects from geometric primitives, marks fingertip-specific allowed and prohibited contact regions with one template per category, and uses a modified Lightning Grasp pipeline to synthesize diverse functional grasps. A privileged teacher then learns with SAPG, an articulation-physics distribution, and a stability-reward curriculum. A temporal student maps 50 steps of proprioceptive history to the teacher’s 16-dimensional latent and reuses the frozen policy.
On held-out simulated instances from four categories, every object has at least one successful grasp, average grasp coverage is 92.3%, and execution-level success ranges from 72.1% to 85.0%. On 12 real objects, 257 of 300 executions complete at least one open–close cycle, for 85.7% success. Increasing the Knife training set from one to 30 instances raises unseen-instance success from 29.4% to 85.0% and also improves geometry- and dynamics-OOD tests.
The 12 test objects receive no policy fine-tuning, but each gets a primitive digital twin; simulation selects five candidate grasps; a person manually recreates each initial grasp; and category-level physics ranges are calibrated with a separate real object. The deployed system has no online vision or touch. Its dynamic evidence comes from proprioceptive history, while static initial geometry and pose descriptors come from the digital twin. I read ArtManip as strong evidence that structured simulation diversity can support category-level articulated control. Autonomous acquisition and recovery after contact loss remain open.
Paper and source version
ArtManip: Category-Level Articulated In-Hand Manipulation is by Yang Yang, Tengyu Liu, Puhao Li, Zeyuan Chen, Yuyang Li, Xingwan Wang, Yingying Wu, Zhaopeng Cui, and Siyuan Huang. The listed affiliations are Zhejiang University, the Beijing Institute for General Artificial Intelligence (BIGAI), Tsinghua University, and Peking University.
These notes follow the 17-page arXiv:2609.12498v1, submitted September 11, 2026. The official project page provides rollout videos and an overview. I read the paper and appendix; the reported results have not been independently reproduced here.
1. The joint and the grasp form one control problem
Given an articulated object instance $m$, a functional initial grasp $\gamma\in\mathcal G_m$, and a relative articulation goal $g$, ArtManip learns a policy that actuates the object’s internal joint while maintaining the grasp. Once the target joint state is reached, the goal switches to the opposite state, producing repeated open–close cycles.
This setting differs from opening a cabinet or drawer because the object base has no environmental support. Contact force on the moving link produces a reaction on the rest of the object. A policy can reach the joint goal and still fail the task by rotating the base, losing the actuation contact, or dropping the object. Initial contact geometry determines whether useful torque and stabilizing force can coexist.
ArtManip treats three sources of variation as a joint distribution:
- Instance geometry: link dimensions, joint locations, and joint limits vary within a category.
- Initial contact: each instance is paired with many functional grasps instead of one canonical hand pose.
- Articulation physics: mass, friction, damping, and spring stiffness vary during training.
The authors train a separate category-level solution for each of four tool categories—utility knives, lighters, staplers, and tongs—on a 22-DoF Sharpa hand. Knives use a prismatic joint; the other three categories use revolute joints. Staplers and tongs include spring-like restoring dynamics.
2. Functional grasps create the useful part of the state distribution
A training asset contains two primitive-box links connected by one joint. Category rules sample link sizes and joint limits. The low-detail model keeps the structures that most directly affect the task: where fingers can contact, how the moving link travels, and how much room the hand has to stabilize the base.
The grasp generator begins with a category-level contact template. For each fingertip, the template marks regions where contact is allowed and regions that should be avoided. The template is specified manually once per category and transfers automatically across procedurally generated instances. A modified Lightning Grasp solver then generates 1,000 candidates per object. Simulation keeps a candidate only if the hand can hold the object for one second without a drop.
The paper generates grasps for 35 objects in each category—30 training objects and five held-out test objects. Grasp synthesis takes roughly 8–24 hours per category, depending on the category. This is automated after the template is defined, though the task semantics still enter through the hand-authored contact regions.
flowchart TD
A["Category rules + contact template"] --> B["Primitive two-link assets"]
B --> C["1,000 constrained grasp candidates per object"]
C --> D["One-second stability filter"]
D --> E["Diverse object–grasp initial states"]
E --> F["Privileged SAPG teacher"]
G["Physics randomization + reward curriculum"] --> F
F --> H["16-D privileged latent"]
H --> I["Temporal student distillation"]
I --> J["Deployable policy with initial descriptors + proprioceptive history"]
J --> K["Digital-twin grasp screening"]
K --> L["Manual initialization and real execution"]
One ablation is decisive. Removing functional-grasp synthesis leaves teacher learning near zero; removing the reward curriculum produces much slower and weaker learning. Exploration begins before RL: the initial-state generator must put the hand where useful contact behaviors can be discovered.
3. The teacher learns progress first, then increasingly stable progress
The privileged teacher receives four inputs. The 57-dimensional initial observation contains the initial hand configuration, the two link poses, fingertip positions, and the primitive bounding-box dimensions of both links. The 59-dimensional proprioceptive observation contains current hand joints, fingertip positions, and the previous action. A scalar goal encodes the target joint displacement. A 21-dimensional privileged observation contains current link poses, joint state and velocity, link masses, friction, damping, and stiffness.
An MLP compresses privileged state into a 16-dimensional latent:
\[z_t=f_{\mathrm{enc}}\!\left(o_t^{\mathrm{priv}}\right), \qquad a_t=\pi\!\left(o_t^{\mathrm{prop}},o^{\mathrm{init}},z_t,g\right).\]The LSTM policy outputs 22 normalized joint commands. They update joint-position targets incrementally:
\[q_t^{\mathrm{target}} =q_{t-1}^{\mathrm{target}}+\alpha a_t, \qquad \alpha=\frac{1}{40}.\]The reward groups articulation progress, base and contact stability, motion regularization, drop penalties, and success bonuses:
\[r=w_1r_{\mathrm{progress}}+w_2r_{\mathrm{stable}} +w_3r_{\mathrm{regulate}}+w_4r_{\mathrm{terminate}}.\]Stability creates an exploration problem. A large early penalty for base motion can make the policy hold still; a weak penalty produces aggressive behaviors that transfer poorly. ArtManip starts the stability weights at 10% of their final values. Beginning at epoch 100, it increases them linearly over 1,000 epochs. The policy first discovers articulation, then learns to execute it with lower base motion and more consistent contact.
The dynamics distribution distinguishes passive and spring-like joints. Knife and lighter joints are actuated through contact against randomized damping. Stapler and tong joints use an internal PD model with randomized stiffness. The remaining randomization covers link masses, friction, and proprioceptive noise. SAPG learns faster and reaches higher cumulative cycle counts than standard PPO in the reported training curves.
Training scale is substantial: 20,000 parallel environments, about 2 billion environment steps, two NVIDIA RTX 5090 GPUs, and roughly 48 hours. The category-level generality is therefore purchased through both structured data generation and large simulation throughput.
4. Closed-loop latent distillation removes the privileged state
The teacher cannot run on hardware because its latent encoder sees exact object state and physical parameters. The student replaces that encoder with a temporal convolutional network:
\[z'_t=f_{\mathrm{stu}}\!\left(o^{\mathrm{prop}}_{t-H:t},o^{\mathrm{init}}\right), \qquad \mathcal L=\left\|z'_t-z_t\right\|_2^2.\]The history length is $H=50$. At the 30 Hz control rate, the temporal window spans about 1.67 seconds. The TCN uses five layers, hidden size 192, and outputs the same 16-dimensional latent expected by the copied policy.
During distillation, the policy backbone is frozen and receives the student’s $z’_t$. Rollouts therefore visit states induced by the student’s own imperfect inference. This closed-loop design reduces the train–deploy state mismatch that would arise if the policy continued to act on the teacher latent during data collection.
“Proprioceptive deployment” needs one qualification. Online adaptation comes from joint state, kinematic fingertip position, and previous-action history. The policy also retains $o^{\mathrm{init}}$: the starting hand configuration, two initial link poses, fingertip positions, and primitive link dimensions. Real deployment obtains those static descriptors from a manually constructed digital twin and a selected grasp. The system removes continuous privileged state estimation; it still depends on an initialized object model.
Teacher–student comparisons show where information is lost. On unseen objects, average grasp coverage falls from 96.4% for the teacher to 92.3% for the student. Average successful-grasp cycle count falls more sharply, from 3.48 to 1.93, even though the best-cycle metric remains similar. History recovers enough latent information for broad success, while exact privileged state still supports more consistently long executions.
5. Coverage metrics separate “a grasp exists” from “most trials work”
A single success rate would hide the method’s sensitivity to initialization. Instance coverage (IC) asks whether every test object has at least one successful grasp. Grasp coverage (GC) measures the fraction of tested grasps that work. CSC summarizes repeated open–close cycles over successful grasps. Execution success rate (SR) counts all randomized rollouts or real trials that complete at least one cycle.
For object $i$, with tested grasps $G_i$ and successful subset $S_i$, the first two metrics are
\[\mathrm{IC}=\frac{1}{N}\sum_i \mathbf 1\!\left(|S_i|>0\right), \qquad \mathrm{GC}=\frac{1}{N}\sum_i\frac{|S_i|}{|G_i|}.\]IC is intentionally permissive: one viable grasp makes an instance count as covered. GC and SR reveal how much of the initial-state and dynamics distribution the policy actually tolerates.
On five unseen simulated objects per category, each grasp receives 100 randomized rollouts:
| Category | Evaluated grasps | IC ↑ | GC ↑ | CSC mean ↑ | CSC max ↑ | SR ↑ |
|---|---|---|---|---|---|---|
| Knife | 54 | 100% | 93.8% | 2.6 | 5.6 | 85.0% |
| Lighter | 101 | 100% | 92.8% | 1.4 | 7.0 | 72.1% |
| Stapler | 213 | 100% | 91.9% | 1.4 | 8.6 | 83.2% |
| Tong | 287 | 100% | 90.6% | 2.3 | 8.6 | 79.8% |
The 100% IC result says that every held-out instance has a workable initialization. The more informative robustness numbers are the 92.3% average GC and the per-execution success rates above. Lighters are the hardest category in randomized simulation despite high grasp coverage, suggesting sensitivity to dynamics after a nominally viable grasp has been found.
6. Training diversity improves interpolation and measured extrapolation
The Knife study holds the evaluation sets fixed and increases the number of training instances:
| Training instances | Anchor SR | Unseen SR | Geometry-OOD SR | Dynamics-OOD SR |
|---|---|---|---|---|
| 1, specialist | 99.2% | 29.4% | 30.8% | 20.7% |
| 5 | 87.0% | 64.6% | 32.3% | 51.5% |
| 10 | 88.8% | 70.5% | 35.7% | 59.6% |
| 30, ArtManip | 93.9% | 85.0% | 55.5% | 74.9% |
One-object specialization wins on its own anchor, then collapses on unseen instances. Thirty-instance training gives up 5.3 points on the anchor and gains 55.6 points on unseen objects. It also more than triples dynamics-OOD success. The geometry-OOD result improves by 24.7 points but remains at 55.5%, leaving a visible gap between interpolation and shape extrapolation.
The OOD construction is controlled and narrow. Geometry-OOD uses box dimensions outside the Knife training intervals; dynamics-OOD uses joint damping outside the training range. These tests show parameter extrapolation inside the same primitive category representation. They do not establish transfer to new articulation topologies or object categories.
7. Real transfer succeeds after simulation-based grasp selection
The hardware study uses three previously unseen objects from each category. For every real object, the authors construct a primitive digital twin, generate candidate functional grasps, roll them out in simulation, and choose the top five. A person then places the real object by matching a rendered simulation view. Each grasp gets five 40-second trials, yielding $12\times5\times5=300$ executions. The open or close target is manually switched after the corresponding state is reached.
| Category | Objects | Selected grasps | IC ↑ | GC ↑ | CSC mean ↑ | CSC max ↑ | SR ↑ |
|---|---|---|---|---|---|---|---|
| Knife | 3 | 15 | 100% | 73.3% | 1.9 | 3.3 | 69.3% |
| Lighter | 3 | 15 | 100% | 86.7% | 2.1 | 3.3 | 80.0% |
| Stapler | 3 | 15 | 100% | 93.3% | 6.5 | 7.0 | 93.3% |
| Tong | 3 | 15 | 100% | 100% | 4.6 | 7.0 | 100% |
Across all categories, 257/300 trials succeed. All 12 objects have at least one successful selected grasp, and average GC is 88.3%. Staplers and tongs sustain the most cycles. Knives are the weakest category, consistent with their higher actuation resistance and the greater sensitivity of marginal fingertip contacts.
The paper’s failure analysis identifies seven selected grasps that fail in all five real trials. Their contacts concentrate near fingertip-pad boundaries, leaving little margin against slip; four belong to the Knife category. Primitive boxes also miss local curvature and thickness changes. Small geometric errors can shift contact and accumulate across cycles. Once the finger loses the articulated part, the policy has no visual or tactile channel with which to localize it again.
Zero-shot here means that the policy is not updated on the 12 test objects. Hardware preparation is still meaningful. A separate real calibration object is used to identify effective damping and stiffness ranges; test objects are not used for this tuning. Digital twins support grasp generation and simulation screening. Manual placement supplies the selected initial state. For lighters, gauze is added at a dorsal contact region to create a smoother, higher-friction surface. These choices define the demonstrated deployment protocol and should travel with the 85.7% number.
8. What the paper establishes—and what remains open
What the experiments establish: diverse instances and grasps materially improve within-category generalization; without contact-aware functional initialization, the policy learns almost no useful behavior; privileged latent distillation can support hardware execution with proprioceptive history; and coarse digital twins can select useful grasps even when they miss local shape detail.
Open questions:
- Functional grasp acquisition is outside the system. A user supplies the initial object pose by manually matching a rendered grasp.
- The object family is limited to two links and one internal DoF. Multi-joint mechanisms introduce coupled goals and more ways to lose contact.
- Each category uses its own geometric rules, contact template, physics ranges, goals, and training process. Cross-category control is not evaluated.
- No online vision or touch corrects model error. Lost contacts, shifted geometry, and inaccurate placement cannot be explicitly re-localized.
- The study has no direct quantitative comparison with another category-level articulated in-hand system. The strongest causal evidence comes from training-diversity, curriculum, optimizer, and grasp-generation ablations.
- Real trials last 40 seconds and start from simulation-screened grasps. Autonomous approach, grasp establishment, regrasping, and longer-horizon wear or drift remain outside the benchmark.
Takeaways for research and practice
What I would keep from ArtManip is the emphasis on the initial-state distribution. It deserves the same attention as the object collection. For articulated tasks, a mesh set without functional contacts leaves the policy searching through many stable but useless grasps. Category-level contact templates inject a small amount of task structure and turn procedural generation into an exploration tool.
ArtManip also offers a practical division of labor. A coarse model proposes object geometry and initial contacts; privileged simulation learns how hidden physics affects control; temporal distillation turns interaction history into an online dynamics cue. Each layer removes some deployment burden, while the remaining failures point directly to the missing layer: online perception for contact recovery.
The next step I would test is a visually or tactilely conditioned recovery policy while keeping the same initialization pipeline. Controlled perturbations could move the articulated link, break one fingertip contact, or offset the object from the rendered pose. Measuring reacquisition separately from ordinary open–close success would show whether category-level control survives outside the carefully selected basin of initial contacts.
