[Paper Notes] ICI-VLA: In-Context Imitation with Spatiotemporally Aligned Demonstrations for Vision-Language-Action Models
Published:
This post supports English / 中文 switching via the site language toggle in the top navigation.
TL;DR
Most VLA adaptation still means collecting task-specific demonstrations and updating model weights. ICI-VLA turns adaptation into retrieval-conditioned inference: the policy stays fixed at test time and receives a few short demonstrations that match the current subtask, phase, and motion geometry.
The framework has three pieces. It decomposes long trajectories into semantically labeled micro-demonstrations, trains a retriever with semantic hard filtering plus Dynamic Time Warping (DTW) supervision, and uses Target Action Masking during policy training to reduce direct copying of contextual action prefixes. With a native text-generation interface, Qwen3-VL-4B predicts continuous actions as numerical text without an extra action head.
ICI-VLA reaches 97.7% average success on LIBERO, 60.4% on RoboTwin 2.0, and 83.2% over four physical dual-arm Aloha tasks. These comparisons follow the paper’s reported protocols; external baseline numbers are benchmark-level references rather than fully paired experiments.
Paper and source version
ICI-VLA: In-Context Imitation with Spatiotemporally Aligned Demonstrations for Vision-Language-Action Models is by Songhua Yang, Ziyu Liu, Xuetao Li, Ruqi Xiao, Kangxin Zhu, and Miao Li from Wuhan University and the Institute of Technological Sciences, Wuhan University. These notes follow arXiv:2609.07581v1, submitted September 7, 2026. See the paper PDF. Results below are author-reported.
1. Adaptation through context instead of test-time updates
A conventional VLA collects demonstrations for a new task and fine-tunes its parameters. ICI-VLA keeps its parameters frozen during deployment and conditions action generation on the current observation plus retrieved reference examples. The setting is few-shot test-time in-context adaptation: the model adapts through context, while all learning happens offline.
The paper preserves a native text-action interface inspired by VLA-0. Continuous actions and proprioceptive states are serialized as numerical text, so Qwen3-VL-4B can generate action chunks without adding a diffusion head, action tokenizer, or specialized control module. This keeps the language model’s original input-output interface available to the retrieval mechanism.
At time $t$, the query is
\[I_q=\langle T,T_{sub},O_t,Z_t\rangle,\]where $T$ is the global instruction, $T_{sub}$ is the active subtask, $O_t$ contains main and wrist camera observations, and $Z_t$ is proprioception. A retrieved micro-demonstration is
\[E=\langle T^e,T^e_{sub},O^e,Z^e,A^e_{e:e+k}\rangle.\]The retriever selects
\[E^*=\arg\max_{E\in\mathcal D}s_\phi(I_q,E),\]and the fixed policy generates the next action chunk conditioned on $I_q$ and $E^*$.
2. Build a library of short, phase-labeled examples
Long demonstrations are a poor retrieval unit for short-horizon control. A full episode can contain reaching, grasping, placing, and returning, while the current policy step may need only one local motion. ICI-VLA therefore uses a Qwen3-VL model to segment approximately 11,200 long-horizon trajectories into subtasks and attach fine-grained instructions to them.
Each micro-demonstration stores the global instruction, subtask instruction, camera observations, initial proprioception, and a $k$-step textualized action chunk. The resulting library contains approximately 139,659 subtask examples. The library combines LIBERO, RoboTwin 2.0, and physical dual-arm Aloha data, including about 1,000 real teleoperation demonstrations.
This decomposition gives retrieval a useful phase vocabulary. A query asking the robot to close a drawer can retrieve a drawer-closing segment even when the full source episode also contains unrelated bowl placement or navigation motions.
3. Semantic filtering plus DTW geometric alignment
Visual or language similarity alone can retrieve an example with the right object and the wrong motion. ICI-VLA trains an RD-Encoder based on Qwen3-VL-Embedding-2B through an iterative two-stage alignment process.
First, semantic hard filtering uses the current embedding to retain a compatible candidate pool. A query such as “the robotic arm pushes the drawer closed” filters out candidates about unrelated actions. Second, Dynamic Time Warping ranks the remaining candidates using labeled end-effector trajectories. The lowest-cost candidate becomes the positive; phase-misaligned examples from the same semantic subtask become hard negatives.
For anchor trajectory $Traj_a$ and candidate $Traj_e$, the mining cost is
\[DTW(Traj_a,Traj_e)=\min_{W}\sum_{(u,v)\in W}\delta(p^a_u,p^e_v),\]where $W$ is a valid warping path and $\delta$ measures waypoint discrepancy across active arms. The RD-Encoder is trained with an InfoNCE-style objective:
\[\mathcal L_{CL}=-\log \frac{\exp(e_A^\top e_{P^+}/\tau)} {\exp(e_A^\top e_{P^+}/\tau)+\sum_i\exp(e_A^\top e_{N_i^-}/\tau)}.\]The key separation is between training and deployment. DTW uses labeled target trajectories only to create offline supervision. At inference, the target action is unknown; the frozen retriever ranks candidates from the observable query alone.
The full five-cycle RD-Encoder raises RoboTwin retrieval Recall@1 from 27.8% for the base embedding to 70.8%, and Recall@5 from 52.4% to 90.1%. Semantic filtering and one DTW cycle provide intermediate gains.
4. Target Action Masking prevents prefix copying
Even a well-matched demonstration can become a misleading action prefix. If the model sees an exact reference action sequence, it may continue the sequence mechanically instead of grounding its next action in the current observation.
During offline policy fine-tuning, ICI-VLA randomly masks a subset $M$ of the contextual target-action tokens. It optimizes only the unmasked positions $U$:
\[\mathcal L_{act}=-\sum_{j\in U} \log\pi_\theta(a_{q,j}\mid I_q,E^*,\tilde a_{q,<j}),\]where $\tilde a_q$ contains corrupted earlier target tokens. Masking is disabled at inference, when the fixed policy generates actions autoregressively.
This objective does not explicitly teach a kinematic residual relative to the demonstration. Its intended effect is behavioral: exact action continuation becomes unreliable during training, so the policy must use the current observation, the task instruction, and the retrieved context together.
5. Simulation results
ICI-VLA evaluates on LIBERO and RoboTwin 2.0 with three retrieved examples at inference. The number of examples, five retriever-mining cycles, and a confidence threshold of 0.65 are selected on validation data. The context is refreshed when policy confidence falls below that threshold.
| Benchmark / split | ICI-VLA success | Reference comparison |
|---|---|---|
| LIBERO Spatial | 98.5% | VLA-0: 98.2% |
| LIBERO Object | 98.7% | OpenVLA-OFT: 99.5% |
| LIBERO Goal | 98.0% | VLA-0: 97.5% |
| LIBERO Long | 96.8% | OpenVLA-OFT: 93.2% |
| LIBERO average | 97.7% | Highest listed baseline: 96.4% |
| RoboTwin 2.0 Easy | 72.4% | Highest listed baseline: 55.2% |
| RoboTwin 2.0 Hard | 46.3% | Highest listed baseline: 24.5% |
| RoboTwin 2.0 average | 60.4% | Highest listed baseline: 41.1% |
The strongest RoboTwin result comes from long-horizon dual-arm coordination. The controlled comparison is especially informative: VLA-0 with the same retrieved context but without Target Action Masking reaches only 10.7%, while ICI-VLA reaches 60.4%.
The ablations separate the contributions:
| Configuration | LIBERO average | RoboTwin average |
|---|---|---|
| Naive ICL, no masking | 71.5% | 10.7% |
| Without DTW ranking | 92.5% | 31.4% |
| Without semantic filtering | 88.4% | 38.1% |
| Full ICI-VLA | 97.7% | 60.4% |
Removing DTW or semantic filtering hurts retrieval alignment. Removing masking causes the largest collapse, which is consistent with a policy that copies a reference prefix without robustly checking the current state.
6. Physical dual-arm evaluation
The physical setup is a dual-arm Aloha system with four tasks: Single-arm Grasp, Dual-arm Grasp, Drawer Placement, and Object Sorting. Evaluation objects, layouts, instructions, and trajectories are disjoint from policy training and the retrieval library. Each task uses 250 rollouts, for 1,000 trials total.
| Task | $\pi_0$ | VLA-0 | ICI-VLA |
|---|---|---|---|
| Single-arm Grasp | 78.4% | 75.2% | 89.6% |
| Dual-arm Grasp | 56.8% | 52.4% | 76.8% |
| Drawer Placement | 62.0% | 58.8% | 81.2% |
| Object Sorting | 68.4% | 65.6% | 85.2% |
| Average | 66.4% | 63.0% | 83.2% |
The reported 95% Wilson interval for ICI-VLA’s average is 80.8%–85.4%. The result suggests that short, phase-relevant references can help under lighting variation, distractors, sensor noise, and contact dynamics. It remains a four-task deployment, so it does not establish broad cross-embodiment adaptation.
7. Sensitivity and limitations
Three retrieved examples work best on the RoboTwin validation split: one or two provide less coverage, while larger contexts introduce irrelevant or conflicting information. Retriever mining improves performance from 34.6% with no optimization to 60.4% after five cycles; additional cycles change the result by at most 0.3 points in the reported sweep.
The approach depends on demonstration coverage and the quality of the subtask planner. It also pays offline costs for full-parameter fine-tuning of Qwen3-VL-4B and the embedding model, plus iterative retrieval mining. Inference avoids gradient updates, but it still requires a library search and confidence-based context refresh. The method’s success does not prove that the VLA has learned an explicit kinematic residual; Target Action Masking is a context-corruption objective whose causal mechanism remains partly open.
My main takeaway is that in-context imitation for robotics needs alignment at the same granularity as action generation. A whole episode is too coarse, and a visually similar frame is too weak. ICI-VLA combines semantic phase labels, trajectory geometry, and action-prefix corruption so the demonstration becomes a local reference instead of a script to replay. The next useful test would be a larger cross-embodiment library where retrieval must match hand morphology and control conventions as well as task phase.
