[Paper Notes] In-Context Robot Learning with VLM Agents

14 minute read

Published:

This post supports English / 中文 switching via the site language toggle in the top navigation.

TL;DR

A human video can show how to lift a thin notebook from a table without supplying a single robot action. GPT-Policy asks how far a fixed, general-purpose vision-language model can carry that information into physical execution. It packages demonstrations and live observations into context, lets the VLM request robot-tool actions, and returns execution feedback for the next decision. The model’s weights remain fixed throughout the trial.

The clearest evidence comes from small real-robot comparisons. Human video changes towel and notebook pickup from 0/3 to 2/3 successful trials. For plug removal and reinsertion, robot video alone remains at 0/3, while video with aligned state and action records reaches 2/3. Context helps the agent choose useful behavior, but precise contact, completion verification, and collision handling remain unresolved. I read this as evidence for a useful interface between general agents and robot controllers, with substantial work still needed to make it dependable.

Paper and source version

In-Context Robot Learning with VLM Agents is by Dongzhou Cheng, Taoran Yi, Ye Fang, Xingwu Zhang, Fan Feng, Yixuan Li, Gengxiong Zhuang, Rongze Wang, Shuai Yang, Wei Song, Weizhi Xue, Minyan Wu, Jie Gui, Jiaqi Wang, and Tong Wu. The first seven authors share first authorship; Tong Wu is the corresponding author. Affiliations include Morphi Robot, Shanghai Innovation Institute, Fudan University, and several other universities.

These notes cover arXiv:2609.19138v1, submitted September 16, 2026, including the PDF appendix. The authors provide a project page and public implementation. Results are taken from the paper; I have not reproduced the robot experiments. Repository availability below reflects its README checked on September 20, 2026.

1. Adaptation lives in the context

The paper defines robotic in-context learning as changing behavior using demonstrations, examples, or interaction experience supplied at deployment, without gradient updates or persistent task-specific parameter changes. This definition covers several kinds of information: a goal image specifies the desired arrangement; a video suggests an interaction procedure; robot state and action records constrain the motion; recent observations and feedback help track what happened.

At decision step $t$, let $T$ be the instruction, $o_t=(I_t,s_t)$ the camera observations and robot state, $c_t$ the available context, and $f_{t-1}$ the preceding tool result. The loop is

\[a_t\sim\pi_\theta(\cdot\mid T,c_t,o_t,f_{t-1}), \qquad (o_{t+1},f_t)=\mathcal E(a_t,o_t).\]

Here $a_t=(u_t,v_t)$ contains a tool name and its arguments. The parameters $\theta$ remain fixed. The execution interface $\mathcal E$ turns a request into robot commands or rejects it, then supplies feedback. A rejected move can therefore inform the next decision without changing the policy weights. There is no preceding result at the first step. Method, §3.1

flowchart TD
    A["Demonstration video or goal image"] --> B["Context compiler: keyframes and aligned records"]
    C["Task instruction + tool definitions"] --> D["Fixed VLM"]
    B --> D
    E["Live camera images + robot state"] --> D
    F["Interaction history + previous tool result"] --> D
    D --> G["Tool name + motion or gripper arguments"]
    G --> H["Robot adapter: path, IK and timing checks"]
    H --> I["Execute accepted request or return rejection"]
    I --> E
    I --> F

That separation also locates the engineering burden. The VLM selects targets and interprets outcomes. The adapter handles coordinate conventions, kinematics, timing, and measured execution status. A model that understands the demonstration can still supply a poor physical target.

2. What survives video compression

The context compiler preserves transitions that affect the procedure: approach, contact, grasp, release, and changes in which arm supports an object. Appendix C describes a vision model selecting candidate moments from overlapping video windows, followed by a global review that removes redundant holds. The reference is capped at 24 keyframes and 48 images, resized within 1,280 pixels per dimension without upscaling. Images are interleaved with timestamps, camera labels, and available stage annotations.

For robot demonstrations, one selected time can contain several camera views. The bottle-opening reference has 13 keyframes and 13 images; plug reinsertion has 14 keyframes and 42 images. The added numerical context is much denser: 205 retained action samples for the bottle and 131 for the plug. This matters because a few images can leave the intervening rotation, support posture, or gripper transition ambiguous. The authors propose that the additional records reduce that ambiguity. Their selected trajectory comparisons support this explanation, but do not isolate it from every other added cue. §4.3 and Appendix C

The alignment rule is concrete. A keyframe uses the nearest measured state within 0.1 s, and extra camera views are matched to the overhead frame with the same tolerance. The image at $t_i$ is paired with its measured state and the command segment leading to $t_{i+1}$. Sampling keeps segment endpoints, approximately one action sample per second, and both sides of gripper-command changes. Missing measurements remain missing. Timestamp matching allows residual sensor misalignment.

These records stay inside the model input; the agent generates new requests from the current scene. Directly replaying the demonstration trajectory would bypass the adaptation being studied.

There is an experimental qualification here: Video + Action adds measured robot states as well as action records. Images are shared between video conditions, but a stricter causal test would also hold annotations fixed and separate video + state from video + state + action. That additional comparison is my proposed follow-up. The reported ablation measures the benefit of the supplied numerical reference package.

Online history has a separate lifecycle. Provider adapters retain reference inputs while limiting older live images; accumulated text may be kept or replaced with host-generated summaries. Memory management can consequently change which evidence reaches later decisions. Appendix B also describes provider-specific execution and completion-review differences, which need attention in model comparisons.

3. From a tool request to a timed trajectory

For the Cartesian interface described in §3.3 and Appendix A, move_to requests one tool-center-point pose and move_eef_chunk requests a sequence. A pose contains position in metres and an xyzw quaternion, expressed in the selected arm’s base frame. In a bimanual sequence, a null arm entry preserves that arm’s preceding pose. Gripper opening changes through a separate set_gripper call and remains fixed during Cartesian motion.

For consecutive targets, the adapter linearly interpolates position and uses quaternion SLERP for orientation. It samples this path and solves IK sequentially:

\[p(s)=(1-s)p_j+sp_{j+1}, \qquad q_k=\operatorname{IK}(\widehat p_k,\widehat R_k;q_{k-1}).\]

The initial seed comes from measured joints. ARX and YAM use execution residual tolerances of 2 mm and approximately 1°; their solver stopping tolerances are tighter. Sequential seeding encourages neighboring solutions, but these two adapters impose no separate hard bound on the joint displacement between samples. Backend acceptance rules differ, so the detailed checks must be read with the robot configuration. Morphi Kino has its own tolerances and continuity checks.

Ruckig supplies a scalar timing profile. For the ARX/YAM path described in Appendix A, sampled joint velocity, acceleration, and jerk are checked against limits, and time is stretched when needed. With $r_v,r_a,r_j$ denoting the largest derivative-to-limit ratios,

\[\alpha_0=\max(1,r_v,\sqrt{r_a},\sqrt[3]{r_j}), \qquad \tau'_k=\alpha\tau_k,\]

where $\alpha=1$ if no stretch is needed and otherwise $\alpha=1.001\alpha_0$. Stretching time scales the computed derivatives by $\alpha^{-1}$, $\alpha^{-2}$, and $\alpha^{-3}$. Both arms are planned before submission, and corresponding segments are synchronized to the longer duration. YAM streams interpolated joint references at 100 Hz. That rate describes low-level playback; VLM decisions occur around tool execution. Appendix A

The limits have a precise scope. These checks constrain the sampled reference, while measured endpoint error and settling are returned separately. The Cartesian planner does not check collisions. The paper reports repeated inter-arm collisions and calls for an independent safety layer. Likewise, reaching a commanded pose or receiving a model completion message does not establish that the plug is seated or that an object remains grasped.

4. What the controlled comparisons show

Table 1 evaluates GPT-6 Astra with three trials per condition. Success requires the final scene to satisfy task-specific geometric and semantic criteria. Decisions and elapsed time are averaged over all trials, including failures. The four tasks with explicit no-demonstration comparisons are:

TaskContextSuccessMean decisionsMean time, min
Pick red towelNone0/396.324.6
Pick red towelHuman video2/376.718.9
Pick up notebookNone0/394.024.6
Pick up notebookHuman video2/366.716.1
Unscrew bottle capNone0/371.016.1
Unscrew bottle capRobot video2/374.315.2
Unscrew bottle capRobot video + action3/354.717.9
Remove and reinsert plugNone0/324.05.3
Remove and reinsert plugRobot video0/333.77.9
Remove and reinsert plugRobot video + action2/348.310.8

Source: paper Table 1. “Video + action” retains the paper’s condition name and includes the measured states discussed above.

The human demonstrations supply no robot action labels, so the towel and notebook results are consistent with transferring an interaction strategy across embodiments. The plug task exposes a harder boundary: watching the procedure alone does not produce a successful insertion in these trials. Aligned numerical references help, although one of three attempts still fails.

Efficiency needs a separate reading. Bottle opening uses fewer decisions with action references but takes longer than video alone. For the plug, both decision count and elapsed time rise as success improves. Because failed episodes enter these averages, early failure can look cheap. A useful extension would report success-conditioned completion time alongside all-trial cost and failure categories.

The remaining six tasks each achieve 3/3 under their selected context: T-shape and fruit arrangement with target images; lemon placement and mobile exploration with self-history; tic-tac-toe and pointed-fruit pickup with human interaction. For tic-tac-toe, wins and draws both count as success. Table 1 supplies no matched no-context result for these six tasks, so their outcomes demonstrate capability under the tested conditions without quantifying each context’s causal contribution.

Table 2 is narrower still: it compares individual towel-pickup runs using task progress, elapsed time, and estimated token usage. Its 100% progress entry is distinct from the 2/3 success rate across the human-video trials. The authors explicitly caution that these examples do not establish a reliable model ranking. Discussion, §5

5. What the public release lets us inspect

The repository README identifies src/gpt_policy/ as the home of input preparation, protocols, planning, recording, and adapters. The published hardware integrations cover ARX X5 and I2RT/YAM; the paper additionally describes Morphi Kino. A profile check is available without opening hardware or starting a model session.

The release excludes site-specific calibration, physical demonstration records, run recordings, and the complete evaluation environment. Its simulation pipeline is listed as future work. Reproducing the framework therefore involves more than installing the package: the input references, camera/TCP calibration, provider behavior, and trial protocol all affect the comparison. This post reviews the paper and README, without claiming an audit or execution of the implementation.

6. The experiment I would run next

My strongest takeaway is that the representation of a demonstration can determine whether general reasoning becomes a useful motion choice. Preserving contact transitions and intermediate commands gives the model evidence it cannot reliably infer from sparse images. The next experiment should separate that benefit into parts: identical keyframes and annotations, then add measured state, then add commands, while varying their temporal density.

For a contact-sensitive task, I would also compare direct VLM pose requests with a fast local controller that refines contact and returns explicit grasp or insertion evidence. That is a research proposal, not a result of this paper. Unknown pretraining exposure also limits claims that the observed behavior constitutes acquisition of a wholly novel skill. I would change my assessment of deployment readiness if broader trials showed reliable recovery, independently verified completion, and low collision/intervention rates across new layouts. Three trials per condition leave those questions open.