[Paper Notes] Transferring the Intelligence of VLMs to Robotic Control
Published:
This post supports English / 中文 switching via the site language toggle in the top navigation.
TL;DR
A VLM may understand that a cup needs to be grasped from the side, yet still need help turning that intention into executable motion. RoboDawn gives a frozen VLM a small command vocabulary for moving, rotating, and opening or closing a gripper. A robot execution layer plans each motion, reports what actually happened, and supplies new images. One task demonstration, translated into that same command language, helps the model use the interface and choose a strategy without updating its weights.
With GPT-6 Astra, the paper reports 53.2% zero-shot and 73.6% one-shot success on RoboTwin 2.0 C2R, and 35.67% and 47.17% on RoboDojo. The practical qualification is substantial: decisions take seconds, motion planning remains external, and real-world performance ranges from 9/10 on placing a block in a basket to 0/10 on cloth folding. I read this as evidence for transferring visual reasoning through a carefully designed control interface, with precision and execution speed still limiting its use.
Paper and sources
Transferring the Intelligence of VLMs to Robotic Control is a technical report by Meng-Hao Guo, Zhe-Han Mo, Jia-Jun Wang, Yi Zhang, Kejin Wang, Yi-Xuan Deng, Jia-Peng Zhang, Yongming Rao, and Shi-Min Hu, from Tsinghua University and Tencent Hunyuan. Hu is the corresponding author.
These notes follow the 16-page arXiv:2609.22966v1, submitted September 19, 2026. The project page and official code repository provide demonstrations and implementation resources. As checked on September 24, the code is public; older search snippets saying “coming soon” are outdated. Experimental values below come from the paper, and I have not independently reproduced them.
1. The interface defines what the VLM is being asked to control
RoboDawn defines the gripper interaction point (GIP) as the midpoint between the fingertips. Visual annotations, reported robot state, and motion commands all use this reference point. That consistency avoids a common ambiguity: moving the wrist to a location does not necessarily put the grasp center there, especially after a rotation.
The model issues commands with a compact grammar:
<arm> move <x|y|z> <distance_cm>
<arm> rotate <roll|pitch|yaw> <angle_deg>
<arm> point <down|forward|down45>
<arm> gripper <open|close|opening_0_to_1>
<arm> home
wait
done
This is an interface description, not a runnable robot program. The arm is left or right, and translation and rotation axes refer to the world frame. A move changes GIP position while preserving orientation; a rotation changes orientation while preserving GIP position. Translation and rotation magnitudes are clipped to 20 cm and 90 degrees per command. Orientation presets handle common poses, while done requests completion checking.
Each command becomes a planned motion to a target GIP pose and runs until the robot is stationary. The VLM can emit a short batch of commands per decision round. Joint-level control and trajectory generation remain in the execution layer. Consequently, “direct VLM control” here means choosing spatial operations within a robot controller’s interface; the model is not generating the servo stream.
2. Adaptation happens in observations, feedback, and memory
The paper writes the decision process as
\[(y_t,a_t)=\pi_\theta(L,E,D;I_t,x_t,F_{t-1},M_t).\]$L$ is the task instruction. $E$ describes the robot and environment, including workspace constraints, cameras, grid-based localization, and gripper properties. $D$ contains demonstrations. The changing inputs are annotated images $I_t$, measured robot state $x_t$, execution feedback $F_{t-1}$, and interaction memory $M_t$. The VLM produces commands $a_t$ and a structured response $y_t$ containing progress assessment, a plan, and a compact scratchpad.
The execution layer parses and plans the commands, runs them, and reports their physical outcome. The next observation contains the resulting scene and measured robot state. Memory combines previous commands, execution results, observations, and the model’s notes. Model weights, the environment profile, and demonstrations remain fixed during the episode. Online control does not receive privileged object poses.
flowchart TD
A["Fixed context: instruction, robot profile, demonstrations"] --> C["Frozen VLM"]
B["Annotated images, robot state, feedback, memory"] --> C
C --> D["Plan and semantic command batch"]
D --> E["Command parsing and motion planning"]
E --> F["Execute robot motion"]
F --> G["New images, measured state, execution outcome"]
G --> H["Update interaction memory"]
H --> B
This loop makes a useful distinction between requested motion and observed motion. A command can fail, only partially execute, or move an object unexpectedly. The next decision has access to that discrepancy. Calling the system training-free is accurate for task-specific weight updates, but its behavior still depends on substantial interface engineering and online state management.
3. “One shot” is a complete, interface-aligned task demonstration
The context has two components:
\[D=D_{\mathrm{prim}}\oplus D_{\mathrm{task}}, \qquad D_{\mathrm{task}}=\{D^{(m)}\}_{m=1}^{N_D}.\]$D_{\mathrm{prim}}$ is a shared primer showing the effects of basic commands. $D_{\mathrm{task}}$ contains complete task demonstrations. Zero-shot means $N_D=0$; the command primer remains. One-shot adds one demonstration of the target task. It does not mean one image, one action, or one task example shared across the whole benchmark.
Raw expert trajectories use continuous controls, so the authors first reduce them to end-effector waypoints and gripper states, then express each transition as commands available to the online model. Each demonstration round contains an image when retained, robot state, a command sequence, its measured effect, and a short rationale:
\[D^{(m)}=\left\{(I_j^{(m)},x_j^{(m)},r_j^{(m)},a_j^{(m)},f_j^{(m)})\right\}_{j=1}^{N_m}.\]The physical-effect record comes from differences in consecutive GIP poses and gripper openings. A VLM writes the rationales after reviewing the recorded episode with a task-agnostic prompt. Those explanations are synthetic annotations, not recorded expert thoughts or evidence that the same reasoning caused the original action.
Simulation demonstrations come from scripted experts in scenes separate from evaluation scenes. They use the same trajectory source as the robot-trained baselines. Long RoboDojo demonstrations preserve the full textual trajectory while sparsifying images, keeping informative stages such as grasping, rotation, and completion. The example therefore teaches both command effects and task ordering in the representation the model will later use.
4. The largest improvement comes from the first task example
RoboTwin 2.0 C2R evaluates 50 bimanual tasks in randomized scenes, with ten evaluation runs per task. Full-set baselines are jointly post-trained on 50 clean demonstrations per task, or 2,500 demonstrations total. RoboDawn uses clean demonstrations only in context. Selected results from the paper’s Table 1 are:
| Method | Benchmark adaptation | Success |
|---|---|---|
| $\pi_{0.5}$ | Full-set post-training | 46.0% |
| LingBot-VLA | Full-set post-training | 50.4% |
| HarnessVLA, Claude Code | Agent using a robot-trained VLA | 58.4% |
| RoboDawn, GPT-6 Astra | No task demonstration | 53.2% |
| RoboDawn, GPT-6 Astra | One task demonstration | 73.6% |
The one-shot gain is 20.4 percentage points over the same model’s zero-shot result. These comparisons establish a strong benchmark result under different adaptation schemes. They do not match pretraining data, model capacity, or inference budgets, so the table cannot isolate an inherent superiority of frozen VLMs over trained action policies.
With Gemini 3.8 Flash, success rises from 47.0% with no task example to 62.2% with one, 63.6% with two, and 65.4% with four; eight examples reduce it to 62.7%. The first example supplies most of the gain. The authors suggest long-context interference for the later decline, though the ablation alone does not establish its cause.
Interface support also matters. In the Gemini zero-shot ablation, removing reasoning lowers success from 47.0% to 34.8%; removing localization grids lowers it to 32.4%; removing the command primer yields 44.0%. The grid result is especially instructive: spatial reference information is a major part of the system’s effectiveness.
5. More interaction helps on RoboDojo, but takes time
RoboDojo covers 42 tasks spanning generalization, memory, long horizons, precision, and open capabilities. The paper reports averages over five runs and separates task progress from complete success. With GPT-6 Astra, RoboDawn moves from 35.67% success and 39.92 progress score zero-shot to 47.17% success and 54.63 score one-shot. The strongest full-set baseline listed, DM0.5, reaches 19.34% success.
Eight open tasks may lack corresponding training data for other methods. On the 34-task subset excluding them, RoboDawn reports 33.96% zero-shot and 43.33% one-shot success. This subset helps qualify the full-benchmark comparison; the paper does not provide a complete matched baseline table for that subset.
Increasing the per-episode command budget from 60 to 240 raises one-shot success from 31.2% to 47.2%, and zero-shot success from 23.7% to 35.7%. Extra interaction can support correction and recovery, but the intervention also gives the robot more physical actions. It therefore measures a joint increase in reasoning and interaction budget, not isolated scaling of internal reasoning compute.
The latency study uses Seed-2.1-Pro, a different backbone from the headline accuracy results. Averaged across RoboTwin tasks, RoboDawn needs 9.74 seconds of inference for a batch averaging 3.4 commands, followed by 2.09 seconds of motion, an inference-to-motion ratio of 4.65. The reported $\pi_{0.5}$ values are 101 ms inference and 2.70 seconds of motion. This leaves a large deployment gap for fast interaction; the paper does not report the same latency measurement for GPT-6 Astra.
6. Real robots expose precision, rotation, and completion errors
The real-world experiments use Gemini 3.8 Flash without task demonstrations or task-specific training:
| Task and robot | Success |
|---|---|
| Block in basket, Franka | 9/10 |
| Block stacking, Franka | 5/10 |
| Cloth folding, Piper | 0/10 |
These are ten-trial demonstrations of transfer for each task, with a pronounced drop as alignment and manipulation demands increase. The cloth results also change robot embodiment, so they cannot isolate rotation difficulty alone. The authors’ explanation that rotations may be less represented in web pretraining remains a hypothesis.
The RoboDojo failures identify three separate problems. The model can choose a sensible strategy yet miss the final positioning needed to insert a coin into a slot. A semantically plausible command can produce an IK motion that collides with surrounding objects. It can also stop pouring before the benchmark’s required amount has been transferred. Better task reasoning, reliable trajectory execution, and a measurable completion condition address different parts of this failure chain.
7. What I would reuse
The most reusable component is the agreement between the action language, demonstrations, visual reference point, and execution feedback. A demonstration becomes much more informative when its commands are exactly those the agent can issue and its effects use the same coordinates as the live robot. The shared primer also makes the meaning of “zero-shot” concrete enough to reproduce.
For a system with long pauses between coarse manipulations, I would test this interface before collecting a large task-specific training set. For insertion, rapid contact corrections, or continuous motion, I would pair the VLM’s task decisions with a faster local controller and evaluate the transition between them. That is a proposed extension, not a result demonstrated here. The next comparison I would want fixes wall-clock time and physical action budget: it would show how much of the success advantage survives when a robot has to finish on a schedule.
