[Paper Notes] RMBench: Memory-Dependent Robotic Manipulation Benchmark with Insights into Policy Design

14 minute read

Published:

This post supports English / 中文 switching via the site language toggle in the top navigation.

TL;DR

Many manipulation policies assume that the current image almost determines the next action. RMBench tests the cases where that assumption fails: a robot must remember a hidden reference object, an earlier block location, how many attempts have failed, or which subtask has already finished. The paper introduces Task Memory Complexity (TMC), a task-level measure of how many task-relevant past observations an optimal policy must retain, and builds a nine-task benchmark on RoboTwin 2.0.

The authors also propose Mem-0, a modular policy with a planning module, an execution module, and a subtask-end classifier. It uses completed-subtask key memory for high-level planning, a persistent anchor for within-subtask reference, and a sliding window for recent motion. With 50 synthesized demonstrations per task and 100 evaluation rollouts, Mem-0 reaches 42.0% average success, compared with 10.4% for Pi0.5 and 9.8% for X-VLA. The gains are concentrated on memory-dependent tasks: 52.8% on M(1) and 28.5% on M(n). The system still struggles with semantic matching, precise orientation, and small button presses.

Paper and source version

Tianxing Chen, Yuran Wang, Mingleyang Li, Yan Qin, Hao Shi, Zixuan Li, Yifan Hu, Yingsheng Zhang, Kaixuan Wang, Yue Chen, Hongcheng Wang, Tianhang Yang, Junjie Wang, Renjing Xu, Ruihai Wu, Yao Mu, Yaodong Yang, Hao Dong, and Ping Luo, with affiliations including MMLab@HKU, Peking University, PsiBot, HKUST (GZ), Tsinghua University, Shenzhen University, and Shanghai Jiao Tong University. These notes follow arXiv:2603.01229v3, revised July 30, 2026. See the paper PDF, project page, and code repository. The source is an arXiv preprint; no accepted venue is assumed. Results below are author-reported and have not been independently reproduced here.

1. Why ordinary long-horizon benchmarks miss memory

A long task is not automatically a memory task. In LIBERO-Long, for example, task-relevant information can remain visible throughout execution, so a policy may succeed by reacting to the current frame. RMBench hides or changes information so that the robot has to carry it forward.

The benchmark is built on RoboTwin 2.0 and SAPIEN, with automated data synthesis, integrated policy evaluation, and fine-grained language annotations for action–observation pairs. The intended comparison is between policies, while the task design controls how much history is genuinely needed.

The distinction is a partially observable one. Let $s_t$ be the latent state, $o_t$ the current observation, $a_t$ the action, and

\[h_t=(o_{1:t},a_{1:t-1})\]

be the complete interaction history. A policy does not need to keep every element of $h_t$ if it can construct a smaller memory state that preserves the information relevant to the next decision.

2. Task Memory Complexity turns “memory” into a task property

RMBench defines Task Memory Complexity as the smallest number $m$ of task-relevant past observations needed by some optimal policy. If $\mathcal M_t^{(m)}$ summarizes at most $m$ such observations, then the task has complexity $m$ when

\[\exists\,\pi^*\ \text{such that}\quad \pi^*(a_t\mid h_t)=\pi^*(a_t\mid \mathcal M_t^{(m)}), \qquad \forall t.\]

The notation is deliberately task-centric:

  • M(0): the current observation is sufficient;
  • M(1): one task-relevant earlier observation must be retained;
  • M(n): several non-local observations, repeated attempts, or completed subtasks matter.

This is a useful separation from architectural memory size. A policy can have a large context window and still fail an M(1) task if it cannot identify which old frame matters. Conversely, a compact memory can solve a task if it retains the right event.

flowchart LR
    A["Current observation is ambiguous"] --> B["Identify task-relevant past event"]
    B --> C["Encode a compact memory state"]
    C --> D["Select next subtask or action"]
    D --> E["New observation updates memory"]
    E --> B

3. What the nine tasks require

RMBench contains five M(1) and four M(n) tasks. Each is designed around a concrete source of partial observability.

TMCTasksMemory demand
M(1)Observe and Pick UpObserve a reference object, hide it, then pick the matching object.
M(1)Rearrange BlocksMove one block, press a button, then use the earlier arrangement to place another block.
M(1)Put Back BlockMove a block to the center, press a button, and return it to its original pad.
M(1)Swap BlocksUse an empty pad to exchange two blocks, then press the button.
M(1)Swap TSwap two T-shaped blocks while preserving their target positions and orientations.
M(n)Battery TryRepeatedly try insertion orders and orientations until both batteries fit.
M(n)Blocks Ranking TryTry block arrangements and press to confirm until the color order is correct.
M(n)Cover BlocksTrack which blocks have been covered and finish the required sequence.
M(n)Press ButtonAccumulate repeated presses with a task-specific count.

M(1) tasks can often be solved by retaining one reference frame or state cue. M(n) tasks require accumulating evidence across attempts or subtasks. The latter category exposes whether a policy can use memory as a counter, a record of completed work, or a history of failures.

4. Mem-0 separates planning memory from execution memory

Mem-0 is designed as an analysis-friendly policy. Its modules can be removed or replaced without changing the benchmark, making it possible to ask which memory mechanism caused a gain.

flowchart TD
    A["Initial image + task instruction + completed-subtask memory"] --> B["Planning VLM"]
    B --> C["Current subtask"]
    D["Current image + subtask"] --> E["Execution VLM"]
    F["Anchor memory"] --> E
    G["Sliding memory window"] --> E
    E --> H["Diffusion Transformer"]
    H --> I["Action chunk"]
    I --> J["Subtask-end MLP"]
    J -->|"8 consecutive end signals"| B
    J -->|"ongoing"| E
    E --> G

Key memory for completed subtasks

At a planning step, the VLM receives the initial observation $o_0$, the global goal $g$, and a memory of completed subtasks:

\[s_t=\mathcal V_{\mathrm{plan}}(o_0,g,\mathcal M_{t-1}), \qquad \mathcal M_{t-1}=\{(s_i,o_i^{\mathrm{end}})\}_{i=1}^{t-1}.\]

Each entry stores the textual description of a finished subtask and the RGB frame at its termination. This gives the planner a structured record of what has happened. It is particularly important for M(n) tasks, where the next action depends on several completed attempts.

Planning happens when a subtask ends, rather than on every frame. If an episode has $N$ subtasks and $N\ll T$ control steps, planning calls fall from $O(T)$ to $O(N)$.

Anchor and sliding memories for execution

The execution module encodes the current image and subtask into latent tokens. The image latent attends to two buffers:

\[\tilde{\mathbf z}_t^l =\operatorname{CrossAttn}(\mathbf z_t^{\mathrm{img}},\mathcal M_t^l) +\mathbf z_t^{\mathrm{img}}, \qquad l\in\{\mathrm{anchor},\mathrm{slide}\}.\]

The conditioning vector concatenates anchor-aware image features, sliding-window features, and text features. At the beginning of a subtask, the first image latent is stored as the anchor and remains fixed. The sliding memory appends the latest image latent and truncates to the most recent $K$ elements:

\[\mathcal S_{t+1}=\operatorname{Trunc}_K \left(\mathcal S_t\cup\{\mathbf z_t^{\mathrm{img}}\}\right).\]

The anchor preserves a stable reference while the sliding buffer captures short-term motion and contact changes. Both buffers reset when the subtask ends. Mem-0 then uses a diffusion transformer with an action horizon of 30 and executes a prefix of each predicted action sequence.

Subtask-end classifier

A small MLP predicts whether the current subtask has finished. To avoid switching plans because of one noisy frame, Mem-0 requires an end prediction for eight consecutive timesteps:

\[\sum_{i=t-7}^{t}\mathcal C_{\mathrm{end}}(\mathbf c_i)=8.\]

This classifier is a control component as much as a memory component. An early transition loses the current subtask; a late transition wastes actions and delays access to the next key memory.

5. Benchmark results

The paper trains DP, ACT, Pi0.5, X-VLA, and Mem-0 with 50 synthesized demonstrations per task, then evaluates each on 100 rollouts. Baselines do not use subtask decomposition. Mem-0 uses its execution module alone on M(1), and uses planning plus execution on M(n).

Task groupDPACTPi0.5X-VLAMem-0
M(1) average6.4%6.8%14.4%11.8%52.8%
M(n) average5.0%4.8%5.5%7.3%28.5%
Overall average5.8%5.9%10.4%9.8%42.0%

Mem-0 is especially strong on Rearrange Blocks (89%), Put Back Block (90%), Swap Blocks (67%), and Cover Blocks (68%). It reaches 28% on Battery Try and 18% on Blocks Ranking Try. Press Button remains at 0%, where the small button motion makes completion detection unreliable. On Observe and Pick Up, Mem-0 reaches 4%; pretrained policies retain an advantage because the task also demands semantic matching. Swap T reaches 14%, reflecting the difficulty of precise orientation and placement.

The relative gains are 38.4 percentage points on M(1) and 21.2 points on M(n) against the best baseline averages. These are benchmark success rates, not per-frame accuracy, and every task uses 100 rollouts. The aggregate should therefore be read together with the task-level failures.

6. What the memory ablations show

The ablations remove one memory component at a time or replace the learned end classifier with simulator ground truth.

VariantM(1) averageM(n) average
Full Mem-052.8%28.5%
Without anchor memory26.8%26.8%
Without sliding memory40.4%25.3%
Without key memory—4.8%
Ground-truth end classifier—45.3%

Anchor memory has a large effect on M(1): removing it allows the sliding window to evict the task-critical reference. Sliding memory supplies recent motion context; removing it often produces unstable or oscillatory behavior even when an anchor remains. The exception is Swap T, where removing sliding memory improves the result, plausibly because transient motion cues interfere with the initial orientation reference.

For M(n), removing key memory collapses the average from 28.5% to 4.8%. A planner that sees only the current frame cannot reliably infer the next subtask after several attempts. Replacing the learned end classifier with ground truth raises the average to 45.3%, which isolates a second bottleneck: planning and memory can be useful, yet poor transition timing prevents them from interacting correctly.

7. Real-world transfer and limitations

The authors also evaluate three tasks on a physical robot. Mem-0 reaches 17.5% on Put Back Block, 37.5% on Rearrange Blocks, and 12.5% on Cover Blocks, for a 22.5% average, compared with 5.83% for Pi0.5 and 0% for ACT. This transfer is encouraging, but the real-world set is smaller than the simulation benchmark and covers only three tasks.

The paper identifies several limits. The benchmark is simulation-first, so visual and contact gaps remain when moving to hardware. Mem-0’s current planner depends on structured subtask descriptions, and its transition classifier is too simple for subtle events such as a small button press. The policy also lacks the semantic strength of large pretrained models on reference matching, and anchor/sliding memory can interfere when the task is highly sensitive to initial orientation.

My main takeaway is that memory should be evaluated at the task level before it is optimized at the architecture level. RMBench’s M(1)/M(n) distinction makes a useful diagnostic: a policy may fail because it cannot preserve one crucial reference, because it cannot accumulate repeated attempts, or because it cannot decide when a subtask ended. Those failures call for different fixes. Adding a longer frame window would not address all three.