[Paper Notes] STAR: Sparse Tactile Representation Learning in Vision–Tactile–Language–Action Models for Dexterous Manipulation

21 minute read

Published:

This post supports English / 中文 switching via the site language toggle in the top navigation.

TL;DR

Most tactile readings from a dexterous hand say very little. At any instant, contact occupies only a few sensing regions; during reaching, the entire tactile stream may be quiet; and when contact does occur, it describes a small patch of the world. STAR builds its training recipe around this sparsity instead of feeding every tactile token into a vision-language-action model and hoping attention will sort it out.

The paper contributes a 200-hour real-robot dataset with 10,576 trajectories across 65 tasks, collected on a bimanual mobile robot with tactile dexterous hands. It then adapts $\pi_{0.5}$ into a vision-tactile-language-action policy through three components: joint masked visual-tactile pretraining, a sparse-global tactile token representation, and prediction of five future tactile states spread over a 1.67-second action horizon.

After 100 task-specific post-training trajectories per evaluation task, STAR reaches 61% mean success on earbud flipping, stacked-book retrieval, postcard retrieval, and multi-object grasping. The decomposition matters: $\pi_{0.5}$ starts at 28%; finetuning it on the 200-hour dexterous dataset without STAR reaches 44%; the full recipe reaches 61%. Tactile feedback produces its clearest gain on earbud flipping, where success rises from 35% without tactile to 65% with it. On stacked-book retrieval, both variants score 55%.

I read STAR as a strong data-and-representation paper. It shows how to spend model capacity on sparse contact events and how to supervise contact dynamics beyond the next few frames. The reported 61% is evidence for task-specific adaptation, with 20 real trials per task. It does not yet establish zero-shot task generalization, transfer across hand hardware, or a broadly released tactile foundation model.

Paper and source version

STAR: Sparse Tactile Representation Learning in Vision–Tactile–Language–Action Models for Dexterous Manipulation is by Xiangcheng Liu, Tianhao Wu, Le Zheng, Yidong Wang, Bowen Jiang, Mingjie Pan, Xinlin Ren, Yi Liu, and Jianlan Luo. Liu, Wu, and Zheng contributed equally; Luo is the corresponding author. The affiliations listed are Shanghai Innovation Institute and Agibot.

These notes follow the 10-page arXiv:2609.12549v1, submitted September 11, 2026. The project page provides the main demonstration video. As of September 17, 2026, I did not find public download links there for the code, model weights, or the 200-hour dataset. I read the paper and appendix; the experiments have not been independently reproduced here.

1. Tactile input is sparse in three different ways

The policy models an action chunk conditioned on images, tactile readings, proprioception, and a language instruction:

\[p(a_t\mid o_t), \qquad o_t=\{I_t,T_t,s_t,l_t\}, \qquad a_t=\{a_t^1,\ldots,a_t^H\}.\]

Adding $T_t$ creates a representation problem. In the collected dataset, the maximum fraction of active taxels in any frame is only 24%, and frames with active tactile readings make up only 35% of the sequence. The paper separates the difficulty into three cases:

  • Spatial sparsity: most sensing locations are inactive in a given frame.
  • Temporal sparsity: long stretches, especially approach motion, contain no contact.
  • Informational sparsity: force at a local patch says little about the global scene or task state.

This taxonomy is useful because each failure calls for a different intervention. Spatial sparsity asks for token selection. Temporal sparsity asks where predictive supervision should be placed. Informational sparsity asks how local contact should be aligned with vision and summarized for the policy.

The paper’s own numbers also explain why naive tactile pretraining can collapse. The tactile signal is rendered as a $224\times224$ image, with zeros in regions that contain no taxel. A masked autoencoder trained only on that image can minimize much of its loss by learning the inactive background. More data would help eventually, but 200 hours is still small next to the corpora used by modern vision-language models. STAR adds inductive bias where brute-force scale is unavailable.

2. The dataset puts dexterous contact into the pretraining distribution

The hardware platform is a bimanual wheeled robot with two 7-DoF arms. Each hand has 10 actuated DoF, 16 total kinematic DoF, and a 268-dimensional piezoresistive tactile array across the palm and inner finger surfaces. Two wrist cameras and one head camera provide RGB observations. Skeleton-based gloves control the hands, trackers control the arms, and all observations and commands are synchronized at 30 Hz.

The teleoperation pipeline modifies DexPilot in two practical ways. It disables an automatic finger-closing rule that interferes with fine manipulation and adds fingertip orientation constraints, so a tilted human thumb does not map to an upright robot thumb. The retargeter is implemented in C++ and placed in a multiprocessing pipeline to keep the system responsive during collection.

The resulting dataset contains 200 hours, 10,576 trajectories, and 65 tasks. The paper classifies 69.5% of the trajectories as multi-finger dexterous manipulation, roughly 20% as other primitives such as pick-and-place or insertion, and about 10% as long-horizon compositions. Every trajectory may include three RGB views, left- and right-hand tactile images, robot state, commanded actions, and a manually assigned task-level language instruction.

The dataset is central to the result. On the four-task benchmark, plain $\pi_{0.5}$ averages 28% success. Finetuning the same backbone on these 200 hours, with tactile tokens supplied but without the STAR recipe, raises success to 44%. The training design then accounts for the remaining reported gain to 61%.

3. Three components handle three forms of sparsity

flowchart TD
    A["200 h synchronized RGB, tactile, state, action, language"] --> B["Joint masked visual-tactile pretraining"]
    B --> C["Pretrained tactile encoder"]
    C --> D["Activated local tactile tokens"]
    C --> E["One global tactile token per hand"]
    D --> F["Sparse-global tactile prefix"]
    E --> F
    G["Three RGB views + prompt + proprioception"] --> H["π0.5 VLM and action expert"]
    F --> H
    H --> I["50-step action chunk"]
    H --> J["Future tactile targets at 10, 20, 30, 40, 50 steps"]

Visual-tactile joint pretraining

STAR pairs each wrist image with the tactile image from the same hand. RGB and tactile inputs use separate ViT encoders, 196 patches per modality, fixed 2D positional embeddings, and independent 75% random masks. Their visible tokens meet in a shared fusion encoder and are reconstructed by modality-specific decoders. The patch-normalized objective is

\[\mathcal L_{\mathrm{joint}} =\sum_{m\in\{\mathrm{img},\mathrm{tac}\}} \frac{1}{|P_m^{\mathrm{mask}}|} \sum_{j\in P_m^{\mathrm{mask}}} \left\|\hat x_m^{(j)}-\bar x_m^{(j)}\right\|_2^2.\]

The paired wrist image makes an all-zero tactile prediction inconsistent with visible contact. This gives the tactile encoder a reason to preserve contact structure and begins the alignment between touch and vision before policy training.

One implementation detail narrows the claim: pretraining selects frames with nonzero tactile signals and their paired RGB images. The encoder learns from informative contact frames; the policy still has to handle inactive frames later.

Sparse-global tactile tokens

Each hand produces 196 spatial patch tokens and one learned global token. A calibrated contact gate marks a patch active when any taxel in it exceeds a sensor-specific threshold. Active spatial tokens remain visible to the backbone, inactive ones are masked, and the global token is always retained.

The global token uses asymmetric attention. Visual, language, and action tokens may attend to it, while it cannot attend back to those modalities. It therefore remains a tactile summary instead of becoming another generic multimodal token. Local tokens preserve where contact occurs; the global token carries a compact view of the whole hand.

This design also has a computational reading. Two dense tactile images would contribute 394 tokens before gating. STAR keeps the fixed global tokens and spends spatial-token capacity only where contact exists. The paper does not report wall-clock speed or the average retained-token count, so its efficiency benefit is plausible but unmeasured.

Sparse future tactile prediction

During multi-task finetuning and task-specific post-training, the policy predicts five future tactile force maps per hand. With action horizon $H=50$, the targets are

\[T_S=\{10,20,30,40,50\}.\]

At 30 Hz, they span approximately 0.33 to 1.67 seconds. A dense-near-term control uses the same prediction budget at $T_D={1,2,3,4,5}$, covering only the next 0.17 seconds. The auxiliary loss operates patch-wise on the normal-force channel:

\[\mathcal L_{\mathrm{future}} =\frac{1}{|T_S|}\sum_{t\in T_S} \frac{1}{|P_{\mathrm{tac}}^t|} \sum_{j\in P_{\mathrm{tac}}^t} \left\|\hat x_{\mathrm{tac}}^{(j)}-x_{\mathrm{tac}}^{(j)}\right\|_2^2.\]

The nearby frames are easy to predict because contact usually changes slowly from one 30 Hz frame to the next. Spreading five labels over the action horizon gives more supervision around contact transitions: when the earbud starts rotating, when a fingertip loses support, or when a book clears the stack. The ablation couples temporal spacing and maximum horizon, so it cannot tell which factor causes the gain.

4. STAR adapts $\pi_{0.5}$ through two training stages

The policy starts from the $\pi_{0.5}$ architecture: a PaliGemma Gemma-2B vision-language backbone and a Gemma-300M action expert. Three $224\times224$ camera images contribute 768 SigLIP tokens. Before sparse masking, the two tactile images contribute 394 tokens. A 64-dimensional proprioceptive state is serialized into the prompt, though only 34 dimensions are active: two 7-dimensional end-effector poses and two 10-dimensional hand-joint vectors. The complete sequence has 1,562 tokens, including a 50-token action suffix.

Action generation uses conditional flow matching. The model learns a velocity field that transports Gaussian noise toward the demonstrated 50-step action chunk and integrates it with ten forward-Euler steps at inference.

Training then proceeds in two stages:

  1. Multi-task finetuning: six epochs on the 200-hour dataset. The VLM and action expert start from $\pi_{0.5}$; the tactile encoder starts from joint pretraining; the action projections are reinitialized for the robot’s 64-dimensional representation.
  2. Task-specific post-training: 100 epochs on 100 demonstrations for each evaluation task.

The evaluation objects are excluded from the 200-hour dataset, then introduced during task-specific post-training. Tests use held-out initial configurations of those objects. The generalization claim therefore concerns adaptation from a multi-task dexterous prior to a held-out task-object setup with 100 demonstrations. It does not test a new task from language alone.

5. The main table contains two separate gains

The four real-world tasks probe different contact patterns. Earbud flipping requires a 180-degree in-hand rotation. Stacked-book retrieval hooks and pulls a tightly packed book before grasping it. Postcard retrieval slides a thin object to the table edge. Multi-object grasping asks both hands to hold previously collected objects while acquiring and discarding more.

Each method receives the same 100 task-specific trajectories and is evaluated on 20 trials per task. Success rate (SR) requires the full task. Task completion rate (TCR) credits completed subtasks.

MethodEarbud SRBook SRPostcard SRMulti-object SRMean SRMean TCR
GR00T N1.70%0%0%0%0%0%
LDA-1B5%0%15%0%5%9%
GR00T N1.7-Dex0%5%25%0%8%15%
LDA-1B-Dex40%0%30%0%18%23%
$\pi_{0.5}$15%35%60%0%28%44%
$\pi_{0.5}$-Dex without STAR25%45%75%30%44%61%
STAR65%55%85%40%61%79%

The strongest baseline is $\pi_{0.5}$, even though its pretraining uses parallel grippers and has no tactile input. The authors attribute this to its real-robot arm-motion prior. GR00T N1.7 produces unsafe arm oscillations under the shared post-training protocol, so its trials are terminated. This result describes the reported adaptation setup; it is too narrow to rank general-purpose VLA models overall.

The comparison I find most informative stays within the same backbone. Dexterous data adds 16 percentage points to mean SR, from 28% to 44%. STAR adds another 17 points, from 44% to 61%. Data and representation design contribute at similar scale in this table.

6. The ablations support the recipe, with one softer component

The full component ablation is run on earbud flipping and stacked-book retrieval, 20 trials each. These tasks intentionally represent tactile-rich and more vision-dominant behavior.

VariantEarbud SRBook SRMean SRMean TCR
Tactile input, no STAR recipe25%45%35%46%
Without visual-tactile pretraining30%45%38%45%
Without sparse local-token selection40%25%33%44%
Without global tactile tokens25%35%30%41%
Without future tactile prediction55%50%53%69%
Dense near-term future prediction55%40%48%62%
Full STAR65%55%60%78%

Joint pretraining and sparse-global tokenization carry the largest measured effects. The local/global split also changes by task: removing local-token sparsification hurts book retrieval more, while removing the global token hurts earbud flipping more. This fits the proposed roles of redundancy reduction and contact aggregation, though only two tasks support that interpretation.

Future prediction is the softer part of the recipe. Removing it lowers mean SR from 60% to 53%; changing to five adjacent targets lowers it to 48%. Those are useful gains, yet each 5-point step represents one trial per task. More repetitions or confidence intervals would make the ranking firmer.

7. Touch helps when the sensor covers the contact that matters

A tempting reading of the headline table is that tactile feedback explains the whole 17-point STAR gain. Table III gives a more precise answer.

Tactile conditionEarbud SRBook SRMean SRMean TCR
Train and test without tactile35%55%45%61%
Train with tactile, mask it at inference55%50%53%69%
Full tactile input65%55%60%78%

Earbud flipping supplies the cleanest evidence. Fingertips remain in contact during rotation, and those surfaces contain sensors; full tactile input adds 30 points over training without tactile. Book retrieval shows no SR gain, though TCR rises from 71% to 83%. The policy can use touch, but the benefit follows the task and the contact geometry.

The multi-object result exposes a hardware boundary. Its 40% SR sits far below its 72% TCR because slight errors occur when the ring and little fingers grasp with their edges. Those edges are not covered by the tactile array. Extra reliance on touch can also draw attention away from visual information, a failure mode the authors acknowledge.

Masking touch only at inference creates a distribution shift, so its 53% mean SR is not a clean estimate of the information carried by tactile signals. Training without touch is the cleaner modality comparison; inference masking still confirms that the trained policy reads the tactile channel when it is available.

8. What the evidence supports

STAR makes three claims that the experiments support well within this platform. Real dexterous data improves adaptation of an existing VLA backbone. Treating inactive tactile regions as ordinary tokens wastes representation capacity. Longer-horizon tactile prediction can be more informative than spending the same label budget on adjacent frames.

Several boundaries travel with those claims:

  • Twenty trials per task: success changes in 5-point increments, and no confidence intervals or significance tests are reported.
  • Four post-trained tasks: every evaluation task uses 100 demonstrations on its evaluation objects. Held-out initial configurations test robustness around that task-object distribution.
  • Two-task ablations: the detailed component and tactile studies cover earbud flipping and book retrieval, leaving postcard and multi-object effects unresolved.
  • One tactile embodiment: transfer to another hand, sensor layout, or sensing technology is not evaluated. Sensor noise, drift, and missing taxels remain open.
  • Normal force only: future tactile prediction uses the sensor’s normal-force channel; shear, slip, and richer contact geometry are absent.
  • Partial multimodal alignment: the tactile encoder is aligned with wrist vision before policy training. Explicit tactile-language alignment is left for future work.

I would also separate collection from release. The 200-hour dataset may be the paper’s most durable contribution, especially because 69.5% of its trajectories contain multi-finger behavior. The current paper documents the dataset, while the project page does not yet provide a public download. Reproducibility will depend heavily on whether the trajectories, tactile calibration, contact thresholds, and preprocessing code become available.

What I would carry forward

STAR offers a good rule for multimodal robot learning: measure the structure of a sensor stream before assigning it tokens. Dense tokenization is natural for RGB because every patch usually carries scene content. A tactile skin behaves differently. Silence is common, contact is local, and the useful event may be the transition that happens one second later.

The sparse-global representation is the part I would reuse first. It combines a simple physical gate with a learned summary token and does not require the VLM to rediscover sensor sparsity from limited data. The asymmetric attention constraint is equally practical; it preserves a tactile-specific summary while keeping that summary readable by the rest of the policy.

My next experiment would keep the four tasks fixed and vary three things separately: the average number of retained tactile tokens, the maximum prediction horizon, and the sampling interval between future targets. I would then inject calibrated drift and missing-taxel patterns at test time. That study would show which gain comes from token economy, which comes from longer contact forecasting, and how quickly the representation fails when the tactile skin stops matching its calibration.