Under review Vision-Language-Action Test-time retrieval OOD robustness

R2A: Retrieve to Act Motion primitive graph retrieval for robust VLA execution

Anonymous authors · paper under double-blind review

89.0+11.9
LIBERO-Plus average, GR00T N1.7 (77.1 without R2A)
75.9+41.7
RoboTwin2.0-Plus, Robot Initial States, π0.5
80+50
Real-world Banana2Pot, shifted start (30 without R2A)
159ms
Per-step latency, vs 138 ms for the base policy
R2A overview: offline Graph-Index over training demonstrations, then test-time rectification, GNN retrieval and harmonization.
Overview of R2A. Offline, Graph-Index builds a motion primitive graph from successful demos. At test time a diffusion model first rectifies an OOD start state. Then at every step the R2A-Retriever locates the rollout in the graph and retrieves a reference motion, and Harmonize blends it into the VLA prediction through a gate.
01 · Abstract

The policy knows what to do. Under a visual shift it forgets how to move.

VLA policies generalize across instructions, but small out-of-distribution (OOD) changes at execution time, such as a moved camera, a new texture or different lighting, push their actions off safe trajectories. Test-time methods that suppress these visual cues ignore the kinematic priors in the policy's own training data, which such shifts leave intact.

Retrieve to Act (R2A) turns successful training demonstrations into an offline motion primitive graph. At test time a query-conditioned GNN locates the rollout in the graph and retrieves compatible future motion, and a learned per-dimension gate blends it into the frozen VLA's actions. On GR00T N1.7, π0.5 and StableVLA it raises success on LIBERO-Plus, RoboTwin2.0-Plus and a real Aloha robot, with no finetuning.

02 · Method

Inside one rollout: query, messages, retrieval, harmonization

A logged rollout of GR00T N1.7 with R2A. Pick a shift: in each, GR00T alone fails. The teal cloud is the task's motion primitive graph, drawn into the camera. At each check the query forms at the gripper, messages travel the graph's edges, the top 8 segments light up with their logged scores, and the policy, retrieved and executed chunks appear.

Camera Viewpoints
step 0
Loading
GR00T N1.7 alonesame scene
Graph node Message Policy Aπ Retrieved Aref Executed Aexec Chunks drawn 4× longer

All data shown is logged, with two exceptions. The message wave follows real edges outward from the nodes nearest the query, but per-layer activations were not logged. The gate bars are fitted from the logged chunks as the S in Aexec = Aπ + S ⊙ (Aref − Aπ).

Offline

Graph-Index

Each demo segment is a node. Temporal edges jump ahead within a demo (Δ = 2, 4, 6, 8). Spatial edges join nearby end-effector states across demos, including other tasks.

Test time

R2A-Retriever

The query combines the VLA's predicted poses with recent states and actions. Message passing conditioned on it spreads relevance along the graph. The future motion of the top-K nodes fuses into a reference chunk.

Test time

Harmonize

A per-dimension gate corrects the spatial pose where prediction and reference disagree. Gripper commands pass through. Before the first step, a diffusion model moves an OOD start pose back toward the training distribution.

Retrieval example under Sensor Noise with GR00T N1.7 at step 32.
Real retrieval, LIBERO-Plus

One query, one retrieved primitive

GR00T N1.7 under Sensor Noise, step 32. The query retrieves demo 49, frames 40 to 55, a segment that goes down to the book. Over 8 steps the policy alone moves 2.6 cm down, the demo 8.3 cm, the harmonized action 3.5 cm. The gate moves the action toward the demo but does not replay it. Background shown clean.

03 · Simulation

LIBERO-Plus: seven perturbation categories, three base policies

Base policy+ R2A
Full Table 1 with all baselines

R2A improves all three policies on average, most under Robot Initial States: +49.1 for GR00T N1.7, +48.0 for StableVLA, +21.9 for π0.5.

Bimanual: RoboTwin2.0-Plus

On π0.5, R2A lifts the average from 56.8 to 66.2 and every category improves. Robot Initial States: 34.2 to 75.9.

Against other test-time OOD methods

LIBERO-10 averages. VLS uses VLM-generated rewards, SDN pushes away from negative references, ICL adds text descriptions of retrieved observations.

Full Table 3, per category
04 · Real world

Aloha with GR00T N1.7, robot started from a shifted pose

Banana2Pot, original initial stateOriginal
Banana2Pot, perturbed initial statePerturbed
TaskMethodBaseRobot init.
Banana2PotGR00T N1.760.030.0
+ SDN75.055.0
+ R2A80.0+20.080.0+50.0
Cube2DrawerGR00T N1.755.010.0
+ SDN70.030.0
+ R2A85.0+30.045.0+35.0

The shifted start breaks the base policy and SDN. R2A holds Banana2Pot at its in-distribution rate and lifts Cube2Drawer from 10 to 45. It also helps without a shift.

05 · Analysis

Ablation and inference time

Ablation, GR00T N1.7 on LIBERO-Plus

Without harmonization, replayed motion collapses to 1.2. ISR (initial state rectification) is measured on Robot Initial States only.

Inference time and success, LIBERO-10

Per-step inference time, GR00T N1.7. SDN needs detection and segmentation. VLS queries a VLM several times per step.

06 · Citation

BibTeX

@inproceedings{r2a2027,
  title     = {Retrieve to Act: Motion Primitive Graph Retrieval for Robust {VLA} Execution},
  author    = {Anonymous},
  booktitle = {Under review},
  year      = {2026}
}