01 / The question
Knowing the parts is only the beginning.
Imagine two robot arms stacking bowls. In the training routine, one arm picks and places its bowl, then the other follows. Now ask them to reverse that order. Or pick together, then place in a different order. The objects and basic actions are familiar. The relationship between the arms has changed.
This is arm-wise compositional generalization: reusing familiar arm-level skills under coordination requirements that were not seen during training. It offers a way to extend a robot’s behavior without collecting demonstrations for every possible combination.
Giving each arm its own policy does not automatically solve coordination. Both arms act in the same workspace: a locally sensible movement may obstruct the other arm, and reaching part of a goal does not guarantee that the full combination succeeds.
02 / ACG-Bench
Keep the skills. Change the composition.
ACG-Bench contains 23 task–condition pairs across eight task families: six in-domain conditions and 17 held-out compositions. The held-out conditions contribute no training demonstrations. We change how familiar skills must work together along four axes.
Reorder
4 conditionsChange which arm acts first, while keeping the same skills and task goal.
Familiar sequence
time →
Unseen composition
time →
Sync
6 conditionsRequire designated events to finish within a time window. Here, the two pick events are aligned.
Familiar sequence
time →
Unseen composition
time →
Sync + Reorder
4 conditionsChange timing and order together. In this example, pick together, then place right before left.
Familiar sequence
time →
Unseen composition
time →
Cross-task
3 conditionsBring together skills learned in different source tasks, such as bowl handling and cube placement.
Schematic event order. Spacing does not represent physical duration; synchronization uses the benchmark’s event-time tolerances.
What counts as success?
The final arrangement is only part of the test. Constraint-compliant success also requires the specified physical milestones, event order, and timing. Arm collisions count as failures. Synchronization is scored using task-specific completion events and time tolerances, rather than requiring identical trajectories.
Every method receives the same per-arm atomic prompts from a common rule-based phase scheduler. This isolates execution of a supplied skill composition. The benchmark does not ask the policy to infer the plan itself.
03 / AE-VLA
One shared policy. Three design choices.
We start from the same pretrained π0.5 backbone used by the baselines. AE-VLA keeps a shared vision-language backbone and action expert, then introduces structure at three levels: action representations, skill-specific parameters, and attention.
Arm-token grouping
Give each arm an explicit group of action tokens. Both groups pass through the same action expert, with each group assigned to its arm’s output.
SkillLoRA
Reuse a bank of ten skill-specific low-rank adapters across arms and tasks. A learned router for each arm selects an adapter from its prompt and state features.
Arm-wise attention
Keep action attention within each arm’s stream while retaining access to shared global context. The global features can still gather observations from both arms.
The aim is to make skills reusable while preserving the scene information needed for coordination. Arm-wise attention blocks direct cross-arm action attention, but the policy remains coupled through shared global features. It also changes temporal attention within each action group, so this is a study of the full attention design.
The combination matters.
Grouping tokens alone brings a small gain. Adding SkillLoRA or arm-wise attention individually also yields modest improvements. Combining all three reaches 21.53% generalization success. In this controlled comparison, the designs work better together.
| Policy design | Success |
|---|---|
| Token Group | 3.82% |
| Token Group + SkillLoRA | 6.00% |
| Token Group + AWA | 5.82% |
| AE-VLA · all three | 21.53% |
Source: paper, Table 2. SkillLoRA adds both parameters and router supervision; the ablation does not isolate those two changes.
04 / From sim to real
Generalization, measured.
We compare Single π0.5, an adapted MA-VLA with the same π0.5 initialization, Dual π0.5 with one policy per arm, and AE-VLA. All use the same source data and evaluation prompts within each setting. The simulation evaluation runs 100 episodes per condition for each method; real-robot evaluation uses 20 trials per condition.
+16.00 percentage points over Dual π0.5
+29.00 percentage points over Dual π0.5
Constraint-compliant success, averaged equally over conditions. Simulation and real-robot sets differ; their percentages are not a direct sim-to-real transfer comparison. Sources: paper, Section 5.2 and Tables 2–3.
On the real robots, AE-VLA succeeds in 12 of 20 Stack Bowls / Sync trials, while all three baselines score 0 of 20. In the held-out Cube in Bowl task it reaches 10 of 20, compared with 4 of 20 for Dual π0.5. The five-condition unseen set consists of two Reorder, two Sync, and one Cross-task condition.
Where the approach still falls short
Pure synchronization remains difficult: simulation success is 8.67%. Generalization gains also come with lower in-domain success: 27.17% versus 41.33% for Dual π0.5 in simulation, and 60% versus 70% on the real robots. On real Push Cubes / Sync, Dual π0.5 scores 6/20 while AE-VLA scores 2/20. The improvement is substantial, but it is not uniform across tasks.
05 / What comes next
Beyond fixed routines.
This work points toward a practical building block for embodied intelligence: a policy that can reuse what each arm knows when their relationship changes. Explicit arm representations, reusable skill parameters, and structured information sharing help move beyond a fixed collaboration routine.
The next step is to connect compositional execution to a high-level planner: decompose a new task into familiar skills, assign them to arms, choose order and timing, and revise the plan from execution feedback. ACG-Bench studies execution under supplied plans; closing that planning–execution loop remains open.
Learn deeply.
Generalize on the fly.