GALAGeometry-Aware Latent Action Modeling
for Vision-Language-Action Model Pretraining
across Embodiments

Yichen Liu1*Puzhen Yuan1*Xiang Zhu1,2*Yanjiang Guo1,2Jianyu Chen1,2†

1 Institute for Interdisciplinary Information Sciences, Tsinghua University, China

2 Shanghai Qi Zhi Institute, China

* Equal contribution† Corresponding author

GALA combines visual transitions and 3D end-effector geometry to learn shared fine-grained latent actions across embodiments.
Learning how end effectors manipulate, alongside where they move.

Abstract

Learning large-scale vision-language-action (VLA) models from multi-embodiment datasets remains challenging due to heterogeneous action spaces across end effectors. Although latent action models (LAMs) can learn embodiment-agnostic action representations from diverse video data, existing image-based LAMs often fail to capture fine-grained end-effector articulation, particularly finger-level geometric changes in human and dexterous robot hands. To address this limitation, we propose GALA, a Geometry-Aware Latent-Action modeling framework that augments image-based latent actions with 3D end-effector geometric motion. However, naively incorporating point clouds yields fine-grained action representations with limited shared semantics, hindering cross-embodiment pretraining. To address this issue, we introduce the Unified End-effector Motion Representation (UEMR), which preserves fine-grained motion information while improving the cross-embodiment generalizability of latent actions. Building upon UEMR, GALA combines visual latent actions that capture scene-level dynamics with geometric latent actions that capture shared fine-grained end-effector articulation, providing effective supervision for VLA pretraining from multi-embodiment data, including action-free ego-centric human videos. Experiments on fine-grained motion probing, cross-embodiment retrieval, and downstream VLA evaluation demonstrate GALA’s effectiveness in modeling generalizable fine-grained motions across embodiments, achieving 68.3% RoboCasa-GR1 success rate and 75.5% real-world success rate.

Overview Video

Method

Visual dynamics and fine-grained geometry, learned together across human hands, dexterous robot hands, and parallel-jaw grippers.

Stage 1 · Geometry-Aware Latent Action Learning

Stage 1 architecture with visual and geometric encoding, shared latent actions and separate reconstruction.
GALA jointly encodes RGB transitions and 3D end-effector point clouds. UEMR combines a unified bimanual motion latent, pair-consistent geometric augmentation, and bidirectional transition learning to capture shared motion semantics without requiring joint or point correspondence.

Stage 2 · Geometry-Aware VLA Co-Training

Stage 2 architecture: frozen GALA supervises VLM bridge tokens and a shared action expert with embodiment-specific heads.
Frozen visual and geometric latent actions supervise dedicated bridge tokens in the vision-language model. A shared flow-matching action expert predicts continuous actions through embodiment-specific heads, preserving each embodiment’s native action space.

Experimental Results

Fine-Grained Motion Probing

A frozen representation and a lightweight probe predict end-effector translation, wrist rotation, and finger articulation.

MethodPosition (cm) ↓Rotation (°) ↓Finger (°) ↓
METIS6.1628.52014.328
Native Kinematics5.8798.0987.768
OPFA6.1028.2607.707
GALA w/o PC6.7908.56715.113
GALA (Ours)5.3597.8987.585

Held-out XHand test set; averaged over three fixed random seeds. Lower is better.

Cross-Embodiment Motion Retrieval

Retrieving grasp, hold, and release transitions across different embodiments.

Human–Robot Track

MethodR@1 ↑P@5 ↑mAP ↑
METIS33.2633.2937.21
GALA w/o PC33.3733.3736.38
GALA w/o UEMR38.6940.3148.15
GALA (Ours)45.6744.5749.49

Robot-Only Track

MethodR@1 ↑P@5 ↑mAP ↑
METIS33.3033.3036.10
Native Kinematics39.3135.4641.94
OPFA41.1836.8343.03
GALA w/o PC33.3833.3836.41
GALA w/o UEMR33.8338.5443.11
GALA (Ours)46.9141.7446.95

All metrics are percentages; higher is better. Macro-averaged across motion classes and directed embodiment pairs.

68.3%RoboCasa-GR1
multi-embodiment success
+12.6%RoboCasa gain from
multi-embodiment co-training
75.5%Real-world
average success

RoboCasa-GR1 · GR-1-Only Training

Latent action modelSuccess rate (%) ↑
UniVLA48.0
METIS43.8
Native Kinematics51.8
OPFA53.5
GALA w/o PC50.9
GALA w/o UEMR54.4
GALA (Ours)55.7

Same VLA architecture, training data, and optimization settings; only the latent action model changes.

RoboCasa-GR1 · Multi-Embodiment Co-Training

Average success across 24 dexterous tabletop manipulation tasks.

Success rate on a 0–100% scale. Publicly reported baselines are reproduced from the paper.

MethodSuccess rate (%) ↑Gain vs. GR-1 only
FLARE55.0
DiT4DiT56.7
JoyAI-RA63.2
UniT66.8
UniVLA53.6+5.6
GALA w/o UEMR58.6+4.2
GALA (Ours)68.3+12.6

RoboCasa-GR1 Demonstrations

One successful rollout for each of the 24 tabletop manipulation tasks.

01Cup → drawer + close
02Potato → microwave + close
03Milk → microwave + close
04Bottle → cabinet + close
05Wine → cabinet + close
06Can → drawer + close
07Cutting board → basket
08Cutting board → cardboard box
09Cutting board → pan
10Cutting board → pot
11Cutting board → tiered basket
12Placemat → basket
13Placemat → bowl
14Placemat → plate
15Placemat → tiered shelf
16Plate → bowl
17Plate → cardboard box
18Plate → pan
19Plate → plate
20Tray → cardboard box
21Tray → plate
22Tray → pot
23Tray → tiered basket
24Tray → tiered shelf

Real-World Manipulation

Success rates across four evaluated tasks, with 50 trials per task.

MethodPick ↑Push ↑Press ↑Flip ↑Average ↑
OpenVLA024181013.0
UniVLA3862322038.0
OpenVLA-OFT5272544255.0
π₀5672563454.5
π₀.₅7280685268.0
HARP-VLA7282745871.5
GALA w/o UEMR7478725068.5
GALA (Ours)8082786275.5

Success rates (%). Pouring demonstrations are supplementary and are not included in this benchmark.

Real-World Demonstrations

Dexterous manipulation across objects and actions.
Each demonstration includes third-person and wrist-camera views.

Pick and Place

Bamboo

Bamboo · 01
Bamboo · 02
Bamboo · 03

Orange

Orange · 01
Orange · 02
Orange · 03

Boxes

Yellow box · 01
Green box · 01
Pink box · 01

Push Object

Blue box · 01
Green box · 02
Red box · 03

Press Button

Button · 01
Button · 02
Button · 03

Flip Cup

Blue cup · 01
Shaker cup · 02
Pink cup · 03

Pour Liquid

Pouring · 01
Pouring · 02
Pouring · 03

Citation

@misc{liu2026galageometryawarelatentaction,
  title={GALA: Geometry-Aware Latent Action Modeling for Vision-Language-Action Model Pretraining across Embodiments},
  author={Yichen Liu and Puzhen Yuan and Xiang Zhu and Yanjiang Guo and Jianyu Chen},
  year={2026},
  eprint={2609.21948},
  archivePrefix={arXiv},
  primaryClass={cs.RO},
  url={https://arxiv.org/abs/2609.21948},
}