OPD for OpenVLA-OFT on LIBERO#
OPD (On-Policy Distillation) distills on the student policy’s own on-policy rollouts:
the student interacts with the environment and stores action tokens; after rollout,
the teacher scores those same student action tokens with dense token-level log
probabilities. In RLinf, the teacher does not generate actions during rollout. It
only scores saved forward_inputs.action_tokens after rollout and is offloaded
immediately afterward to reduce rollout latency and GPU memory pressure.
Overview#
This example trains an OpenVLA-OFT student with an OpenVLA-OFT teacher on LIBERO-Spatial using OPD.
LIBERO-Spatial
OPD
OpenVLA-OFT student · OpenVLA-OFT teacher
1 node · 8 GPUs
unnorm_key -> launch libero_spatial_opd_openvlaoft -> watch env/success_once and OPD actor metrics.Method#
OPD uses the relative log probability assigned by teacher and student to the same student action-token sequence:
Implementation details:
Student rollout is the only source of environment interaction; the teacher does not sample actions.
After rollout, the teacher computes dense token logprobs over stored
forward_inputs.action_tokens.OPD loss uses
loss_maskto exclude invalid tokens after success or termination.
Checkpoints#
Role |
Hugging Face checkpoint |
Config field |
|
|---|---|---|---|
Student |
|
|
|
Teacher |
|
|
Download the checkpoints with either method:
# Method 1: git-lfs
git lfs install
git clone https://huggingface.co/RLinf/RLinf-OpenVLAOFT-LIBERO-130-Base-Lora
git clone https://huggingface.co/RLinf/RLinf-OpenVLAOFT-LIBERO-130
# Method 2: huggingface-hub (set HF_ENDPOINT=https://hf-mirror.com in mainland China)
pip install huggingface-hub
hf download RLinf/RLinf-OpenVLAOFT-LIBERO-130-Base-Lora --local-dir RLinf-OpenVLAOFT-LIBERO-130-Base-Lora
hf download RLinf/RLinf-OpenVLAOFT-LIBERO-130 --local-dir RLinf-OpenVLAOFT-LIBERO-130
Configuration#
The example config is examples/embodiment/config/libero_spatial_opd_openvlaoft.yaml.
Key fields are:
algorithm:
adv_type: opd
loss_type: opd
loss_agg_func: token-mean
actor:
global_batch_size: 4096
micro_batch_size: 32
model:
model_path: /path/to/RLinf-OpenVLAOFT-LIBERO-130-Base-Lora
unnorm_key: libero_130_no_noops_trajall
optim:
lr: 2.0e-6
rollout:
model:
model_path: /path/to/RLinf-OpenVLAOFT-LIBERO-130-Base-Lora
expert_model:
model_path: /path/to/RLinf-OpenVLAOFT-LIBERO-130
unnorm_key: libero_130_no_noops_trajall
env:
train:
rollout_epoch: 8
max_episode_steps: 240
max_steps_per_rollout_epoch: 240
Note
Student and teacher must use the same unnorm_key. Both checkpoints in this
recipe use libero_130_no_noops_trajall.
Run#
Launch the full training run:
source switch_env openvla-oft
bash examples/embodiment/run_embodiment.sh libero_spatial_opd_openvlaoft
Logs are written under runner.logger.log_path and can be inspected with TensorBoard:
tensorboard --logdir ./logs --port 6006
Key metrics:
env/success_once: unnormalized episodic success rate.actor/opd_reverse_kl: teacher-vs-student logprob gap on student action tokens.actor/opd_reward: token-level distillation reward used by OPD.actor/policy_lossandactor/total_loss: actor update stability.
Results#
The table below reports LIBERO-Spatial results. The base student is
RLinf/RLinf-OpenVLAOFT-LIBERO-130-Base-Lora;
the teacher is
RLinf/RLinf-OpenVLAOFT-LIBERO-130.
Model / training stage |
Success rate |
Notes |
|---|---|---|
OpenVLA-OFT LIBERO-130 Base-Lora student |
~67% |
Student baseline before OPD training. |
OPD training, 30 steps |
96%+ |
Early OPD result with |