MolmoAct2 Evaluation#
MolmoAct2 is AllenAI’s open vision-language-action model, served through a LeRobot policy that predicts continuous actions with a flow-matching action expert. RLinf runs the official MolmoAct2-LIBERO checkpoint through its LIBERO runner. This integration is evaluation only — the upstream policy is inference-only, so there is no training path.
Overview#
Evaluate the official MolmoAct2-LIBERO checkpoint across the four LIBERO suites.
LIBERO
Evaluation only
Spatial · Object · Goal · Long
1 node · 1–8 GPUs
eval/success_once.Tasks#
One config ships per suite. Each runs 20 parallel environments for 25 episodes
each — the full 500-trajectory suite — with a step budget of
max_episode_steps Ă— 25.
Suite |
Config |
|
Trajectories |
|---|---|---|---|
Spatial |
|
240 |
500 |
Object |
|
240 |
500 |
Goal |
|
320 |
500 |
Long |
|
520 |
500 |
Observation and Action#
Field |
Description |
|---|---|
Observation |
Two camera views — |
Action |
Continuous 7-DoF actions (6-DoF delta EE + gripper) from the action expert. One action is executed per rollout step. |
Reward |
LIBERO task success. |
Prompt |
Natural-language task instruction from |
Installation#
First, clone the RLinf repository:
# Mainland China users can use a mirror for faster cloning:
# git clone https://ghfast.top/github.com/RLinf/RLinf.git
git clone https://github.com/RLinf/RLinf.git
cd RLinf
Then set up the dependencies with one of the two methods below — a prebuilt
Docker image (recommended) or a custom environment. The general setup
(prerequisites, GPU drivers, the in-image switch_env helper, mirrors, and
troubleshooting) is documented once in Installation;
the commands in this recipe only differ in the Docker image tag and the
--env value.
Option 1: Docker image — image tag agentic-rlinf0.3-libero:
docker run -it --rm --gpus all \
--shm-size 20g \
--network host \
--name rlinf \
-v .:/workspace/RLinf \
rlinf/rlinf:agentic-rlinf0.3-libero
# Mainland China mirror: docker.1ms.run/rlinf/rlinf:agentic-rlinf0.3-libero
# Inside the container, switch to the MolmoAct2 virtual environment:
source switch_env molmoact2
Option 2: Custom environment — install bundle --model molmoact2 --env libero:
# Add --use-mirror for faster downloads in mainland China.
bash requirements/install.sh embodied --model molmoact2 --env libero
source .venv/bin/activate
Set MOLMOACT2_LEROBOT_PATH before installing to reuse an existing checkout of
RLinf/lerobot.
Download the Model#
Download the official allenai/MolmoAct2-LIBERO checkpoint:
hf download allenai/MolmoAct2-LIBERO \
--local-dir /path/to/model/MolmoAct2-LIBERO
Then set rollout.model.model_path in the eval config to that directory. The
config has no actor section, so there is no second path to keep in sync.
Run It#
bash evaluations/run_eval.sh libero libero_10_molmoact2_eval \
rollout.model.model_path=/path/to/model/MolmoAct2-LIBERO
What this command does:
Loads the official checkpoint through the MolmoAct2 model adapter.
Runs the LIBERO-Long suite with the settings in
evaluations/libero/libero_10_molmoact2_eval.yaml.Writes terminal output and
eval/success_onceto a timestamped log.
Swap in any config from the Tasks table above to evaluate another suite.
Warning
A full suite covers 500 trajectories and can take several hours. Lower
env.eval.max_steps_per_rollout_epoch for a smoke test — one episode per
environment is max_episode_steps.
Configure further
Inference settings (
num_steps,norm_tag, action mode) → themolmoact2block inexamples/embodiment/config/model/molmoact2.yaml.rollout.model.precisionhas no effect: MolmoAct2 loads its weights in fp32 upstream.Keep
rollout.pipeline_stage_num: 1; the policy keys its per-environment action queues by batch index.Keep
env.eval.max_episode_stepsa multiple of the policy’s 10-step action queue (240 / 320 / 520 all are), or an episode starts on the previous one’s leftover actions.Placement and throughput → Placement and Execution modes.
Visualization and Results#
The terminal reports eval/success_once. Logs are written to:
logs/<timestamp>-libero_10_molmoact2_eval/eval_embodiment.log
See LIBERO Evaluation for the benchmark protocol and Evaluation Results for metric interpretation.