RL with RoboCasa365 Benchmark#
This document describes the benchmark-native RoboCasa365 integration in RLinf.
Unlike the legacy RoboCasa recipe, this setup keeps the original
RoboCasa task wrapper untouched and adds a separate robocasa365 environment that
selects tasks through the official RoboCasa dataset registry.
The main goal is to train and evaluate vision-language-action models on the official RoboCasa365 benchmark splits while keeping the legacy RoboCasa recipe stable. Compared with the single-task RoboCasa examples, this recipe focuses on:
Benchmark split control: select
pretrainortargettasks through the official registry.Task-soup evaluation: evaluate benchmark slices such as
atomic_seenorcomposite_unseen.Task metadata at runtime: expose split and task-soup metadata in environment observations.
Environment#
RoboCasa365 Benchmark
Environment: RoboCasa365 kitchen simulation benchmark
Selection API: official RoboCasa dataset registry via
split+task_soupRobot: mobile Panda configuration (
PandaOmronby default)Observation: multi-view RGB images + configurable proprioceptive state extractor
Action Space: configurable RoboCasa mobile-manipulation action schema
The default RLinf recipe uses:
split=pretrainfor trainingsplit=pretrainfor the lightweight evaluation configtask_soup=atomic_seenas the first benchmark slice
You can switch to other official task soups, for example composite_seen or
composite_unseen, by editing the YAML config.
Observation Structure
Main camera image:
robot0_agentview_left_imageby defaultWrist camera image:
robot0_eye_in_hand_imageby defaultProprioceptive state: configurable state vector assembled from the keys in
observation.state_layout
Action Structure
The default OpenPI recipe uses a 12-dimensional action schema. The wrapper can disable base control and map valid OpenPI action slices before stepping the underlying RoboCasa environment.
Configuration#
The benchmark-specific environment config lives in:
examples/embodiment/config/env/robocasa365.yaml
Key fields:
task_source: should staydataset_registryfor RoboCasa365dataset_source: data source registered by RoboCasa, typicallyhumansplit: benchmark split such aspretrainortargettask_soup: official task soup name such asatomic_seentask_filter: optional include / exclude filter for narrowing the selected taskstask_mode: optionalatomicorcompositeguardrailobservation: camera keys and state-layout mapping used by RLinfaction_space: action schema and OpenPI slice mapping used before env stepping
Dependency Installation#
1. Clone RLinf Repository#
git clone https://github.com/RLinf/RLinf.git
cd RLinf
2. Prepare RoboCasa Assets#
export ROBOCASA_ASSETS_PATH=/path/to/robocasa_assets
mkdir -p "$ROBOCASA_ASSETS_PATH"/{_downloads,objects}
hf download robocasa/robocasa-assets \
--repo-type dataset \
--include textures.zip generative_textures.zip fixtures.zip objaverse.zip aigen_objs.zip \
--local-dir "$ROBOCASA_ASSETS_PATH/_downloads"
hf download nvidia/PhysicalAI-Robotics-Manipulation-Objects-Kitchen-MJCF \
--repo-type dataset \
--include "fixtures_lightwheel/*.zip" \
--local-dir "$ROBOCASA_ASSETS_PATH/_downloads"
hf download nvidia/PhysicalAI-Robotics-Manipulation-Objects-Kitchen-MJCF \
--repo-type dataset \
--include "objects_lightwheel/*.zip" \
--local-dir "$ROBOCASA_ASSETS_PATH/_downloads"
for item in \
textures.zip:textures \
generative_textures.zip:generative_textures \
objaverse.zip:objects/objaverse \
aigen_objs.zip:objects/aigen_objs; do
zip_name="${item%%:*}"
target_rel="${item#*:}"
parent_rel="$(dirname "$target_rel")"
if [[ "$parent_rel" == "." ]]; then
unzip_dir="$ROBOCASA_ASSETS_PATH"
else
unzip_dir="$ROBOCASA_ASSETS_PATH/$parent_rel"
fi
mkdir -p "$unzip_dir"
unzip -q "$ROBOCASA_ASSETS_PATH/_downloads/$zip_name" \
-d "$unzip_dir"
done
mkdir -p "$ROBOCASA_ASSETS_PATH/fixtures"
find "$ROBOCASA_ASSETS_PATH/_downloads/fixtures_lightwheel" -name "*.zip" \
-exec unzip -q {} -d "$ROBOCASA_ASSETS_PATH/fixtures" \;
mkdir -p "$ROBOCASA_ASSETS_PATH/objects/lightwheel"
find "$ROBOCASA_ASSETS_PATH/_downloads/objects_lightwheel" -name "*.zip" \
-exec unzip -q {} -d "$ROBOCASA_ASSETS_PATH/objects/lightwheel" \;
3. Install Dependencies#
export ROBOCASA_ASSETS_PATH=/path/to/robocasa_assets
bash requirements/install.sh embodied --model openpi --env robocasa365
source .venv/bin/activate
Dataset Selection#
RLinf now delegates task selection to RoboCasa’s official registry. The following command pattern from the RoboCasa docs is what the RLinf wrapper mirrors internally:
from robocasa.utils.dataset_registry import get_ds_soup
task_names = get_ds_soup(
task_set="atomic_seen",
split="pretrain",
source="human",
)
RoboCasa names this argument task_set; it corresponds to RLinf’s
task_soup YAML field.
Useful references:
RoboCasa dataset usage: https://robocasa.ai/docs/build/html/datasets/using_datasets.html
Model Checkpoint#
Download the official RoboCasa365 Pi0 checkpoint:
hf download RLinf/RLinf-pi0-robocasa-pretrain-human300 \
--local-dir /path/to/pi0-robocasa-pretrain-human300
Point the config to the downloaded directory:
rollout:
model:
model_path: "/path/to/pi0-robocasa-pretrain-human300"
actor:
model:
model_path: "/path/to/pi0-robocasa-pretrain-human300"
Training#
The OpenDrawer training recipe is:
bash examples/embodiment/run_embodiment.sh robocasa365_opendrawer_ppo_openpi
This config trains a single RoboCasa365 task with OpenPI and PPO:
env.train.split=pretrainenv.train.task_soup=atomic_seenenv.train.task_filter=["OpenDrawer"]env.train.task_mode=atomic
OpenDrawer Batch Divisibility
Use this check when changing GPU placement, rollout size, or actor batch size in
examples/embodiment/config/robocasa365_opendrawer_ppo_openpi.yaml.
chunk_steps = env.train.max_steps_per_rollout_epoch / actor.model.num_action_chunks
rollout_size_per_actor_rank = (
env.train.total_num_envs * env.train.rollout_epoch / actor_world_size
) * chunk_steps
rollout_size_per_actor_rank % (actor.global_batch_size / actor_world_size) == 0
actor.global_batch_size % (actor.micro_batch_size * actor_world_size) == 0
End-to-End Test#
The shortened PPO e2e config is robocasa365_ppo_openpi:
bash tests/e2e_tests/embodied/run.sh robocasa365_ppo_openpi
Evaluation#
The lightweight evaluation config is robocasa365_eval_openpi:
bash evaluations/run_eval.sh robocasa365_eval_openpi
This config evaluates on:
env.eval.split=pretrainenv.eval.task_soup=atomic_seenenv.eval.task_mode=atomicno
env.eval.task_filterby default, so the selected pretrain slice is not narrowed
To evaluate a different benchmark slice, override the YAML fields directly:
env:
eval:
split: target
task_soup: composite_unseen
task_mode: composite
For example:
bash evaluations/run_eval.sh robocasa365_eval_openpi \
env.eval.task_soup=composite_unseen \
env.eval.task_mode=composite