srlwam β stack_v1 policies
LiLa-WAM and Cosmos Edge policies for two-cube stacking on a Franka Panda, trained on the stack_v1 simulation
corpus and on real teleoperated recordings, packaged for deployment on the real robot. Training
data: dataset repo aabyaneh/srlwam
(revision 4ee4768c, stack/sim/stack_v1 and stack/real/stack_v0).
Deploying? Read
AGENTS.mdfirst, then followdeploy/DEPLOYMENT.mdgate by gate. The checkpoints do not all expect the same robot state; the deploy code handles it from each checkpoint'sdeploy.json, and the sanity checks verify it before the arm moves.
Cosmos Edge stack_v2
Cosmos3 Edge checkpoints trained on the stack_v2 simulation data, six training runs (two fine-tune on 50% / 20% of the
real train episodes). Each run is released as its
final checkpoint and its best checkpoint (lowest held-out normalized action MAE, open-loop, evaluated every
250 steps; one bundle when they coincide). Shared assets are the same cosmos_edge/shared/ used by stack_v1.
Sim-unit checkpoints (sim_v2_*, real_video_*) use gain 0.1975 m/unit: use command.delta, not
command.real_controller_delta.
| run | checkpoint | role | observation frame | dir. cosine (step 0) | translation L1 (h16) | gripper acc. (step 0) |
|---|---|---|---|---|---|---|
| sim (stack_v2 sim from Cosmos3-Edge-Policy-DROID) | sim_v2_droid_iter1750 | best (MAE 0.0593) | sim_ee_frame |
+0.273 | 0.0059 | 0.914 |
| sim (stack_v2 sim from Cosmos3-Edge-Policy-DROID) | sim_v2_droid_iter2000 | final (MAE 0.0624) | sim_ee_frame |
+0.280 | 0.0058 | 0.923 |
| real supervised from sim | real_supervised_from_sim_v2 | best = final (MAE 0.1966) | raw_robot |
+0.342 | 0.0058 | 0.966 |
| real supervised from sim, 50% real data | real_supervised_from_sim_v2_real50 | best = final (MAE 0.1867) | raw_robot |
+0.364 | 0.0055 | 0.969 |
| real supervised from sim, 20% real data | real_supervised_from_sim_v2_real20_iter2000 | best (MAE 0.2194) | raw_robot |
+0.376 | 0.0063 | 0.967 |
| real supervised from sim, 20% real data | real_supervised_from_sim_v2_real20_iter4000 | final (MAE 0.2257) | raw_robot |
+0.365 | 0.0061 | 0.967 |
| real only from DROID | real_only_from_droid_iter3500 | best (MAE 0.1967) | raw_robot |
+0.367 | 0.0059 | 0.966 |
| real only from DROID | real_only_from_droid_iter4000 | final (MAE 0.2007) | raw_robot |
+0.356 | 0.0059 | 0.964 |
| real video from sim | real_video_from_sim_v2_iter2500 | best (MAE 0.3143) | sim_ee_frame |
+0.389 | 0.0051 | 0.947 |
| real video from sim | real_video_from_sim_v2_iter4000 | final (MAE 0.3218) | sim_ee_frame |
+0.394 | 0.0051 | 0.951 |
Held-out results use the same 906 real windows (episodes 0, 1, 15, 17) and seeds as stack_v1. They are recorded-action agreement scores, not closed-loop robot success rates.
Cosmos Edge stack_v1
Five Cosmos3 Edge checkpoints are packaged with shared processor, vision-encoder and Wan2.2 VAE assets.
Start with real_supervised_from_sim_v1 for supervised real adaptation.
Follow Cosmos deployment instructions for a local GPU or the provided remote server/client.
Each checkpoint is about 6.74 GB; shared assets add about 3.83 GB. The complete Cosmos bundle is about 37.5 GB.
| checkpoint | observation frame | adaptation |
|---|---|---|
| real_supervised_from_sim_v1 | raw_robot |
real stack recordings (16 train episodes), joint action + video objective, 2000 steps, from sim_v1_droid; sim normalization |
| real_video_from_sim_v1_iter2000 | sim_ee_frame |
as iter1250, end of the 2000-step schedule; lowest aggregate held-out action MAE of the video runs |
| real_video_from_sim_v1_iter1250 | sim_ee_frame |
sim_v1_droid's generation expert adapted on real camera video only (no real state or actions), step 1250 = lowest held-out video flow loss; action heads frozen |
| sim_v1_droid | sim_ee_frame |
stack_v1 sim, 1800 train demos, 2000 steps x global batch 2048, from Cosmos3-Edge-Policy-DROID |
| sim_v1_limiteddata_droid | sim_ee_frame |
stack_v1 sim, 200-demo subset (seed 1701), same 2000-step recipe, own normalization |
Deployment and timing
Cosmos predicts 16 future actions and video frames jointly. The adapter executes the 16-action chunk at 10 Hz, holds position while inference runs, and then replans. Measured H200 inference is approximately 1.2-1.4 s per chunk with 30 sampling steps; peak model GPU allocation is approximately 7.7 GiB. This is not a continuous 10 Hz inference service. Budget additional GPU memory for the runtime and other processes. Use gate 1 to measure the deployment machine.
Held-out real results (open-loop)
Predictions use 906 windows from real episodes 0, 1, 15 and 17, with a full 16-step future.
Sampling uses 30 steps and seed 1701 + window index. Values below use each checkpoint's declared observation frame.
These are recorded-action agreement scores, not closed-loop robot success rates. These real recordings also informed scene calibration;
the step-1250 variant was selected using held-out video loss, so this is not an untouched test set.
| checkpoint | direction cosine (step 0) | translation L1 (h16) | gripper accuracy (step 0) | dz sign agreement (h16) | net dz pred / truth (sim units) |
|---|---|---|---|---|---|
| real_supervised_from_sim_v1 | +0.363 | 0.0064 | 0.965 | 0.777 | -0.0698 / -0.0593 |
| real_video_from_sim_v1_iter2000 | +0.376 | 0.0072 | 0.953 | 0.807 | -0.1158 / -0.0593 |
| real_video_from_sim_v1_iter1250 | +0.369 | 0.0072 | 0.953 | 0.807 | -0.1075 / -0.0593 |
| sim_v1_droid | +0.163 | 0.0103 | 0.945 | 0.575 | -0.0476 / -0.0593 |
| sim_v1_limiteddata_droid | +0.165 | 0.0109 | 0.949 | 0.746 | -0.1175 / -0.0593 |
The supervised checkpoint has translation L1 about 0.0064 and predicted net descent about -0.070 versus -0.059 in the recordings. Video-only adaptation improves direction agreement relative to the sim-only checkpoint, but predicts substantially more downward motion (about -0.108 to -0.116 sim units). The limited-data sim baseline also overpredicts descent (about -0.117). Start with shadow mode and the staged motion checks; do not interpret a higher direction score as proof of better robot control.
Each checkpoint's eval_real_holdout.json contains both observation frames, including the wrong-frame control.
The exported bf16 weights approximately reproduce the original full-precision evaluations; small numerical changes can flip gripper signs
near decision boundaries. Fixture checks compare the exported deployment weights against their own fixed-seed references.
Release verification
All five Cosmos checkpoints passed 8/8 offline checks with an empty Hugging Face cache. The isolated robot-side HTTP client passed 8/8 checks, and the wrong-checkpoint handshake was rejected. Both fault-injection replays passed the clean baseline and flagged the six injected mistakes. All 11 existing LiLa-WAM weight variants passed their 8/8 regression checks. These are offline deployment checks; no robot was commanded. See verification report and check output.
Files and provenance
cosmos_edge/stack_v1/<checkpoint>/ contains model/ (including hidden .metadata), deploy.json,
checkpoint.json, normalization.json, split.json, reference predictions, and evaluation results.
deploy/cosmos/ contains the loader, HTTP server/client, offline checker, setup instructions, and robot dependency pins.
Use this loader with the pinned NVIDIA framework; these are distributed checkpoints, not Transformers AutoModel bundles.
See component licenses and Cosmos notices.
Which LiLa-WAM checkpoint
| checkpoint | trained on | robot state it expects (obs_frame) |
use it for |
|---|---|---|---|
real_finetune_from_sim_v1 |
sim_v1 pretrain β 16 real episodes | raw_robot |
deploy first |
real_only |
16 real episodes, no sim | raw_robot |
baseline: what sim pretraining adds |
real_finetune_from_sim_v1_limiteddata |
200-demo sim pretrain β 16 real episodes | raw_robot |
ablation: amount of sim data |
sim_v1_pretrain |
1800 sim demos | sim_ee_frame |
zero-shot sim-to-real baseline |
sim_v1_limiteddata_pretrain |
200 sim demos | sim_ee_frame |
zero-shot ablation |
real_video_from_sim_v1 |
sim_v1 pretrain + real camera video only | sim_ee_frame |
video-only adaptation (no real labels) |
Why real_finetune_from_sim_v1 first: on held-out real data it is as accurate as the other
real-trained checkpoints, and it is the one whose predicted descent matches the operator's
(-0.063 vs -0.059 sim units over 16 steps; the others over-descend by 20-35%). Its sim
pretraining also covers a wider workspace than the 16 real episodes. This is an open-loop
judgement; the robot has the final word (DEPLOYMENT.md, gate 5).
Each directory has model.pt (selected by act_l1_h16 on validation, the default) and
model_final.pt (last epoch), except sim_v1_limiteddata_pretrain, where the selected epoch is
the last one.
Repository layout
AGENTS.md instructions for the agent/operator on the robot computer
deploy/ shared by every checkpoint
βββ DEPLOYMENT.md the procedure: gates 0-5, interface, evaluation protocol
βββ lilawam_deploy.py robot <-> policy adapter; every convention lives here
βββ lilawam_policy.py loads a checkpoint directory, predicts action chunks
βββ check_deployment.py gate 1: offline self-test (8 checks), nothing moves
βββ check_live_log.py gates 2-5: checks a step log recorded on the robot
βββ robot_loop_template.py control loop with shadow/execute modes; implement RobotInterface
βββ fixture_observations.npz 16 real lab frames: raw 640x480 cameras, raw robot readings, operator actions
βββ model/ inference code (models/, utils/, robotwin_infer.py)
βββ dinov3-vitl16-pretrain-lvd1689m/ frozen vision encoder (1.2 GB)
lilawam/stack_v1/<checkpoint>/
βββ model.pt, model_final.pt action-model weights (inference only; optimizer state not included)
βββ deploy.json obs_frame, action conventions, training state range, sha256, lineage
βββ config.yaml model/dataset config (paths relative to the bundle)
βββ norm_stats.json the normalisation the checkpoint was trained with
βββ split.json train/val episode split
βββ fixture_reference.npz this machine's predictions on the fixture, per weight file
βββ eval_real_holdout.json open-loop scores on held-out real episodes, both frames
tools/ how this release was built and verified (not needed on the robot)
The observation-frame difference
The real recordings store the robot's readings unconverted: the panda_hand origin and the
panda_link8 flange orientation. The simulator records the ee_frame (0.1034 m further along the
hand's z) and the hand orientation (flange Γ Rz(-45Β°)). A checkpoint that learned from real state
must be fed raw readings (raw_robot); one that only saw sim state must be fed converted readings
(sim_ee_frame). Getting it wrong raises no error β the stack_v0 real deployment failed this way
(the policy would not descend). lilawam_deploy applies the right frame from deploy.json,
check_deployment.py and check_live_log.py verify it, and the wrong-frame control below shows
the effect on held-out data.
Held-out real results (open-loop)
Predict at every frame of the 4 held-out real episodes (IDs 0, 1, 15, 17; 906 windows with a full
16-step future) and compare with the operator's next actions, in the policy's convention (sim units,
+1 = open). Seed 0, no smoothing, the exact deployment pipeline. This is not robot success.
Caveats: the real checkpoints selected model.pt on these same episodes, and the dataset release
notes say these recordings were also used for scene calibration.
| checkpoint | weights | epoch | obs_frame | dir. cos | transl. L1 h16 | gripper acc | dz sign h16 | net dz h16 pred / truth [sim units] |
|---|---|---|---|---|---|---|---|---|
real_finetune_from_sim_v1 |
model.pt |
160 | raw_robot | +0.329 | 0.0063 | 0.954 | 0.757 | -0.0632 / -0.0593 |
real_finetune_from_sim_v1 |
model_final.pt |
200 | raw_robot | +0.337 | 0.0064 | 0.956 | 0.764 | -0.0635 / -0.0593 |
real_finetune_from_sim_v1_limiteddata |
model.pt |
90 | raw_robot | +0.355 | 0.0063 | 0.966 | 0.780 | -0.0733 / -0.0593 |
real_finetune_from_sim_v1_limiteddata |
model_final.pt |
200 | raw_robot | +0.352 | 0.0061 | 0.967 | 0.791 | -0.0802 / -0.0593 |
real_only |
model.pt |
560 | raw_robot | +0.335 | 0.0058 | 0.965 | 0.770 | -0.0734 / -0.0593 |
real_only |
model_final.pt |
600 | raw_robot | +0.333 | 0.0058 | 0.965 | 0.776 | -0.0718 / -0.0593 |
sim_v1_pretrain |
model.pt |
110 | sim_ee_frame | +0.190 | 0.0090 | 0.944 | 0.639 | -0.0433 / -0.0593 |
sim_v1_pretrain |
model_final.pt |
120 | sim_ee_frame | +0.202 | 0.0089 | 0.939 | 0.661 | -0.0540 / -0.0593 |
sim_v1_limiteddata_pretrain |
model.pt |
120 | sim_ee_frame | +0.092 | 0.0095 | 0.945 | 0.595 | -0.0235 / -0.0593 |
real_video_from_sim_v1 |
model.pt |
43 | sim_ee_frame | +0.196 | 0.0091 | 0.944 | 0.648 | -0.0396 / -0.0593 |
real_video_from_sim_v1 |
model_final.pt |
50 | sim_ee_frame | +0.198 | 0.0091 | 0.945 | 0.648 | -0.0402 / -0.0593 |
dir. cos: cosine between predicted and operator translation at the executed step. transl. L1 h16: mean |error| over the 16 executed steps. gripper acc: executed-step sign agreement. dz sign / net dz: summed vertical motion over 16 steps β whether the policy descends when the operator did, and by how much. 1 sim unit β 0.186 m of tool motion per step.
Findings: the three real-trained checkpoints are close to each other and clearly ahead of the sim-only ones (direction cosine ~0.33-0.36 vs ~0.19). Video-only adaptation leaves the sim policy essentially unchanged (0.196 vs 0.190), consistent with the stack_v0 finding that its gain lands in the discarded video predictor. Using the full 1800 sim demos instead of 200 roughly doubles zero-shot direction agreement.
Wrong-frame control (selected weights, same windows): what happens if a checkpoint is fed the other frame.
| checkpoint | frame | dir. cos | transl. L1 h16 | gripper acc | net dz h16 pred / truth |
|---|---|---|---|---|---|
real_finetune_from_sim_v1 |
sim_ee_frame (wrong) | +0.309 | 0.0060 | 0.962 | -0.0144 / -0.0593 |
real_finetune_from_sim_v1 |
raw_robot (deployed) | +0.329 | 0.0063 | 0.954 | -0.0632 / -0.0593 |
real_finetune_from_sim_v1_limiteddata |
sim_ee_frame (wrong) | +0.313 | 0.0063 | 0.959 | -0.0780 / -0.0593 |
real_finetune_from_sim_v1_limiteddata |
raw_robot (deployed) | +0.355 | 0.0063 | 0.966 | -0.0733 / -0.0593 |
real_only |
sim_ee_frame (wrong) | +0.266 | 0.0057 | 0.959 | -0.0881 / -0.0593 |
real_only |
raw_robot (deployed) | +0.335 | 0.0058 | 0.965 | -0.0734 / -0.0593 |
sim_v1_pretrain |
sim_ee_frame (deployed) | +0.190 | 0.0090 | 0.944 | -0.0433 / -0.0593 |
sim_v1_pretrain |
raw_robot (wrong) | -0.008 | 0.0085 | 0.524 | -0.1299 / -0.0593 |
sim_v1_limiteddata_pretrain |
sim_ee_frame (deployed) | +0.092 | 0.0095 | 0.945 | -0.0235 / -0.0593 |
sim_v1_limiteddata_pretrain |
raw_robot (wrong) | +0.068 | 0.0085 | 0.526 | -0.0408 / -0.0593 |
real_video_from_sim_v1 |
sim_ee_frame (deployed) | +0.196 | 0.0091 | 0.944 | -0.0396 / -0.0593 |
real_video_from_sim_v1 |
raw_robot (wrong) | -0.001 | 0.0085 | 0.526 | -0.1265 / -0.0593 |
For sim-only checkpoints the wrong frame destroys direction and drops the gripper to chance. For
real-trained checkpoints it is subtler β direction degrades a little and descent goes wrong (for
real_finetune_from_sim_v1, a quarter of the operator's descent) β which is why the frame is
checked explicitly rather than trusted to show up as obviously bad behaviour.
Verification done before this release
On an H200 with the files as published:
deploy/check_deployment.py --all: 11/11 weight files pass 8/8 checks (files and sha256, load, raw-frame pipeline reproduces the recorded observations exactly, frame inside training range, reference predictions, operator agreement and gripper sign, latency, adapter + step log).- The adapter's
raw_robotstate reproduces the real training pack bit for bit on the 16 fixture frames; itssim_ee_framestate reproduces the stack_v0 deployment reference (< 1e-5). - Fault injection (
tools/fault_injection.py, results intools/fault_injection_results.txt): real held-out episode 0 replayed through the adapter and policy, forreal_finetune_from_sim_v1andsim_v1_pretrain. The clean replay gives no FAIL or WARN. Every injected mistake is caught: tool/TCP-frame readings, (w, x, y, z) quaternion order and a frozen camera -> FAIL; gripper width in mm -> refused by the adapter; BGR frames and swapped cameras -> WARN (exit 2, which blocks the next gate until a person acknowledges it).
What changed from the previous release
The stack_v0 bundles (policy_stack_v0, policy_stack_v0_fulldata, runs/stack_real_v0_*,
archive/policy_stack_v0_ep6.pt) were removed; they remain in this repo's history at revision
cba7827.
Their deployment knowledge β camera crops, the ee_frame offset and flange roll, the gripper sign,
the control loop, calibration and evaluation protocol β is carried into deploy/, with two
corrections: the observation frame is now per checkpoint (the v0 adapter converted every reading
to the sim frame, although the v0 real-trained runs had been trained on raw readings), and the
action gain is the v1-measured 0.1861 (v0 used 0.192).