You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

srlwam β€” stack_v1 policies

LiLa-WAM and Cosmos Edge policies for two-cube stacking on a Franka Panda, trained on the stack_v1 simulation corpus and on real teleoperated recordings, packaged for deployment on the real robot. Training data: dataset repo aabyaneh/srlwam (revision 4ee4768c, stack/sim/stack_v1 and stack/real/stack_v0).

Deploying? Read AGENTS.md first, then follow deploy/DEPLOYMENT.md gate by gate. The checkpoints do not all expect the same robot state; the deploy code handles it from each checkpoint's deploy.json, and the sanity checks verify it before the arm moves.

Cosmos Edge stack_v2

Cosmos3 Edge checkpoints trained on the stack_v2 simulation data, six training runs (two fine-tune on 50% / 20% of the real train episodes). Each run is released as its final checkpoint and its best checkpoint (lowest held-out normalized action MAE, open-loop, evaluated every 250 steps; one bundle when they coincide). Shared assets are the same cosmos_edge/shared/ used by stack_v1. Sim-unit checkpoints (sim_v2_*, real_video_*) use gain 0.1975 m/unit: use command.delta, not command.real_controller_delta.

run checkpoint role observation frame dir. cosine (step 0) translation L1 (h16) gripper acc. (step 0)
sim (stack_v2 sim from Cosmos3-Edge-Policy-DROID) sim_v2_droid_iter1750 best (MAE 0.0593) sim_ee_frame +0.273 0.0059 0.914
sim (stack_v2 sim from Cosmos3-Edge-Policy-DROID) sim_v2_droid_iter2000 final (MAE 0.0624) sim_ee_frame +0.280 0.0058 0.923
real supervised from sim real_supervised_from_sim_v2 best = final (MAE 0.1966) raw_robot +0.342 0.0058 0.966
real supervised from sim, 50% real data real_supervised_from_sim_v2_real50 best = final (MAE 0.1867) raw_robot +0.364 0.0055 0.969
real supervised from sim, 20% real data real_supervised_from_sim_v2_real20_iter2000 best (MAE 0.2194) raw_robot +0.376 0.0063 0.967
real supervised from sim, 20% real data real_supervised_from_sim_v2_real20_iter4000 final (MAE 0.2257) raw_robot +0.365 0.0061 0.967
real only from DROID real_only_from_droid_iter3500 best (MAE 0.1967) raw_robot +0.367 0.0059 0.966
real only from DROID real_only_from_droid_iter4000 final (MAE 0.2007) raw_robot +0.356 0.0059 0.964
real video from sim real_video_from_sim_v2_iter2500 best (MAE 0.3143) sim_ee_frame +0.389 0.0051 0.947
real video from sim real_video_from_sim_v2_iter4000 final (MAE 0.3218) sim_ee_frame +0.394 0.0051 0.951

Held-out results use the same 906 real windows (episodes 0, 1, 15, 17) and seeds as stack_v1. They are recorded-action agreement scores, not closed-loop robot success rates.

Cosmos Edge stack_v1

Five Cosmos3 Edge checkpoints are packaged with shared processor, vision-encoder and Wan2.2 VAE assets. Start with real_supervised_from_sim_v1 for supervised real adaptation. Follow Cosmos deployment instructions for a local GPU or the provided remote server/client. Each checkpoint is about 6.74 GB; shared assets add about 3.83 GB. The complete Cosmos bundle is about 37.5 GB.

checkpoint observation frame adaptation
real_supervised_from_sim_v1 raw_robot real stack recordings (16 train episodes), joint action + video objective, 2000 steps, from sim_v1_droid; sim normalization
real_video_from_sim_v1_iter2000 sim_ee_frame as iter1250, end of the 2000-step schedule; lowest aggregate held-out action MAE of the video runs
real_video_from_sim_v1_iter1250 sim_ee_frame sim_v1_droid's generation expert adapted on real camera video only (no real state or actions), step 1250 = lowest held-out video flow loss; action heads frozen
sim_v1_droid sim_ee_frame stack_v1 sim, 1800 train demos, 2000 steps x global batch 2048, from Cosmos3-Edge-Policy-DROID
sim_v1_limiteddata_droid sim_ee_frame stack_v1 sim, 200-demo subset (seed 1701), same 2000-step recipe, own normalization

Deployment and timing

Cosmos predicts 16 future actions and video frames jointly. The adapter executes the 16-action chunk at 10 Hz, holds position while inference runs, and then replans. Measured H200 inference is approximately 1.2-1.4 s per chunk with 30 sampling steps; peak model GPU allocation is approximately 7.7 GiB. This is not a continuous 10 Hz inference service. Budget additional GPU memory for the runtime and other processes. Use gate 1 to measure the deployment machine.

Held-out real results (open-loop)

Predictions use 906 windows from real episodes 0, 1, 15 and 17, with a full 16-step future. Sampling uses 30 steps and seed 1701 + window index. Values below use each checkpoint's declared observation frame. These are recorded-action agreement scores, not closed-loop robot success rates. These real recordings also informed scene calibration; the step-1250 variant was selected using held-out video loss, so this is not an untouched test set.

checkpoint direction cosine (step 0) translation L1 (h16) gripper accuracy (step 0) dz sign agreement (h16) net dz pred / truth (sim units)
real_supervised_from_sim_v1 +0.363 0.0064 0.965 0.777 -0.0698 / -0.0593
real_video_from_sim_v1_iter2000 +0.376 0.0072 0.953 0.807 -0.1158 / -0.0593
real_video_from_sim_v1_iter1250 +0.369 0.0072 0.953 0.807 -0.1075 / -0.0593
sim_v1_droid +0.163 0.0103 0.945 0.575 -0.0476 / -0.0593
sim_v1_limiteddata_droid +0.165 0.0109 0.949 0.746 -0.1175 / -0.0593

The supervised checkpoint has translation L1 about 0.0064 and predicted net descent about -0.070 versus -0.059 in the recordings. Video-only adaptation improves direction agreement relative to the sim-only checkpoint, but predicts substantially more downward motion (about -0.108 to -0.116 sim units). The limited-data sim baseline also overpredicts descent (about -0.117). Start with shadow mode and the staged motion checks; do not interpret a higher direction score as proof of better robot control.

Each checkpoint's eval_real_holdout.json contains both observation frames, including the wrong-frame control. The exported bf16 weights approximately reproduce the original full-precision evaluations; small numerical changes can flip gripper signs near decision boundaries. Fixture checks compare the exported deployment weights against their own fixed-seed references.

Release verification

All five Cosmos checkpoints passed 8/8 offline checks with an empty Hugging Face cache. The isolated robot-side HTTP client passed 8/8 checks, and the wrong-checkpoint handshake was rejected. Both fault-injection replays passed the clean baseline and flagged the six injected mistakes. All 11 existing LiLa-WAM weight variants passed their 8/8 regression checks. These are offline deployment checks; no robot was commanded. See verification report and check output.

Files and provenance

cosmos_edge/stack_v1/<checkpoint>/ contains model/ (including hidden .metadata), deploy.json, checkpoint.json, normalization.json, split.json, reference predictions, and evaluation results. deploy/cosmos/ contains the loader, HTTP server/client, offline checker, setup instructions, and robot dependency pins. Use this loader with the pinned NVIDIA framework; these are distributed checkpoints, not Transformers AutoModel bundles.

See component licenses and Cosmos notices.

Which LiLa-WAM checkpoint

checkpoint trained on robot state it expects (obs_frame) use it for
real_finetune_from_sim_v1 sim_v1 pretrain β†’ 16 real episodes raw_robot deploy first
real_only 16 real episodes, no sim raw_robot baseline: what sim pretraining adds
real_finetune_from_sim_v1_limiteddata 200-demo sim pretrain β†’ 16 real episodes raw_robot ablation: amount of sim data
sim_v1_pretrain 1800 sim demos sim_ee_frame zero-shot sim-to-real baseline
sim_v1_limiteddata_pretrain 200 sim demos sim_ee_frame zero-shot ablation
real_video_from_sim_v1 sim_v1 pretrain + real camera video only sim_ee_frame video-only adaptation (no real labels)

Why real_finetune_from_sim_v1 first: on held-out real data it is as accurate as the other real-trained checkpoints, and it is the one whose predicted descent matches the operator's (-0.063 vs -0.059 sim units over 16 steps; the others over-descend by 20-35%). Its sim pretraining also covers a wider workspace than the 16 real episodes. This is an open-loop judgement; the robot has the final word (DEPLOYMENT.md, gate 5).

Each directory has model.pt (selected by act_l1_h16 on validation, the default) and model_final.pt (last epoch), except sim_v1_limiteddata_pretrain, where the selected epoch is the last one.

Repository layout

AGENTS.md                         instructions for the agent/operator on the robot computer
deploy/                           shared by every checkpoint
β”œβ”€β”€ DEPLOYMENT.md                 the procedure: gates 0-5, interface, evaluation protocol
β”œβ”€β”€ lilawam_deploy.py             robot <-> policy adapter; every convention lives here
β”œβ”€β”€ lilawam_policy.py             loads a checkpoint directory, predicts action chunks
β”œβ”€β”€ check_deployment.py           gate 1: offline self-test (8 checks), nothing moves
β”œβ”€β”€ check_live_log.py             gates 2-5: checks a step log recorded on the robot
β”œβ”€β”€ robot_loop_template.py        control loop with shadow/execute modes; implement RobotInterface
β”œβ”€β”€ fixture_observations.npz      16 real lab frames: raw 640x480 cameras, raw robot readings, operator actions
β”œβ”€β”€ model/                        inference code (models/, utils/, robotwin_infer.py)
└── dinov3-vitl16-pretrain-lvd1689m/   frozen vision encoder (1.2 GB)
lilawam/stack_v1/<checkpoint>/
β”œβ”€β”€ model.pt, model_final.pt      action-model weights (inference only; optimizer state not included)
β”œβ”€β”€ deploy.json                   obs_frame, action conventions, training state range, sha256, lineage
β”œβ”€β”€ config.yaml                   model/dataset config (paths relative to the bundle)
β”œβ”€β”€ norm_stats.json               the normalisation the checkpoint was trained with
β”œβ”€β”€ split.json                    train/val episode split
β”œβ”€β”€ fixture_reference.npz         this machine's predictions on the fixture, per weight file
└── eval_real_holdout.json        open-loop scores on held-out real episodes, both frames
tools/                            how this release was built and verified (not needed on the robot)

The observation-frame difference

The real recordings store the robot's readings unconverted: the panda_hand origin and the panda_link8 flange orientation. The simulator records the ee_frame (0.1034 m further along the hand's z) and the hand orientation (flange Γ— Rz(-45Β°)). A checkpoint that learned from real state must be fed raw readings (raw_robot); one that only saw sim state must be fed converted readings (sim_ee_frame). Getting it wrong raises no error β€” the stack_v0 real deployment failed this way (the policy would not descend). lilawam_deploy applies the right frame from deploy.json, check_deployment.py and check_live_log.py verify it, and the wrong-frame control below shows the effect on held-out data.

Held-out real results (open-loop)

Predict at every frame of the 4 held-out real episodes (IDs 0, 1, 15, 17; 906 windows with a full 16-step future) and compare with the operator's next actions, in the policy's convention (sim units, +1 = open). Seed 0, no smoothing, the exact deployment pipeline. This is not robot success. Caveats: the real checkpoints selected model.pt on these same episodes, and the dataset release notes say these recordings were also used for scene calibration.

checkpoint weights epoch obs_frame dir. cos transl. L1 h16 gripper acc dz sign h16 net dz h16 pred / truth [sim units]
real_finetune_from_sim_v1 model.pt 160 raw_robot +0.329 0.0063 0.954 0.757 -0.0632 / -0.0593
real_finetune_from_sim_v1 model_final.pt 200 raw_robot +0.337 0.0064 0.956 0.764 -0.0635 / -0.0593
real_finetune_from_sim_v1_limiteddata model.pt 90 raw_robot +0.355 0.0063 0.966 0.780 -0.0733 / -0.0593
real_finetune_from_sim_v1_limiteddata model_final.pt 200 raw_robot +0.352 0.0061 0.967 0.791 -0.0802 / -0.0593
real_only model.pt 560 raw_robot +0.335 0.0058 0.965 0.770 -0.0734 / -0.0593
real_only model_final.pt 600 raw_robot +0.333 0.0058 0.965 0.776 -0.0718 / -0.0593
sim_v1_pretrain model.pt 110 sim_ee_frame +0.190 0.0090 0.944 0.639 -0.0433 / -0.0593
sim_v1_pretrain model_final.pt 120 sim_ee_frame +0.202 0.0089 0.939 0.661 -0.0540 / -0.0593
sim_v1_limiteddata_pretrain model.pt 120 sim_ee_frame +0.092 0.0095 0.945 0.595 -0.0235 / -0.0593
real_video_from_sim_v1 model.pt 43 sim_ee_frame +0.196 0.0091 0.944 0.648 -0.0396 / -0.0593
real_video_from_sim_v1 model_final.pt 50 sim_ee_frame +0.198 0.0091 0.945 0.648 -0.0402 / -0.0593

dir. cos: cosine between predicted and operator translation at the executed step. transl. L1 h16: mean |error| over the 16 executed steps. gripper acc: executed-step sign agreement. dz sign / net dz: summed vertical motion over 16 steps β€” whether the policy descends when the operator did, and by how much. 1 sim unit β‰ˆ 0.186 m of tool motion per step.

Findings: the three real-trained checkpoints are close to each other and clearly ahead of the sim-only ones (direction cosine ~0.33-0.36 vs ~0.19). Video-only adaptation leaves the sim policy essentially unchanged (0.196 vs 0.190), consistent with the stack_v0 finding that its gain lands in the discarded video predictor. Using the full 1800 sim demos instead of 200 roughly doubles zero-shot direction agreement.

Wrong-frame control (selected weights, same windows): what happens if a checkpoint is fed the other frame.

checkpoint frame dir. cos transl. L1 h16 gripper acc net dz h16 pred / truth
real_finetune_from_sim_v1 sim_ee_frame (wrong) +0.309 0.0060 0.962 -0.0144 / -0.0593
real_finetune_from_sim_v1 raw_robot (deployed) +0.329 0.0063 0.954 -0.0632 / -0.0593
real_finetune_from_sim_v1_limiteddata sim_ee_frame (wrong) +0.313 0.0063 0.959 -0.0780 / -0.0593
real_finetune_from_sim_v1_limiteddata raw_robot (deployed) +0.355 0.0063 0.966 -0.0733 / -0.0593
real_only sim_ee_frame (wrong) +0.266 0.0057 0.959 -0.0881 / -0.0593
real_only raw_robot (deployed) +0.335 0.0058 0.965 -0.0734 / -0.0593
sim_v1_pretrain sim_ee_frame (deployed) +0.190 0.0090 0.944 -0.0433 / -0.0593
sim_v1_pretrain raw_robot (wrong) -0.008 0.0085 0.524 -0.1299 / -0.0593
sim_v1_limiteddata_pretrain sim_ee_frame (deployed) +0.092 0.0095 0.945 -0.0235 / -0.0593
sim_v1_limiteddata_pretrain raw_robot (wrong) +0.068 0.0085 0.526 -0.0408 / -0.0593
real_video_from_sim_v1 sim_ee_frame (deployed) +0.196 0.0091 0.944 -0.0396 / -0.0593
real_video_from_sim_v1 raw_robot (wrong) -0.001 0.0085 0.526 -0.1265 / -0.0593

For sim-only checkpoints the wrong frame destroys direction and drops the gripper to chance. For real-trained checkpoints it is subtler β€” direction degrades a little and descent goes wrong (for real_finetune_from_sim_v1, a quarter of the operator's descent) β€” which is why the frame is checked explicitly rather than trusted to show up as obviously bad behaviour.

Verification done before this release

On an H200 with the files as published:

  • deploy/check_deployment.py --all: 11/11 weight files pass 8/8 checks (files and sha256, load, raw-frame pipeline reproduces the recorded observations exactly, frame inside training range, reference predictions, operator agreement and gripper sign, latency, adapter + step log).
  • The adapter's raw_robot state reproduces the real training pack bit for bit on the 16 fixture frames; its sim_ee_frame state reproduces the stack_v0 deployment reference (< 1e-5).
  • Fault injection (tools/fault_injection.py, results in tools/fault_injection_results.txt): real held-out episode 0 replayed through the adapter and policy, for real_finetune_from_sim_v1 and sim_v1_pretrain. The clean replay gives no FAIL or WARN. Every injected mistake is caught: tool/TCP-frame readings, (w, x, y, z) quaternion order and a frozen camera -> FAIL; gripper width in mm -> refused by the adapter; BGR frames and swapped cameras -> WARN (exit 2, which blocks the next gate until a person acknowledges it).

What changed from the previous release

The stack_v0 bundles (policy_stack_v0, policy_stack_v0_fulldata, runs/stack_real_v0_*, archive/policy_stack_v0_ep6.pt) were removed; they remain in this repo's history at revision cba7827. Their deployment knowledge β€” camera crops, the ee_frame offset and flange roll, the gripper sign, the control loop, calibration and evaluation protocol β€” is carried into deploy/, with two corrections: the observation frame is now per checkpoint (the v0 adapter converted every reading to the sim frame, although the v0 real-trained runs had been trained on raw readings), and the action gain is the v1-measured 0.1861 (v0 used 0.192).

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading