RLVR-World
updated
RLVR-World: Training World Models with Reinforcement Learning
Paper
• 2505.13934
• Published • 16
thuml/rt1-frame-tokenizer
0.1B • Updated • 35
thuml/rt1-world-model-single-step-base
0.1B • Updated • 13
thuml/rt1-world-model-single-step-rlvr
0.1B • Updated • 23
thuml/rt1-compressive-tokenizer
0.1B • Updated • 23
thuml/rt1-world-model-multi-step-base
0.1B • Updated • 13
thuml/rt1-world-model-multi-step-rlvr
0.1B • Updated • 42
thuml/webarena-world-model-cot
Viewer
• Updated • 6.41k • 152
thuml/webarena-world-model-sft
2B • Updated • 14
thuml/webarena-world-model-rlvr
2B • Updated • 29
• 1
thuml/bytesized32-world-model-cot
Viewer
• Updated • 304k • 288
• 3
thuml/bytesized32-world-model-sft
2B • Updated • 12
thuml/bytesized32-world-model-rlvr-binary-reward
2B • Updated • 12
thuml/bytesized32-world-model-rlvr-task-specific-reward
2B • Updated • 10