Papers
arxiv:2609.33748

AnyStep-WAM: Budget-Aligned Distillation and Adaptive Inference for World Action Models

Published on Sep 27
· Submitted by
Rui Wang
on Sep 30
Authors:
,
,
,
,
,
,

Abstract

World-action models (WAMs) couple predictive visual modeling with action generation, typically relying on iterative denoising with a fixed denoising steps. However, manipulation tasks contain actions chunks with varying sensitivity to generation errors: critical actions require precision, while less sensitive actions allow faster generation with fewer denoising steps. Here we introduce AnyStep World Action Model, a general framework for tunable-budget prediction and scene-dependent computation allocation. Our budget-aligned teacher-trajectory distillation trains interval-conditioned flow maps using explicit frozen-teacher transitions and shared low-rank adapters, supporting action generation from one-step prediction to multi-step refinement. Building on this capability, a lightweight risk-benefit scheduler predicts teacher-curvature-based difficulty and budget-specific student-teacher fidelity from a single one-step preview, selecting the smallest budget predicted to satisfy risk-adaptive fidelity requirements. We evaluate our framework on three widely used WAMs Motus, FastWAM, and LingBotVA using RoboTwin 2.0. Our method reduces average denoising steps by 60.2%, 49.8%, and 85.28%, respectively, while maintaining baseline task success rates. In particular, our AnyStep training substantially improves model performance under a one-step denoising budget, increasing task success rates by 7.07%, 12.08%, and 8.94% on Motus, FastWAM, and LingBotVA, respectively. Experiments on six real-world manipulation tasks further validate its effectiveness.

Community

Paper submitter

World-action models (WAMs) typically use fixed-step denoising, despite varying precision requirements across manipulation stages. We introduce AnyStep WAM, a framework for cross-budget prediction and scene-dependent computation allocation. Budget-aligned teacher-trajectory distillation with shared low-rank adapters enables action generation from one-step prediction to multi-step refinement. A lightweight risk-benefit scheduler predicts difficulty and budget-specific student–teacher fidelity from a single preview, selecting the smallest budget predicted to meet adaptive fidelity requirements. On RoboTwin 2.0, AnyStep reduces average denoising steps by 60.2%, 49.8%, and 85.28% on Motus, FastWAM, and LingBotVA, respectively, while maintaining baseline success rates. It also improves one-step success rates by 7.07, 12.08, and 8.94 percentage points, respectively. Experiments on six real-world manipulation tasks further validate its effectiveness.

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.33748
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.33748 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.33748 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.33748 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.