Status
Started 2026-06-22, still reading.
What it’s about
World Action Models (WAMs) are embodied predictive-action models: they forecast the future in a form that can be turned into action. The survey argues the field has gotten blurry — broad world models, video generation models, action-grounded video world models, VLA policies, and WAMs all overlap — and tries to give everyone a shared vocabulary.
How it organizes the field (two views)
- What must be generated: rendered futures vs. latent futures vs. video-generation-free action reasoning.
- How a method is built: decomposed by predictive substrate, backbone, action coupling, and deployment regime.
The takeaway I want to remember
WAMs aren’t just video generators with an action head. Their design choices trade representational richness against compute, memory, latency, and action-label cost — and the trend is “dream less, act more”: generate less of the future, preserve what control requires.
My notes (to fill in)
-