Core idea
JEPA-style world models like LeWM plan via expensive one-step autoregressive latent rollouts, which accumulate error and are slow over long horizons. Fast-LeWM trains with prefix-level supervision: encode a whole action prefix and directly predict the latent reached after executing it, at multiple horizons in parallel — at planning time only the final prefix token is needed to score a future latent, no intermediate traversal. Reports higher average success, much lower planning time, and open-loop latent loss that grows far slower with horizon.
(Direct follow-up to my LeWM experiments — worth reading closely.)