Core idea
VLAs fail to generalize to novel execution contexts (new viewpoints, new robot bodies) because they condition only on the current observation and instruction, implicitly assuming the training setup. ICWM runs a short history of self-generated, task-agnostic interactions through the model’s context window — in-context demonstrations of system properties rather than of the task — letting it implicitly learn the current dynamics and adapt without weight updates. Reported to beat standard VLA baselines on unseen viewpoints in sim and on real robots.