← All paper notes
Read June 29, 2026

World Models for Robotic Control — Paper Triage & Notes

Tags
roboticsworld-modelsjepadiffusionin-context-learningstudy-notes
Reading Session — Paper Triage ← All paper notes

Reading Session · Triage

Today’s Papers — Classified

Ten papers grouped into four themes. For each we’ll work through abstract, method, conclusion, limitations, and related work.

Session 2026-06-29  ·  10 papers  ·  4 groups

World models & robot control4 papers

Learning or using models of world dynamics for prediction, planning, and control.

LLM safety, agents & reasoning2 papers

Whether “reasoning” and agentic tool use actually behave safely.

Post-training, generalization & adaptation2 papers

How training and adaptation choices reshape in-domain vs. out-of-domain generalization.

Interactive & agentic multimodal generation2 papers

Foundation models that interact or generate across modalities with streaming or planning.

World models & robot control

4 · sub-split: generative/visual vs. latent-predictive (JEPA)

All four learn or exploit world/dynamics models — two are generative (predict pixels/meshes), two are latent-predictive (predict in feature space).

P4

PhysiFormer: Learning to Simulate Mechanics in World Space

Chen, Lan, Vedaldi · Visual Geometry Group, Oxford

A diffusion transformer that simulates physically-plausible 3D object motion as meshes in world coordinates — casting vertex-trajectory prediction as one denoising process, with no hand-coded rigidity/causality priors.

AbstractMethodConclusionLimitationsRelated work
world modeldiffusion3D / physics
P4 PhysiFormer · Abstract

1 · Pixel-space video world models vs. coordinate-space meshes — a new paradigm

Most video world models predict the future as pixels in video frames — view-dependent pixel space, tied to a particular camera. That conflates an object’s physics with its rendering: change the camera and the prediction changes. PhysiFormer instead represents objects as 3D meshes (vertices) in world coordinates, modeling the physical state directly, independent of any view. This is a new paradigm the paper calls coordinate-space diffusion — a step toward view-invariant, geometry-aware world modeling rather than 2D pixel prediction.

2 · Neural physics, “ad-hoc latent spaces,” and dropping inductive biases

Neural physics approaches are learned simulators (e.g. Graph Network Simulators, MeshGraphNets) that predict how a physical system evolves. Prior work tends to do two things PhysiFormer avoids: (1) roll out dynamics inside an ad-hoc latent space — a bespoke, method-specific feature space rather than interpretable world coordinates; and (2) explicitly enforce rigidity and causality as hard-coded inductive biases. PhysiFormer’s claim is that none of that is necessary: casting vertex-trajectory prediction as one denoising diffusion process directly in world coordinates already works — and beats autoregressive baselines. Note the “single” process: it denoises the whole trajectory at once rather than rolling out step-by-step, which is why it avoids autoregressive error accumulation.

3 · Factorised attention & permutation invariance

Factorised attention means the model does not run one giant attention over all (time × space × object) tokens at once; it factorises into separate attentions along each axis — time, space (vertices), and objects — for efficiency, mixing across axes over stacked layers (axial / factorised spatiotemporal attention). Permutation-invariant means the output does not depend on the ordering of objects: relabel object 1 ↔ object 2 and the result is unchanged. Objects have no canonical order, and attention is a set operation, so attending over the object axis without positional encoding gives permutation-invariant multi-object reasoning — without assigning each object an explicit ID or slot.

4 · How the diffusion process works (what is learnable)

The object being generated is a vertex trajectory $x_0$ — future 3D vertex positions over time, shape $\approx [T \times N \times 3]$ in world coordinates.

Forward (noising), no learnable parameters — add Gaussian noise on a fixed schedule:

$$x_k = \sqrt{\bar\alpha_k}\,x_0 + \sqrt{1-\bar\alpha_k}\,\epsilon,\qquad \epsilon\sim\mathcal{N}(0, I)$$

Reverse (denoising), learnable — the PhysiFormer transformer $\epsilon_\theta$ predicts the noise from the noisy trajectory $x_k$, the diffusion step $k$, and the conditioning $c$; iterate from noise back to a clean trajectory. The training objective is the standard diffusion MSE:

$$\mathcal{L} = \mathbb{E}_{x_0,\epsilon,k}\big\lVert \epsilon - \epsilon_\theta(x_k, k, c)\big\rVert^2$$

Conditioning $c$ = initial vertex positions + velocities and material type (rigid / elastic); these are inputs, not generated. Learnable parameters are the transformer weights: vertex embeddings, the factorised time/space/object attention layers, the timestep-$k$ embedding, the material and initial-state encoders, and the output head. The noise schedule is typically fixed. At inference you start from a noise trajectory and run $K$ reverse steps (DDPM) or fewer (DDIM); different seeds give diverse plausible futures, matching the abstract’s probabilistic framing.

x_K ~ N(0, I) pure-noise trajectory conditioning c initial vertex pos + vel material: rigid / elastic (input, not generated) PhysiFormer ε_θ(x_k, k, c) diffusion transformer (denoiser) time space objects attention factorised → permutation-invariant (no explicit object encoding) reverse diffusion ×K x_0 (clean) vertex trajectory [T × N × 3], world coords
PhysiFormer as a conditional diffusion model: a fixed forward process noises the vertex trajectory; the learnable transformer $\epsilon_\theta$ denoises it over $K$ reverse steps, conditioned on initial state and material. Factorised attention over time/space/objects keeps it efficient and permutation-invariant.

Architecture specifics — the exact prediction parameterization, noise schedule, and how $c$ is injected — are paper internals to confirm when we read the Method section.

P4 PhysiFormer · Method
PhysiFormer architecture, Figure 2 (highlighted)
Figure 2: PhysiFormer architecture, with the reader’s highlights.

Refines the abstract-level guesses: the generative framework is flow matching (not vanilla DDPM); hidden size $D=1024$; conditioning is additive; the backbone is a factorised DiT-L with 16 register tokens.

Markovian

The Markov property means the future depends only on the present state, not the full history (memoryless). Two senses are relevant:

  • Physics is Markovian in (position, velocity). Newtonian dynamics is a second-order ODE, i.e. first-order in the state $(x, v)$: given initial positions $X_0$ and velocities $V_0$, the future trajectory is determined. This is why PhysiFormer conditions on both first-frame position and velocity — a single frame of positions cannot reveal velocity, so $(x, v)$ together form the Markov state.
  • The flow / diffusion process is a Markov chain. Each denoising step depends only on the current noisy sample, not earlier steps.

What “second-order ODE” means. An ODE relates a function to its own derivatives, and its order is the highest derivative that appears. Newton’s law $F = m\ddot{x}$ involves acceleration $\ddot{x}$ (a second derivative of position), so it is second-order. A second-order ODE needs two initial conditions for a unique solution: the initial position and the initial velocity. Position alone cannot tell whether the object is moving, so the future is not determined by position by itself. Writing $v=\dot{x}$ turns it into a first-order system in the state $(x,v)$:

$$\dot{x} = v, \qquad \dot{v} = \tfrac{F}{m} = a.$$

The rate of change of $(x,v)$ then depends only on the current $(x,v)$, which is exactly the Markov property, so $(X_0, V_0)$ determines the whole trajectory. Analogy: for a thrown ball, position alone says nothing about where it goes next, but position together with velocity fixes the entire arc.

Inferred from the figure; if the Method text uses “Markovian” in a specific sentence, paste it and I’ll pin the exact meaning.

Architecture walkthrough (left → right in Fig. 2)

  1. Input. Mesh vertex trajectory: N vertices over T frames, $\mathbb{R}^{T\times N\times 3}$.
  2. x_embed. A linear embedder lifts each 3-D coordinate to hidden size $D=1024$, giving vertex tokens $\mathbb{R}^{T\times N\times D}$.
  3. Flow-matching noising. Vertex tokens are noised along the flow-matching schedule (at inference, start from noise).
  4. Conditioning (additive). To each noised vertex token, add the sum of: a material embedding, the first-frame position embedding x_embed_cond(X₀), and the first-frame velocity embedding v_embed(V₀).
  5. Register tokens. 16 global register tokens are prepended to aggregate global context (learnable “scratch” slots, from the ViT-registers idea).
  6. Factorised DiT-L backbone (×6). Each block runs four attention layers along different axes, each with its own RoPE; the token tensor is reshaped between them. 4 attentions × 6 = 24 layers.
  7. Linear head. Final tokens are projected back to $\mathbb{R}^{T\times N\times 3}$ (clean vertex coordinates).
  8. Render. Iterative denoising yields clean vertex trajectories, assembled into triangle meshes with the provided topology, then rendered from any viewpoint (view-invariant), arbitrary material — “4D” = 3D geometry + time.

Factorised attention (the heart of the backbone)

Instead of one attention over all $T\times N$ tokens, the tensor is reshaped so attention runs along one axis at a time:

BlockReshape (batch, seq, D)Attends overRoPE
Spatial Attention(T, N, D)vertices within a frameSpatial
Temporal Attention(N, T, D)time, per vertexTemporal
Object Attention(TK, N/K, D)grouped per-object vertices—
Temporal Attention(N, T, D)time, per vertexTemporal

Spatial attention mixes geometry within a frame; temporal attention propagates motion across time per vertex; object attention partitions the N vertices into K groups for object-level structure. Together this gives efficient, permutation-invariant multi-object reasoning without explicit object encoding. DiT = Diffusion Transformer (transformer backbone for diffusion, replacing the U-Net); RoPE = rotary position embedding (relative position by rotating Q/K), applied per axis.

Key settings

  • Hidden $D = 1024$; backbone factorised DiT-L, $4\times 6$ layers.
  • 16 prepended global register tokens.
  • Generative framework: flow matching.
  • Conditioning: first-frame position + velocity (Markov state) + material, added to each token.
  • Output: vertex trajectories → triangle meshes via fixed topology → view-invariant, arbitrary-material 4D rendering.
P4 PhysiFormer · Conclusion, Limitations & Brainstorm

Correction — what [T × N × 3] means

The three axes are not positions / velocities / material. $T$ = number of timesteps, $N$ = number of mesh vertices, and $3$ = the $(x, y, z)$ coordinates of each vertex. One vertex at one timestep is a 3-vector. Velocities and material type are conditioning, not part of this tensor — initial velocities are given alongside initial positions, and the material label is mapped to its own embedding vector. The model generates positions; it is given the initial state and material.

Does the 100k simulated trajectories explain the success?

Partly, but the framing matters. The autoregressive baselines are presumably trained on the same data, so data scale alone cannot explain the gap. Data makes strong performance possible; the joint single-diffusion formulation (generate the whole trajectory at once, not step-by-step) and the world-coordinate + factorised-attention design are what make it better than equally-trained baselines. Also, “predicts better in one step” is slightly off for diffusion: inference still runs $K$ reverse steps — what is “single” is joint trajectory generation, so the real contrast is joint generation vs. autoregressive rollout (which avoids accumulated error).

What “generalization” means here (§4.6 & conclusion)

Read against the limitations, the OOD tested is mostly extended values of known axes — mixed materials, unseen geometries, larger object counts. This is OOD in the value distribution, not the discovery of new physical factors (fluids, fracture, friction regimes, articulation are not represented). Fixed trajectory length and fixed mesh resolution are further hard limits of the representation. Tellingly, the admitted artifacts — spurious contacts, interpenetration, rare orientation discontinuities — show that with a generic diffusion objective and no collision/consistency constraints, the model learns the statistics of seen trajectories, not the underlying physics of contact.

Is the dataset the bottleneck for other domains?

It is a major one, but not alone. Data coverage and the objective/representation are co-bottlenecks: a simulator that never shows fluids or articulation caps what can be represented, but even with perfect data a constraint-free objective will still occasionally interpenetrate — the paper itself points the fix at the objective (contact-focused training, physical-consistency losses), not only at more data.

Is the data from video? Could we build it from multi-scene video for manipulation?

Almost certainly not from video: it is trained on simulated trajectories, and exact world-space vertex tracks plus material labels come from a physics simulator — you cannot read those off raw RGB (video gives view-dependent pixels, which this method deliberately avoids).

Building a video-derived, scene-factorised version for robot manipulation inverts the setup, which is exactly where the difficulty lives: PhysiFormer consumes clean 3D meshes in world coordinates, so a real-video version must first recover them — per-object 3D reconstruction, mesh tracking, world-frame registration, and material estimation across scenes. The binding difficulty is the perception / system-identification layer, not the dynamics model. That “factorise across scenes” idea is really about disentangling scene-invariant dynamics from scene-specific appearance — closer to object-centric video models and to ICWM (P10, system identification as in-context adaptation) than to PhysiFormer itself.

Where it could genuinely help manipulation: if object meshes/states are available (depth sensors, multi-view capture, or a sim digital-twin), a coordinate-space, view-invariant, permutation-invariant dynamics predictor is an attractive world model for planning — predict how objects move under candidate actions, robust to camera changes, in cluttered multi-object scenes. The realistic bridge is sim-to-real plus a 3D-perception front-end that lifts real video/depth into the mesh/state representation, with contact-aware objectives — connecting directly to ICWM (P10) and the world-model-for-control framing in P7 / P9.

Limitations are grounded in the conclusion text provided; the §4.6 generalization specifics (exact axes and numbers) can be confirmed from the PDF when available.

P4 PhysiFormer · Related Work

Five lines of work; PhysiFormer positions itself against all of them. The reader’s highlighted page is embedded below.

PhysiFormer related work, highlighted page
The PhysiFormer Related Work page, with the reader’s highlights.

1 · Traditional physical simulation

  • Motion & material modelling: equations of how objects move (dynamics) plus constitutive models of how materials deform/respond (elastic, plastic, viscoelastic; cloth, thin shells, fluids).
  • Contact resolution: detecting and resolving collisions/contacts — non-penetration, contact forces, friction, impacts, stacking.
  • Time integration: stepping state (positions, velocities) forward in time (explicit / implicit / symplectic), trading stability vs accuracy vs energy/momentum conservation.

Physically grounded and realistic, but expensive, complex, and hard to generalize.

2 · Per-scene optimized physical dynamics

  • NeRF: implicit scene (an MLP maps position + view direction to color + density; images via volume rendering). 3D Gaussian Splatting: explicit scene (a cloud of 3D Gaussians, real-time rasterization).
  • Inverse-physics optimization: from observed video motion, optimize physical parameters (stiffness, mass, forces, initial velocities) so a (often differentiable) simulator reproduces it.
  • Why “overfitting” is emphasized: these fit one specific scene/video and do not generalize without re-optimizing — the opposite of a feed-forward model trained across many trajectories.
  • “Strong simulator assumptions”: built-in priors (spring-mass models, simplified contact) cap fidelity; they also often need dense multi-view capture or tracking supervision.

NeRF in depth — implementation, current work, applications

Implementation. A NeRF is an MLP $F_\Theta(\mathbf{x},\mathbf{d}) \to (\mathbf{c},\sigma)$ from a 3D point $\mathbf{x}$ and view direction $\mathbf{d}$ to color $\mathbf{c}$ and density $\sigma$. Inputs are positionally encoded (Fourier features) so the MLP fits high-frequency detail. A pixel is rendered by casting a camera ray, sampling points along it, querying the MLP, and volume-rendering (alpha-compositing color weighted by density). Training is per-scene: a photometric loss between rendered and real pixels over posed images; geometry emerges implicitly from the density field.

Current work. Speed/quality (Instant-NGP hash grids, Mip-NeRF 360, TensoRF, Plenoxels); surfaces (NeuS, VolSDF); dynamic/4D (D-NeRF, Nerfies, HyperNeRF); generalizable / feed-forward (pixelNeRF, Large Reconstruction Models); generative and text-to-3D (EG3D, DreamFusion). The big shift: 3D Gaussian Splatting (2023) has largely overtaken NeRF for real-time rendering, and feed-forward reconstruction (no per-scene optimization) is the current trend — exactly the “overfitting” limitation PhysiFormer points at.

Applications. Novel-view synthesis / free-viewpoint video; 3D reconstruction and digitization; AR/VR, VFX and content creation; digital humans/avatars; robotics (scene representation and dense SLAM — iMAP, NICE-SLAM); autonomous driving (reconstruction and simulation); medical imaging; text-to-3D; and, most relevant here, as a scene representation coupled with physics (the per-scene optimized physical dynamics line).

3 · Learning-based physics simulation

CNNs on regular grids; GNNs for particle simulation; MeshGraphNets / FIGNet / HopNet / HCMT for mesh dynamics. Mesh-aware biases help but need extra geometric machinery (explicit connectivity, topology preprocessing, hierarchy, learned shape representations). PhysiFormer instead diffuses raw 3D vertex trajectories, multi-object and multi-material, without specialized contact modeling.

4 · Autoregressive prediction in VFM feature space

  • Deterministic AR outputs a single future (no uncertainty, blurs under ambiguity, accumulates error); probabilistic models (diffusion) sample diverse plausible futures. PhysiFormer is probabilistic.
  • VFM pipeline: encode context frames with a frozen Vision Foundation Model (e.g. DINO) → predict future latent features → decode to depth/segmentation. Example: DINO-Foresight (4 frames → latent futures → depth/seg).
  • Advantage over pixel-space: semantic, lower-dimensional, stabler to predict; one latent decodes to several tasks; less blur. (DINO-Foresight is the deterministic-AR version of “predict in latent, not pixels”; our brainstorm was the probabilistic-diffusion version.)

5 · Diffusion models for world simulation

Video Diffusion Models simulate convincingly but often violate Newtonian dynamics and are view-dependent in pixel space; diffusion over 3D representations is mostly single, static objects. PhysiFormer’s gap: multi-object, dynamic, view-invariant diffusion directly in 3D coordinate space.

★ Cross-paper brainstorm — P4 × P9 × LeWM

Action-Prefix-Conditioned Latent-Trajectory Diffusion

A JEPA world model that denoises whole latent rollouts, conditioned on action prefixes.

The shared enemy: autoregressive rollout

PhysiFormer and the LeWM line are two reactions to the same failure mode — predicting the future one step at a time, which accumulates error over the horizon and is slow (each step is sequential). PhysiFormer attacks it on the generative side (denoise the entire vertex trajectory jointly, in world coordinates); Fast-LeWM attacks it on the latent-predictive side (replace repeated one-step latent rollout with action-prefix prediction, computing prefix-reached latents in parallel). So “don’t roll out autoregressively” isn’t merely transferable to LeWM — Fast-LeWM is already one instantiation of it. A diffusion world-model and a JEPA world-model converging on the same anti-autoregressive idea, from opposite directions, is the real signal.

Where it turns novel — fuse the three ingredients

  1. denoise an entire latent trajectory jointly (PhysiFormer’s joint diffusion),
  2. in JEPA’s reconstruction-free latent space (LeWM),
  3. conditioned on action prefixes (Fast-LeWM’s prediction unit).
Thesis: a world model that, given the current latent and a candidate action sequence, samples the whole future latent trajectory in one (few-step) diffusion pass, conditioned on action prefixes — joint generation (no accumulated error) + JEPA latent (no perception ground-truth) + prefix conditioning (multi-horizon action effects).

Precedent in spirit: trajectory-level diffusion planners — Diffuser and Decision Diffuser — diffuse over state-action trajectories. The new twist is doing it over JEPA latents with prefix conditioning, not raw states.

Easy borrows (low tension, architecture-level)

Factorised attention over time/space/objects and permutation-invariant multi-object handling are orthogonal to the JEPA-vs-diffusion question; they drop cleanly into any LeWM handling multiple objects or structured latents.

Synergy with the earlier P4 brainstorm

We found PhysiFormer’s bottleneck for real-robot video is perception / system-ID — it needs clean meshes in world coordinates. JEPA/LeWM sidesteps exactly that (reconstruction-free, learns its own latent from observations, no ground-truth geometry), so for manipulation-from-video a LeWM-style latent is the more practical substrate than explicit meshes — a concrete argument for pushing the idea toward LeWM rather than PhysiFormer.

Tensions to respect

  • JEPA deliberately avoids generative modeling; diffusion reintroduces a generative objective — but diffusing a small latent is far cheaper than pixel diffusion, and modeling latent-dynamics uncertainty is reasonable, so not fatal.
  • Diffusion’s $K$ reverse steps fight Fast-LeWM’s speed goal — use few-step / consistency-distilled diffusion to keep planning fast.
  • Making the latent “geometry-aware” (view-invariance) means injecting 3D/object structure JEPA’s minimalism avoids — possible (object-centric / 3D-grounded JEPA), but a tradeoff.

Next step: when we read P9, check whether Fast-LeWM’s prefix mechanism already covers part of this (deterministic prefix prediction) or leaves clear room for the probabilistic diffusion variant.

References (DOIs) PhysiFormer (P4) — Chen, Lan, Vedaldi · arXiv:2606.27364 · DOI 10.48550/arXiv.2606.27364
Fast-LeWorldModel (P9) — Gao, Xu · arXiv:2606.26217 · DOI 10.48550/arXiv.2606.26217
LeWorldModel (LeWM) — Maes, Le Lidec, Scieur, LeCun, Balestriero · arXiv:2603.19312 · DOI 10.48550/arXiv.2603.19312
Diffuser — Janner et al. 2022 · arXiv:2205.09991 · DOI 10.48550/arXiv.2205.09991
Decision Diffuser — Ajay et al. 2023 · arXiv:2211.15657 · DOI 10.48550/arXiv.2211.15657
I-JEPA — Assran et al. 2023 · arXiv:2301.08243 · DOI 10.48550/arXiv.2301.08243
P7

World Action Models Enable Continual Imitation Learning with Recurrent Generative Replays

Govind, Reilly, Patel, Le, Das · UNC Charlotte · cs.RO

REGEN uses a World Action Model’s generative ability to synthesize pseudo-replay trajectories, letting a robot policy rehearse old tasks without storing original demonstrations — cutting catastrophic forgetting.

AbstractMethodConclusionLimitationsRelated work
world modelcontinual learningimitation / robot
P9

Fast LeWorldModel

Gao, Xu · Xi’an Jiaotong University · arXiv:2606.26217 · cs.LG

A JEPA-style latent world model that replaces repeated one-step latent rollout with action-prefix prediction, modeling effects over multiple horizons in parallel — faster planning, slower-growing latent error.

AbstractMethodConclusionLimitationsRelated work
world modelJEPAplanning
P9 Fast-LeWM · Method & key details
Fast-LeWM training pipeline, Figure 2 (highlighted)
Figure 2: Fast-LeWM training pipeline, with the reader’s highlights.

Confirms the earlier P4×P9 brainstorm: Fast-LeWM is exactly the deterministic action-prefix + parallel-predictor mechanism; our idea adds probabilistic diffusion on top. Dense loss: $\frac{1}{H}\sum_{k=1}^{H}\lVert \hat z_{t+k}-z_{t+k}\rVert_2^2 + \lambda\,\mathrm{SIGReg}(\mathcal Z)$, where SIGReg is LeWM/LeJEPA’s collapse regularizer.

Why prepend a state token, then discard the 0-th output

An action prefix alone doesn’t fix the outcome: the same open-loop actions lead to different results depending on the current scene. So the prefix encoding is conditioned on the current latent $z_t$, mapped through a lightweight MLP into a state token prepended as the 0-th token. Under the causal mask, the output at action position $k$ attends to tokens $0..k$ (state token + first $k$ actions), so it summarizes the prefix $(a_t,\dots,a_{t+k-1})$ conditioned on $z_t$: that is $p_{t,k}=E_\psi^{(k)}(a_t,\dots,a_{t+k-1}\mid z_t)$ (eq. 14). One forward pass yields all prefixes $p_{t,1:H}$. The 0-th output is discarded because position 0 sees no actions (a length-0 prefix), i.e. just $z_t$, which is already known.

CEM planning (Cross-Entropy Method)

A sampling-based, derivative-free optimizer for planning/MPC: (1) keep a Gaussian over action sequences; (2) sample N candidates; (3) roll out the world model and score each against the goal; (4) keep the top-K elites; (5) refit the Gaussian to the elites; (6) repeat a few iterations; (7) execute the first action, then replan. The model is queried N×iterations×H times, so dynamics-evaluation dominates the cost — exactly what Fast-LeWM cuts (one prefix pass + parallel prediction per candidate instead of H sequential steps).

Green-highlighted terms

  • Sinusoidal positional encodings (sine/cosine at exponentially spaced frequencies): the original Transformer position code; wavelengths form a geometric progression, low dims encode fine position and high dims coarse. Tells the encoder the order of state/action tokens.
  • AdaLN-zero modulation: Adaptive LayerNorm, zero-initialized (from DiT). Adaptive means the LayerNorm scale $\gamma$ and shift $\beta$ are not fixed constants but are predicted from the conditioning (here the prefix representation / $z_t$), so the normalization adapts per input (this is FiLM-style modulation). A small MLP emits three things: scale (a conditioning-dependent gain on each feature), shift (a conditioning-dependent bias), and a residual gate $\alpha$ that scales the sublayer output before it rejoins the residual stream, $x \leftarrow x + \alpha\,\mathrm{Sublayer}(\mathrm{AdaLN}(x))$. Zero-init sets $\alpha=0$, so each block starts as identity, which stabilizes training. This is how the predictor modulates $z_t$ with the action prefix.
  • “10 epochs, matching the LeWM protocol”: trained 10 epochs to match LeWM exactly, for a fair, controlled comparison.
  • “same image resolution and latent dimensionality”: one dynamics evaluation costs an encode + a predict, and that compute is set by two knobs — image resolution drives the encoder FLOPs, latent dim drives the predictor FLOPs. Every task is run at the same resolution and the same latent dim, so the per-evaluation cost is nearly identical across tasks. Only the scene / goal / difficulty changes, which affects how many evaluations CEM needs, not the cost per evaluation. Hence the Two-Room timing represents all tasks.

Conditioning: FiLM / AdaLN vs cross-attention

AdaLN is the standard way to inject conditioning into a transformer / MLP without concatenating it — it is FiLM (feature-wise linear modulation) applied to normalized features.

“Inject conditioning” means making the computation depend on an external signal $c$ (class, timestep, action prefix). FiLM does it by turning $c$ into per-feature scale/shift, $\hat x\cdot\gamma(c)+\beta(c)$: a cheap affine modulation, no extra tokens, no attention. Cross-attention instead lets the main tokens (queries) attend to a set of conditioning tokens (keys/values).

AspectAdaLN / FiLMCross-attention
Mechanismcondition → MLP → per-feature scale/shift/gate on normalized featuresmain tokens (Q) attend to condition tokens (K, V)
Conditioning shapeone global vector (class, timestep, prefix summary)a set / sequence of tokens (variable length)
Granularityuniform — same $\gamma,\beta$ for every tokenper-token, content-dependent routing
Costcheap: one MLP + elementwise, ~$O(d)$costly: $O(n\,m\,d)$ attention + params
Best forcompact global conditionrich / long / structured condition
ExamplesDiT, StyleGAN (AdaIN), Fast-LeWMStable Diffusion (text), Flamingo, enc–dec

Intuition: FiLM/AdaLN is a broadcast (one global condition reshapes the whole feature space uniformly — cheap, stable); cross-attention is routing (each token selectively reads different parts of a structured condition — expressive, costly). Fast-LeWM summarizes the prefix into a compact representation, so global AdaLN suffices and stays fast; a language instruction would lean toward cross-attention.

FLOPs & latency metrics

  • FLOPs (count): floating-point operations for one forward pass (a matmul $(n\times k)(k\times m)\approx 2nkm$). Hardware-agnostic measure of compute work. Trap: FLOPs = a count per inference; FLOPS = operations per second, a hardware rate.
  • MACs: multiply-accumulates, $\approx$ FLOPs$/2$.
  • Latency: measured wall-clock time per inference on given hardware — what this paper reports (31.4s → 8.0s) and what users feel.
  • Throughput: items/second under batching ($\approx$ inverse latency for big batches).
  • #Parameters / peak memory (VRAM): model size and footprint — storage proxies, not directly speed.

FLOPs ≠ latency: a low-FLOP model can be slow if memory-bound or poorly parallelized. That is Fast-LeWM’s story exactly — FLOPs per candidate are similar, but replacing $H$ sequential predictor steps with one parallel pass collapses the wall-clock latency.

Does 10 epochs converge? And how do LeWM and Fast-LeWM relate?

The 10-epoch budget is LeWM’s protocol, adopted for fair comparison, not a claim of optimality; convergence depends on dataset size and LR schedule, and is judged by the loss plateauing plus downstream planning, not zero loss. The “same resolution/latent” sentence is across tasks, not Fast-LeWM vs LeWM. Fast-LeWM is a drop-in replacement of LeWM’s dynamics module: same encoder and latent space, matched parameters (17.9M vs 18.0M), same protocol, all other hyperparameters following LeWM — the only change is autoregressive one-step rollout → action-prefix encoder + parallel predictor + dense multi-horizon loss. This isolates the effect of the prefix/parallel design.

Planning horizon & action skip

  • Planning horizon $H$: the look-ahead length — how many future steps are predicted/optimized. Here the training horizon is set adaptively per trajectory and clamped to $[1,5]$.
  • Action skip (action repeat / frame skip): holding the same action for $k$ environment steps so the agent decides less often, cutting compute and lengthening effective temporal reach. (Not in the shown excerpts; standard meaning.)

Key settings

  • State / action-prefix token dim 192; action-prefix Transformer: 3 layers, 6 heads, head dim 32; sinusoidal PE.
  • Predictor: 6-layer action-modulated residual MLP, latent dim 192, hidden 2048, fusion 768, AdaLN-zero, dropout 0.1; 17.9M params.
  • Planning cost vs LeWM: dynamics-eval 31.4s → 8.0s; CEM solve 54.4s → 28.3s. History size 1 (no visual history needed).
P9 Fast-LeWM · Highlighted source excerpts
Eq 14 and the prepend-state-token highlight
Action-prefix conditioning, eq. (14): the current latent $z_t$ is mapped through a lightweight MLP into a state token and prepended as the 0-th token.
Hyperparameters and training details, highlighted
Hyperparameters and training: token dim 192, 3-layer / 6-head prefix Transformer, sinusoidal PE, 6-layer action-modulated residual MLP with AdaLN-zero, 17.9M params, 10 epochs matching LeWM.
Planning-cost results, highlighted
Planning cost: dynamics-evaluation 31.4s → 8.0s, full CEM solve 54.4s → 28.3s; per-task dynamics-eval costs are uniform because all tasks share resolution and latent dim.

★ Research direction — P9 (and P4)

Open-box physics probing & force factorization for latent world models

Beyond latent MSE: test what physics the world model actually learned.

The gap

Fast-LeWM reports latent MSE (the dense prefix loss) and planning success — black-box, aggregate metrics that show latents are predicted accurately but not what physics the model internalized. Contrast P4 PhysiFormer, which measured momentum-consistency and rigidity-preservation: physics-aware metrics worth borrowing.

Directions

  1. Open-box probing: train linear probes on the latents $z$ to recover physical quantities (velocity, force magnitude/direction, mass, friction, contact). Recoverability = the model internalized them.
  2. Force factorization: expose force-like intermediate variables in the predictor, or analyze the AdaLN modulation to decompose the predicted latent change into components (applied force vs gravity vs contact); test with counterfactual action perturbations.
  3. More intuitive metrics: physical-consistency on rollouts (momentum/energy, penetration, contact) plus probe-recoverability — an open-box benchmark for latent world models.

An “interpretability + physics” axis on top of the speed/accuracy axis, spanning LeWM / Fast-LeWM / the latent-diffusion variant. Connects to P4 (physics metrics) and the earlier P4×P9 brainstorm.

Idea seed (not in the paper). Could become a deeper-research project: a physics-probing benchmark + a force-factorized predictor.

★ NEW IDEA Brainstorm

Self-organized modulation (core-periphery)

Observation. AdaLN’s $\gamma,\beta$ are hetero-conditioned (“他律”): an external condition $c$ (another modality, the action prefix, a timestep) is pushed through a learned MLP that decides, top-down, which feature dimensions to amplify or suppress. Core-periphery (CP) coreness is the opposite, self-organized (“自律”): a node’s coreness $c_i$ emerges bottom-up from the group’s internal relations (e.g. its mean similarity to the others) — the set itself decides who is core and who is periphery.

Idea. Add a self-organized CP signal alongside the hetero-conditioned modulation: from the token set (prefix tokens, vertex/object tokens, or latent dims) compute a differentiable coreness from mutual similarity, and use it to gate / weight tokens. The prefix says what to do; CP says which parts of the state matter for it. Schematically, $\text{mod} = \text{AdaLN}(c)\,(\text{hetero}) \;\odot\; \text{coreness}(\{z\})\,(\text{self})$.

What it could solve in this scenario.

  • Adaptive compute: full predictor on core (dynamically important) tokens, cheap-path the periphery — serves Fast-LeWM’s speed goal.
  • Open-box interpretability: emergent coreness reveals which vertices / objects / latent dims drive the predicted dynamics — a free probe signal, feeding the force-factorization agenda above.
  • Contact & multi-object reasoning: contacts are where dynamics couple; CP could surface contact-relevant nodes emergently, replacing PhysiFormer’s fixed $N/K$ object grouping or hand-designed contact modeling.
  • Richer inductive bias than isotropy: SIGReg pushes the latent toward isotropic Gaussian (all dims equal); a CP-structured latent (core carries dynamics, periphery regularized) is an expressive alternative to weigh against it.

Open questions: a cheap, differentiable coreness; whether self + hetero beats hetero alone empirically; and the tension with SIGReg’s isotropy. Janelle’s core-periphery background is the natural lever here.

P10

In-Context World Modeling for Robotic Control

Wang, Shi, Fei, Fu, Ji, Gong, Qiu · Fudan / OpenMOSS · arXiv:2606.26025 · cs.RO

ICWM treats system identification as in-context adaptation: a policy infers essential system variables from a short history of self-generated, task-agnostic interactions, adapting to novel configurations without parameter updates.

AbstractMethodConclusionLimitationsRelated work
world modelin-contextVLA / robot
P10 ICWM · Abstract / core idea
ICWM abstract, highlighted
ICWM core paragraph, with the reader’s highlights.

Core contribution (two coupled halves)

  • Method: repurpose the context window for system identification, not behavior specification. Before a task, the robot runs short random exploratory moves, records the visual transitions, and prepends these self-probed clips as context; the model implicitly recovers the current system configuration (viewpoint, geometry, dynamics) and adapts its actions.
  • Surprising finding (the red block): task-agnostic random movements alone suffice for implicit system identification, with no task-specific demonstrations, and beat standard VLA training on novel viewpoints in simulation and on real robots.

So the red “Surprisingly…” sentence is the headline result; the cyan/yellow part is the method. They are inseparable.

System identification from random movements?

Yes. System identification = inferring the hidden configuration/parameters (viewpoint, layout, dynamics, calibration) from action→observation responses. Classically one designs informative excitation; ICWM shows even random moves expose how the scene responds, and an in-context model reads the configuration off those transitions implicitly. The probe is the excitation; the model infers the latent config and adapts. It is especially strong for novel viewpoints, because the probe clip shows how the current camera sees the scene. (Same flavor as inferring a Markov state from observed transitions.)

In-context learning vs test-time adaptation

In-context learning (ICL)Test-time adaptation (TTA)
Adapt byconditioning on context, no weight updateslightweight parameter / statistic updates on test data
Origin / methodsGPT-3 few-shot prompting; in-context RL (Algorithm Distillation, Decision-Pretrained Transformer)TENT, Test-Time Training (TTT), BN-stat adaptation, CoTTA / EATA / MEMO / SHOT
Signalexamples / probe clips placed in the inputunlabeled test inputs + a self-supervised objective

ICWM is ICL: it reaches the TTA goal (handle novel viewpoints) by conditioning alone — test-time adaptation without test-time training.

On-the-fly skill acquisition

“On-the-fly” = at runtime, with no separate training phase: acquiring or adapting a skill immediately at deployment from a little in-situ experience, instead of the slow train-then-deploy loop. In ICWM, when the robot meets a new configuration it does a quick random probe, conditions on it, and adapts in real time — on-the-fly adaptation via in-context conditioning.

LLM safety, agents & reasoning

2

Both probe whether “deliberation” or autonomous tool choice actually yields safer behavior.

P1

Do Thinking Tokens Help with Safety?

Ri, Panigrahi, Arora · Princeton Language & Intelligence

Evidence that reasoning models’ safety isn’t very deliberative: the refusal/compliance outcome is already predictable from the first token’s hidden state (0.84–0.95 AUROC) before any visible thinking; thinking resembles prefix completion more than revision.

AbstractMethodConclusionLimitationsRelated work
safetyreasoninginterpretability
P2

When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents

Yang, Bu, Yi, Wang, Zhou, Dai, Hu, Yang · CAS / BAAI / CUHK / PKU

Studies over-privileged tool selection — agents escalating to higher-privilege tools when a lower-privilege one suffices — via ToolPrivBench, measuring initial choice and escalation after transient tool failures.

AbstractMethodConclusionLimitationsRelated work
agent safetytool usebenchmark

Post-training, generalization & adaptation

2

Both ask how training/adaptation choices reshape ID vs. OOD performance over time.

P3

How Post-Training Shapes Biological Reasoning Models

Fesser, Zhang, Li, Wang, Perozzi, Azizi, Kakade, Zitnik · Harvard / Google DeepMind

Trains 100+ biological reasoning models under controlled CPT / SFT / RL to show each post-training stage reshapes ID vs. OOD generalization differently — gains are non-monotonic; stage composition matters.

AbstractMethodConclusionLimitationsRelated work
post-trainingID / OODscientific reasoning
P5

Distill Once, Adapt Life-Long: Dataset Distillation for Continual Test-Time Adaptation

Jang, Kim, Kweon, Yoon · KAIST / Chung-Ang · arXiv:2606.20196 · cs.CV

DO-ALL uses dataset distillation to keep a compact, privacy-conscious set of source anchors; each target sample is matched to its closest anchor for stable continual test-time adaptation (source replay, representation alignment, manifold smoothing).

AbstractMethodConclusionLimitationsRelated work
test-time adaptationdataset distillationcontinual learning

Interactive & agentic multimodal generation

2

Foundation models that interact or generate across modalities — one streaming, one agentic.

P6

Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models

Wan Team, Alibaba · arXiv:2606.25041 · cs.CV

A native-streaming, end-to-end interactive foundation model unifying language/audio/video in one Transformer with block-causal attention — ~200 ms model-side latency for full-duplex audio-visual interaction, no separate VAD/ASR/TTS modules.

AbstractMethodConclusionLimitationsRelated work
multimodal FMstreamingblock-causal attention
P8

Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation

Z. Zhang, Li, J. Zhang, Gao, Yan … Wu · Qwen

Frames real-world T2I failures as a “context gap” and closes it with an agentic loop — context-aware planning + grounding via reason/search/memory/feedback; introduces IA-Bench for plan/reason/search/memory.

AbstractMethodConclusionLimitationsRelated work
image generationagentbenchmark

Cross-cutting threads (for horizontal comparison)

World models → P4, P7, P9, P10  ·  Continual learning → P5, P7

Agents → P2, P8  ·  Safety → P1, P2  ·  Robotics / VLA → P7, P10  ·  Diffusion → P4