Reading Notes · AI
Paper Reading Notes
A running log of questions and answers, one section per paper. Each entry keeps the paper's abstract and key passages on record, followed by the Q&A as it accumulates.
Industry Pulse · 2026-07-06
The AI landscape continues expanding across multiple frontiers: multimodal models like Qwythos-9B and GLM-5.2 are gaining traction for complex reasoning tasks, while specialized models for OCR and tabular data classification signal deepening vertical application.
Major labs are advancing practical capabilities — OpenAI's ChatGPT adoption continues climbing, Google is scaling foundation models for zero-shot tabular tasks and heat resilience prediction, and Microsoft Research is exploring trainable agent skills and memory architectures.
Developer tooling is democratizing as open-source projects gain momentum, particularly video-aware LLM systems (claude-real-video), local inference guides, and GPU worker networks (Talos) that enable distributed inference.
The sector is simultaneously wrestling with timing pressures (Zuckerberg noting slower-than-expected agent progress) and ethical questions (rich families using AI tutors, contentious AI commercials). Red teaming platforms and security-focused tools are accelerating in response to the need for safer AI deployment at scale.
Referenced projects & models · verified links
- GLM-5.2 open weights · MIT Zhipu / Z.ai — github.com/zai-org/GLM-5 · weights: huggingface.co/zai-org/GLM-5.2
- Qwythos-9B open weights · Apache-2.0 Empero AI — huggingface.co/empero-ai/Qwythos-9B-Claude-Mythos-5-1M · empero.org (weights live on Hugging Face, not a single GitHub repo)
- claude-real-video open source · MIT github.com/HUANGCHIHHUNGLeo/claude-real-video (core free version; a paid "crv Pro" tier also exists)
- Talos open source · MPL-2.0 Likely Talos Linux (Sidero Labs) — github.com/siderolabs/talos; the immutable Kubernetes OS commonly used to build GPU-worker clusters for distributed inference
Notes: “local inference guides” is a category (llama.cpp / Ollama / vLLM tutorials), not one repo. The OCR / tabular models, OpenAI ChatGPT, and the Google / Microsoft Research efforts named above are products, labs, or unnamed categories — no single canonical repo to link. Links verified via web search on 2026-07-06.
Concepts & Background · from the Pulse
Local inference and the .cpp projects — is this hardware deployment?
Local inference means running a model on your own device instead of calling a cloud API. The cpp in projects like llama.cpp and whisper.cpp — and Embodied.cpp from the 07-05 digest — is C++: they re-implement the inference engine in C/C++ (instead of Python + PyTorch) to run faster, use less memory, drop heavy dependencies, and work across hardware — plain CPUs, Apple Silicon, various GPUs, even phones or a Raspberry Pi. The key trick is quantization (compressing weights from FP16 down to ~4-bit, the GGUF format) so a large model fits in consumer VRAM.
So it is closely related to hardware deployment, but more precisely it is a runtime for running models efficiently on local / edge hardware. “Local inference guides” are tutorials on installing, choosing quantization, sizing VRAM, and picking a backend (llama.cpp / Ollama / vLLM) — a category of content, not one piece of software. Embodied.cpp applies the same idea to embodied models running on robots.
Distributed vs decentralized inference — GPU scheduling, or moving compute onto users?
These two words point at two different layers:
- Distributed inference. A model too big for one GPU is split across several that cooperate on a single forward pass (tensor / pipeline / expert parallelism). ELDR operates at this layer, routing decode requests. This usually happens inside one data center with fast interconnects (NVLink / InfiniBand); the focus really is distributing across GPUs plus scheduling / routing.
- Decentralized inference. Emphasizes no single owner — compute is crowd-sourced across many people's machines (geographically spread, over the public internet). This is the “move compute onto users” idea: volunteers plug in their own GPUs to earn crypto (e.g. a GPU marketplace like Infernet Protocol; Talos Linux stitches such workers into clusters). It only suits loosely-coupled jobs (models that fit on one card, batch embeddings, LoRA fine-tunes); the hard parts are slow networking, unreliable nodes, result verification, and privacy.
In one line: cluster-internal parallelism ≈ distributed; compute moved onto users' machines ≈ decentralized. There is also a third, related mode — on-device / hybrid offload (part on the phone, part in the cloud; or, like Embodied.cpp, directly on the robot) — driven by latency, privacy, or cost.
Video-aware vs image-aware — what does it add, and what does it solve?
Image-aware models see a single frame: recognize objects, OCR, read charts, describe a scene — what is missing is the time dimension. Video-aware adds time, enabling things that only work across frames: motion / action (waving, falling, speed), temporal & causal order (“first… then…”, “what happens at 0:30”), tracking the same target across frames, plus editing rhythm and audio timelines.
The core thing it solves: a single frame cannot show a process — an image tells you what, a video tells you how it changes. The main engineering challenge is information explosion (30 frames per second, each treated as an image → token blow-up), which is why tools like claude-real-video keep only the frames that actually change (scene-change + dedup) instead of sampling uniformly. Notably this is the same philosophy as PixelEyes in the spatial dimension (don't feed every pixel, only look at the important region) — one saves along time, the other across space.
Could a continuous medical-imaging tracker work?
Technically feasible — and continuous / temporal scenarios are common in medical imaging:
- Ultrasound — real-time tracking of valve motion, blood flow, probe guidance.
- Endoscopy / colonoscopy — tracking a polyp across frames in continuous video (avoiding double-counting or misses).
- Surgical video — instrument and anatomy tracking, surgical-phase recognition.
- Follow-up of fundus, skin lesions, or pathology over time.
These map onto video-aware strengths: consistent cross-frame tracking (rather than per-frame detection that makes IDs jump), temporal smoothing to suppress single-frame noise / artifacts, and detecting change over time (a lesion growing, abnormal motion).
Some honest constraints:
- Don't make the LLM the doctor. A sounder architecture: a dedicated vision model does perception (video segmentation / tracking, e.g. SAM 2 or a specialized medical model) and the VLM/LLM does the summarizing reasoning — exactly PixelEyes' decouple perception from reasoning. A “sample frames, feed a general LLM” approach (like claude-real-video) is usually not precise enough for fine medical localization.
- Real-time. Clinical continuous tracking needs low latency and a stable frame rate — closer to Embodied.cpp's closed-loop, edge, low-latency runtime than a cloud request-response service.
- Privacy & regulation. Medical images are sensitive → local / on-device inference fits better; diagnostic use typically needs clinical validation and regulatory approval (FDA / CE, etc.). A research prototype or assistive tool is low-barrier; clinical decision-making is high-barrier.
I'm not a physician; this is an engineering / research feasibility discussion, not medical advice.
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning
★ KeyTianjin University · Alibaba
Project page: anitaleungxx.github.io/MIPU · methods: MIPI (principle) / MIPU (framework)
Key Passages · from the paper
Gist · paraphrased from the abstract
RL for LLM post-training is fragile — it can destabilize or collapse. A key cause is the training-inference mismatch: LLMs use separate engines for generation (inference, e.g. vLLM/SGLang) and for gradient computation (training, e.g. FSDP/Megatron), so the same trajectory gets different probabilities on the two sides even with synchronized parameters — a built-in off-policyness that quietly poisons training.
Prior work mostly reduces or stabilizes the mismatch on the training side. This paper flags a deeper objective misalignment: improving the training policy π does not guarantee improving the inference policy μ — the one actually deployed. They propose MIPI (Monotonic Inference Policy Improvement), a principle defined on μ, and MIPU, a two-step realization: Step 1 makes a sampler-referenced candidate update; Step 2 accepts or rolls back the synchronized candidate using an inference-side gap proxy. Under high mismatch (FP8-quantized rollout) on Qwen3-4B and Qwen3-1.7B, MIPU improves average reasoning performance and training stability.
Questions & Answers
Explain “off-policyness” here.
On-policy RL learns from data generated by the current policy; off-policy means you update using data generated by a different policy. In LLM RL the rollouts are produced by the inference engine (call its policy μ), while the gradients are computed by the training engine (policy π). Even with identical weights, differences in precision, decoding, and serving backend make π and μ assign different probabilities to the same trajectory.
So the data you train on was not really produced by π — it came from μ. That is a form of off-policyness, and unlike the usual “old policy vs new policy” kind, it is engine-induced and present even at the same training step with synchronized parameters. It biases the gradient and, left unchecked, can destabilize or collapse training.
What new objective does the paper propose, and how does it target the mismatch?
Existing work reduces the mismatch but still judges an update by whether it improves the training policy π (the canonical RL objective). The authors argue the deployed policy is μ (inference), so the real question is whether a synchronized update actually yields a better μ.
They define MIPI (Monotonic Inference Policy Improvement): require J(μ₋₁) − J(μₖ) ≥ 0 — monotonic improvement along the inference-policy trajectory. That gain decomposes into three terms: ① post-update inference gap J(μ₋₁)−J(π₋₁), ② training-side update J(π₋₁)−J(πₖ), ③ pre-update inference gap J(πₖ)−J(μₖ). MIPU realizes it in two steps: Step 1 (sampler-referenced update) optimizes ②+③; Step 2 (inference-gap-aware acceptance) estimates & validates ① with a proxy T̂_post, accepting the synchronized candidate if T̂_post ≥ −c and otherwise rolling back. So it is a principled accept/reject layered on top of the update, aimed at the deployed policy.
What are the components of a modern LLM RL pipeline?
Modern LLM RL deliberately separates rollout generation from gradient computation:
- Inference / rollout engines (sample responses fast):
vLLM,SGLang. - Training engines (compute log-probs and gradients precisely):
FSDP,Megatron.
The split exists for efficiency — but it is exactly what creates the problem: precision, decoding, and serving-backend differences make the inference side (μ) and training side (π) assign different probabilities to the same tokens, which is the training-inference mismatch this paper studies.
Table: compare the remedies for the mismatch.
Grouped by where they intervene:
| Method (ref) | Side | What it does |
|---|---|---|
| TIS (Yao 2025) | Algorithm · sampler-aware | Uses the training-to-sampler probability ratio (π/μ) as a clipped correction weight on the update |
| MIS (Liu 2025) | Algorithm · sampler-aware | Filters out tokens / sequences whose mismatch signal is extreme |
| LR decay (Zhang 2026) | Algorithm · optimization | Shrinks the update magnitude when mismatch-related instability appears |
| Qi et al. (2025) | Infrastructure | Identifies numerical precision as a direct cause → uses FP16 rollout to reduce the train/infer discrepancy |
| Li et al. (2026) | Infrastructure | Analyzes low-precision / quantized rollout and its effect on reasoning-oriented RL |
| MIPU (this paper) | Objective-level | Doesn't just reduce mismatch — accepts or rejects a synchronized update by whether it improves the inference policy μ |
First three correct/filter the update from the algorithm side; the next two shrink mismatch from the system side. MIPU is complementary: it acts at the objective level, deciding whether a synchronized update should become the next inference policy.
Organize the current remedies — and what is “sampler-side information”?
Two families (see the table in Q4):
- Algorithm-side (sampler-aware) — fold sampler information into the training update: TIS (clipped π/μ ratio as a correction), MIS (filter extreme-mismatch tokens/sequences), LR decay (shrink updates when unstable).
- Infrastructure-side — reduce the mismatch at the system level: Qi et al. (precision is a direct cause → FP16 rollout), Li et al. (analyze low-precision/quantized rollout).
“Sampler-side information” = information from the inference/sampler engine — concretely, the probabilities (log-probs) the sampler μ assigned to the tokens it actually generated. So yes: it is derived from the samples the inference engine produced, i.e. the sampler's own probability for each token. Sampler-aware algorithms then use the ratio between the training-side probability π and the sampler-side probability μ (π/μ) to correct or filter the update. This paper is complementary: rather than only shrinking the mismatch, it asks whether a synchronized update should be accepted as the next inference policy.
They use FP8 rollout to amplify the mismatch — is this an experiment-design choice? Would FP16 make the remedy weaker?
Good read — yes, that is deliberate. They evaluate under FP8-quantized rollout precisely because inference-side quantization amplifies the training-inference gap, making both the problem and their fix easy to see. It is a stress test that exaggerates the effect.
The implication you suspect is right: under milder mismatch (FP16, or carefully-aligned bf16) the gap is smaller, so MIPU's accept/reject step would fire less often and buy less — its value scales with the severity of the mismatch. That doesn't make the method wrong; it means the benefit depends on how large your deployment's mismatch actually is. Honest caveat: the excerpts I have don't include an FP16 ablation quantifying how much the gain shrinks, so to judge the real-world payoff at FP16 you'd want that ablation from the experiments section. Your skepticism is reasonable — high-mismatch settings flatter the method.
It claims “more stable training dynamics” — how would you actually show that, and with what metrics?
Stability is usually demonstrated with a bundle of signals rather than one number:
- Reward / accuracy curve over steps — smoother and more monotonic, fewer dips, higher final value, and (importantly) lower variance across random seeds.
- No collapse — policy entropy doesn't crash toward 0 (mode collapse) or blow up; response length doesn't degenerate.
- Bounded KL / gradient magnitude — KL between successive policies stays controlled; no gradient spikes.
- Shrinking train-inference gap — the ①/③ gaps stay small over training.
- Seed variance — a tighter shaded band across runs.
- Rollback rate (MIPU-specific) — how often Step 2 rejects, as direct evidence it is catching bad synchronized updates.
In short: smoother monotone reward curves, controlled entropy/KL, no collapses, and lower run-to-run variance.
The MDP preliminary is clear — explain the MDP with clear annotations.
LLM generation is framed as an MDP M = (S, A, P, R, γ):
- State s_t = (q, y₁:ₜ) — the prompt plus the tokens generated so far.
- Action a_t — the next token, chosen from the vocabulary V.
- Transition P(s₋₁ | s_t, a_t) is deterministic: appending the chosen token gives the next state. The only randomness is the policy's token choice, not the environment.
- Episode starts at a prompt s₀ and ends at an end-of-sequence token or the max length H.
- Reward is sparse / outcome-level: 0 at every non-terminal step; the terminal reward R(τ) = 1 if the whole sequence is correct and well-formatted, else 0. γ is the discount.
Introduce the rule-based reward function and the learned reward model.
- Rule-based (verifiable) reward. A deterministic checker assigns reward by rules — does the answer match ground truth, pass unit tests, or satisfy a required format? Cheap, objective, and hard to game; it underpins RL-with-verifiable-rewards (RLVR) for math and code. This paper uses exactly this: R(τ)=1 if the output is correct and well-formatted, else 0.
- Learned reward model (RM). A neural network trained on human (or AI) preference data to score outputs; it powers RLHF. It captures fuzzy qualities (helpfulness, tone, style) that no rule can express — but it needs preference data, is only as good as that data, and can be reward-hacked (the policy finds high-scoring outputs the RM wrongly likes).
Trade-off: rule-based rewards are precise but narrow (verifiable domains); reward models are broad but noisier and hackable. This paper stays in the verifiable-reward regime, which is why the reward is a clean 0/1 at the end of the sequence.
Why is the policy in this MDP the LLM itself?
Because the policy is by definition the thing that chooses the next action given the state — here, the next token given the prompt-plus-tokens-so-far. That mapping is the LLM: π(a_t | s_t) is literally the model's next-token distribution. The LLM's parameters are the policy parameters, and generating text is rolling out the policy. There is no separate agent module; RL-fine-tuning the LLM just means optimizing this policy so the sequences it samples earn more reward.
Walk me through it thoroughly: policy difference (Eq. 1) → TRPO → PPO → GRPO. Why is KL used? What is the clipped ratio? What is the value function?
Why not vanilla policy gradient? ∇J(π) is very sensitive to step size — too big a step and performance craters. So we want a principled way to take the largest safe step.
Policy difference (Eq. 1, Kakade & Langford). For two policies π′ and π: J(π′) − J(π) = E over states/actions drawn from π′ of the advantage Aπ(s,a). Reading: the gain from switching to π′ equals the expected advantage of π′'s behavior, measured against π. Monotonic improvement means choosing π₋₁ so this is ≥ 0.
Why it's hard. Eq. 1 needs the new policy's state-visitation distribution d^{π′}, which you don't have until after you update — a chicken-and-egg problem.
TRPO's fix. Approximate by using the old policy's state distribution d^{πₖ} and an importance-sampling ratio π₋₁(a|s) / πₖ(a|s) (Eq. 2). This surrogate is only valid while the new policy stays close to the old one, so TRPO maximizes it subject to a KL trust-region constraint between old and new policy.
Why KL is everywhere in LLM RL. KL divergence measures how much one distribution diverges from another (it is a divergence, not a symmetric true metric). It shows up in two roles: (1) as a trust region keeping the update small enough that the importance-sampling surrogate stays valid → stability; and (2) as a penalty toward a reference/SFT model so the policy doesn't drift too far, which curbs reward hacking and preserves fluency. So yes — it's the standard way to quantify divergence between two token distributions and to bound how far you move per step.
PPO. Solving a hard KL-constrained problem is annoying, so PPO replaces it with a clipped probability-ratio objective. Let r_t = π_new(a|s) / π_old(a|s) — how much more (or less) likely the new policy makes the taken action. PPO multiplies r_t by the advantage but clips r_t to [1−ε, 1+ε]: if the advantage is positive it still can't push the probability up beyond 1+ε, and if negative it can't push below 1−ε. That removes any incentive for a huge one-step jump — a cheap, first-order way to “discourage excessively large policy updates.”
Value function (and why GRPO drops it). The value function V(s) is the expected future return from state s under the policy — a critic. PPO uses it as a baseline to form the advantage A = return − V (via GAE), which lowers gradient variance. But it means training a second network, extra memory, and it's hard to fit well when reward is a single sparse 0/1 at the end. GRPO (Shao et al., 2024) removes the critic: for one prompt it samples a group of G responses, scores each, and sets each response's advantage to (rᵢ − group-mean) / group-std. The group is the baseline — no value network, less memory, and it fits outcome-reward RL naturally.
Tell me more about RL in agentic model training — why are long-horizon rollouts and complex optimization dynamics the pain points?
Recently RL has moved into agentic training: models learn from environment feedback in tool-use, coding, and multi-turn interaction (examples the paper cites: GLM-4.5, Kimi K2.5, ROME). It's powerful but especially fragile, for three reasons:
- Long-horizon rollouts. Agentic tasks span many turns and tool calls, so trajectories are long. Credit assignment over a long horizon is hard, reward is sparse and delayed, small errors compound, and rollouts are slow/expensive — and the training-inference mismatch accumulates over length.
- Complex optimization dynamics. Tool outputs inject external, non-stationary tokens the policy didn't generate; variance is high; entropy and KL are harder to keep in a stable band; and the behavior distribution shifts as the agent improves.
- System-level implementation details. To make long rollouts affordable you quantize/optimize the inference engine — which enlarges exactly the mismatch this paper targets.
That's the paper's positioning: as we push toward agentic, long-horizon RL, the mismatch (and the question of whether a synchronized update truly improves the deployed policy) matters more, not less.
How is the mismatch “engine-induced” even at the same step with synchronized parameters?
“Engine-induced” means the gap comes not from the math of RL (an old-vs-new policy lag) but from the implementations of the two engines running the same weights θ.
- The inference engine (vLLM/SGLang) computes token probabilities with fused/optimized kernels, paged attention, custom batching, and often lower precision (FP8/FP16).
- The training engine (FSDP/Megatron) computes log-probs in higher precision with a different kernel and parallelism layout.
Floating-point arithmetic is not associative — (a+b)+c ≠ a+(b+c) in finite precision — so summing logits in a different order gives tiny differences, which the softmax turns into slightly different token probabilities, and over a long sequence they compound into a meaningfully different trajectory probability. So even at the same step k with byte-identical parameters, μ(τ) ≠ π(τ). No policy lag is needed; the engines alone induce it.
Why inference can afford lower precision. Generation (decode) is memory-bandwidth-bound — each step mostly shuttles weights and the KV cache through memory — so shrinking them to FP8/FP16 makes tokens come out faster and cheaper. Generation also only needs “good enough” probabilities to sample a token: a tiny numerical error rarely flips which token is chosen, and the text stays fluent. So you trade a little precision for a lot of throughput.
Why training keeps higher precision. Gradients are computed from log-probs and accumulated across billions of parameters and thousands of steps; small errors there compound and can destabilize optimization — so the training engine uses higher precision (e.g. FP32 master weights, BF16 compute with FP32 accumulation). The two sides therefore run at different precisions by design → different logits → the mismatch. Quantizing the rollout to FP8 (the paper's stress test) just widens that gap.
| Engine | Side | Key ideas | Typical precision |
|---|---|---|---|
| vLLM | Inference (sampler) | PagedAttention (KV-cache paging), continuous batching, high-throughput OpenAI-compatible serving | BF16 / FP16 · FP8 optional |
| SGLang | Inference (sampler) | RadixAttention (prefix caching), structured / JSON generation, fast serving | BF16 / FP16 · FP8 optional |
| FSDP | Training | PyTorch Fully Sharded Data Parallel — shards params, grads, optimizer states across GPUs | BF16 compute + FP32 accumulate |
| Megatron-LM | Training | NVIDIA tensor / pipeline / sequence parallelism for very large models | mixed precision, FP32 grads |
Inference engines optimize for throughput/latency (memory-bound decode → lower precision is fine); training engines optimize for gradient accuracy (higher precision). Running the same weights at different precisions is a core source of the mismatch.
Is the post-update inference gap ① a measure of whether the training update improved the inference policy — and is that Step 2?
Term ① = J(μ₋₁) − J(π₋₁) is, after the update, the performance gap between the deployed inference policy μ₋₁ and the training policy π₋₁. Your reading is basically right — it captures whether the synchronized (deployed) version carries over the training update — but it is the gap, not the improvement itself.
The full improvement of the inference policy is J(μ₋₁)−J(μₖ) = ①+②+③. Step 2 (inference-gap-aware acceptance) is the check on ①: it estimates a proxy T̂_post ≈ ① and accepts the synchronized candidate if T̂_post ≥ −c, else rolls back. So yes — Step 2 validates whether syncing produced a good inference policy, specifically by controlling the post-update gap ①.
I don't get why Step 1 optimizes ②+③ — why add those two?
It's a telescoping identity — just add and subtract J(π₋₁) and J(πₖ):
J(μ₋₁) − J(μₖ) = [J(μ₋₁)−J(π₋₁)] + [J(π₋₁)−J(πₖ)] + [J(πₖ)−J(μₖ)] = ① + ② + ③
The inner terms cancel, so the three pieces sum back to the true inference-policy improvement. Now group Step 1's part:
② + ③ = [J(π₋₁)−J(πₖ)] + [J(πₖ)−J(μₖ)] = J(π₋₁) − J(μₖ)
So ②+③ isn't a strange sum — it telescopes to “get from the old deployed policy μₖ to the new training policy π₋₁.” That is exactly what a training update can control while referencing the old sampler μₖ (hence “sampler-referenced update”). Step 2 then handles the leftover ① (from π₋₁ to μ₋₁).
The policy learning / update is actually changing the weights, right?
Yes. The policy is the LLM, and its weights θ are the policy parameters, so πₖ → π₋₁ is literally a gradient step that changes the weights. “Sync” (π → μ) means loading those updated weights into the inference engine so the sampler uses them; “rollback” means discarding the candidate weights and keeping the previous ones.
The training-to-sampler ratio π/μ — does “sampler” mean it comes from inference? What is the sampler exactly?
The sampler is the inference engine — the component that samples (generates) rollouts from the current policy. RL needs trajectories to learn from, and the sampler produces them by autoregressively sampling tokens. In LLM RL pipelines this is the fast inference engine (vLLM/SGLang), so “sampler” and “inference engine” are the same thing, with policy μ.
So the ratio π/μ = (training-engine probability of a token) ÷ (sampler/inference-engine probability of that same token). “Sampler-side” therefore means the generation/inference side — and yes, it's exactly where the rollouts come from.
“Autoregressively” means each token is predicted conditioned on all the tokens before it — the prompt plus everything generated so far — not on itself. At step t the model sees s_t = (q, y₁:ₜ) and outputs a distribution over the next token; that token is appended and fed back in for step t+1. So “auto-regressive” = regressing on its own previous outputs: it sees previous + current context and predicts the next; it never sees the token it is about to produce.
Policy entropy is interesting — what are the common algorithms around it?
Policy entropy is the entropy of the next-token distribution π(·|s) — how spread-out / uncertain the choice is. High entropy = exploratory and diverse; low entropy = confident and near-deterministic.
It's watched as a health metric because entropy collapse (→ 0) means the policy becomes over-confident, stops exploring, and often degenerates (repetition, mode collapse) right as reward plateaus; too-high entropy means it's basically random. Common ways to manage it:
- Entropy bonus — add +β·H(π) to the objective to encourage exploration (A3C; maximum-entropy RL like SAC).
- KL penalty to a reference / SFT model — indirectly keeps entropy from collapsing by staying near a sane distribution.
- Sampling temperature, and recent LLM-RL tricks like clip-higher (DAPO) and entropy-aware advantage / clipping.
Clip-higher / DAPO. DAPO (Decoupled clip and Dynamic sAmpling Policy Optimization, 2025) tweaks PPO's clipping. Standard PPO clips the ratio symmetrically to [1−ε, 1+ε]. DAPO decouples the two sides and uses a larger upper clip (ε_high > ε_low) — “clip-higher” — so the policy may raise the probability of good but currently low-probability tokens more aggressively. That preserves exploration and fights entropy collapse (symmetric clipping tends to crush rare-but-promising tokens). DAPO also adds dynamic sampling (drop prompts whose samples are all-correct or all-wrong, since they give zero gradient) and a token-level loss.
The rollback — does the previous trajectory become useless? And under RLVR a wrong answer is reward 0; if we re-label that wrong rollout, does the problem change?
Rollback discards the candidate weight update, not the experience. When Step 2 rejects, the model reverts to (πₖ, μₖ); the rollouts you collected still did their job (they produced and evaluated the candidate) and can inform the next attempt. So the trajectory isn't wasted — only the unsafe weight step is undone.
On your RLVR idea: yes, a wrong answer gets reward 0, a very sparse signal. Re-labeling wrong rollouts is a real research direction:
- Process / partial reward — credit correct intermediate steps or format, not only the final 0/1.
- Hindsight relabeling (HER-style) — treat what the model did produce as success for a relabeled goal.
- Preference among wrongs — rank “less wrong” vs “more wrong” via a judge/verifier for a denser signal.
- Exploration bonuses for novel attempts.
Caveat: relabeling changes the reward semantics and can reintroduce reward hacking or bias — RLVR's whole appeal is the clean, unhackable 0/1. But for hard problems where reward is almost always 0, denser/relabeled signals are exactly what people are exploring. Good instinct: this is the “reward shaping under sparse verifiable rewards” frontier.
State s_t = (q, y₁:ₜ) — is the prompt fed in via prefill?
Yes — s₀ = the prompt q is ingested via the prefill phase: the model runs one parallel forward pass over all prompt tokens, building the KV cache, before any token is generated. Then decode proceeds token-by-token, each new token appended to form s₁, s₂, …
So in MDP terms, prefill establishes s₀ (populates the KV cache with the prompt context) and decode extends the state. Prefill isn't a special RL “intervention” — it's just how the LLM conditions on the prompt — but it is where inference-engine optimizations live, so prefill vs decode is part of where the training-inference mismatch creeps in.
Example. For “What is 2+2?”, prefill runs all four prompt tokens through the model together in one pass (each position attends to the earlier ones via the causal mask) and stores a key/value pair per position — the KV cache. Decode then generates “4” using that cache, appends its own K,V, and continues — so the prompt is never recomputed.
More on: the transition is deterministic, the only randomness is the policy's token choice.
In a general MDP an action can lead to several possible next states with probabilities P(s′|s,a) — the environment is stochastic (a robot's wheel might slip). In the LLM token MDP there is no such environment randomness: once you pick token a_t, the next state is exactly s₋₁ = (q, y₁:ₜ, a_t) — you just append the token. P is a delta (probability 1 on that one state).
The only stochasticity in the whole process is the policy's own token choice (π is a distribution over the vocabulary). Why it matters: (1) no environment model is needed — the dynamics are trivial; (2) all learning is about the action distribution, not about predicting transitions; (3) it breaks in agentic settings, where tool outputs are external tokens the policy didn't choose — there the transition becomes genuinely stochastic, part of why agentic RL (Q12) is harder.
Your three questions, directly:
- “Environment is stochastic” = the next state is a distribution? Yes — P(s′|s,a) is a probability distribution over several possible next states; the same action can land you in different states, each with some probability.
- Why is the LLM transition determined? Because appending a chosen token yields exactly one next state — concatenation is deterministic. The choice of which token (sampling from the vocabulary distribution π) is the stochastic/learned part; but given that choice, the state update has no randomness.
- Does “transition” mean the latent space / representation? No. The transition is the MDP's state-update dynamics — how the state changes given an action — not the model's internal latent features. In the LLM, state = the token sequence so far and the transition = “append the token.” In agentic RL the state also absorbs tool / environment outputs, which are external and unpredictable, so the transition genuinely becomes stochastic (the environment injects tokens you didn't choose and can't perfectly predict).
Policy = the LLM's next-token distribution — for VLMs, does the next token include visual and text tokens, individually or combined?
It depends on the VLM's architecture:
- Most current VLMs (LLaVA-style, Qwen-VL) are multimodal on input but text-only on output: the image is encoded into visual tokens fed into the context (part of the state), but the model only generates text tokens. So the action space stays the text vocabulary; visual tokens are conditioning, not actions.
- Unified / any-to-any models (Chameleon-style, with a discrete image tokenizer) have a combined vocabulary of text + image tokens and generate both from one next-token distribution — there the action does include image tokens.
So individually vs combined: in most VLMs the visual tokens are input-only and the policy predicts text tokens conditioned on visual+text context; in unified generative models text and image tokens share one combined next-token distribution. For reasoning-RL like this paper the usual setup is: visual tokens condition the state, actions = text tokens — exactly what PixelEyes (07-05 log) does, reasoning in text while a tool handles pixels.
Expand on “reward-hacked” — the policy finds high-scoring outputs the RM wrongly likes.
Reward hacking = the policy earns high reward without actually doing the task well, by exploiting flaws in the reward signal. With a learned reward model (RM) — an imperfect proxy for human preference — hard optimization drifts into inputs the RM over-rates. Classic examples:
- Length bias — the RM prefers longer answers, so the policy pads.
- Sycophancy — the RM rewards agreeable answers, so the policy agrees even when wrong.
- Format / tone gaming — a confident tone or markdown scores well, so style beats substance.
The root cause is Goodhart's law: when a measure becomes a target, it stops being a good measure. Mitigations: KL penalty to a reference model (stay in-distribution), RM ensembles, iterative RM retraining on fresh policy outputs, conservative optimization — and, where possible, verifiable rewards. That's the point of this paper's rule-based 0/1: you can't fake passing a unit test, so it's essentially unhackable — at the cost of only working in verifiable domains (Q9).
Let's start building on the agentic-RL angle — organize the existing methods with GitHub links.
If you want to build on the agentic-RL direction (Q12), here is a curated starting set of open-source frameworks, grouped. Links verified via web search on 2026-07-06 — still, confirm each repo before relying on it, and note the space moves fast.
| Framework | Focus | Repo |
|---|---|---|
| verl | High-performance RL post-training; the de-facto base many others extend | verl-project/verl |
| RAGEN | Multi-turn, trajectory-level agent RL (StarPO) + training-stability diagnostics | RAGEN-AI/RAGEN |
| SkyRL | Full-stack, long-horizon real-world / Docker environments | NovaSky-AI/SkyRL |
| Agent Lightning | Add RL to ANY existing agent framework, decoupled (Microsoft) | microsoft/agent-lightning |
| MARTI | Multi-agent RL training + inference; tree-search RL | TsinghuaC3I/MARTI |
| AReaL | Large-scale asynchronous RL for reasoning (Ant) | inclusionAI/AReaL |
| OpenRLHF | Scalable RLHF on Ray (general base) | OpenRLHF/OpenRLHF |
| TRL | Hugging Face RL / alignment (PPO, GRPO, DPO…) | huggingface/trl |
| slime | RL-scaling post-training (THUDM) | THUDM/slime |
| ROLL | Stable multi-GPU parallel RL (Alibaba) | alibaba/ROLL |
First six are agentic / multi-turn oriented; the rest are general RLHF/RL bases that agentic frameworks build on. Verified via web search on 2026-07-06 — still confirm each repo before relying on it.
Two models the paper name-drops: GLM-4.5 (Zhipu / Z.ai — github.com/zai-org) and Kimi K2 (Moonshot — github.com/MoonshotAI/Kimi-K2). I could not find a confirmed public repo for ROME (Wang et al., 2026) as cited, so I'm flagging that rather than guessing a link.
Where I'd start a first experiment: verl (the de-facto base most others extend) or RAGEN (multi-turn, trajectory-level, with training-stability diagnostics that connect directly to this paper's mismatch/stability themes). If you already have an agent, Agent Lightning adds RL with almost no code change.
Causal mask vs autoregressive vs KV cache — how do they relate? And does “per layer” mean every layer's cache updates?
Three related-but-distinct ideas. Let me separate them, since this is exactly where they blur together.
1) Causal mask — inside self-attention. Each position i forms a query and scores it against the keys at all positions; the causal mask sets the scores for positions > i to −∞, so position i attends only to positions 1..i — former tokens and itself (the diagonal is included). It's a lower-triangular mask on the Q·Kᵀ score matrix. So yes: “only attend to itself and earlier tokens” is precisely the causal mask.
2) Autoregressive decoding — about prediction. The representation at position i (which, via the mask, has attended to 1..i including i) is used to predict the next token, i+1. That resolves the apparent contradiction you spotted:
- “Attends to itself” (mask diagonal) refers to the current token i — already in the input, so the model can look at it.
- “Never sees the token it's about to produce” refers to the target token i+1 — not in the input yet.
Different tokens (i vs i+1), so both hold. The causal mask is how each position gathers context (self + past); autoregressive is the objective of predicting the next token from that context. Causal masking is what makes autoregression valid — no peeking at the future.
3) KV cache — related to the mask, but not the same thing. Because attention is causal, the K and V vectors of past positions never change as you generate more tokens. So you cache them and, for each new token, compute only its own K,V and append — instead of recomputing the whole prefix. The causal structure is why caching is valid; the KV cache is the stored K,V tensors (a speed optimization), while the causal mask is the constraint on the scores. During decode you often don't even need an explicit mask — you just attend to the cached past plus the current token.
“Per layer.” Yes — a transformer stacks L layers, each with its own self-attention, so each layer keeps its own K,V cache. For an n-token prompt you store K,V for each of the n positions in each of the L layers (roughly shape [layers × positions × heads × head_dim]). During decode, every new token appends one K,V entry to every layer's cache. It's not that one shared cache is rewritten at each layer; rather each layer independently has and fills its own cache, and the vectors differ per layer because each layer transforms the representation.
Method · Section 4 — step-by-step formulas
Rendered as math via MathJax. Each block: the original passage, the formula, then every symbol.
Basics — the RL objective \(J(\pi)\)
Everything starts from the objective RL maximizes:
\[ J(\pi) \;=\; \mathbb{E}_{\tau \sim \pi}\!\left[R(\tau)\right], \qquad \pi^{*} \;=\; \arg\max_{\pi_\theta} J(\pi_\theta). \]- \(\pi\) (often \(\pi_\theta\)) — the policy: the LLM with parameters \(\theta\); \(\pi(a\mid s)\) is its next-token distribution.
- \(\tau\) — a trajectory: one full generated sequence (prompt + all output tokens).
- \(\tau\sim\pi\) — the trajectory is produced by rolling out \(\pi\).
- \(R(\tau)\) — the reward of the whole trajectory (here the verifiable 0/1 terminal reward).
- \(\mathbb{E}_{\tau\sim\pi}[\cdot]\) — the average over trajectories the policy generates.
- \(J(\pi)\) — the expected reward; training maximizes it, giving the optimal policy \(\pi^{*}\).
- \(d^{\pi}_{s},\,d^{\pi}_{s,a}\) — the (discounted) state and state-action visitation distributions: how often \(\pi\) visits \(s\) or takes \((s,a)\). They are the measures under which the next expectations are taken.
Eq 1 — the policy difference \(\Delta(\pi',\pi)\)
The policy difference (Kakade & Langford, 2002) writes the gap between two policies exactly:
\[ \Delta(\pi',\pi) \;:=\; J(\pi') - J(\pi) \;=\; \mathbb{E}_{s,a\sim d^{\pi'}}\!\left[A^{\pi}(s,a)\right]. \tag{1} \]- \(\Delta(\pi',\pi)\) — the performance difference of \(\pi'\) over \(\pi\).
- \(A^{\pi}(s,a)=Q^{\pi}(s,a)-V^{\pi}(s)\) — the advantage: how much better action \(a\) at state \(s\) is than \(\pi\)'s average there.
- \(d^{\pi'}\) — the visitation of the new policy \(\pi'\). This is the catch: it depends on the policy you don't have yet.
Reading: the gain of moving \(\pi\to\pi'\) equals the expected advantage (scored under the old \(\pi\)) of the actions the new \(\pi'\) tends to take. Monotonic improvement asks for \(\Delta(\pi_{k+1},\pi_k)\ge 0\).
Eq 2 — the TRPO local surrogate
Because Eq 1 needs the new policy's visitation \(d^{\pi_{k+1}}\), TRPO replaces it with the old one \(d^{\pi_k}\) and corrects the action mismatch with an importance ratio:
\[ \widetilde{\Delta}(\pi_{k+1},\pi_k) \;:=\; \mathbb{E}_{s\sim d^{\pi_k},\,a\sim\pi_k(\cdot\mid s)}\!\left[\frac{\pi_{k+1}(a\mid s)}{\pi_k(a\mid s)}\,A^{\pi_k}(s,a)\right]. \tag{2} \]- \(\dfrac{\pi_{k+1}(a\mid s)}{\pi_k(a\mid s)}\) — the importance-sampling ratio: how much more/less likely the new policy makes the same action.
- \(A^{\pi_k}(s,a)\) — the advantage under the old policy.
This local surrogate underlies the practical algorithms: TRPO maximizes it under a KL trust region; PPO clips the ratio; GRPO keeps the ratio but replaces the value-function advantage with a group estimate (Q11). The paper builds Step 1 on this exact template.
Eq 3 — the objective misalignment
The whole paper hinges on this non-implication:
\[ J(\pi_{k+1}) - J(\pi_k) \;\ge\; 0 \quad\not\Rightarrow\quad J(\mu_{k+1}) - J(\mu_k) \;\ge\; 0. \tag{3} \]- \(\pi\) — the training policy (moved by gradients).
- \(\mu\) — the inference policy (what runs after syncing weights to the inference engine).
Improving \(\pi\) does not guarantee that, after synchronization, the deployed \(\mu\) improves. This is the objective misalignment the method fixes.
Eq 4 — instantiating in GRPO \(\mathcal{J}_{\mathrm{GRPO}}\)
Now put the mismatch inside the GRPO objective:
\[ \mathcal{J}_{\mathrm{GRPO}}(\theta) = \mathbb{E}_{x,\,\{y_i\}\sim\mu_k}\!\Big[\tfrac{1}{G}\textstyle\sum_{i}\min\!\big(r_i\hat A_i^{\mu_k},\ \operatorname{clip}(r_i,1{-}\epsilon,1{+}\epsilon)\hat A_i^{\mu_k}\big)\Big], \tag{4} \] \[ r_i(\theta)=\frac{\pi_\theta(y_i\mid x)}{\pi_k(y_i\mid x)}. \]- \(x\sim\mathcal{D}\) — a prompt from the dataset \(\mathcal{D}\).
- \(\{y_i\}_{i=1}^{G}\sim\mu_k\) — a group of \(G\) responses sampled by the inference policy \(\mu_k\).
- \(r_i(\theta)=\pi_\theta/\pi_k\) — the training-side probability ratio.
- \(\hat A_i^{\mu_k}\) — the group-relative advantage: \(y_i\)'s reward normalized by the mean/std of rewards in the \(\mu_k\)-sampled group. The superscript \(\mu_k\) stresses it is induced by inference rollouts, not the old training-policy advantage \(A^{\pi_k}\).
- \(\operatorname{clip}(\cdot,1-\epsilon,1+\epsilon)\) — PPO clipping with width \(\epsilon\).
So under mismatch the standard update carries two flaws: a ratio-level mismatch (it clips \(\pi_\theta/\pi_k\) though samples come from \(\mu_k\)) and an advantage-level bias (\(\hat A_i^{\mu_k}\) is misaligned with \(A^{\pi_k}\) when \(\mu_k\neq\pi_k\)). Deeper: the training transition \(\pi_k\to\pi_{k+1}\) can differ from the inference transition \(\mu_k\to\mu_{k+1}\) — which is why the objective is redefined on \(\mu\).
Eq 5 — the MIPI decomposition
MIPI takes the inference-policy improvement \(J(\mu_{k+1})-J(\mu_k)\) as the target and decomposes it exactly:
\[ \begin{aligned} J(\mu_{k+1})-J(\mu_k) = \ & \underbrace{J(\mu_{k+1})-J(\pi_{k+1})}_{\text{(1) post-update inference gap}}\\ &+\ \underbrace{J(\pi_{k+1})-J(\pi_k)}_{\text{(2) training-side update}}\\ &+\ \underbrace{J(\pi_k)-J(\mu_k)}_{\text{(3) pre-update inference gap}}. \end{aligned} \tag{5} \]- (1) \(J(\mu_{k+1})-J(\pi_{k+1})\) — after update & sync, how far the deployed \(\mu_{k+1}\) sits from the trained \(\pi_{k+1}\).
- (2) \(J(\pi_{k+1})-J(\pi_k)\) — the ordinary training-side improvement.
- (3) \(J(\pi_k)-J(\mu_k)\) — the pre-update gap between old training and old inference policies.
A telescoping identity (inner terms cancel). MIPU splits the work: Step 1 constructs a candidate by handling (2)+(3); Step 2 verifies (1) and accepts or rolls back. (See Q14, Q15.)
Step 1 / Eqs 6-7 — sampler-referenced update
Step 1 targets (2)+(3), which telescope to \(J(\pi_{k+1})-J(\mu_k)=\Delta(\pi_{k+1},\mu_k)\). Following Eq 2 but referencing the sampler \(\mu_k\) gives the weight \((\pi_\theta/\mu_k)\,A^{\mu_k}\). The full trainer-to-sampler ratio factorizes:
\[ \rho_i(\theta) \;=\; \underbrace{\frac{\pi_k(y_i\mid x)}{\mu_k(y_i\mid x)}}_{w_i^{k}\;:\ \text{pre-update mismatch}} \;\cdot\; \underbrace{\frac{\pi_\theta(y_i\mid x)}{\pi_k(y_i\mid x)}}_{r_i(\theta)\;:\ \text{current update}}. \tag{6} \]- \(w_i^{k}=\pi_k/\mu_k\) — the pre-update mismatch weight (fixed given the checkpoint).
- \(r_i(\theta)=\pi_\theta/\pi_k\) — the current update ratio (what the gradient moves).
Clipping the whole \(\rho_i\) (“PPO-IS”) over-constrains the update when \(w_i^{k}\) is already outside the clip band, even if \(r_i(\theta)\approx1\). “Vanilla-IS” keeps the full \(w_i^{k}\) and clips only \(r_i\), but the unbounded weight adds variance. The paper adopts TIS: truncate \(\bar w_i^{k}=\min(w_i^{k},\,w_{\max})\) and clip only \(r_i(\theta)\). The Step-1 surrogate:
\[ \mathcal{J}_{\mathrm{S1}}(\theta) = \mathbb{E}_{x,\,\{y_i\}\sim\mu_k}\!\Big[\tfrac{1}{G}\textstyle\sum_{i}\bar w_i^{k}\,\min\!\big(r_i\hat A_i^{\mu_k},\ \operatorname{clip}(r_i,1{-}\epsilon,1{+}\epsilon)\hat A_i^{\mu_k}\big)\Big]. \tag{7} \]Optimizing Eq 7 yields the candidate \(\pi_{k+1}:=\pi_{\theta^\star}\). It proposes a better \(\pi\) relative to the sampler, but does not yet check whether the synchronized \(\mu_{k+1}\) realizes that gain — that is Step 2.
Step 2 / Eqs 8-9 — inference-gap-aware acceptance
After syncing \(\pi_{k+1}\to\mu_{k+1}\), the leftover term (1) is the post-update inference gap \(T_{\text{post}}=J(\mu_{k+1})-J(\pi_{k+1})\). A negative value means the deployed policy underperforms its training-side counterpart — an unreliable update. The direct form needs the advantage under \(\pi_{k+1}\) (unavailable in GRPO), so a reverse identity anchors it at \(\mu_{k+1}\):
\[ \begin{aligned} T_{\text{post}} &= -\Delta(\pi_{k+1},\mu_{k+1}) = -\,\mathbb{E}_{s,a\sim d^{\pi_{k+1}}}\!\big[A^{\mu_{k+1}}(s,a)\big]\\[2pt] &\rightsquigarrow\ \widehat{T}_{\text{post}} = -\,\mathbb{E}_{x\sim\mathcal{D}_{\mathrm{val}},\,y_i\sim\mu_{k+1}}\!\big[\rho_i\,\hat A_i^{\mu_{k+1}}\big]. \end{aligned} \tag{8} \]with a length-normalized sequence importance weight
\[ \rho_i = \exp\!\left(\frac{1}{T_i}\sum_{t=1}^{T_i}\log\frac{\pi_{k+1}(y_{i,t}\mid x,\,y_{i,\lt t})}{\mu_{k+1}(y_{i,t}\mid x,\,y_{i,\lt t})}\right). \tag{9} \]- \(\widehat{T}_{\text{post}}\) — the proxy for the post-update gap, estimated on a validation set \(\mathcal{D}_{\mathrm{val}}\) with responses from \(\mu_{k+1}\).
- \(\hat A_i^{\mu_{k+1}}\) — group-relative advantage from validation rollouts of \(\mu_{k+1}\).
- \(\rho_i\) — the per-token log-ratio averaged over the \(T_i\) tokens, then exponentiated: a stabilized (length-normalized) importance weight from \(\mu_{k+1}\) to \(\pi_{k+1}\).
Acceptance test: \(\widehat{T}_{\text{post}}\ge -c\) ⇒ accept (set \(\pi_k\leftarrow\pi_{k+1},\ \mu_k\leftarrow\mu_{k+1}\)); otherwise reject and roll both the trainer and the inference engine back to the previous checkpoint. \(c\ge 0\) tolerates proxy noise. No formal monotonicity guarantee, but it filters updates whose gains the inference policy won't realize.
Algorithm 1 — the full MIPU loop
The loop ties Steps 1 and 2 together with checkpointing and rollback:
- Init: trainer \(\pi_{\theta_0}=\pi_{\text{base}}\), optimizer state \(\Omega_0\), inference \(\mu_0=\mathrm{Sync}(\pi_{\theta_0})\).
- Each iteration \(k\): save checkpoint \((\theta_k,\Omega_k,\mu_k)\).
- Step 1 (proposal): sample \(G\) responses per prompt from \(\mu_k\); compute rewards, group advantages \(\hat A^{\mu_k}\), and mismatch weights \(w^{k}=\pi_k/\mu_k\); maximize the Step-1 surrogate (Eq 7) → \(\pi_{k+1}\); sync \(\mu_{k+1}=\mathrm{Sync}(\pi_{k+1})\).
- Step 2 (acceptance): roll out a validation batch from \(\mu_{k+1}\); compute \(\hat A^{\mu_{k+1}}\) and ratios \(\rho_i\); estimate \(\widehat{T}_{\text{post}}\).
- Gate: if \(\widehat{T}_{\text{post}} \lt -c\), reject — restore \(\theta_{k+1}\leftarrow\theta_k,\ \Omega_{k+1}\leftarrow\Omega_k,\ \mu_{k+1}\leftarrow\mu_k\).
So each accepted step is one the inference engine is verified to benefit from — the concrete realization of “optimize the policy you deploy, not just the one you train.”
Method & concept follow-ups
M1 — what are \(A^\pi\), \(Q\), \(V\), and “visitation”? Why replace with the old policy?
Three quantities, plainest terms:
\[ \begin{aligned} V^\pi(s) &= \mathbb{E}_{\pi}\!\big[\textstyle\sum_{t}\gamma^{t}r_t \mid s_0=s\big],\\ Q^\pi(s,a) &= \mathbb{E}_{\pi}\!\big[\textstyle\sum_{t}\gamma^{t}r_t \mid s_0=s,\,a_0=a\big],\\ A^\pi(s,a) &= Q^\pi(s,a)-V^\pi(s). \end{aligned} \]- \(V^\pi(s)\) — the value: the average score of being in situation \(s\) and playing on with \(\pi\).
- \(Q^\pi(s,a)\) — the action-value: the average score if you commit to move \(a\) now, then play on with \(\pi\).
- \(A^\pi(s,a)=Q-V\) — the advantage: the move's edge over just playing normally. Positive ⇒ better than \(\pi\)'s average here; negative ⇒ worse. (For an LLM: at a half-written sequence \(s\), is token \(a\) better than the average token \(\pi\) would pick?)
Visitation \(d^\pi\). Run the policy many times; some states come up often, others rarely. \(d^\pi\) is that (discounted) frequency distribution over states (or state-actions). It's the weighting in the expectation — you care more about states you actually reach.
Why swap \(d^{\pi'}\to d^{\pi_k}\). Eq 1's expectation is over the new policy's visitation \(d^{\pi'}\) — but you don't have \(\pi'\) until after you update. It's circular. So TRPO cheats gently: assume the new policy visits roughly the same places as the old one (use \(d^{\pi_k}\), which you can sample now) and fix up the actions with the importance ratio \(\pi'/\pi_k\). Valid only while \(\pi'\) stays near \(\pi_k\) — hence the trust region / clipping.
M2 — show the TRPO / PPO / GRPO objectives side by side
Same surrogate core \(r\,A\); they differ in how they bound the step and where \(A\) comes from.
\[ \textbf{TRPO:}\quad \max_\theta\ \mathbb{E}\!\left[\tfrac{\pi_\theta}{\pi_k}A^{\pi_k}\right]\ \text{s.t.}\ \mathbb{E}\big[\mathrm{KL}(\pi_k\|\pi_\theta)\big]\le\delta. \] \[ \textbf{PPO:}\quad \max_\theta\ \mathbb{E}\!\left[\min\!\big(r\,A,\ \operatorname{clip}(r,1-\epsilon,1+\epsilon)\,A\big)\right],\quad r=\tfrac{\pi_\theta}{\pi_k},\ A=A^{\pi_k}. \] \[ \begin{aligned} \textbf{GRPO:}\ \ &\max_\theta\ \mathbb{E}\!\Big[\tfrac{1}{G}\textstyle\sum_i\min\!\big(r_i\hat A_i,\ \operatorname{clip}(r_i,1{-}\epsilon,1{+}\epsilon)\hat A_i\big)\Big],\\ &\hat A_i=\tfrac{R_i-\mathrm{mean}(R)}{\mathrm{std}(R)}. \end{aligned} \]| Method | How the step is controlled | Advantage source | Critic? |
|---|---|---|---|
| TRPO | hard KL trust region: \(\mathbb{E}[\mathrm{KL}(\pi_k\|\pi_\theta)]\le\delta\) | \(A^{\pi_k}\) (critic / GAE) | yes |
| PPO | clip the ratio to \([1-\epsilon,\,1+\epsilon]\) | \(A^{\pi_k}\) (critic / GAE) | yes |
| GRPO | clip the ratio (same as PPO) | group-normalized reward \(\hat A_i=\dfrac{R_i-\mathrm{mean}(R)}{\mathrm{std}(R)}\) | no |
GRPO — is the group of \(G\) responses sampled simultaneously?
Yes, effectively. For one prompt \(x\) you draw \(G\) independent completions from \(\mu_k(\cdot\mid x)\). The inference engine generates them as a batch: the prompt is prefilled once, then \(G\) parallel decode streams sample \(G\) different continuations. So it is \(G\) inferences for the same prompt, run together in parallel — not a sequence where one depends on the last. That parallel group is exactly what lets GRPO form a baseline by comparing the \(G\) rewards to their own mean/std.
Eq 5 — is the decomposition a special design? Can it transfer (e.g. to MLLMs)?
Mathematically the decomposition is just a telescoping identity (add and subtract \(J(\pi_{k+1})\) and \(J(\pi_k)\)) — trivially true. What's clever is the choice of anchor points, which makes each term separately meaningful and separately actionable: a part you can optimize (②+③) and a part you can only check (①).
That framing is a reusable pattern — use it whenever your optimization target differs from your evaluation/deployment target:
- MLLMs — same training-inference mismatch (quantized vision encoder/connector/decoder at serve time), so MIPI applies almost directly; decompose “improvement of the deployed multimodal policy” and gate on the vision-aware inference gap.
- Quantization-aware deployment / distillation — “improvement of the quantized model” = “improvement of the full model” + “quantization gap”; accept only if the gap is tolerable.
- Sim-to-real RL — “real-policy improvement” = “sim improvement” + “sim-to-real gap.”
Lesson: when train ≠ deploy, decompose the deployment improvement into an optimizable piece and a verification gate. Good instinct to flag it.
Eq 6 — how should I read \(\pi_k\)?
\(\pi_k\) is the training policy at the start of iteration k — the frozen checkpoint weights, before this step's gradient update. Three policies are in play:
- \(\pi_k\) — old trainer (checkpoint, frozen reference for this step).
- \(\pi_\theta\) — the trainer being updated; at the end \(\pi_{k+1}:=\pi_{\theta^\star}\).
- \(\mu_k=\mathrm{Sync}(\pi_k)\) — the inference policy that generated the rollouts.
So \(\rho_i=\pi_\theta/\mu_k\) splits as \((\pi_k/\mu_k)\cdot(\pi_\theta/\pi_k)\). Using \(\pi_k\) as the pivot cleanly separates the gap you inherited \(w_i^{k}=\pi_k/\mu_k\) (fixed, both frozen this step) from the change you're making \(r_i=\pi_\theta/\pi_k\). That's the whole point of introducing \(\pi_k\).
M6 — table of PPO-IS / Vanilla-IS / TIS with formulas
All three keep PPO clipping on the current ratio; they differ in how they treat the pre-update mismatch weight \(w_i^{k}\):
| Variant | Weight / clip | Behaviour & when to use |
|---|---|---|
| PPO-IS | clip the full ratio \(\rho_i=w_i^{k}\,r_i(\theta)\) | over-constrains when the pre-update gap \(w_i^{k}\) is already outside the band, even if \(r_i\approx1\). OK only under mild mismatch. |
| Vanilla-IS | keep full \(w_i^{k}\), clip only \(r_i(\theta)\) | correct in expectation but unbounded \(w_i^{k}\) → high variance. OK when mismatch is small. |
| TIS (used here) | \(\bar w_i^{k}=\min(w_i^{k},\,w_{\max})\), clip only \(r_i(\theta)\) | bounded variance + keeps the correction — robust under large mismatch (the FP8 setting). The paper's choice. |
Intuition: PPO-IS is simplest but brittle when the sampler and trainer already disagree a lot; Vanilla-IS is faithful but noisy; TIS truncates the weight to get the correction and bounded variance — which is why it's chosen for the high-mismatch FP8 experiments.
Eq 8 — the reverse identity, \(\mathcal{D}_{\mathrm{val}}\), importance weights, length-normalization
Reverse identity. The performance-difference identity \(\Delta(\pi',\pi)=\mathbb{E}_{d^{\pi'}}[A^\pi]\) puts the advantage on the old policy and the expectation on the new one. The gap they want, \(T_{\text{post}}=\Delta(\mu_{k+1},\pi_{k+1})\), would need \(A^{\pi_{k+1}}\) — not available in GRPO. So they flip the roles: \(T_{\text{post}}=-\Delta(\pi_{k+1},\mu_{k+1})=-\mathbb{E}_{d^{\pi_{k+1}}}[A^{\mu_{k+1}}]\), anchoring the advantage at \(\mu_{k+1}\) (measurable from \(\mu_{k+1}\) rollouts) and correcting the expectation from \(\mu_{k+1}\) to \(\pi_{k+1}\) with an importance weight. That swap is the “reverse” part.
\(\mathcal{D}_{\mathrm{val}}\). A held-out validation set of prompts, separate from the training set \(\mathcal{D}\). You roll out \(\mu_{k+1}\) on it to estimate \(\widehat{T}_{\text{post}}\). Yes — it's a probe subset used to simulate/measure the post-update gap without reusing the training batch.
Importance weights & expectation. Importance sampling estimates an expectation under \(p\) from samples of \(q\): \(\mathbb{E}_{p}[f]=\mathbb{E}_{q}\!\big[\tfrac{p}{q}f\big]\). Here \(\rho_i=\pi_{k+1}/\mu_{k+1}\) is that reweighting factor, and the whole quantity is an expectation estimated by averaging over validation samples — so yes, it lives inside an expectation.
Length-normalization. A full-sequence ratio \(\prod_t \pi/\mu\) over hundreds of tokens has explosive variance. Eq 9 instead takes the average per-token log-ratio \(\tfrac{1}{T_i}\sum_t\log(\pi/\mu)\) then exponentiates — a per-token geometric mean. So yes, it's computing an average (of log-ratios) to keep the weight stable regardless of sequence length.
Which formulas are the authors' own contribution?
Splitting background from proposed (based on the citations in the excerpts):
- Background (not theirs): \(J(\pi)\) (standard RL); Eq 1 policy difference (Kakade & Langford, 2002); Eq 2 TRPO surrogate (Schulman, 2015); Eq 4 GRPO objective (Shao, 2024).
- Proposed (theirs): Eq 3 (the misalignment framing / claim); Eq 5 (the MIPI decomposition — their organizing identity); Eq 7 (the Step-1 truncated sampler-referenced surrogate); Eq 8 (the \(T_{\text{post}}\) reverse-identity proxy); Eq 9 (the length-normalized importance weight); and Algorithm 1 (MIPU).
- Eq 6 is their analysis (the \(w\cdot r\) factorization), but the TIS truncation it leads to is adopted from Yao et al. (2025), so it's adapted rather than wholly new.
So yes, Eq 8 is one of their proposed pieces. The novel core is the two-step MIPI/MIPU: the decomposition (Eq 5) plus the acceptance test (Eqs 8-9). (Attribution is inferred from the cited names in the excerpts; the full related-work section would pin down the exact boundaries.)
Summarize the whole method — the MIPU pipeline
One sentence: optimize the policy you deploy, not just the one you train — propose an update, then accept it only if the synchronized inference policy actually benefits.
In plain language. When you RL-train an LLM you use a fast engine to generate answers and a precise engine to compute gradients; even with the same weights they disagree a little, and — worse — making the training model better doesn't guarantee the deployed (inference) model gets better once you copy the weights over. Standard RL blindly accepts every step, which can quietly hurt the model you actually ship. MIPU's fix: judge success by whether the deployed model improved, and enforce it in two steps each round — (1) build a candidate update that accounts for the data having come from the inference sampler; (2) after copying weights to the inference engine, quickly validate whether the deployed model really improved, and if not, throw the step away and restore the previous checkpoint.
Algorithm 1, step by step (strict):
- Input: base model \(\pi_{\text{base}}\), group size \(G\), clip range \(\epsilon\), tolerance \(c\).
- Init: trainer \(\pi_{\theta_0}=\pi_{\text{base}}\), optimizer state \(\Omega_0\), inference \(\mu_0=\mathrm{Sync}(\pi_{\theta_0})\).
- for \(k=0,1,2,\dots\) do
- Save checkpoint \((\theta_k,\Omega_k,\mu_k)\).
- [Step 1] Collect a training batch \(\mathcal{B}\sim\mathcal{D}\); for each prompt sample \(G\) responses from \(\mu_k\).
- Compute rewards, group-relative advantages \(\hat A^{\mu_k}\), and mismatch weights \(w^k=\pi_k/\mu_k\) on the training rollouts.
- Update trainer params + optimizer state by maximizing the Step-1 surrogate (Eq 7).
- Set \(\pi_{k+1}\leftarrow\pi_{\theta^\star}\) and \(\mu_{k+1}\leftarrow\mathrm{Sync}(\pi_{k+1})\).
- [Step 2] Sample a validation batch \(\mathcal{B}_{\mathrm{val}}\sim\mathcal{D}_{\mathrm{val}}\); roll out \(G\) responses per prompt from \(\mu_{k+1}\).
- Compute rewards, advantages \(\hat A^{\mu_{k+1}}\), and ratios \(\rho_i\) on the validation rollouts.
- Estimate the post-update inference gap \(\widehat{T}_{\text{post}}\) (Eqs 8–9).
- if \(\widehat{T}_{\text{post}} \lt -c\) then
- Reject the update: restore \(\theta_{k+1}\leftarrow\theta_k,\ \Omega_{k+1}\leftarrow\Omega_k,\ \mu_{k+1}\leftarrow\mu_k\).
- end if
- end for
(This mirrors Algorithm 1 line-for-line. Send the full paper and I'll cross-check every symbol against Appendix A.2's implementation details.)
The rollback “checkpoint” — is that the stored KV cache?
No — different thing. The checkpoint is a saved snapshot of the training state: the trainer weights \(\theta_k\), the optimizer state \(\Omega_k\) (e.g. Adam moments), and the synced inference weights \(\mu_k\). Rolling back = restoring those, i.e. undoing the gradient step.
The KV cache is a transient, per-request inference buffer (the K,V of one generation's tokens); it's rebuilt and discarded every generation and isn't something you “roll back.” So: checkpoint = weights + optimizer + inference copy (persistent training artifact); KV cache = ephemeral attention memory during decoding.
M3 — idea: constrain the parameter space? Turn the structure into a “space” problem?
Good direction, and it dovetails with the paper. MIPU already imposes a constraint, but in behavior/output space (the acceptance test \(\widehat{T}_{\text{post}}\ge -c\) rejects updates that harm the deployed policy). Your idea — constrain in parameter space — is the smooth, proactive complement.
Concretely, replace the post-hoc gate with a mismatch-aware trust region: prefer parameter directions where the trainer policy and its quantized inference realization agree. For example add a penalty
\[ \max_\theta\ \mathcal{J}(\theta)\ -\ \lambda\, D\big(\pi_\theta,\ \mu(\pi_\theta)\big), \]where \(\mu(\pi_\theta)\) is the inference (quantized) realization of \(\pi_\theta\) and \(D\) a divergence (KL / per-token). Geometrically: treat the “deploy-consistent” parameters as a constraint surface and project updates onto it (a trust region defined w.r.t. the \(\mu\)-realization rather than the old \(\pi\)).
Honest caveats: evaluating \(D(\pi_\theta,\mu(\pi_\theta))\) each step means running the inference realization (costly), and quantization is non-differentiable, so you'd need a proxy / straight-through estimator. But “inference-aware trust region for RL” is a real, publishable extension of this paper. (Speculative research direction, not an established result.)
Q25 follow-ups — why “often” no explicit mask? Is the KV-cache size identical across layers?
Why “often” not “always.” In a pure single-token decode step, the one new query only ever attends to past cached positions plus itself — all legitimately allowed — so there's nothing to mask. But an explicit mask is needed whenever you process more than one position at once or batch uneven lengths:
- Prefill / training — many positions in one pass need the triangular causal mask so position \(i\) can't see \(j\gt i\).
- Padding in a batch — mask out pad tokens of shorter sequences.
- Chunked prefill / speculative decoding — several new tokens at once need a causal mask among them.
- Sliding-window / block-sparse attention — a windowed mask, not full-causal.
Cache size per layer. In a vanilla transformer, yes — identical: each layer stores K and V of shape \([\text{batch},\ n_{\text{kv-heads}},\ \text{seq},\ d_{\text{head}}]\), same across all \(L\) layers, so total \(=L\times\) that. Caveats from modern efficiency tricks: GQA/MQA shrink \(n_{\text{kv-heads}}\) (but equally across layers); sliding-window layers cache only a window (smaller); cross-layer KV sharing reuses one cache for several layers; and hybrid models (linear-attention layers — like FlashMorph from the 07-05 log) keep no KV cache at all for some layers. So “identical per layer” holds for a homogeneous transformer, but the efficiency-oriented architectures deliberately break it.
Follow-ups · round 2
Is “visitation” the state distribution? Same as the inference distribution? A weights-space paper?
Yes — \(d^\pi_s\) is a distribution over states (how often \(\pi\) visits each state); \(d^\pi_{s,a}\) over state-actions. It's the policy's occupancy measure.
Same as the inference distribution? No, two different objects: the policy \(\mu\) is a distribution over next tokens given a state (action-conditional); the visitation \(d^\mu\) is the distribution over states that results from rolling out \(\mu\). They're linked — \(d^\mu\) is induced by \(\mu\). Here the rollouts come from \(\mu_k\), so the empirical state distribution of the samples is \(d^{\mu_k}\); that's exactly why Step 1 “references the sampler.”
Weights-space idea. Visitation lives in state space, not weight space. Relating a weight change to the resulting occupancy shift is what trust-region theory does (bounding KL in policy space bounds the occupancy shift). Predicting how weight updates move the visitation is real but hard — the “policy-improvement bound” literature; explorable if you find a tractable surrogate for weights → occupancy.
Is the importance ratio just proportional scaling? Does it assume linear correlation?
Not “equal-proportion scaling.” It's a per-sample multiplicative weight, from the identity
\[ \mathbb{E}_{p}[f] = \mathbb{E}_{q}\!\left[\tfrac{p}{q}\,f\right]. \]Each sample drawn from \(q\) is reweighted by its own density ratio \(p/q\), so the average estimates the expectation under \(p\). It is exact in expectation for any \(f\) — no linearity or linear-correlation assumption — as long as \(q>0\) wherever \(p>0\). The catch is variance: if \(p/q\) swings a lot, the estimate is noisy, which is why the paper clips \(r_i\) and truncates \(w_i^{k}\) (TIS). So: per-sample density reweighting, exact but high-variance — not a uniform scale.
MQ2 — what does “critic” mean?
In actor-critic RL the actor is the policy (picks actions); the critic is a separate network estimating the value function \(V(s)\) (or \(Q\)). The critic “critiques” the actor by predicting expected future return, giving a baseline to form the advantage \(A=\text{return}-V\), which lowers gradient variance. PPO trains such a critic alongside the policy; GRPO drops it, using the group's mean reward as the baseline instead — the “free of learning the value function” line.
Explain: “improvement of the quantized model = improvement of the full model + quantization gap”
Same telescoping trick, applied to quantization. The deployed policy is \(\mu=\mathrm{quantize}(\pi)\); its score differs from full-precision \(J(\pi)\) by a quantization gap \(J(\mu)-J(\pi)\). So the deployed improvement splits:
\[ \underbrace{J(\mu_{k+1})-J(\mu_k)}_{\text{quantized improvement}} = \underbrace{J(\pi_{k+1})-J(\pi_k)}_{\text{full-model improvement}} + \big[\text{change in quantization gap}\big]. \]You can improve the full model while the quantization gap worsens — so the deployed model doesn't actually improve. The lesson (identical to MIPI): gate on the quantized model's real improvement — accept a step only if, after quantizing, the model still improved. It's MIPU with “inference engine” = “quantizer.”
MQ4 — could we optimize only a learnable residual path?
Practical and neat. Instead of moving all weights, train only a small residual/adapter path (e.g. LoRA) and freeze the base. Why it helps here:
- Smaller mismatch surface — fewer parameters change per step, so the train-vs-inference divergence is smaller and easier to control.
- Precision control — keep the low-rank residual in higher precision while the base stays quantized, shrinking the quantization gap on exactly the part that learns.
- Cheaper sync and rollback (only the residual moves).
It composes with MQ11: constrain/regularize only the residual for deploy-consistency. Caveats: LoRA-style residuals have limited capacity, and the residual is still quantized at serve time unless kept separate. But “residual-only, mismatch-aware RL” is a clean direction.
MQ5 — what does “\(:=\)” mean?
\(:=\) means “is defined as.” It introduces a definition: the left side is newly defined to equal the right side, rather than an equality you derived. So \(\pi_{k+1}:=\pi_{\theta^\star}\) reads “define \(\pi_{k+1}\) to be the optimized policy \(\pi_{\theta^\star}\).” (Plain \(=\) asserts an equality that could be proven; \(:=\) is an assignment/definition.)
MQ6 — is the unified role of PPO-IS / Vanilla-IS / TIS the alignment between trained and inference?
Yes — that's the unifying view. All three solve the same problem: samples were produced by the inference/sampler policy \(\mu_k\), but you're updating the training policy \(\pi_\theta\). Each uses the trainer-vs-sampler ratio (\(\pi/\mu\)) to realign the off-policy samples so the gradient reflects the training objective despite the mismatch — i.e., alignment between the trained policy and the inference policy that generated the data. They differ only in how they bound that weight to control variance: PPO-IS clips the whole ratio, Vanilla-IS leaves it unbounded, TIS truncates it. Same job (importance-sampling correction of the train↔inference gap), different variance knob.
The KV cache is “rebuilt and discarded every generation” — does every inference update it?
Two timescales:
- Within one generation: yes — the cache is updated every token: each new token computes its own K,V and appends them, so the cache grows step by step.
- Across generations / requests: the cache is per-request — a new generation starts with a fresh cache (or reuses a shared prefix cache) and is freed when the generation ends.
So “rebuilt and discarded every generation” = each independent generation has its own cache lifecycle; it's a transient inference buffer, appended within a generation and thrown away after — not persistent model state (unlike the weights/optimizer checkpoint, which is what gets rolled back).
How exactly do queries attend to the cached KV — in the head? a matrix lookup?
Per head, per layer: the new token forms a query \(q\), scores it against every cached key by dot product, softmaxes, and takes the weighted sum of the cached values:
\[ \text{out} = \sum_{j\le t}\ \mathrm{softmax}_j\!\Big(\tfrac{q\cdot k_j}{\sqrt{d}}\Big)\; v_j. \]So “attend to the cached KV” = compute \(q\,K^{\top}\) against the stored keys (a matrix-vector product), softmax the scores, then combine the stored values \(\cdot V\). It runs independently in each head (each head has its own \(q,K,V\)); heads are then concatenated and projected. It's not a hash/lookup table — it's soft attention: a weighted average via matrix multiplies. The KV cache just spares recomputing \(k_j,v_j\) for past positions.
M11 — connecting the math to spatial geometry (geometric loss, balance, metrics)
A rich, legitimate lens — much ML progress is geometry in disguise:
- Information geometry — the natural gradient treats parameter space as a Riemannian manifold with the Fisher metric; TRPO's KL trust region is a local metric ball. The mismatch \(\pi\) vs \(\mu\) reads as a geodesic distance, and MIPU's gate as “stay inside a metric ball around the sampler.”
- Optimal transport — Wasserstein losses to compare rollout or latent-trajectory distributions.
- Loss-landscape geometry — flat vs sharp minima (SAM); tied to update stability.
- Search as navigation — PixelEyes' region search (07-05) becomes navigation on a manifold; a “geometric balance” loss can trade exploration (spread/volume in embedding space) against exploitation (concentration).
- Latent reasoning (AMVL, 07-05) already lives in a continuous latent space — distances/curvature there are natural.
How I'd start: pick one problem and define the metric first. For this paper the cleanest handle is information geometry — define the train↔inference discrepancy as a KL/Fisher distance and recast the accept/reject gate as a trust-region constraint in that geometry (MQ11, made rigorous). Keep the metric computable (KL/Fisher, cosine, entropic-regularized OT) so it yields a real gradient. Productive direction — the risk is elegance without a tractable objective, so anchor each geometric idea to a loss you can actually optimize. (Research direction, not a solved result.)
Follow-ups · round 3
MQ15 — is the critic like a teacher? Is there much to improve? Does it matter much?
Like a teacher? Only loosely. A supervised teacher hands you target answers (labels or soft distributions) — an external supervisor with the “right” output. A critic doesn't know the right action; it learns the value \(V(s)\) (how good a state is on average), trained from the same reward the actor sees (often bootstrapped via TD/Bellman). So it's more a self-taught coach / referee that supplies a baseline to lower gradient variance than a teacher with answers. (The reward is the real supervisor.) The genuinely teacher-like case is distillation-style RL where a teacher gives distributions — e.g. DemoPSD from the 07-05 log.
Much to improve? Does it matter? Yes, and “it depends.” The critic is a classic pain point in LLM RL: it's a second network (~doubles memory), and with sparse, terminal rewards its value estimates are noisy → biased advantages. In classic RL (control, games) a good critic is crucial for variance reduction and sample efficiency; in LLM reasoning RL its payoff is questionable — which is exactly why critic-free methods (GRPO's group baseline, RLOO's leave-one-out) now dominate. So the main “improvement” has been to remove or replace it. Other directions: better value targets (GAE, V-trace), Monte-Carlo / per-token value estimates (VinePPO), ensemble or distributional critics (curb over-estimation), sharing the critic trunk with the policy. Net: historically important, but for verifiable-reward LLM RL the field is moving away from an explicit critic.
Let's build on GRPO with MQ17's residual direction — and maybe optimize the objective
Yes — GRPO (critic-free) + MQ17's residual idea + this paper's mismatch-awareness compose into a concrete, testable design:
- Freeze the base, train only a residual. \(W=W_0+\Delta W\) with \(\Delta W\) low-rank (LoRA); optimize only \(\Delta W\).
- Keep the residual precise. Serve the base \(W_0\) quantized but the low-rank \(\Delta W\) in higher precision — so the learned part barely suffers quantization, pushing the pre-update mismatch \(w_i^{k}=\pi_k/\mu_k\) toward 1 on exactly the part that changes.
- Mismatch-aware objective. Restrict Eq 7 to \(\Delta W\) and add a deploy-consistency penalty:
where \(\mu(\cdot)\) is the quantized inference realization and \(D\) a KL / per-token divergence. Intuition: fewer, precise, deploy-consistent parameters → smaller \(T_{\text{post}}\) gaps → fewer Step-2 rejections → faster, steadier training.
How to test. Same setup as the paper (Qwen3-1.7B/4B, FP8 rollout). Baselines: GRPO, TIS, full MIPU. Ablations: LoRA-only, LoRA-only + precise-residual, + penalty. Metrics: average accuracy, reward-curve stability (seed variance), rollback rate (how often Step 2 rejects), wall-clock (LoRA sync/rollback is cheaper). I can't run training here, but I can write the full experiment plan + pseudocode / a loop scaffold on an open framework (verl / TRL). Want that as a research-direction card?
In the checkpoint, are “weights” the same as the “optimizer” state?
No — different things, both saved in the checkpoint. Weights \(\theta\) are the model parameters (the function itself). The optimizer state \(\Omega\) is the bookkeeping the optimizer keeps to compute the next update — for Adam, the per-parameter first/second moment estimates \((m,v)\) and step count; for SGD-momentum, the velocity. It's “how to move,” not “where you are.”
That's why Algorithm 1 restores both \(\theta_{k+1}\leftarrow\theta_k\) and \(\Omega_{k+1}\leftarrow\Omega_k\) on rollback: undoing the weights but keeping stale Adam moments would corrupt the next step (wrong momentum / adaptive scaling). Think position (weights) + velocity/inertia (optimizer state) — you need both to faithfully undo a move.
Could a world model predict the “spatial evolution” — define the geometry change as a world-model problem?
Ambitious but coherent. A world model learns the environment's dynamics — how a (latent) state evolves under actions — so you can predict and plan (Dreamer-style) rather than only react. Mapping onto the geometry idea:
- Where it earns its keep. In the plain LLM token MDP the dynamics are trivial (deterministic append), so a world model adds little. But in agentic RL (tool use, multi-turn) the environment is genuinely stochastic — there a world model of tool/environment responses lets you plan in a learned latent space, and that space has geometry: “geometry change” = a trajectory through its manifold.
- Modeling the optimization itself. You could learn a world model of how the policy's behavior distribution evolves under weight updates — predict the occupancy shift or reward-landscape change after a step. That's exactly the hard map weights → occupancy from MQ13, and it links to learned optimizers and model-based RL over the improvement process.
- Search as model-based planning. For region search (PixelEyes, 07-05), a latent \(z\) could encode explored-coverage; a transition model \(z'=f(z,\text{action})\) with a geometric loss (preserve distances / coverage) lets you plan the next region instead of fixed BFS.
So yes — “define the geometry change as a world-model prediction problem” is a real fusion of model-based RL + information geometry: learn a predictive latent dynamics whose metric you control, then optimize/plan in that space. Caveats: world models are hard to learn accurately and errors compound, so keep the latent low-dimensional and the predictive loss tractable, and start where the dynamics are genuinely non-trivial (agentic / search). (Research direction, not a solved result.)
Foundations & cross-domain
TD, the Bellman equation, and Q-learning — the machinery behind the critic
This is how a value function is actually learned — the engine under the critic (MQ23). Source you shared: EITCA — Bellman equation in TD & Q-learning.
Bellman equation — value is self-consistent: a state's value = immediate reward + discounted value of the next state.
\[ V^\pi(s) = \mathbb{E}_{a\sim\pi,\,s'}\big[\,r + \gamma\,V^\pi(s')\,\big],\qquad Q^\pi(s,a)=\mathbb{E}_{s'}\big[\,r+\gamma\,\mathbb{E}_{a'\sim\pi}Q^\pi(s',a')\,\big]. \]The optimal forms take a max over actions:
\[ V^{*}(s)=\max_a \mathbb{E}\big[r+\gamma V^{*}(s')\big],\qquad Q^{*}(s,a)=\mathbb{E}\big[r+\gamma \max_{a'} Q^{*}(s',a')\big]. \]Temporal-Difference (TD) learning — don't wait for the whole episode (that's Monte Carlo); bootstrap from your current estimate of the next state. TD(0):
\[ V(s)\leftarrow V(s) + \alpha\,\underbrace{\big[\,r+\gamma V(s')-V(s)\,\big]}_{\delta\,:\ \text{TD error}}. \]The TD error \(\delta\) is the “surprise” — how much better/worse reward-plus-expected-future turned out than predicted; \(\alpha\) is the learning rate, \(\gamma\) the discount.
Q-learning — off-policy TD control: learn optimal action-values directly, regardless of the behaviour policy:
\[ Q(s,a)\leftarrow Q(s,a)+\alpha\big[\,r+\gamma \max_{a'} Q(s',a') - Q(s,a)\,\big]. \]Deep RL replaces the table with a neural net \(Q_\theta\) (function approximation) — DQN, stabilized by a target network and a replay buffer (the linked lesson's topic).
Tie-back: a PPO critic is trained by exactly this — regressing \(V\) toward the TD/Bellman target \(r+\gamma V(s')\). GRPO sidesteps all of it with a group baseline, which is why this paper never needs a value network.
Could TD learning apply to fMRI / the medical domain?
Strong instinct — and not a stretch: TD learning has one of the most celebrated links to the brain in all of neuroscience.
The established bridge (dopamine ≈ TD error). The reward-prediction-error hypothesis of dopamine (Schultz, Dayan & Montague, 1997) found that phasic dopamine-neuron firing behaves like the TD error \(\delta\). This launched model-based fMRI: fit an RL/TD model to a subject's trial-by-trial choices, generate the latent signals (value \(V\), prediction error \(\delta\)) per trial, and use them as parametric regressors in the fMRI GLM to find regions whose BOLD signal tracks them. Robust findings: TD prediction error → ventral striatum; expected value → vmPFC / OFC. So the common direction is TD → interpret fMRI.
Medical / more novel directions.
- Computational psychiatry — fit TD/RL models to patients (depression, addiction, Parkinson's); altered learning rates or prediction-error signalling become biomarkers, with fMRI localizing the deficit. Parkinson's (a dopamine disorder) directly implicates TD.
- Real-time fMRI neurofeedback — a closed loop where RL selects stimuli/feedback to steer a target brain state, with a TD-learned value of brain states.
- Adaptive / closed-loop stimulation (DBS, tACS) tuned by a value function over brain states.
Honest caveats. fMRI is slow (hemodynamic lag ~4–6 s, low temporal resolution), noisy, and indirect (a blood-flow proxy, not spikes), so credit assignment over these signals is hard; and any clinical use needs proper validation, ethics, and regulatory approval. So: well-established as a model of brain reward learning; promising-but-hard for closed-loop control. This is a research-area overview, not medical advice — I can pull recent references (O'Doherty, Daw, computational-psychiatry reviews) if you want to go deeper.
DataComp-VLM: Improved Open Datasets for Vision-Language Models
multimodal benchmarks · VLM trainingLarge multi-institution collaboration (ML Foundations / DataComp; institutions per superscripts in the paper)
Code: github.com/mlfoundations/dcvlm · Site: datacomp.ai/dcvlm · arXiv 2606.28551 · baseline: FineVision
Gist · paraphrased from the abstract
Building strong VLMs hinges on how you curate the training data, but the field lacked a controlled benchmark for comparing curation strategies. DataComp-VLM (DCVLM) fills that gap: a data-centric benchmark that gathers 160 datasets across four types — image-caption pairs, interleaved multimodal documents, text-only, and instruction-tuning — into a 6T-token corpus, and lets you test filtering, mixing, formatting, sampling across 1B–8B models and 6.25B–200B token budgets, evaluated on up to 52 downstream benchmarks in 9 domains.
Headline finding: data mixing, not filtering, drives quality — instruction-heavy mixtures scale better than caption-heavy ones, with the gap widening at larger scale. The resulting DCVLM-Baseline trains an 8B VLM to 63.6% on their 33-task core suite at 200B tokens — +5.4pp over FineVision, the prior SOTA open VLM training set. All artifacts to be released publicly.
Questions & Answers — none yet.
Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning
perception / reasoning decouplingZhejiang University · Alibaba Group
arXiv 2607.01191 · GitHub & Hugging Face (linked in the paper) · methods: P2R / PRA-GRPO
Gist · paraphrased from the abstract
Fine-grained visual reasoning is hard for VLMs when a small but decisive cue is buried in a high-resolution image. Prior work injects local evidence by repeated cropping or test-time visual search but rarely separates perception from reasoning. Perceive-to-Reason (P2R) makes the split explicit and simple: the model first acts as a Perceiver that localizes the question-relevant evidence (emitting bounding boxes), then as a Reasoner that answers from the annotated image plus the cropped regions — one model, two stages, a clean single-pass pipeline.
Training uses PRA-GRPO (Perception-Reasoning Alternating GRPO): a role-aware RL scheme that alternates perception-focused and reasoning-focused updates using only final-answer supervision (no bounding-box labels). On Qwen3-VL-Instruct 2B/4B/8B it improves across scales — P2R-4B hits 93.2% on V-Star and 81.9%/80.5% on HR-Bench-4K/8K — and the gains extend to broader multimodal reasoning. Diagnostic motivation: giving Qwen3-VL-4B oracle boxes lifts V-Star from 81.7% to 90.6%, showing perception is the bottleneck.
Main results · Table 1
| Method | V-Star | HR-4K | HR-8K | Avg. |
|---|---|---|---|---|
| GPT-4o | 66.0 | 59.0 | 55.0 | 60.0 |
| o3 (OpenAI) | 95.7 | – | – | – |
| Qwen3-VL-2B | 74.9 | 70.4 | 64.6 | 70.0 |
| Qwen3-VL-4B | 81.7 | 73.8 | 67.0 | 74.2 |
| Qwen3-VL-8B | 83.8 | 74.8 | 70.1 | 76.2 |
| Qwen3-VL-32B | 86.9 | 78.6 | 73.6 | 79.7 |
| ZoomEye-7B | 90.6 | 69.6 | 69.3 | 76.5 |
| DeepEyes-7B | 90.1 | 75.1 | 72.6 | 79.3 |
| Thyme-7B | 82.2 | 77.0 | 72.0 | 77.1 |
| P2R-2B | 84.3 | 75.1 | 74.8 | 78.1 |
| P2R-4B | 93.2 | 81.9 | 80.5 | 85.2 |
| P2R-8B | 93.7 | 81.5 | 82.6 | 85.9 |
“Overall” columns from the paper's Table 1 (per-column Attr/Spatial and FSP/FCP breakdown is in the embedded table above); P2R rows highlighted. Note: PixelEyes is not in this table — the two are concurrent (PixelEyes: 30 Jun 2026, arXiv 2607.00115; P2R shortly after, 2607.01191), so neither benchmarks against the other.
Compared with PixelEyes (07-05 digest)
Both papers reach the same diagnosis — perception, not reasoning, is the bottleneck in fine-grained visual reasoning, so decouple it — and both build on Qwen3-VL with GRPO-style RL. The interesting split is how far they push complexity.
| Aspect | PixelEyes (07-05) | Perceive-to-Reason (P2R) |
|---|---|---|
| Core thesis | decouple perception (where) from reasoning (what) | same — a Perceiver localizes, a Reasoner answers |
| Decoupled into… | two models: the VLM reasoner + an external SAMTok tool | two roles of one model: perceive stage → reason stage |
| Localization output | pixel-precise mask (referring segmentation) | the model emits bounding boxes (JSON bbox_2d) + annotated / cropped images |
| Pipeline | multi-turn agentic search: Semantic-Region BFS, mask-guided search, switchable-tool fallback | single-pass two-stage; explicitly “clean context, simple pipeline, easy to optimize” |
| Training / RL | SFT on expert trajectories (PixelEyes-6K, Gemini-3-Flash teacher) → vanilla GRPO | PRA-GRPO: role-aware RL alternating perception/reasoning updates, only final-answer supervision (no box labels) |
| Base model | Qwen-3-VL | Qwen3-VL-Instruct 2B / 4B / 8B |
| External dependency | needs an external segmentation model (SAMTok) | none — self-contained single model |
| Benchmarks | Pinpoint-Bench (own; targets ~0.07% of image; LSR + Turn-to-Answer, zero-hint) | V-Star, HR-Bench-4K/8K (existing) + broader multimodal; P2R-4B: 93.2 V-Star |
| Failure framing | “inattentional blindness” (LSR − accuracy gap) | perception bottleneck: oracle boxes lift Qwen3-VL-4B 81.7 → 90.6 on V-Star |
Takeaway. PixelEyes goes more complex — an external mask tool plus agentic BFS search — to nail pixel-precise, ultra-tiny targets (~0.07% of the image) through iterative search. P2R deliberately goes simpler: one model in two roles, bounding boxes instead of masks, a single perceive-then-reason pass — and it explicitly argues that agentic-search pipelines (ZoomEye, and by extension PixelEyes-style methods) are “complex and hard to optimize.” Its cleverest piece is PRA-GRPO: it learns both roles from only final-answer reward, with no localization labels, whereas PixelEyes needs SFT on teacher-generated expert trajectories first.
So they sit at two ends of a precision-vs-simplicity spectrum: PixelEyes likely wins where targets are minuscule and pixel-exact localization + iterative search matter; P2R trades some localization precision for a self-contained, easily-optimized pipeline that also generalizes to broader multimodal reasoning. (Comparison is from each paper's abstract/figures — the benchmarks differ, so the headline numbers aren't directly comparable; flagging that rather than ranking them.)
PixelEyes results (Table 2) & head-to-head
| Method | V* | HR-4K | HR-8K |
|---|---|---|---|
| PixelEyes-4B | 91.62 | 81.75 | 79.88 |
| P2R-4B | 93.2 | 81.9 | 80.5 |
| PixelEyes-8B | 94.24 | 85.00 | 83.15 |
| P2R-8B | 93.7 | 81.5 | 82.6 |
Reported in separate, concurrent papers — likely different eval protocols/checkpoints (their reported Qwen3-VL-4B on V* even differs: 80.1 vs 81.7), so read cross-paper deltas as indicative. Reading: at 4B the two are ~tied (P2R edges V* and HR-8K by a hair); at 8B PixelEyes leads, notably on HR-Bench-4K (85.0 vs 81.5). PixelEyes additionally reports VisualProbe, Pinpoint-Bench, and the localization metrics TAE/LSR, which P2R doesn't.
Questions & Answers — none yet. Send questions to start this paper's log, or ask me to read the full PDF for a deeper analysis.
VLA-Corrector: Lightweight Detect-and-Correct Inference for Adaptive Action Horizon
VLA · detect-and-correctZhejiang University (OmniAI Group, ACES Lab) · Alibaba DAMO Academy
Code: github.com/ZJU-OmniAI/vla-corrector · June 2026 · methods: LVM / OGG
Gist · paraphrased
Action-chunked VLA policies predict a batch of future actions and run them open-loop under a fixed action horizon to cut policy-call frequency. The cost: a “predict-then-blindly-execute” blind spot — in contact-rich tasks a small perturbation amplifies unseen and compounds into failure. VLA-Corrector adds a lightweight detect-and-correct loop without touching the backbone weights: a Latent-space Vision Monitor (LVM) continuously compares predicted vs actual visual-feature evolution to spot drift; on persistent drift it fires a truncation event (dropping stale actions) and re-plans via Online Gradient Guidance (OGG).
This yields an event-triggered adaptive action horizon: long-horizon when the chunk is reliable, short corrective replanning when it drifts — softening the static-horizon trade-off between robustness and policy-call frequency. It drops into different VLAs with no backbone retraining and improves long-horizon, contact-rich manipulation.
Section 3 framework · summary
The core move: decouple action generation from execution monitoring. The frozen VLA generates chunks; a small external corrector watches whether execution stays on track and intervenes only on drift. Four stages:
1) Corrector training (§3.1). Freeze the VLA, use its visual encoder to get latents. For a transition (o_t, a_t, o_{t+k}) the target is the residual latent evolution ΔZ*; train a lightweight (~40M MLP) external corrector M_φ to predict it:
\[ \Delta\hat Z_{t+k} = M_\phi\big(Z^{\text{real}}_t,\ a_t\big), \tag{3} \] \[ \mathcal{L}_{\text{corr}} = \big\|\Delta\hat Z_{t+k} - \Delta Z^{*}_{t+k}\big\|_2^2 + \beta\big[\,1 - \operatorname{CosSim}(\Delta\hat Z_{t+k},\ \Delta Z^{*}_{t+k})\,\big]. \tag{4} \]Predicting the residual (not the absolute future latent) suppresses static background and focuses on task-relevant motion. Only the small corrector trains — modular and cheap.
2) LVM detection (§3.2). Online, compare the corrector's expected residual with the actual one from fresh camera input:
\[ E_t = 1 - \operatorname{CosSim}\big(\Delta Z^{\text{exp}}_{t+k},\ \Delta Z^{\text{real}}_{t+k}\big). \tag{5} \]Large \(E_t\) = the world is evolving differently than predicted → drift.
3) Event-triggered truncation (§3.3). Thresholding \(E_t\) directly is jumpy, so use robust adaptive thresholds from a sliding window (median \(M_e\) + MAD), with hysteresis:
\[ T_{\text{on}} = M_e + \lambda_{\text{on}}\,\text{MAD},\qquad T_{\text{off}} = M_e + \lambda_{\text{off}}\,\text{MAD},\qquad \lambda_{\text{on}} > \lambda_{\text{off}}. \tag{7} \]An interrupt fires only when \(E_t > T_{\text{on}}\) for \(p\) consecutive steps (isolated spikes ignored). It then drops the remaining stale actions → the realized horizon shrinks to \(H_{\text{adaptive}} = h < H\).
4) OGG correction (§3.4). The replan right after the interrupt is steered. Build a corrective latent direction (intended dynamics minus accumulated drift), then nudge the flow-matching velocity so the candidate action's predicted effect aligns with it:
\[ \Delta Z_{\text{corr}} = \Delta Z_{\text{exp}} - \Delta Z_{\text{dev}}, \tag{9} \] \[ \mathcal{L}_{\text{OGG}} = 1 - \operatorname{CosSim}(\Delta\hat Z_{\text{act}},\ \Delta Z_{\text{corr}}),\qquad v^{\text{guide}}_\tau = v_\tau - \eta\,\nabla_{v_\tau}\mathcal{L}_{\text{OGG}}. \tag{10-11} \]Because it perturbs the velocity field (not raw action coordinates), it stays compatible with flow matching and gives smoother recovery.
Plain-language walkthrough
In everyday terms — a VLA robot predicts a batch of future moves (a “chunk”) and executes them blindly for a while to save compute. The danger: if something goes slightly wrong mid-batch, it keeps going blind and the error snowballs. VLA-Corrector adds a cheap watcher:
- Learn what should happen. Offline, on demos, with the big policy frozen, train a tiny 40M predictor of how the scene's visual features should change right after each action.
- Watch while executing. During blind execution, keep predicting the expected visual change and compare it to what the camera actually shows. Same direction → fine; diverging (low cosine similarity) → a drift signal \(E_t\).
- Don't panic on a blip. Only if drift stays high for several steps in a row does it hit the brakes — throwing away the rest of the pre-planned moves (shortening the horizon on the fly).
- Re-plan, not blindly. It nudges the new plan in the direction that undoes the accumulated drift (OGG), steering the robot back on track.
So it runs long-horizon and efficient when things are fine, and switches to short-horizon corrective replanning exactly when it drifts — an adaptive action horizon — with no retraining of the expensive policy.
Key passages & results
| Backbone · Method | Easy | Med | Hard | V.Hard | Avg |
|---|---|---|---|---|---|
| π0.5 · Baseline | 70.5 | 45.0 | 38.3 | 41.0 | 48.70 |
| π0.5 · +Corrector | 83.2 | 61.7 | 47.5 | 65.0 | 64.35 |
| SmolVLA · Baseline | 81.3 | 53.6 | 51.7 | 61.0 | 61.90 |
| SmolVLA · +Corrector | 83.4 | 56.0 | 64.2 | 63.0 | 66.65 |
| X-VLA · Baseline | 72.5 | 46.4 | 48.3 | 55.0 | 55.55 |
| X-VLA · +Corrector | 74.4 | 50.0 | 50.0 | 64.0 | 59.60 |
| Model | Object | Spatial | Goal | Long | Avg |
|---|---|---|---|---|---|
| π0.5 (full FT) | 99.4 | 98.2 | 97.8 | 92.4 | 96.95 |
| π0.5 (few-shot FT) | 97.8 | 95.4 | 96.2 | 86.6 | 94.00 |
| π0.5 (few-shot) +Corrector | 99.8 | 100.0 | 98.0 | 93.4 | 97.80 |
Corrector rows highlighted. Biggest gains on the hardest / longest-horizon splits (MetaWorld Very-Hard +24.0 on π0.5; LIBERO-Long +6.8) — exactly where the open-loop blind spot hurts most.
Connection to our research thread
Worth flagging: VLA-Corrector is essentially the robotics cousin of the MIPU / MQ26 idea we've been developing. Its 40M corrector \(M_\phi\) is exactly a lightweight world model of the local dynamics (MQ26); the detect → gate → correct loop mirrors MIPU's propose → verify → accept/rollback; and its concrete tricks — a cosine-directional loss, MAD-based robust thresholds with hysteresis, persistence-checking, and velocity-space corrective guidance — are directly borrowable for the MIPU-Predict plan (predict expected behaviour, gate on deviation, shape the correction). Different domain, same skeleton — good validation that the direction is real, and a source of ready-made methods.
Questions & Answers
The persistence counter (Eq 12) — what are the cases, and how is each classified?
The counter \(c_t\) turns the noisy per-step drift score \(E_t\) into a persistence signal, so the system truncates only on sustained drift, not on blips. First, two thresholds are built from the recent window's robust statistics (median \(M_e\) and MAD):
\[ T_{\text{on}} = M_e + \lambda_{\text{on}}\,\text{MAD},\qquad T_{\text{off}} = M_e + \lambda_{\text{off}}\,\text{MAD},\qquad \lambda_{\text{on}} \gt \lambda_{\text{off}}. \]Because \(\lambda_{\text{on}} \gt \lambda_{\text{off}}\), these split \(E_t\) into three zones, and the counter (Eq 12) has one case per zone:
| Where E_t falls | Meaning | Counter update |
|---|---|---|
| E_t > T_on | clearly abnormal now | increment: +1 |
| E_t < T_off | clearly normal now | reset to 0 |
| T_off ≤ E_t ≤ T_on | ambiguous (dead-band) | hold (unchanged) |
Trigger (Eq 13): an interrupt fires when \(c_t \ge p\) (the patience parameter). Then the queue's remaining actions are discarded, the counter resets, and the next replan runs in corrective mode — so the realized horizon becomes \(H_{\text{adaptive}} = h \lt H\).
Why classify it this way — it's a hysteresis (Schmitt-trigger) design:
- Increment only when clearly abnormal (\(E_t \gt T_{\text{on}}\)) — you accumulate evidence of real drift.
- Reset only when clearly normal (\(E_t \lt T_{\text{off}}\)) — a brief spike that quickly returns to normal wipes the count, so isolated spikes can't trigger.
- Hold in the dead-band (\(T_{\text{off}} \le E_t \le T_{\text{on}}\)) — borderline noise neither falsely accumulates nor erases progress; the gap between the two thresholds stops the counter from flip-flopping when \(E_t\) hovers near a single line.
Example (\(p=3\)). Isolated spike: \(E_t\gt T_{\text{on}}\) (c=1) → \(E_t\lt T_{\text{off}}\) (c=0, reset) → no interrupt. Sustained drift: \(E_t\gt T_{\text{on}}\) (c=1) → (c=2) → dead-band (c=2, hold) → \(E_t\gt T_{\text{on}}\) (c=3 \(\ge p\)) → interrupt & truncate.
The State-Prediction Separation Hypothesis
architecture · LLM efficiency to readCornell University · Harvard University
arXiv 2607.01218
Gist · paraphrased from the abstract
Transformers use a single forward stream both to predict the next token and to store state for future tokens. The state-prediction separation hypothesis: disentangling these two roles improves language modeling. They build a State-Prediction Separation Transformer with two computation streams (one for prediction, one for carried state) and compare across scales.
Separation consistently improves data- and compute-efficiency: at 1.6B parameters it matches a standard Transformer's validation loss using 2.6× fewer tokens, and beats standard Transformers by 2–3 points on average downstream. Extensive analysis rules out confounders and shows the two designs' gradients differ fundamentally.
Questions & Answers — none yet.
OrbitQuant: Data-Agnostic Quantization for Image and Video Diffusion Transformers
diffusion · quantization to readCantina Labs · University of Southern California · UIUC
Project: saurabhcantina.github.io/orbitquant · arXiv 2607.02461
Gist · paraphrased from the abstract
Diffusion transformers (DiTs) are SOTA for image/video generation but costly at inference; post-training quantization (PTQ) is the natural fix, yet DiT activations shift across timesteps, prompts, and guidance branches, forcing prior methods to re-fit calibration data per checkpoint/modality. OrbitQuant is a data-agnostic weight-activation quantizer that bypasses range estimation by quantizing in a normalized, rotated basis.
A randomized permuted block-Hadamard (RPBH) rotation concentrates each coordinate around one fixed, known marginal regardless of input, so a single Lloyd-Max codebook serves all timesteps/prompts/layers; the rotation cancels inside each linear layer (only a forward rotation on activations remains at runtime), and the same recipe transfers image→video with no per-modality tuning. Across FLUX.1, Z-Image-Turbo, Wan 2.1, and CogVideoX it sets SOTA low-bit PTQ — pushing image DiTs to W2A4 with usable quality.
Questions & Answers — none yet.
Optimizing Visual Generative Models via Distribution-wise Rewards
visual generation · RL rewards to readUSTC · Shanghai Innovation Institute · Hunyuan Frontier Lab, Tencent · NUS
arXiv 2607.02291 · ICML 2026
Gist · paraphrased from the abstract
RL for visual generation typically uses sample-wise rewards, which invites reward hacking that erodes diversity and adds artifacts. This paper finetunes generative models with distribution-wise rewards that account for the whole batch's distribution, mitigating the mode collapse that arises when all samples are pushed the same way independently.
To tame the expensive reward estimation, a subset-replace strategy updates only a small subset of a generated reference set; they also RL-optimize post-hoc model-merging coefficients to counter the train-inference inconsistency introduced by SDE in standard RL. Results: FID-50K improves across bases (SiT 8.30→5.77, EDM2 3.74→3.52) with better perceptual quality while preserving diversity. (ICML 2026.)
Questions & Answers — none yet.