Paper Reading
Notes on the papers I read. One entry a day, ideally. · 11 so far
-
2026.07.06
Paper Reading Notes - 2026-07-06 Session
Running Q&A log over the 2026-07-06 session - 7 papers. Deep-read: 01 MIPU (Monotonic Inference Policy Update - optimize the deployed inference policy, not just the training policy, under training-inference mismatch; full method + Q&A) and 04 VLA-Corrector (lightweight detect-and-correct inference giving action-chunked VLA policies an adaptive action horizon). Also logged: 02 DataComp-VLM (data-centric benchmark for VLM training) and 03 Perceive-to-Reason / P2R (decoupling perception from reasoning for fine-grained visual reasoning, compared head-to-head with PixelEyes). To read: 05 The State-Prediction Separation Hypothesis, 06 OrbitQuant (data-agnostic quantization for diffusion transformers), 07 Optimizing Visual Generative Models via Distribution-wise Rewards.
-
2026.07.05
Paper Reading Notes - 2026-07-05 Digest
Running Q&A log over the 2026-07-05 paper digest. Paper 01 - PixelEyes: decoupling perception (an external SAMTok mask) from reasoning (the VLM) for pinpoint visual evidence seeking.
-
2026.07.02
ABot-M0.5 — Reading Notes
Close reading of ABot-M0.5 (AMAP CV Lab, Alibaba): a unified mobility-and-manipulation World Action Model aligning temporal granularity, action space, and train-test consistency.
-
2026.07.01
World Models, Agents & VLA — Paper Reading Notes
Close readings of five AI papers: DreamForge (real-time controllable world model), LUMOS (semantic OS layer for agents), QVAL (evaluating dense-supervision signals via Q-alignment), MemLearner (learned context memory for video world models), and Drop-Then-Recovery (redundancy analysis of vision-language-action models).
-
2026.06.29
World Models for Robotic Control — Paper Triage & Notes
A 10-paper triage centered on world models for robotic control, with deep reads of PhysiFormer (flow-matching diffusion over raw 3D vertex trajectories, view-invariant, factorized DiT), Fast-LeWM (action-prefix encoding + parallel latent prediction over a JEPA latent, AdaLN conditioning, CEM planning), and ICWM (in-context system identification from self-probed random clips); plus dives into the Markov state and second-order ODEs, AdaLN vs. cross-attention, FLOPs vs. latency, and two new-idea brainstorms — action-prefix-conditioned latent-trajectory diffusion and self-organized core-periphery modulation.
-
2026.06.28
OctoSense — Multimodal Robot Perception
Notes on OctoSense — a late-fusion masked autoencoder with per-modality tokenizers that fuses RGB, event, LiDAR, thermal, IMU, GPS and proprioception at once; plus a deep dive into SSL, DINO / SigLIP / Hiera, and contrastive vs. masked methods.
-
2026.06.27
Reading Notes — Attention & KV-Cache Mechanics
How recent work limits what a query attends to — InfoKV's info-aware cache compression, block-causal masks, R-SWA, plus MHA/GQA/MLA and KV-serving internals.
-
2026.06.26
AI Research Q&A — Agentic RL, World Models & Embodied AI
A consolidated Q&A study log from the 6.26 daily-papers list — on-policy distillation, VLA/WAM/world models, JEPA, Goodhart, RoPE attention, and KV-cache quantization, with diagrams.
-
2026.06.25
Causal-rCM: Study Notes
Deep-dive notes on Causal-rCM — building up autoregressive video diffusion distillation concept by concept, with diagrams.
-
2026.06.24
Unlimited OCR Works: Welcome the Era of One-shot Long-horizon Parsing
Replaces decoder attention with a Reference Sliding Window Attention that keeps the KV cache constant, letting an OCR model transcribe dozens of pages in a single forward pass without the usual slowdown.
-
2026.06.22
World Action Models: A Survey — Dream Less, Act More
Surveys World Action Models — predictive-action models that make a forecast of the future usable for control — and argues the field is moving toward generating less of the future while keeping what control actually needs.