Reading Queue · Triaged Backlog
Want to Read
Good papers I want to read but haven’t had time for yet — sorted by theme, each with a reason to read and its paper first page (click to expand). Stars mark the ones sitting right on the world-model / robotic-control thread I’m already reading.
Priority picks (closest to my current thread)
- PhysiFormer — diffusion transformer that predicts 3D-mesh mechanics in world space; the explicit-3D anchor of the thread.
- Fast-LeWM — replaces autoregressive one-step rollout with parallel action-prefix prediction; the planning-latency / error-accumulation fix from my Fast-LeWM notes.
- REGEN — world action models generate pseudo-replay so a policy rehearses old tasks without stored demos; continual learning by imagination.
Embodied AI, world models & robot learning
5Directly adjacent to the PhysiFormer / Fast-LeWM / world-model thread.
PhysiFormer: Learning to Simulate Mechanics in World Space
Yiming Chen, Yushi Lan, Andrea Vedaldi · Visual Geometry Group, University of Oxford · arXiv:2606.27364
Why readThe anchor of my current world-model thread — the explicit-3D, world-coordinate route to physical plausibility that I’m tracking against the video-diffusion approaches.
Key ideaA diffusion transformer that samples future vertex trajectories of 3D meshes in world coordinates (not view-dependent pixels), given initial positions/velocities and material type — with no ad-hoc latent space and no hand-coded rigidity/causality. Attention factorised over time/space/objects gives permutation-invariant multi-object reasoning without explicit object encoding.
ResultTrained on 100k+ simulated trajectories; handles rigid + elastic mechanics and generalises to mixed materials, unseen real geometries, and larger object counts; beats autoregressive baselines on trajectory accuracy, rigidity preservation, and momentum consistency.
Paper first page (click to collapse)
Fast LeWorldModel (Fast-LeWM)
Yuntian Gao, Xiangyu Xu · Xi’an Jiaotong University · arXiv:2606.26217 · code
Why readThe other paper already in my notes — it attacks exactly the autoregressive-rollout cost and accumulated-latent-error problem from the Fast-LeWM latency thread.
Key ideaA JEPA-style latent world model that swaps repeated one-step local rollout for action-prefix prediction: encode a candidate’s action prefixes and predict, in parallel, the future latents reached after each prefix. Prefix-level supervision teaches how states evolve under different prefixes rather than fitting only one-step transitions; at planning time it reads the last prefix token instead of rolling through every imagined state.
ResultHigher average success than LeWM with substantially less planning time; lower open-loop latent loss whose growth slows markedly as the rollout horizon increases.
Paper first page (click to collapse)
World Action Models Enable Continual Imitation Learning with Recurrent Generative Replays (REGEN)
Manish Kumar Govind, Dominick Reilly, Smit Patel, Hieu Le, Srijan Das · UNC Charlotte · arXiv:2606.27374 · project
Why readUses a world model’s generative capability for continual learning without stored demos — resonates with the replay / on-the-fly-adaptation angles in my notes.
Key ideaRecurrent Generative Replay: query a World Action Model to synthesize pseudo-replay trajectories — conditioned only on prior task instructions and current-task observations — so a policy can rehearse previously learned tasks without keeping the original human demonstrations.
ResultCuts catastrophic forgetting by up to 50% vs sequential fine-tuning, approaching privileged experience-replay that needs real replay data; flags long-horizon visual degradation and action-observation inconsistency as the main bottlenecks.
Paper first page (click to collapse)
Flatness Preserves Instruction Following in Vision-Language-Action Models
Haochen Zhang, Yonatan Bisk · Carnegie Mellon University · arXiv:2606.23641
Why readA one-line fix for “instruction blindness” in VLA finetuning — cheap, no retraining, and complementary to guidance methods; slots into the VLA-robustness threads.
Key ideaLimited-data finetuning degrades pretrained VLA representations so policies follow visual shortcuts and ignore language. They argue sparse gradients → sharp, high-curvature minima, and apply flatness-preserving optimization (sharpness-aware minimization) on the exact same data so the model tolerates weight-space perturbations.
ResultSAM during finetuning improves instruction following by 60%+ across sim and real benchmarks, with no extra data, architecture change, or retraining; includes analysis of selective sharpness.
Paper first page (click to collapse)
OctoSense: Self-Supervised Learning for Multimodal Robot Perception
Anthony Bisulco, Jeremy Wang, Kostas Daniilidis, Randall Balestriero, Pratik Chaudhari · UPenn GRASP + Brown · arXiv:2606.27317 · project
Why readOpen sensor platform + dataset + a fast late-fusion MAE that stays robust under degraded sensors — useful infra on the perception side of the robot-learning stack.
Key ideaAn open-source multi-sensor rig (stereo RGB, event, LiDAR, thermal, IMU, RTK-GPS, proprioception) with a 59-hour time-synced driving dataset. A “late-fusion” masked autoencoder uses modality-specific tokenizers and caches per-modality tokens at inference to ingest new measurements as they arrive.
ResultFast representation (6.68 ms on RTX 5090 / 112 ms on Orin NX); beats image-only foundation models on optical flow, depth, segmentation, and ego-motion; predicts robustly at night and under degraded sensors.
Paper first page (click to collapse)
Agents, tool use & reasoning safety
2Safety-relevant tool selection and whether “thinking” actually buys deliberation.
When Lower Privileges Suffice: Investigating Over-Privileged Tool Selection in LLM Agents
Kaiyue Yang, Yuyan Bu, Jingwei Yi, Yuchi Wang, Biyu Zhou, Juntao Dai, Songlin Hu, Yaodong Yang · CAS / BAAI / CUHK / Peking · arXiv:2606.20023
Why readA crisp, safety-relevant agent failure mode with its own benchmark and a post-training fix — pairs with the tool-orchestration / reliability threads.
Key ideaIntroduces TOOLPRIVBENCH to test whether agents pick a higher-privilege tool when a sufficient lower-privilege alternative exists — measuring both initial selection and escalation after transient tool failures, across 8 domains and 5 recurring risk patterns.
ResultOver-privileging is common and amplified by transient failures; general safety alignment does not transfer to least-privilege choice and prompt controls help only a little. Their privilege-aware post-training defense cuts unnecessary high-privilege use while preserving general capability.
Paper first page (click to collapse)
Do Thinking Tokens Help with Safety?
Narutatsu Ri, Abhishek Panigrahi, Sanjeev Arora · Princeton Language and Intelligence · arXiv:2606.25013
Why readPunctures the “deliberation = safer” assumption for reasoning models — useful counter-evidence before leaning on chain-of-thought for alignment.
Key ideaTests whether thinking tokens enable real safety deliberation across GPT-OSS, Qwen, Olmo, and Phi. A trained head on the first token’s hidden state already predicts refusal/compliance before any visible thinking.
Result0.84–0.95 AUROC / ~88% balanced accuracy at first token; the outcome rarely changes after the first ~20% of thinking, and ~74% of text-level “deliberation” happens after the outcome is already locked — i.e. more prefix-completion than revision. Existing inference-/training-time safety interventions mostly push toward over-refusal while suppressing the already-scarce deliberation signal.
Paper first page (click to collapse)
Training, adaptation & efficiency
3Latency↔capability, adaptation under distribution shift, and how post-training stages compose.
Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models
Lianghua Huang, Zhifan Wu, Wei Wang, Yupeng Shi et al. · Wan Team, Alibaba Group · arXiv:2606.25041 · site
Why readHottest paper in the batch (buzz 93) — a single-Transformer, full-duplex audio-visual streamer; a fresh latency↔capability point resonant with the FLOPs-vs-latency / streaming threads in my notes.
Key ideaOne Transformer models language, audio, and video as both input and output via interleaved tokens with block-causal attention for incremental streaming — no separate VAD / ASR / TTS / animation modules. The whole stack (causal encoders/decoders, low-latency token scheduling) is redesigned for streamability, with streaming units as short as 160 ms at 25 fps.
Result~200 ms model-side response latency; ~550 ms total interaction latency (with 350 ms network), supporting sub-second full-duplex audio-visual communication.
Paper first page (click to collapse)
Distill Once, Adapt Life-Long (DO-ALL): Dataset Distillation for Continual Test-Time Adaptation
Hyun-Kurl Jang, Jihun Kim, Hyeokjun Kweon, Kuk-Jin Yoon · KAIST + Chung-Ang University · arXiv:2606.20196 · code
Why readUses dataset distillation to retain source knowledge in a compact, privacy-conscious form for stable long-term adaptation — a clean recipe for adaptation under distribution shift.
Key ideaBefore deployment, distill the source distribution into a small set of synthetic anchors. During CTTA, match each target sample to its most semantically aligned anchor and use it for source replay, representation alignment, and manifold-smoothing regularization — plug-and-play into existing CTTA methods, with no raw source data retained.
ResultConsistent long-term robustness gains across CIFAR100-C, ImageNet-C, and the CCC benchmark; counters the compounding self-training errors and forgetting of source-free CTTA.
Paper first page (click to collapse)
How Post-Training Shapes Biological Reasoning Models
Lukas Fesser, Hanlin Zhang, Michelle M. Li, Eric Wang, Bryan Perozzi, Shekoofeh Azizi, Sham M. Kakade, Marinka Zitnik · Harvard + Google DeepMind / Research · arXiv:2606.16517 · code
Why readA careful ID-vs-OOD study of how CPT / SFT / RL compose — the training-recipe lessons generalise well beyond biology and are worth stealing for any post-training pipeline.
Key ideaTrain and evaluate 100+ biological reasoning models over genomics, transcriptomics, and proteins with controlled variation in backbone, continued pre-training (CPT), SFT, and RL, measuring both in-domain and out-of-domain performance.
ResultEach stage reshapes generalization differently: CPT aligns to biological language; SFT lifts ID but makes OOD peak early then decline as it fits the training distribution; RL on strong SFT checkpoints with aligned rewards improves OOD and partially recovers generalization. Under a fixed budget, the best ID-OOD trade-off comes from brief SFT, larger RL allocation, and asymmetric adaptation capacity across stages.
Paper first page (click to collapse)
Industry pulse
Condensed from the same digest · 2026-06-28
- Specialized architectures + inference optimization are consolidating. New releases led by OpenAI’s GPT-5.6 Sol and Baidu’s Unlimited-OCR; the landscape is organizing around task-specific models rather than one-size-fits-all scale.
- Speculative decoding is the performance multiplier of the moment. DeepSpec is being adopted fast on GitHub and DSpark is trending on Hacker News — decode-side speedups are now a first-class lever.
- Agent systems are exploding as a category. From Qwen’s AgentWorld framework to autonomous trading agents like Eleven, the shift is toward multi-step reasoning and tool use over single-turn generation.
- Efficiency infra is maturing. Google and others push frozen Multi-Token Prediction and linear elastic caching, while vLLM and peer inference servers become essential deployment infrastructure for increasingly complex models.
- The market is bifurcating. Frontier providers (OpenAI, Google, Baidu) on one side; an open-source ecosystem (Qwen, Liquid AI, community fine-tunes) on the other, the latter leaning into practical optimization and domain-specific capability rather than raw scale.