Reading Queue · Triaged Backlog
Want to Read
Good papers I want to read but haven’t had time for yet — sorted by theme, with relevance and buzz signals. Stars mark the ones closest to the world-model / robotic-control thread I’m already reading.
Priority picks (closest to my current thread)
- PhysisForcing — physics-reinforced video world simulator; the diffusion counterpart to PhysiFormer’s explicit-physics approach. Buzz 30 — hottest in the batch.
- SimFoundry — automated real→sim digital twins + “digital cousins”; the sim-to-real bridge I flagged in the PhysiFormer notes.
- Learning to Fold — LeHome 2026 winner; a unified VLA head that doubles as its own value/failure detector.
Embodied AI, world models & sim-to-real
4Directly adjacent to the PhysiFormer / Fast-LeWM / ICWM thread.
PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation
Peiwen Zhang, Yufan Deng, Shangkun Sun, Juncheng Ma et al. · arXiv:2606.28128
Why readThe video-diffusion route to physical consistency, complementing PhysiFormer’s explicit 3D approach — same problem (physically implausible rollouts), different substrate.
Key ideaSupervise DiT diffusion features at pixel-level (point-trajectory alignment) and semantic-level (relational alignment via a frozen video encoder) to focus training on physics-informative contact regions.
Result+22.3% / +9.2% on R-Bench; closed-loop WorldArena success 16.0% → 24.0%; improves downstream policy success.
SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation
Nadun Ranawaka, Josiah Wong, Wei-Lin Pai, Wei-Teng Chu et al. · arXiv:2606.28276
Why readThe concrete sim-to-real bridge I sketched in the PhysiFormer notes: reconstruct sim-ready twins from video, then vary them to generalize.
Key ideaZero-shot real→sim digital-twin reconstruction from video + procedural “digital cousins” (affordance-preserving object/scene/task variations) for policy generalization.
ResultSim–real correlation Pearson r=0.911 across 7 tasks × 5 architectures; +17% / +21% / +40% real success from object/scene/task cousins.
Learning to Fold: prizewinning solution at LeHome Challenge 2026
Ilia Larchenko · 1st online / 2nd offline of 62 teams · arXiv:2606.27163
Why readA clean RL+VLA recipe where the policy head is its own value/failure detector — resonates with the auxiliary-prediction and on-the-fly-adaptation ideas in my notes.
Key ideaRepurpose the VLA head as a multi-task predictor (actions + success/progress/future states) for self-supervised advantage estimation and live failure detection; AWR + RECAP flow-matching, Thompson-sampling inference-time tuning.
Object-Centric Residual RL for Zero-Shot Sim-to-Real VLA Enhancement
Kinam Kim, Namiko Saito, Heecheol Kim, Katsushi Ikeuchi et al. · arXiv:2606.18953
Why readObject-centric + residual RL for VLA sim-to-real robustness — pairs with the core-periphery / object-structure idea I’m chasing. (Digest had no generated summary; read the abstract first.)
Agentic systems, tool use & reasoning
4Production reliability, multi-agent optimization, and adaptive tool orchestration.
Constraint Tax in Open-Weight LLMs: Tool-Calling Suppression Under Structured Output Constraints
Fangzheng Li, Aimin Zhang, Chen Lv · arXiv:2606.25605
Why readA crisp, fixable production failure mode — the kind of implementation-level gotcha worth knowing before it bites.
Key ideaGrammar-based JSON-schema masking removes tool-call tokens from the reachable set (a priority inversion); fix = Transparent Two-Pass Execution (tool selection unconstrained, then schema-compliant generation), no retraining.
GBC: Gradient-Based Connections for Optimizing Multi-Agent Systems
Xiaocheng Yang, Abdulrahman Alrabah, Dilek Hakkani-Tür, Gokhan Tur · arXiv:2606.28187
Why readGradient-based credit assignment for multi-agent LLM coordination — fine-grained optimization across agents. (Digest had no generated summary.)
Towards Automating Scientific Review with Google’s Paper Assistant Tool (PAT)
Rajesh Jayaram, Drew Tyler, David Woodruff, Corinna Cortes et al. · arXiv:2606.28277
Why readAgentic review + inference scaling (test-time compute) for a high-stakes verification task; piloted at STOC and ICML.
Key ideaMulti-step agent ingests full manuscripts, checks theorems / experiments / methodology; inference scaling yields +34% over zero-shot recall on the SPOT benchmark.
ProMSA: Progressive Multimodal Search Agents for Knowledge-Based VQA
ZhengXian Wu, Hangrui Xu, Kai Shi, Zhuohong Chen et al. · arXiv:2606.27974
Why readAdaptive multi-tool retrieval with budget awareness, trained by a tool-aware RL objective — a tidy agentic-RAG design.
Key ideaAgent iteratively chooses image-search / text-search / stop under tool-call budgets; two-stage training (rejection-sampling SFT, then sequence-level RL via TN-GSPO normalizing by generation length and tool depth).
Inference efficiency & generative
2Latency↔capability tradeoffs and RL for generative models.
Thinking While Speaking: Inference-Time Knowledge Transfer for Conversational Voice Agents (ConvFill)
Vidya Srinivas, Zachary Englhardt, Shwetak Patel, Vikram Iyer et al. · arXiv:2511.07397
Why readA new latency↔capability Pareto point — directly resonant with the FLOPs-vs-latency and sequential-vs-parallel threads from the Fast-LeWM notes.
Key ideaConversational infill: a small “talker” (135M–1.7B) emits an immediate response while a larger “reasoner”’s tokens are streamed and fluently infilled, hiding reasoner latency.
ResultMillisecond time-to-first-response within 6.3% of frontier reasoner accuracy; n=18 study on Apple M2 rates it on par overall, preferred for retrieval-heavy tasks.
Qwen-Image-2.0-RL Technical Report
Yixian Xu, Kaiyuan Gao, Yuxiang Chen, Yilei Chen et al. · arXiv:2606.27608
Why readRLHF + on-policy distillation (OPD) applied to diffusion image models at industrial scale — on-policy distillation is worth tracking. (Digest had no generated summary.)
Industry pulse
Condensed from the same digest · 2026-06-29
- Inference optimization is consolidating. Speculative decoding is the lever of the moment (DeepSpec going viral, multiple labs cutting latency); new model releases from Baidu (Unlimited-OCR) and Zhipu (GLM 5.2).
- Agents are moving research→production. OpenAI workforce-mapping, Google’s computer use in Gemini 3.5 Flash, and open-source routers like OpenTag (dispatching across Claude and Codex) point to autonomous-agent infrastructure maturing.
- Enterprise AI is going domain-specific. HP×OpenAI, HackerRank open-sourcing its ATS, and Microsoft’s Talos rare-disease diagnosis signal a shift from chatbot novelty to operational tools.
- The replacement narrative is tempering. Ford rehired engineers after AI underperformed on manufacturing tasks — a useful counterweight.
- Evaluation rigor is becoming the differentiator. Benchmarking resources (e.g. awesome-evals) are gaining traction as teams realize assessment, not just scale, drives production value.