← Reading list
June 29, 2026

Reading Queue — Want to Read

Tags
reading-queueembodied-aiworld-modelsagentsinference-efficiencystudy-notes
Reading Queue — Want to Read (2026-06-29)

Reading Queue · Triaged Backlog

Want to Read

Compiled 2026-06-29  ·  10 papers  ·  3 groups  ·  from hf_daily / arXiv pulse

Good papers I want to read but haven’t had time for yet — sorted by theme, with relevance and buzz signals. Stars mark the ones closest to the world-model / robotic-control thread I’m already reading.

Priority picks (closest to my current thread)

  • PhysisForcing — physics-reinforced video world simulator; the diffusion counterpart to PhysiFormer’s explicit-physics approach. Buzz 30 — hottest in the batch.
  • SimFoundry — automated real→sim digital twins + “digital cousins”; the sim-to-real bridge I flagged in the PhysiFormer notes.
  • Learning to Fold — LeHome 2026 winner; a unified VLA head that doubles as its own value/failure detector.

Embodied AI, world models & sim-to-real

4

Directly adjacent to the PhysiFormer / Fast-LeWM / ICWM thread.

07rel 0.90🔥 buzz 30arxiv+hf☐ to read

PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation

Peiwen Zhang, Yufan Deng, Shangkun Sun, Juncheng Ma et al. · arXiv:2606.28128

Why readThe video-diffusion route to physical consistency, complementing PhysiFormer’s explicit 3D approach — same problem (physically implausible rollouts), different substrate.

Key ideaSupervise DiT diffusion features at pixel-level (point-trajectory alignment) and semantic-level (relational alignment via a frozen video encoder) to focus training on physics-informative contact regions.

Result+22.3% / +9.2% on R-Bench; closed-loop WorldArena success 16.0% → 24.0%; improves downstream policy success.

world modelvideo diffusioncontact physicsmanipulation
05rel 0.97🔥 6hf_daily☐ to read

SimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluation

Nadun Ranawaka, Josiah Wong, Wei-Lin Pai, Wei-Teng Chu et al. · arXiv:2606.28276

Why readThe concrete sim-to-real bridge I sketched in the PhysiFormer notes: reconstruct sim-ready twins from video, then vary them to generalize.

Key ideaZero-shot real→sim digital-twin reconstruction from video + procedural “digital cousins” (affordance-preserving object/scene/task variations) for policy generalization.

ResultSim–real correlation Pearson r=0.911 across 7 tasks × 5 architectures; +17% / +21% / +40% real success from object/scene/task cousins.

sim-to-realdigital twinsdata generationpolicy learning
03rel 1.00🔥 4hf_daily☐ to read

Learning to Fold: prizewinning solution at LeHome Challenge 2026

Ilia Larchenko · 1st online / 2nd offline of 62 teams · arXiv:2606.27163

Why readA clean RL+VLA recipe where the policy head is its own value/failure detector — resonates with the auxiliary-prediction and on-the-fly-adaptation ideas in my notes.

Key ideaRepurpose the VLA head as a multi-task predictor (actions + success/progress/future states) for self-supervised advantage estimation and live failure detection; AWR + RECAP flow-matching, Thompson-sampling inference-time tuning.

VLARLbimanualfailure detectionsim-to-real
04rel 1.00🔥 3no abstract summary☐ to read

Object-Centric Residual RL for Zero-Shot Sim-to-Real VLA Enhancement

Kinam Kim, Namiko Saito, Heecheol Kim, Katsushi Ikeuchi et al. · arXiv:2606.18953

Why readObject-centric + residual RL for VLA sim-to-real robustness — pairs with the core-periphery / object-structure idea I’m chasing. (Digest had no generated summary; read the abstract first.)

residual RLobject-centricVLAzero-shot

Agentic systems, tool use & reasoning

4

Production reliability, multi-agent optimization, and adaptive tool orchestration.

09rel 0.90🔥 2hf_daily☐ to read

Constraint Tax in Open-Weight LLMs: Tool-Calling Suppression Under Structured Output Constraints

Fangzheng Li, Aimin Zhang, Chen Lv · arXiv:2606.25605

Why readA crisp, fixable production failure mode — the kind of implementation-level gotcha worth knowing before it bites.

Key ideaGrammar-based JSON-schema masking removes tool-call tokens from the reachable set (a priority inversion); fix = Transparent Two-Pass Execution (tool selection unconstrained, then schema-compliant generation), no retraining.

tool useconstrained decodingagentsreliability
06rel 0.93🔥 3arxiv+hf☐ to read

Towards Automating Scientific Review with Google’s Paper Assistant Tool (PAT)

Rajesh Jayaram, Drew Tyler, David Woodruff, Corinna Cortes et al. · arXiv:2606.28277

Why readAgentic review + inference scaling (test-time compute) for a high-stakes verification task; piloted at STOC and ICML.

Key ideaMulti-step agent ingests full manuscripts, checks theorems / experiments / methodology; inference scaling yields +34% over zero-shot recall on the SPOT benchmark.

agentictest-time computepeer reviewverification
08rel 0.90🔥 7hf_daily☐ to read

ProMSA: Progressive Multimodal Search Agents for Knowledge-Based VQA

ZhengXian Wu, Hangrui Xu, Kai Shi, Zhuohong Chen et al. · arXiv:2606.27974

Why readAdaptive multi-tool retrieval with budget awareness, trained by a tool-aware RL objective — a tidy agentic-RAG design.

Key ideaAgent iteratively chooses image-search / text-search / stop under tool-call budgets; two-stage training (rejection-sampling SFT, then sequence-level RL via TN-GSPO normalizing by generation length and tool depth).

multimodalretrievaltool budgetsRL

Inference efficiency & generative

2

Latency↔capability tradeoffs and RL for generative models.

01rel 1.00🔥 4hf_daily☐ to read

Thinking While Speaking: Inference-Time Knowledge Transfer for Conversational Voice Agents (ConvFill)

Vidya Srinivas, Zachary Englhardt, Shwetak Patel, Vikram Iyer et al. · arXiv:2511.07397

Why readA new latency↔capability Pareto point — directly resonant with the FLOPs-vs-latency and sequential-vs-parallel threads from the Fast-LeWM notes.

Key ideaConversational infill: a small “talker” (135M–1.7B) emits an immediate response while a larger “reasoner”’s tokens are streamed and fluently infilled, hiding reasoner latency.

ResultMillisecond time-to-first-response within 6.3% of frontier reasoner accuracy; n=18 study on Apple M2 rates it on par overall, preferred for retrieval-heavy tasks.

on-devicestreaminglatencyknowledge transfer
10rel 0.87🔥 21no abstract summary☐ to read

Qwen-Image-2.0-RL Technical Report

Yixian Xu, Kaiyuan Gao, Yuxiang Chen, Yilei Chen et al. · arXiv:2606.27608

Why readRLHF + on-policy distillation (OPD) applied to diffusion image models at industrial scale — on-policy distillation is worth tracking. (Digest had no generated summary.)

diffusionRLHFon-policy distillationtech report

Industry pulse

Condensed from the same digest · 2026-06-29

  • Inference optimization is consolidating. Speculative decoding is the lever of the moment (DeepSpec going viral, multiple labs cutting latency); new model releases from Baidu (Unlimited-OCR) and Zhipu (GLM 5.2).
  • Agents are moving research→production. OpenAI workforce-mapping, Google’s computer use in Gemini 3.5 Flash, and open-source routers like OpenTag (dispatching across Claude and Codex) point to autonomous-agent infrastructure maturing.
  • Enterprise AI is going domain-specific. HP×OpenAI, HackerRank open-sourcing its ATS, and Microsoft’s Talos rare-disease diagnosis signal a shift from chatbot novelty to operational tools.
  • The replacement narrative is tempering. Ford rehired engineers after AI underperformed on manufacturing tasks — a useful counterweight.
  • Evaluation rigor is becoming the differentiator. Benchmarking resources (e.g. awesome-evals) are gaining traction as teams realize assessment, not just scale, drives production value.