Reading Queue · Triaged Backlog
Want to Read
Good papers I want to read but haven’t had time for yet. Theme of the day: agent memory & reliability are the new frontier — raw capability is table-stakes, trustworthiness under deployment is the bottleneck. Six of these had no digest summary, so I pulled the details from the papers directly.
Priority picks (closest to my current thread)
- WorldDirector — LLM-orchestrated 3D trajectories decoupled from rendering, with persistent object memory across occlusions; controllable long-horizon world simulation.
🔥 Hottest in the batch (by buzz) — agent memory is the story
- AgenticSTS — buzz 37: treats memory as a bounded, typed contract on what each decision may see (Slay-the-Spire-2 testbed).
- MemSyco-Bench — buzz 21: benchmarks memory-induced sycophancy — agents over-aligning with the user at the cost of accuracy.
World models & simulation
1The one squarely on the world-model / robotic-control thread.
WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory
Hanlin Wang, Hao Ouyang, Qiuyu Wang, Wen Wang et al. · arXiv:2607.02517
Why readA world model that decouples semantic motion orchestration from rendering, with persistent object identity across occlusions — right on the controllable-world-sim thread, and a cleaner abstraction for enforcing physical consistency at scale.
Key ideaAn LLM coordinates 3D object trajectories and camera movement, then passes these structured trajectories to a video-generation module as explicit control signals — enforcing physical logic and appearance stability instead of baking dynamics into pixels end-to-end.
ResultSynthesizes complex, extended events with strict physical logic, stable appearances, and exact preservation of dynamic entity identity even after prolonged out-of-view periods — persistent memory + unrestricted viewpoint control that end-to-end methods can’t achieve.
Paper first page (click to collapse)
Agents: memory, reliability & evaluation
5The frontier this batch — how agents maintain, evaluate, and improve reasoning over long horizons.
AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents
Xiangchen Cheng, Yunwei Jiang, Jianwen Sun, Zizhen Li et al. · arXiv:2607.02255 · code
Why readTop of the batch (buzz 37) and a principled reframing: memory design as a contract, not an ever-growing context dump — core to reliable long-horizon agents.
Key ideaMemory for a long-horizon agent is a contract about what each future decision is allowed to see. AgenticSTS provides a bounded, typed, ablatable memory contract and uses the roguelike Slay the Spire 2 as a long-horizon testbed, so you can isolate which memory actually matters instead of feeding unbounded context.
ResultA testbed + analysis rather than one headline number: it targets the information overload and forgetting that unbounded context creates; trajectories and code are released (under EMNLP 2026 review).
Paper first page (click to collapse)
MemSyco-Bench: Benchmarking Sycophancy in Agent Memory
Zhishang Xiang, Zerui Chen, Yunbo Tang, Zhimin Wei et al. · Xiamen University + Jilin University · arXiv:2607.01071 · code
Why readNames a reliability failure that store/retrieve benchmarks miss: retrieved memories can make agents sycophantic. A timely robustness lens as agents become long-term collaborators.
Key ideaRetrieved memories often induce sycophancy — agents over-align with the user at the cost of factual accuracy or objective reasoning. MemSyco-Bench evaluates when memory should influence a decision and how valid memory should be used, via five tasks: rejecting memory as factual evidence, respecting its applicable scope, resolving conflicts between memory and objective evidence, and more.
ResultShifts memory evaluation from “is it stored/retrieved/updated correctly” to how retrieved memory shapes downstream reasoning and decisions.
Paper first page (click to collapse)
SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use
Jiayin Zhu, Kelong Mao, Yudong Guo, Dengbo He et al. · arXiv:2607.01874
Why readProcess-level (not outcome-only) supervision for agentic skill-use — exposes failures that final-accuracy metrics hide, and doubles as a training signal.
Key ideaAuto-generate skill-grounded process rubrics from real agent rollouts and score trajectories across four dimensions — skill selection, following, composition, reflection — independent of the outcome; rubrics evolve iteratively and are reused as process supervision for trajectory filtering.
ResultEvolved rubrics improve evaluation quality, surface process failures masked by final accuracy, and give stronger supervision than outcome-only filtering (specific deltas not quantified in the abstract).
Paper first page (click to collapse)
SWE-Interact: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions
Mohit Raghavendra, Anisha Gunjal, Aakash Sabharwal, Yunzhong He · arXiv:2606.30573
Why readFlips SWE evaluation from upfront-spec autonomous coding to realistic multi-turn, user-driven sessions — tests intent discovery and building on prior work, closer to how agents are actually used.
Key ideaA user simulator starts with vague or incomplete instructions, progressively reveals requirements, inspects the agent’s workspace, and gives targeted feedback, revisions, and new constraints until the goal is handed off — grounded in large-scale studies of real coding-agent interactions.
ResultStrong single-turn SWE performance doesn’t reliably transfer: the best models (incl. Opus 4.8 and GPT 5.5) solve ~50% of single-turn baseline tasks but only ~25% of the matching SWE-Interact tasks — measuring an orthogonal axis of interactive goal discovery and iterative refinement with a user in the loop.
Paper first page (click to collapse)
Autonomous Scientific Discovery via Iterative Meta-Reflection (DiscoPER)
Bingchen Zhao, Sara Beery, Oisin Mac Aodha · University of Edinburgh + MIT · arXiv:2607.01131
Why readOpen-ended autonomous research with a meta-reflection loop — an agent that synthesizes its own accumulated findings rather than generating one-off hypotheses; the iterative-research direction to track.
Key ideaDiscoPER dynamically generates and executes code to explore datasets without pre-specified research objectives; every proposed discovery must pass statistical testing. A second-order (meta-reflection) mechanism periodically analyzes its own accumulated discoveries to surface complex, interconnected phenomena.
ResultOn iNatDisco (a new multimodal ecological-knowledge benchmark with pattern-level ground truth), DiscoPER recovers 8 of 9 known patterns at a 72.7% hypothesis-support rate, outperforming classical causal discovery and LLM-guided baselines; it scales with more data.
Paper first page (click to collapse)
RL & reasoning training
2Mechanistic clarity on GRPO-family methods, and smarter multi-domain RLVR curricula.
GRPO, Dr. GRPO, and DAPO Are Three Operations on One Number: The Group-Standard-Deviation Identity
Yong Yi Bay, Kathleen A. Yearick · arXiv:2607.00152
Why readTheoretical clarity on three trendy RL-for-reasoning methods — they turn out to be one dial, not three tricks. A cleaner mental model for RLVR / on-policy-distillation pipelines.
Key ideaGRPO (divides by group std-dev), Dr. GRPO (omits the division), and DAPO (discards zero-std-dev groups) are unified by a group-standard-deviation identity — three settings on one control dial. The insight: split groups (high disagreement) drive learning; unanimous groups (zero std-dev) contribute nothing.
ResultProves the identity governs both the magnitude of updates and which problems deserve weight and retry budget; confirmed on Big-Math and controlled runs (no perf deltas in the abstract).
Paper first page (click to collapse)
Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR (TAC)
Yongjin Yang, Jiarui Liu, Yinghui He, Lezhen Zhang, Bernhard Schölkopf, Zhijing Jin · Toronto/Vector, CMU, Princeton et al. · arXiv:2606.25178 · code
Why readA near-free curriculum that samples domains by how much their updates help the rest of the suite — squarely in the multi-domain RLVR trend, with a clean signal-reuse trick.
Key ideaTransfer-Aware Curriculum (TAC), a bandit-style online curriculum that prioritizes domains whose gradient steps broadly benefit other domains. It reuses signals RL already produces: per-domain advantages (local learnability) + projected gradients from the GRPO step (cross-domain transferability via gradient-geometry alignment), at <1% wall-clock overhead.
ResultBest macro-averaged accuracy on a six-domain suite for Qwen3-1.7B and Llama3.2-3B; beats proportional random sampling, a hand-designed schedule, and a learnability-only bandit — by up to 2.8 points (10% relative) over the last.
Paper first page (click to collapse)
Medical AI
2Clinical reasoning crossing from reasoning-only into deployable, editable workflows.
Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning
Junha Jung, Minbyul Jeong, Suhyeon Lim, Sungwook Jung et al. · arXiv:2606.31825
Why readStep-level credit assignment for medical multimodal reasoning — targets the early-stage cascading errors that outcome-only RL can’t fix.
Key ideaOutcome-centric post-training gives sparse credit assignment, and their analysis shows early-stage reasoning failures propagate into cascades that drive wrong answers in medical VQA. They propose MRPO (Medical Reasoning-aware Policy Optimization), which reshapes the GRPO advantage to assign exponentially larger penalties to tokens in earlier invalid reasoning steps — breaking cascades without hurting already-successful traces.
ResultOn Qwen3-VL-8B, MRPO beats standard GRPO and a recent RL baseline and surpasses much larger medical MLLMs like HuatuoGPT-Vision-34B by 2.79 points; it cuts early-stage reasoning failures from 64.0% to 13.0%.
Paper first page (click to collapse)
Discrete Diffusion Language Models for Interactive Radiology Report Drafting
Max Van Puyvelde, Halil Ibrahim Gulluk, Wim Van Criekinge, Olivier Gevaert · arXiv:2607.01436
Why readDiscrete diffusion (not autoregressive) for radiology drafting — enables any-order infill so radiologists edit fragments instead of regenerating, and it’s 3.5–4.4× faster. A rare inference-efficiency + clinical-workflow win.
Key ideaAdapt discrete diffusion LMs (bidirectional token denoising) to medical VQA and radiology reporting. Bidirectional denoising gives inherent any-order infill — fix report fragments and have the model fill the gaps — a capability AR models can’t natively provide. They fine-tune DiffusionGemma-26B (MoE) against a same-size AR Gemma-4-26B under identical LoRA recipes.
ResultDiffusion matches or exceeds AR on all medical VQA benchmarks; competitive with frontier VLMs at 3.8B active params; decoding 3.5–4.4× faster; infill is inherent to diffusion but structurally absent from AR.
Paper first page (click to collapse)
Industry pulse
Condensed from the same digest · 2026-07-03
- Consolidation around multimodal + agentic. Anthropic’s Claude Fable 5 and Mythos 5 are building commercial momentum after U.S. export controls were lifted, while open-source alternatives like Qwythos-9B and GLM-5.2 gain traction in image-text and generation tasks.
- Practical agent tooling is where devs are. Self-learning skill frameworks, video understanding for LLMs, and browser-automation agents are seeing strong adoption — a shift from raw scale toward real-world agent autonomy and task automation.
- Labs push specialized applications. Voice-enabled Gemma 4 models, enterprise-focused Java-migration benchmarks, and tabular-data foundation models.
- Regulatory headwinds are broadening. From Japan’s stance on AI patent inventorship to Spain’s Palantir restrictions and ongoing export-control debates.
- An enthusiasm gap is widening. Some skepticism is emerging about AI’s pace and hype cycles as use cases move from headline-grabbing abstractions toward grounded enterprise and consumer productivity.