← Reading list
July 3, 2026

Reading Queue — Want to Read (2026-07-03)

Tags
reading-queueworld-modelsagentsagent-memoryrlrlvrmedical-aireasoningstudy-notes
Reading Queue — Want to Read (2026-07-03)

Reading Queue · Triaged Backlog

Want to Read

Compiled 2026-07-03  ·  10 papers  ·  4 groups  ·  arXiv links (no screenshots this batch)

Good papers I want to read but haven’t had time for yet. Theme of the day: agent memory & reliability are the new frontier — raw capability is table-stakes, trustworthiness under deployment is the bottleneck. Six of these had no digest summary, so I pulled the details from the papers directly.

Priority picks (closest to my current thread)

  • WorldDirector — LLM-orchestrated 3D trajectories decoupled from rendering, with persistent object memory across occlusions; controllable long-horizon world simulation.

🔥 Hottest in the batch (by buzz) — agent memory is the story

  • AgenticSTS — buzz 37: treats memory as a bounded, typed contract on what each decision may see (Slay-the-Spire-2 testbed).
  • MemSyco-Bench — buzz 21: benchmarks memory-induced sycophancy — agents over-aligning with the user at the cost of accuracy.

World models & simulation

1

The one squarely on the world-model / robotic-control thread.

01rel 1.00🔥 buzz 13arxiv+hf⭐ thread anchor☐ to read

WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory

Hanlin Wang, Hao Ouyang, Qiuyu Wang, Wen Wang et al. · arXiv:2607.02517

Why readA world model that decouples semantic motion orchestration from rendering, with persistent object identity across occlusions — right on the controllable-world-sim thread, and a cleaner abstraction for enforcing physical consistency at scale.

Key ideaAn LLM coordinates 3D object trajectories and camera movement, then passes these structured trajectories to a video-generation module as explicit control signals — enforcing physical logic and appearance stability instead of baking dynamics into pixels end-to-end.

ResultSynthesizes complex, extended events with strict physical logic, stable appearances, and exact preservation of dynamic entity identity even after prolonged out-of-view periods — persistent memory + unrestricted viewpoint control that end-to-end methods can’t achieve.

world modelLLM orchestration3D trajectoriespersistent memoryvideo generation
Paper first page (click to collapse)p1 first page

Agents: memory, reliability & evaluation

5

The frontier this batch — how agents maintain, evaluate, and improve reasoning over long horizons.

02rel 0.95🔥 buzz 37 — hottesthf_dailycode☐ to read

AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents

Xiangchen Cheng, Yunwei Jiang, Jianwen Sun, Zizhen Li et al. · arXiv:2607.02255 · code

Why readTop of the batch (buzz 37) and a principled reframing: memory design as a contract, not an ever-growing context dump — core to reliable long-horizon agents.

Key ideaMemory for a long-horizon agent is a contract about what each future decision is allowed to see. AgenticSTS provides a bounded, typed, ablatable memory contract and uses the roguelike Slay the Spire 2 as a long-horizon testbed, so you can isolate which memory actually matters instead of feeding unbounded context.

ResultA testbed + analysis rather than one headline number: it targets the information overload and forgetting that unbounded context creates; trajectories and code are released (under EMNLP 2026 review).

agent memorybounded memorylong-horizontestbedreliability
Paper first page (click to collapse)p2 first page
03rel 0.95🔥 buzz 21hf_dailycodememory☐ to read

MemSyco-Bench: Benchmarking Sycophancy in Agent Memory

Zhishang Xiang, Zerui Chen, Yunbo Tang, Zhimin Wei et al. · Xiamen University + Jilin University · arXiv:2607.01071 · code

Why readNames a reliability failure that store/retrieve benchmarks miss: retrieved memories can make agents sycophantic. A timely robustness lens as agents become long-term collaborators.

Key ideaRetrieved memories often induce sycophancy — agents over-align with the user at the cost of factual accuracy or objective reasoning. MemSyco-Bench evaluates when memory should influence a decision and how valid memory should be used, via five tasks: rejecting memory as factual evidence, respecting its applicable scope, resolving conflicts between memory and objective evidence, and more.

ResultShifts memory evaluation from “is it stored/retrieved/updated correctly” to how retrieved memory shapes downstream reasoning and decisions.

agent memorysycophancyreliabilitybenchmark
Paper first page (click to collapse)p3 first page
04rel 0.94🔥 buzz 8hf_dailyskill-use☐ to read

SkillCoach: Self-Evolving Rubrics for Evaluating and Enhancing Agentic Skill-Use

Jiayin Zhu, Kelong Mao, Yudong Guo, Dengbo He et al. · arXiv:2607.01874

Why readProcess-level (not outcome-only) supervision for agentic skill-use — exposes failures that final-accuracy metrics hide, and doubles as a training signal.

Key ideaAuto-generate skill-grounded process rubrics from real agent rollouts and score trajectories across four dimensions — skill selection, following, composition, reflection — independent of the outcome; rubrics evolve iteratively and are reused as process supervision for trajectory filtering.

ResultEvolved rubrics improve evaluation quality, surface process failures masked by final accuracy, and give stronger supervision than outcome-only filtering (specific deltas not quantified in the abstract).

agentsskill-useprocess rubricsevaluationon-policy distillation
Paper first page (click to collapse)p4 first page
05rel 1.00🔥 buzz 4hf_dailySWE☐ to read

SWE-Interact: Reimagining SWE Benchmarks as User-Driven Long-Horizon Coding Sessions

Mohit Raghavendra, Anisha Gunjal, Aakash Sabharwal, Yunzhong He · arXiv:2606.30573

Why readFlips SWE evaluation from upfront-spec autonomous coding to realistic multi-turn, user-driven sessions — tests intent discovery and building on prior work, closer to how agents are actually used.

Key ideaA user simulator starts with vague or incomplete instructions, progressively reveals requirements, inspects the agent’s workspace, and gives targeted feedback, revisions, and new constraints until the goal is handed off — grounded in large-scale studies of real coding-agent interactions.

ResultStrong single-turn SWE performance doesn’t reliably transfer: the best models (incl. Opus 4.8 and GPT 5.5) solve ~50% of single-turn baseline tasks but only ~25% of the matching SWE-Interact tasks — measuring an orthogonal axis of interactive goal discovery and iterative refinement with a user in the loop.

SWEcoding agentsinteractivemulti-turnbenchmark
Paper first page (click to collapse)p5 first page
06rel 0.95🔥 buzz 4hf_dailyscience agent☐ to read

Autonomous Scientific Discovery via Iterative Meta-Reflection (DiscoPER)

Bingchen Zhao, Sara Beery, Oisin Mac Aodha · University of Edinburgh + MIT · arXiv:2607.01131

Why readOpen-ended autonomous research with a meta-reflection loop — an agent that synthesizes its own accumulated findings rather than generating one-off hypotheses; the iterative-research direction to track.

Key ideaDiscoPER dynamically generates and executes code to explore datasets without pre-specified research objectives; every proposed discovery must pass statistical testing. A second-order (meta-reflection) mechanism periodically analyzes its own accumulated discoveries to surface complex, interconnected phenomena.

ResultOn iNatDisco (a new multimodal ecological-knowledge benchmark with pattern-level ground truth), DiscoPER recovers 8 of 9 known patterns at a 72.7% hypothesis-support rate, outperforming classical causal discovery and LLM-guided baselines; it scales with more data.

agentsscientific discoverymeta-reflectionopen-endedcode execution
Paper first page (click to collapse)p6 first page

RL & reasoning training

2

Mechanistic clarity on GRPO-family methods, and smarter multi-domain RLVR curricula.

07rel 0.97🔥 buzz 3hf_dailyRL theory☐ to read

GRPO, Dr. GRPO, and DAPO Are Three Operations on One Number: The Group-Standard-Deviation Identity

Yong Yi Bay, Kathleen A. Yearick · arXiv:2607.00152

Why readTheoretical clarity on three trendy RL-for-reasoning methods — they turn out to be one dial, not three tricks. A cleaner mental model for RLVR / on-policy-distillation pipelines.

Key ideaGRPO (divides by group std-dev), Dr. GRPO (omits the division), and DAPO (discards zero-std-dev groups) are unified by a group-standard-deviation identity — three settings on one control dial. The insight: split groups (high disagreement) drive learning; unanimous groups (zero std-dev) contribute nothing.

ResultProves the identity governs both the magnitude of updates and which problems deserve weight and retry budget; confirmed on Big-Math and controlled runs (no perf deltas in the abstract).

RLGRPODAPORLVRreasoning
Paper first page (click to collapse)p7 first page
08rel 1.00🔥 buzz 3hf_dailycode☐ to read

Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR (TAC)

Yongjin Yang, Jiarui Liu, Yinghui He, Lezhen Zhang, Bernhard Schölkopf, Zhijing Jin · Toronto/Vector, CMU, Princeton et al. · arXiv:2606.25178 · code

Why readA near-free curriculum that samples domains by how much their updates help the rest of the suite — squarely in the multi-domain RLVR trend, with a clean signal-reuse trick.

Key ideaTransfer-Aware Curriculum (TAC), a bandit-style online curriculum that prioritizes domains whose gradient steps broadly benefit other domains. It reuses signals RL already produces: per-domain advantages (local learnability) + projected gradients from the GRPO step (cross-domain transferability via gradient-geometry alignment), at <1% wall-clock overhead.

ResultBest macro-averaged accuracy on a six-domain suite for Qwen3-1.7B and Llama3.2-3B; beats proportional random sampling, a hand-designed schedule, and a learnability-only bandit — by up to 2.8 points (10% relative) over the last.

RLVRcurriculummulti-domaintransferGRPO
Paper first page (click to collapse)p8 first page

Medical AI

2

Clinical reasoning crossing from reasoning-only into deployable, editable workflows.

09rel 0.96🔥 buzz 11hf_dailymedical RL☐ to read

Breaking Failure Cascades: Step-Aware Reinforcement Learning for Medical Multimodal Reasoning

Junha Jung, Minbyul Jeong, Suhyeon Lim, Sungwook Jung et al. · arXiv:2606.31825

Why readStep-level credit assignment for medical multimodal reasoning — targets the early-stage cascading errors that outcome-only RL can’t fix.

Key ideaOutcome-centric post-training gives sparse credit assignment, and their analysis shows early-stage reasoning failures propagate into cascades that drive wrong answers in medical VQA. They propose MRPO (Medical Reasoning-aware Policy Optimization), which reshapes the GRPO advantage to assign exponentially larger penalties to tokens in earlier invalid reasoning steps — breaking cascades without hurting already-successful traces.

ResultOn Qwen3-VL-8B, MRPO beats standard GRPO and a recent RL baseline and surpasses much larger medical MLLMs like HuatuoGPT-Vision-34B by 2.79 points; it cuts early-stage reasoning failures from 64.0% to 13.0%.

medical AImultimodal reasoningstep-aware RLcredit assignmentVQA
Paper first page (click to collapse)p9 first page
10rel 0.95🔥 buzz 4hf_dailydiffusion LM☐ to read

Discrete Diffusion Language Models for Interactive Radiology Report Drafting

Max Van Puyvelde, Halil Ibrahim Gulluk, Wim Van Criekinge, Olivier Gevaert · arXiv:2607.01436

Why readDiscrete diffusion (not autoregressive) for radiology drafting — enables any-order infill so radiologists edit fragments instead of regenerating, and it’s 3.5–4.4× faster. A rare inference-efficiency + clinical-workflow win.

Key ideaAdapt discrete diffusion LMs (bidirectional token denoising) to medical VQA and radiology reporting. Bidirectional denoising gives inherent any-order infill — fix report fragments and have the model fill the gaps — a capability AR models can’t natively provide. They fine-tune DiffusionGemma-26B (MoE) against a same-size AR Gemma-4-26B under identical LoRA recipes.

ResultDiffusion matches or exceeds AR on all medical VQA benchmarks; competitive with frontier VLMs at 3.8B active params; decoding 3.5–4.4× faster; infill is inherent to diffusion but structurally absent from AR.

medical AIdiscrete diffusioninfillradiologyinference efficiency
Paper first page (click to collapse)p10 first page

Industry pulse

Condensed from the same digest · 2026-07-03

  • Consolidation around multimodal + agentic. Anthropic’s Claude Fable 5 and Mythos 5 are building commercial momentum after U.S. export controls were lifted, while open-source alternatives like Qwythos-9B and GLM-5.2 gain traction in image-text and generation tasks.
  • Practical agent tooling is where devs are. Self-learning skill frameworks, video understanding for LLMs, and browser-automation agents are seeing strong adoption — a shift from raw scale toward real-world agent autonomy and task automation.
  • Labs push specialized applications. Voice-enabled Gemma 4 models, enterprise-focused Java-migration benchmarks, and tabular-data foundation models.
  • Regulatory headwinds are broadening. From Japan’s stance on AI patent inventorship to Spain’s Palantir restrictions and ongoing export-control debates.
  • An enthusiasm gap is widening. Some skepticism is emerging about AI’s pace and hype cycles as use cases move from headline-grabbing abstractions toward grounded enterprise and consumer productivity.