← Reading list
July 4, 2026

Reading Queue — Want to Read (2026-07-04)

Tags
reading-queueembodied-aivlaagentsagent-memorybenchmarksinference-efficiencydistillationstudy-notes
Reading Queue — Want to Read (2026-07-04)

Reading Queue · Triaged Backlog

Want to Read

Compiled 2026-07-04  ·  10 papers  ·  4 groups  ·  arXiv links + first-page screenshots

Good papers I want to read but haven’t had time for yet. Theme of the day: agentic benchmarking & capability measurement and memory as a trainable skill — plus a couple of cautionary results on naive test-time scaling and self-distillation. Six had no digest summary, so I pulled the details from the screenshots directly.

Priority picks (closest to my current thread)

🔥 Hottest in the batch (by buzz)

  • EvoPolicyGym — buzz 41: makes autonomous policy self-improvement a first-class, budget-constrained eval with trajectory diagnostics.
  • AgenticDataBench — buzz 25: fine-grained, skill-level benchmark for data agents.
  • Seed2.0 Model Card — buzz 23: ByteDance’s Pro/Lite/Mini frontier family aimed at real-world complexity.

Embodied AI & VLA

1

The one squarely on the world-model / robotic-control thread.

01rel 0.97🔥 buzz 4arxiv+hf⭐ thread anchor☐ to read

Learning to Move Before Learning to Do: Task-Agnostic Pretraining for VLAs (TAP)

Junhao Shi, Siyin Wang, Xiaopeng Yu, Li Ji, Jingjing Gong, Xipeng Qiu · Fudan University + Shanghai Innovation Institute · arXiv:2607.02466 · homepage · code

Why readTask-agnostic VLA pretraining that attacks demonstration scarcity head-on — learn how to move from cheap unlabeled data first, ground in language later. Right on the embodied / VLA thread.

Key ideaA “Decomposition Hypothesis”: VLA learning conflates physical competence (how to move) with semantic alignment (what to do). TAP is two-stage — stage 1 learns transferable motor priors from cheap, unlabeled interaction (including discarded off-task trajectories and autonomous robot play) via a self-supervised Inverse Dynamics objective; a lightweight stage 2 grounds those priors in language using minimal expert data.

ResultOn SIMPLER, matches models trained on 1M+ expert trajectories while using orders of magnitude less labeled data (+10% absolute over standard behavior cloning). On real WidowX, retains 25% success under camera perturbations where internet-scale baselines collapse to 0%.

VLAtask-agnostic pretraininginverse dynamicsmotor priorsembodied AI
Paper first page (click to collapse)TAP first page

Agentic benchmarking & capability measurement

4

The hottest convergence this batch — instrumenting how agents improve and what actually predicts real-world performance.

02rel 0.90🔥 buzz 41 — hottestarxiv+hfagent eval☐ to read

EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments

Zhilin Wang, Han Song, Runzhe Zhan, Jusen Du et al. · arXiv:2607.02440 · code

Why readHottest in the batch (buzz 41). Treats autonomous policy improvement as a first-class, budget-constrained evaluation — not a collapsed final score — with trajectory-level diagnostics of how agents actually spend feedback.

Key idea“Autonomous Policy Evolution”: a harness-model repeatedly edits an executable policy under a fixed interaction budget. EvoPolicyGym instantiates this across 16 compact interactive RL environments, logging edits, feedback allocation, and refinements to measure both aggregate performance and how agents convert feedback into parametric improvements.

ResultGPT-5.5 achieves the strongest aggregate rank and top-two performance across all 16 environments; trajectory analysis shows strong policy evolution depends on discovering task-appropriate mechanisms, not isolated task wins.

agent evaluationpolicy evolutionself-improvementbudget-constraineddiagnostics
Paper first page (click to collapse)EvoPolicyGym first page
03rel 0.95🔥 buzz 25hf_dailydata agents☐ to read

AgenticDataBench: A Comprehensive Benchmark for Data Agents

Zhaoyan Sun, Shan Zhong, Daizhou Wen, Jiaxing Han et al. · Tsinghua University + Ant Group · arXiv:2607.01647

Why readA data-agent benchmark with fine-grained, skill-level ground truth — moves beyond coarse final-answer scoring to diagnose where data-science agents actually fail. Buzz 25.

Key ideaA benchmark for data agents spanning diverse domains with fine-grained ground-truth labels. Built from 15 real datasets (including 5 real B2B use cases from a leading fintech company) plus data-science skills (e.g., handling missing data) mined from Stack Overflow via skill-aligned hierarchical clustering; LLM-based generation fills domains lacking real tasks.

ResultEnables evaluation of the full data-science workflow — planning → iterative execution → termination — at both task and skill granularity across current data agents.

data agentsbenchmarkskill-leveldata scienceevaluation
Paper first page (click to collapse)AgenticDataBench first page
04rel 0.87🔥 buzz 6hf_dailycode☐ to read

PACE: A Proxy for Agentic Capability Evaluation

Yueqi Song, Lintang Sutawika, Jiarui Liu, Lindia Tjuatja et al. · CMU + Salesforce AI Research · arXiv:2607.02032 · code

Why readA cheap proxy that predicts expensive agentic-benchmark performance from small atomic evals — deployable today for model selection/triage. The “agentic capability is compressible” idea is worth stealing.

Key ideaPACE selects a small, curated subset of atomic (non-agentic) evaluation instances from 19 non-agentic benchmarks and fits a regression mapping their scores to full agentic-benchmark performance. The subset is chosen by dual instance-selection: target-relevance local ranking + globally informative global selection.

ResultPACE-Bench hits leave-one-out MAE <4%, Spearman >0.80, and ~85% pairwise model-ranking accuracy across 14 models and 4 agentic targets — at <1% of full agentic-eval cost.

agent evaluationproxy benchmarkmodel selectionregressioncost
Paper first page (click to collapse)PACE first page
05rel 1.00🔥 buzz 3hf_dailycode☐ to read

HealthAgentBench: Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents

Qianchu Liu, Sheng Zhang, Guanghui Qin, Jeya Maria Jose Valanarasu et al. · Microsoft Research · arXiv:2606.31179 · code

Why readMedical AI as a serious agentic frontier — end-to-end clinical workflows, not LLM+RAG toys, and a hard, realistic bar (best agent only ~42%).

Key ideaA unified suite of 54 agentic healthcare tasks across 7 categories, each its own environment spanning the patient journey and many modalities. Given minimal instructions, an agent must explore raw healthcare data, operate in a complex environment, and execute multi-step solutions (e.g., querying large clinical databases, interpreting gigapixel pathology images); all scored by binary success/failure against expert labels.

ResultOverall task success stays low — the strongest and most cost-effective agent, Codex GPT-5.5, reaches only ~42%; frontier agents show promise on EHR research-modeling pipelines but medical imaging remains especially hard.

medical AIagentic benchmarkclinical workflowshealthcareevaluation
Paper first page (click to collapse)HealthAgentBench first page

Agent memory as a trainable skill

2

Memory crystallizing as a learnable cognitive skill, not an architectural afterthought.

06rel 0.90🔥 buzz 9hf_dailysite☐ to read

AutoMem: Automated Learning of Memory as a Cognitive Skill

Shengguang Wu, Hao Zhu, Yuhui Zhang, Xiaohan Wang, Serena Yeung-Levy · Stanford University · arXiv:2607.01224 · site

Why readMemory as a trainable cognitive skill (“metamemory”), not an afterthought — automates both memory structure and proficiency, and the gains are large.

Key ideaTreat memory management (what to encode, when to retrieve, how to organize) as a learnable skill using file-system operations as first-class actions. Two loops: (1) a strong LLM reviews full trajectories and iteratively revises the memory structure (prompts, schemas, action vocab); (2) the agent’s own good memory decisions from successful episodes train its memory proficiency.

ResultOptimizing memory alone (no task-action changes) improves base-agent performance 2×–4× across Crafter / MiniHack / NetHack, bringing a 32B open-weight model to parity with Claude Opus 4.5 and Gemini 3.1 Pro Thinking on these long-horizon tasks.

agent memorymetamemorytrainable skilllong-horizonself-supervised
Paper first page (click to collapse)AutoMem first page
07rel 0.93🔥 buzz 3hf_dailyon-device☐ to read

DuoMem: Towards Capable On-Device Memory Agents via Dual-Space Distillation

Peyman Hosseini, Ondrej Bohdal, Ahmed Alajrami, Andrea Maracani et al. · Samsung R&D UK + Queen Mary University of London · arXiv:2606.29961

Why readGets capable memory agents onto small on-device models via dual-space distillation — a practical resource-constrained recipe.

Key ideaDistill procedural problem-solving from a large teacher into a compact student in two complementary spaces: (1) context-space distillation replaces student-generated memories with higher-quality teacher procedural memories prepended to the input; (2) parameter-space distillation fine-tunes lightweight LoRA adapters on successful teacher trajectories.

ResultOn ALFWorld, a 4B model jumps from 4.3% → 77.9% success, closing most of the gap to a 72B teacher (87.1%) with fewer than 10M trainable params; the enhanced 4B runs ~3× faster than the 72B in wall-clock — viable for real-time on-device deployment.

agent memoryon-devicedual-space distillationLoRAprocedural memory
Paper first page (click to collapse)DuoMem first page

Inference, distillation & scaling

3

Two cautionary results on naive scaling, plus a frontier model card.

08rel 0.89🔥 buzz 4hf_dailytest-time scaling☐ to read

When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling

Yong Yi Bay, Kathleen A. Yearick · University of Illinois at Urbana-Champaign · arXiv:2606.28661

Why readA cautionary result for test-time scaling — coverage keeps rising but selection is capped, so past a point extra samples just make the model surer of a confident mistake. Same authors as the GRPO-identity paper from 07-03.

Key ideaCoverage (fraction with ≥1 correct try) climbs with samples, but selection (returning one answer without knowing which is right) is capped; the identifiability gap is the answer a model can produce but can’t pick. They define a modal ceiling (the vote settles within a few dozen draws) and a correlation ceiling (for scoring a benchmark, sooner still), turning the cutoff into a single “effective number of samples” any run already reveals.

ResultBeyond that cutoff, extra draws cost compute and add nothing — and can even make the answer worse; the bottleneck is recognizing a right answer, not generating one.

test-time scalinginference computesamplingselectionceilings
Paper first page (click to collapse)When More Sampling Hurts first page
09rel 0.86🔥 buzz 4hf_dailycode☐ to read

Denser ≠ Better: Limits of On-Policy Self-Distillation for Continual Post-Training

Meng Wang, Haohan Zhao, Wenzhuo Liu, Lu Yang et al. · HKISI-CAS / Institute of Automation, CAS · arXiv:2607.01763 · code

Why readA corrective to the optimism that dense self-distillation stabilizes continual learning — it can amplify drift and forgetting. Pairs with the on-policy-distillation threads (AsyncOPD, the GRPO identity).

Key ideaRevisits self-distillation policy optimization (SDPO) for continual post-training. SDPO accelerates in-domain specialization when teacher signals are stable but struggles to generalize OOD; it induces larger parameter- and response-space drift than GRPO, amplifying high-frequency formatting artifacts via a self-reinforcing teacher–student loop.

ResultSDPO exhibits stronger forgetting and potential collapse, while GRPO adapts more conservatively and better preserves prior capabilities — on-policy data alone isn’t a sufficient stabilizer without stable, well-aligned teacher signals.

self-distillationcontinual post-trainingforgettingGRPOstability
Paper first page (click to collapse)Denser not Better first page
10rel 0.85🔥 buzz 23hf_dailymodel card☐ to read

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity

ByteDance Seed · arXiv:2607.00248

Why readByteDance’s frontier push toward “real-world complexity” — a Pro/Lite/Mini family emphasizing multimodal understanding, flexible inference, and reliable complex-instruction execution. Buzz 23.

Key ideaSeed2.0 Series (Pro / Lite / Mini) is designed around user experience under large-scale deployment, prioritizing robust visual + multimodal understanding, fast/flexible inference across three sizes, and reliable complex-instruction execution (benchmarked by DeR² and CL-bench). It aims beyond Olympiad-style problems toward research-level reasoning, tackling Erdös problems and Scientific Coding.

ResultA model card rather than a methods paper — positions Seed2.0 as a deployment-oriented frontier family for long-horizon, multi-step real-world workflows (details in the card).

foundation modelmultimodalfrontiermodel cardByteDance
Paper first page (click to collapse)Seed2.0 first page

Industry pulse

Condensed from the same digest · 2026-07-04

  • Local inference & open-source keep rising. Trending models like Qwythos-9B and GLM-5.2 gain traction amid widespread discussion of running SOTA LLMs locally without cloud reliance.
  • Richer input modalities. Hugging Face and Cerebras are integrating real-time voice with Gemma 4, while community tools like claude-real-video let LLMs process video natively.
  • Agents maturing into production. New frameworks — ScarfBench (agent eval), SkillOpt (trainable agent skills), ASPIRE (self-improving robotics) — signal a shift from experimentation to production-ready systems.
  • Dev tooling + foundational research. Momentum around self-learning skills for coding agents and location-spoofing utilities, alongside work on tabular data (TabFM), memory (Memora), and Lean theorem proving (Leanstral 1.5).
  • Labs expanding beyond traditional AI. OpenAI reports ChatGPT adoption growth, Google is collaborating with A24 on creative-AI research, and Anthropic is exploring pharmaceutical development as a new application domain.