Reading Queue · Triaged Backlog
Want to Read
Good papers I want to read but haven’t had time for yet. Theme of the day: scale the agent horizon, not the parameters, and consolidate reasoning + memory + action into one backbone. Stars mark the ones closest to my embodied / world-model thread; screenshots expand on click.
Priority picks (closest to my current thread)
- Qwen-RobotManip — a VLA manipulation foundation model that solves the “align heterogeneous data → scale” problem; ~38K hours, 1st in RoboChallenge.
- Vesta (NVIDIA) — one generalist embodied reasoner that beats specialist ensembles; the unify-reasoning-memory-planning direction.
- Qwen-RobotNav — a navigation backbone you reconfigure at inference (task mode + observation budget); a modular block for hierarchical agents.
🔥 Hottest in the batch (by buzz)
- Agentic Abstention — buzz 80, top of the batch: when should an agent stop instead of act? CONVOLVE distills stopping rules from trajectories, no retraining.
- Scaling the Horizon (Agents-A1) — buzz 57: a 35B MoE agent matching 1T models by scaling the agent horizon, not the parameters.
Embodied AI, world models & robot learning
3Unified embodied / VLA foundation models — right on the world-model / robotic-control thread.
Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models
Qwen Team · arXiv:2606.17846 · blog · code
Why readA VLA manipulation foundation model that cracks the “align heterogeneous manipulation data, then scale” problem — and wins a competition doing it. Squarely on the embodied/VLA thread.
Key ideaBuilt on Qwen-VL, a unified alignment framework across the representation, motion, and behavioral dimensions of manipulation harmonizes messy multi-source data into coherent large-scale pretraining. A human-to-robot synthesis pipeline turns egocentric hand demos into robot trajectories across 15 platforms → a ~38,100-hour corpus, all from open-source datasets + human video (no proprietary collection).
ResultRanks 1st in RoboChallenge (+20% relative over π0.5); emergent zero-shot instruction following, perturbation robustness, reactive error recovery, cross-embodiment transfer; validated on AgileX ALOHA, Franka, UR, ARX. Evaluated on OOD suites (RoboCasa365, LIBERO-Plus, EBench, RoboTwin-IF/XE) since standard benchmarks miss pretraining quality.
Paper first page (click to collapse)
Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation System
Qwen Team · arXiv:2606.18112 · blog · code
Why readA single navigation backbone you reconfigure at inference time — a clean modular building block for hierarchical / agentic planning.
Key ideaBuilt on Qwen3-VL, a parameterised interface exposes two knobs: task modes (instruction following, object search, tracking, driving) and controllable observation parameters (token budget, per-camera weights). Training-time randomization over all configs makes it robust to any inference-time setting with zero architectural change; co-training with vision-language data prevents collapse into reactive action mappers.
ResultSOTA across navigation benches — 76.5% VLN-CE RxR, 90.0% tracking on EVT-Bench, 91.4 PDMS on NAVSIM; an agentic system built on it sets new SOTA on Embodied QA (+10.8% HM-EQA, +15.4% EXPRESS-Bench) with 77% fewer navigation steps; scales 2B→8B with strong zero-shot real-robot transfer.
Paper first page (click to collapse)
Vesta: A Generalist Embodied Reasoning Model
Johan Bjorck, Zhiqi Li, Yunze Man, Jing Wang et al. · NVIDIA · arXiv:2606.20905
Why readA single generalist embodied reasoner that beats specialist ensembles — direct evidence for the “unify world-model + memory + planning in one backbone” direction I’m tracking.
Key ideaConsolidates localization, spatial reasoning, navigation, and long-horizon planning into one foundation model instead of a specialist stack. Uses a massive curated corpus to induce spatial grounding plus a multimodal memory harness for extended-horizon temporal reasoning.
ResultBeats individual per-task SOTA by >20% on average and an ensemble of per-category-best baselines by >10%; on real-world robotic tasks needing memory + reasoning, >35% higher task success — a scalable alternative to combining specialists.
Paper first page (click to collapse)
Agents, long-horizon autonomy & evaluation
5Horizon scaling, knowing when to stop, and deployment-like continual evaluation.
Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent
Agents-A1 Team, Shanghai Artificial Intelligence Laboratory · arXiv:2606.30616
Why readTrend-setting (buzz 57): a 35B MoE agent matches 1T models by scaling the agent horizon rather than parameters. The multi-teacher on-policy distillation recipe is worth stealing.
Key ideaScale trajectory length and domain heterogeneity, not size. Three stages on a 35B MoE base: (1) full-domain SFT, (2) train domain-level teacher models, (3) multi-teacher domain-routed on-policy distillation with salient-vocabulary alignment across ~45K-token agentic trajectories, unifying six heterogeneous domains into one student.
ResultAgents-A1 (35B) matches or exceeds 1T models (Kimi-K2.6, DeepSeek-V4-pro) on long-horizon benches — SEAL-0 56.4, IFBench 80.6, HiPhO 46.4, FrontierScience-Olympiad 79.0, MolBench-Bind 56.8; competitive on SciCode, HLE, BrowseComp — at ~35× smaller size.
Paper first page (click to collapse)
Agentic Abstention: Do Agents Know When to Stop Instead of Act?
Han Luo, Bingbing Wen, Lucy Lu Wang · arXiv:2606.28733
Why readHottest in the batch (buzz 80). Reframes “when should an agent stop?” as a learned sequential decision — a lightweight, model-agnostic reliability lever for long-horizon tool use.
Key ideaAgentic Abstention treats stopping as a sequential decision under uncertainty (answer / abstain / gather more info), distinct from single-turn abstention. CONVOLVE is a context-engineering method that distills full interaction trajectories into reusable stopping rules — no parameter updates.
ResultAcross 13 LLM systems × 2 scaffolds over web-shopping, terminal, and QA (28k+ tasks), many agents never abstain or over-abstain, and larger/reasoning models sometimes do worse. CONVOLVE lifts timely abstention on WebShop from 26.7% → 57.4% for Llama-3.3-70B without retraining.
AgentOdyssey: Open-Ended Long-Horizon Text Game Generation for Test-Time Continual Learning Agents
Zheyuan Zhang, Zehao Wen, Alvin Zhang, Andrew Wang et al. · arXiv:2606.24893
Why readShifts agent evaluation from fixed train/test splits to deployment-like continuous adaptation — useful lens on what actually enables test-time continual learning.
Key ideaA procedurally generated text-game framework that interleaves learning and inference over long horizons, with diagnostic tests isolating world knowledge, episodic memory, exploration, and action diversity.
ResultTop agents remain far below human with big headroom; short-term memory emerges as a critical component across paradigms; performance scales with base-model strength but exposes fundamental long-horizon reasoning limits.
TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents
Shoufa Chen, Luyuan Wang, Xuan Yang, Zhiheng Liu et al. · arXiv:2606.28480
Why readFills a real gap: evaluating general-purpose terminal-use agents (command-line / computer-use beyond GUIs). Buzz 37 — a benchmark to watch as terminal agents move to production.
Key ideaA benchmark for agents that operate directly in the terminal across general-purpose computer-use tasks, targeting the real-world CLI skills that GUI-centric suites miss. (No abstract summary in the digest and I couldn’t retrieve the paper — details are from the digest blurb; read the abstract first.)
OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
Mengqi Yuan, Zilong Zhou, Xinzhuang Xiong, Weiming Wu et al. · arXiv:2606.29537 · code · site
Why readThe frontier computer-use eval just got much harder and more realistic — a sober measuring stick for how far deployment-grade agents actually are.
Key idea108 long-horizon computer-use workflows across everyday and professional domains, each a realistic end-to-end task taking a human a median ~1.6 hours and averaging 318 tool calls (vs ~30 in OSWorld 1.0). Targets under-tested phenomena: streaming interaction, dynamic environments, cross-source reasoning, implicit-state inference, visual-spatial precision.
ResultReported figures put the best frontier agent at ~20.6% full task completion (54.8% partial credit) at a 500-step budget — scores stay far below OSWorld 1.0 even with more reasoning effort.
Distillation & inference efficiency
2On-policy distillation systems and long-context attention — the plumbing behind the above.
AsyncOPD: How Stale Can On-Policy Distillation Be?
Wonjun Kang, Kevin Galim, Seunghyuk Oh, Minjun Kang et al. · FuriosaAI / Ajou / UC Berkeley / MSR / KRAFTON · arXiv:2606.24143 · code
Why readThe systems side of on-policy distillation (which shows up in Agents-A1 too): how much staleness can you tolerate to get throughput? Practical for any post-training pipeline.
Key ideaFirst systematic study of stale data in asynchronous OPD (decoupling rollout generation from learner updates). KL direction matters: teacher-weighted forward KL is robust to stale rollouts, student-weighted reverse KL is vulnerable. For the vulnerable case, a simple OPD-specific surrogate — recompute the reverse-KL under the current student at learner time — beats async-RL stabilizers; finite teacher-score caches create a bias-variance tradeoff, motivating multi-sample Monte Carlo.
ResultAsyncOPD, a fully asynchronous pipeline built from these estimator choices, improves training throughput 1.6×–3.8× over strict synchronous training at comparable accuracy.
Paper first page (click to collapse)
Simplified Sparse Attention via Gist Tokens
Yuzhen Mao, Michael Y. Li, Emily B. Fox · Stanford University · arXiv:2604.20920 · code
Why readA no-architecture-change sparse attention for long-context inference — directly on the KV-cache / latency threads, and it even beats full attention in RAG.
Key ideaGist tokens — compressed chunk summaries learned via continued pretraining under restricted attention masks. At inference the query scores only the small set of gist tokens (low memory bandwidth), then selectively unfolds the top-k chunks’ raw tokens; a hierarchical gist-of-gist variant (H-SSA) gives log-linear decoding.
ResultBeats compression and inference-time sparse-attention baselines on LongBench at matched compression; in RAG, exceeds full attention by >5.7 points; up to 3.37× end-to-end decoding speedup over Flash-Decoding; H-SSA holds/improves accuracy at 32× compression.
Paper first page (click to collapse)
Industry pulse
Condensed from the same digest · 2026-06-30
- Benchmark-war noise & OCR dominance. GLM 5.2 is generating buzz after claiming benchmark wins over Claude, while Baidu’s Unlimited-OCR keeps dominating image-text-to-text with strong adoption.
- Open-source is accelerating agents. Standout repos like DeepSpec (speculative decoding) and awesome-evals (agent evaluation) are drawing attention, alongside autonomous trading agents and digital-human platforms.
- Big labs push across modalities. Google added computer-use to Gemini 3.5 Flash and demoed frozen multi-token prediction; Microsoft and Meta shipped brain-to-text decoding and harmonic-memory architectures.
- Commercial + governance signals. OpenAI expanded its HP partnership amid enterprise-adoption talk; emerging debates on agent economics (OKX exploring agent-to-agent hiring) and music royalties hint at maturing AI-infrastructure governance.
- The shift is toward practical agentic deployment. Cheaper/faster inference (vLLM), better evaluation, and autonomous-agent frameworks are converging — the field is moving from capability races to agentic deployment at scale.