Reading List
Papers I've come across and want to read — each with a one-line idea. · 15 so far
-
2026.07.15
Paper Digest — Reading Queue (2026-07-15)
A triaged 2026-07-15 reading queue with arXiv links and first-page screenshots, organized as one signal supply chain: sourcing training signal (TerraZero — procedural driving self-play, zero demos), transferring it between models (Direct-OPD, PUST), routing it so it gets used (ABot-N1 pixel-goal interface, the Knowing–Using Gap), and the weak link that everything assumes — exploration (MACE). Two industry-led papers (Applied Intuition, Alibaba AMAP) are the focus. Includes a worked deep dive on porting the reuse layer into embodied policies. Theme of the day: signal supply chain & industry infrastructure.
-
2026.07.04
Reading Queue — Want to Read (2026-07-04)
A triaged 2026-07-04 backlog with arXiv links and first-page screenshots: task-agnostic VLA pretraining (TAP), agentic benchmarking & capability measurement (EvoPolicyGym, AgenticDataBench, PACE, HealthAgentBench), agent memory as a trainable skill (AutoMem, DuoMem), and inference / distillation / scaling limits (When More Sampling Hurts, Denser != Better, Seed2.0), plus an industry pulse. Theme of the day: agentic benchmarking and memory-as-a-skill.
-
2026.07.03
Reading Queue — Want to Read (2026-07-03)
A triaged 2026-07-03 backlog with arXiv links: controllable world simulation (WorldDirector), agent memory / reliability / evaluation (AgenticSTS, MemSyco-Bench, SkillCoach, SWE-Interact, DiscoPER), RL & reasoning training (the GRPO/Dr.GRPO/DAPO identity, Transfer-Aware Curriculum for multi-domain RLVR), and medical AI (step-aware RL for medical reasoning, discrete-diffusion radiology drafting), plus an industry pulse. Theme of the day: agent memory & reliability are the new frontier.
-
2026.06.30
Reading Queue — Want to Read (2026-06-30)
A triaged 2026-06-30 backlog with arXiv links and first-page screenshots: embodied foundation models (Qwen-RobotManip, Qwen-RobotNav, Vesta), agents / long-horizon autonomy & evaluation (Agents-A1 horizon scaling, Agentic Abstention, AgentOdyssey, TUA-Bench, OSWorld2.0), and distillation / inference efficiency (AsyncOPD, Simplified Sparse Attention), plus an industry pulse. Theme of the day: scaling agent horizons, not parameters, and unifying reasoning+memory+action in one backbone.
-
2026.06.29
Reading Queue — Want to Read
A triaged backlog of good papers I want to read but haven't yet: embodied AI / world models / sim-to-real (PhysisForcing, SimFoundry, Learning to Fold, Object-Centric Residual RL), agentic systems and tool use (GBC, PAT, ProMSA, Tool-Suppression), and inference efficiency / generative (ConvFill, Qwen-Image-2.0-RL), plus a condensed industry pulse.
-
2026.06.28
Reading Queue — Want to Read (2026-06-28)
A triaged 2026-06-28 backlog with arXiv links, first-page screenshots, and relevance/buzz signals: embodied AI / world models / robot learning (PhysiFormer, Fast-LeWM, REGEN, VLA flatness, OctoSense), agents & reasoning safety (over-privileged tool selection, do thinking tokens help with safety), and training / adaptation / efficiency (Wan-Streamer, DO-ALL, biological post-training), plus an industry pulse.
-
2026.06.27
Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation
Wraps text-to-image in an agent loop (plan → reason → search → memory) that diagnoses the 'context gap' in underspecified prompts and gathers the missing intent and domain context before generating. Ships a new IA-Bench.
-
2026.06.27
In-Context World Modeling for Robotic Control
Treats system identification as in-context adaptation: feed a VLA a short window of task-agnostic self-generated interactions so it infers the current system's dynamics and adapts zero-shot to new camera viewpoints and robot morphologies — no fine-tuning.
-
2026.06.27
Fast LeWorldModel
Speeds up LeWM visual planning by swapping step-by-step latent rollouts for action-prefix prediction — encode an action prefix and predict the resulting future latent in parallel, cutting planning time and slowing latent-error growth over long horizons.
-
2026.06.27
When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents
Argues a 'co-failure ceiling' caps routing / voting / mixture-of-agents gains across 67 frontier models: when models fail on the same inputs, combining them recovers little — bounding how far multi-model systems can scale.
-
2026.06.26
RoboAtlas: Contextual Active SLAM
Uses a contextual multi-armed bandit to adaptively balance frontier-based exploration against VLM-guided semantic navigation over a scalable 3D semantic map — strong real-robot navigation results.
-
2026.06.26
ReNIO: Reweighting Negative Trajectory Importance for LLM On-Policy Distillation
Finds that incorrect reasoning traces are more informative than correct ones for on-policy distillation, and up-weights likely-negative trajectories via the student-to-teacher probability ratio — up to ~9–10% gains on math/code.
-
2026.06.26
The Hitchhiker's Guide to Agentic AI: From Foundations to Systems
A full-stack practitioner's reference for agentic AI — treating LLM substrate, alignment, reasoning, tool use, multi-agent coordination, and deployment as one interdependent pipeline rather than siloed topics.
-
2026.06.26
Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models
Pairs teacher-forcing (offline) and self-forcing (on-policy) as complementary phases to distill autoregressive video diffusion — reaching SOTA VBench-T2V with just 2-step inference, enabling interactive world models.
-
2026.06.26
Autodata: An Agentic Data Scientist to Create High-Quality Synthetic Data
AI agents acting as data scientists that generate, evaluate, and iteratively refine synthetic datasets — then meta-optimize the agent itself to produce progressively stronger data.