← All paper notes
Read July 5, 2026

Paper Reading Notes - 2026-07-05 Digest

Tags
reading-notes
Paper Reading Notes

Reading Notes · AI

Paper Reading Notes

Digest 2026-07-0510 papersPixelEyes — perception / reasoning decoupling

A running log of questions and answers, one section per paper. Each entry keeps the paper's abstract and key passages on record, followed by the Q&A as it accumulates.

Summary

A running study log over the 2026-07-05 arXiv digest — 10 papers. Paper 01 (PixelEyes) is read in depth with a full question-and-answer thread; papers 02–10 have their abstracts on record with a short paraphrased gist each. Six papers are flagged ★ Key: 01, 02, 03, 05, 06, 08.

By theme:

01

PixelEyes: Decoupling Perception and Reasoning for Pinpoint Visual Evidence Seeking

★ Key

Dengxian Gong, Yuanzheng Wu, Haobo Yuan, Zhengdong Hu, Tao Zhang, Yikang Zhou, Shihao Chen, Quanzhu Niu, Kai Wang, Jason Li, Haochen Wang, Lu Qi, Shunping Ji, Ming-Hsuan Yang

Wuhan University · UC Merced · UTS · NUS · NTU · CASIA

Project page: godx-7.github.io/PixelEyeSite  ·  arXiv 2607.00115

Abstract
PixelEyes first page: title, authors, figure, and start of abstract PixelEyes abstract continuation

Key Passages · from the paper

A · Method overview (approach)
PixelEyes method overview excerpt
Original screenshot. Key claims: general VLM decides what; SAMTok returns where as a pixel mask; Semantic-Region BFS + Switchable Tool Use keep trajectories short.
B · Training data & policy (PixelEyes-6K)
PixelEyes-6K training excerpt
Original screenshot. 5.8K expert trajectories from Gemini-3-Flash + mask_based_crop → SFT on Qwen-3-VL → vanilla GRPO.
C · Benchmarks & Pinpoint-Bench
Pinpoint-Bench excerpt
Original screenshot. 433 samples, target masks averaging 0.07% of image area, zero-hint protocol, reports LSR and Turn-to-Answer efficiency.
D · Related work
Related work excerpt
Original screenshot. Two threads: (1) VLMs vs specialized perception; (2) multi-turn visual reasoning agents.

Questions & Answers

1

Multi-turn visual reasoning — how many turns, and which models do this?

Multi-turn means the agent does not answer in one shot; it iteratively crops, zooms, and re-observes the image over several rounds to gather evidence — the "Thinking with Images" paradigm popularized by OpenAI o3.

Turn count is bounded, not fixed. Each task runs under a turn budget. The paper's whole thesis is that the structure of the search matters more than the raw number of turns: two agents show the extremes — Mini-o3 scales turns aggressively (dozens of bbox crops per trajectory), while ZwZ collapses zooming into a single forward pass. PixelEyes keeps trajectories short by making each turn precise.

2

“Localize the target” — which task is this? Perception?

Yes. Localization (grounding) is a perception task: given a query, find where the target sits in the image, expressed as a bounding box or a pixel mask. It is distinct from reasoning, which interprets the content once found. PixelEyes' core move is to hand this perception step to a dedicated tool (SAMTok) instead of asking the reasoning model to do it.

4

Is the “entanglement” assumption actually validated by evidence?

The paper backs it with three observations rather than asserting it:

  • A measured trade-off: one jointly-trained model is consistently weaker at grounding than perception specialists and weaker at reasoning than strong general VLMs. Fine-tuning grounding into a reasoner degrades VQA ability — the two objectives fight.
  • Two observed failure modes: weak perception → long "blind crop" trajectories; and correct crop but weak reasoning → inattentional blindness.
  • A direct measurement: the LSR−accuracy gap on Pinpoint-Bench quantifies how often a model reaches the target but still fails — the entanglement made visible.

So the assumption is supported empirically, though it is presented as the paper's motivating diagnosis rather than a controlled ablation of "entangled vs not" on the same base model.

5

Beyond reasoning and perception, are there other components in the inference loop?

Yes — an agentic VLM loop has several moving parts beyond those two:

  • Control / policy: decides the next action — call a tool, crop, or answer. This is the orchestrator that sequences turns.
  • Tool use / action: the interface that invokes SAMTok, performs the crop, and feeds results back.
  • Memory / context management: the trajectory history (past crops, masks) held in context — the thing that gets "polluted" by noisy crops.
  • Stop / answer decision: when to terminate the search and commit.

At the raw-LLM level you would also list tokenization/embedding, attention, and decoding (plus expert routing in MoE), but in this paper's framing the relevant extras are the controller, tool interface, and trajectory memory.

Structure of an agentic VLM loop (Q5)
CONTROLLER / POLICY — decides each turn: call tool · crop · answer Reasoner (VLM) · formulate query · recognize + answer Tool Interface call · crop · return Perception (SAMTok) referring segmentation → pixel mask Trajectory Memory crops + masks history Answer / Stop 1 query 2 call 3 mask 4 crop 5 context answer
Raw-LLM inference structure (F3) — the parts under the agent loop
Input text / image Tokenizer + Embedding Transformer block × N Self-Attention (+ KV cache) MoE Router → Experts (or a dense FFN) Final Norm + LM Head Decoding (sampling) For a VLM: a Vision Encoder → Projector turns the image into visual tokens that enter alongside text tokens before the stack. Decoding runs autoregressively, reusing the KV cache.
6

Is “reasoner = what to look for, tool = where it is” a precise description of the two roles?

It is a clean and mostly accurate division of labor, but a slight simplification.

  • Perception tool = purely where: pixel-level localization of a named target. Accurate.
  • Reasoner = more than just what to look for. It formulates the referring query, proposes regions, judges whether the tool succeeded, decides the next region, and — crucially — performs the final content recognition and answer. "What to look for" captures the query-formulation half but understates that the reasoner also does the concluding reasoning. (See Round 2, F4 — does this mean the reasoner still has perception ability?)

So: a great mnemonic for the control split, slightly under-describing the reasoner's full job.

7

Mask-guided Visual Search: why masks? which segmentation model? Is the expert doing the reasoning?

Why a mask: a pixel-level mask is far more precise than a coarse bounding box, which matters enormously when targets occupy under 1% of a high-resolution image. Precise localization means the reasoner no longer has to spend turns compensating for sloppy grounding.

What "sloppy grounding" means: a localization that is coarse or slightly off — a box that is too large, misplaced, or clips a tiny target — so the resulting crop doesn't cleanly isolate the object. The reasoner then has to patch it up with extra crops and extra reasoning turns, which is precisely the trajectory bloat PixelEyes removes.

Which model: SAMTok [52], a referring-segmentation tool that unifies mask generation inside the language-model interface.

Important correction: the external expert does perception, not reasoning. SAMTok only answers where. Reasoning stays entirely with the VLM. So PixelEyes introduces an external perception specialist and keeps a single reasoner — the opposite of outsourcing reasoning. (Could outsourcing reasoning help instead? See Round 2, F6.)

8

Semantic-region BFS — how exactly does the breadth-first search work?

The trick is coverage before depth. Concretely:

  • Anchor in the original image. Every candidate region's coordinates are referenced to the full original image, not to the previous crop — so errors don't compound as you zoom.
  • Propose a low-IoU region on failure. Whenever SAMTok fails to ground the target, the reasoner proposes a new region that barely overlaps the ones already tried (low IoU = genuinely new area), instead of re-cropping the same wrong spot.
  • Expand siblings first. It explores neighboring semantic regions at the same level (breadth) before descending into any single one (depth).

This is precisely what kills the redundant DFS-style loops of prior agents that keep cropping deeper into an already-wrong sub-region. See Q15 for "low-IoU."

9

What is “inattentional blindness”? And besides instance-level, are there other annotation levels?

Inattentional blindness is borrowed from cognitive psychology (the "invisible gorilla" effect): you look right at something but fail to notice it. Here it names a specific gap — the agent visits the correct region (it is in the crop) but fails to recognize / answer it. The paper measures it directly as the gap between LSR (reached the target) and accuracy (answered correctly).

Annotation levels. Pinpoint-Bench uses instance-level masks and boxes, which is what cleanly separates localization failures from reasoning failures. Other granularities exist and could be used: category/semantic-level, panoptic or region-level, part-level, and box vs pixel-mask precision. Instance-level is the sweet spot for attributing a failure to "didn't find it" vs "found it but got it wrong."

10

The “passive observers → active visual reasoners” framing — note.

Agreed — a good framing. The setting: VLMs shift from passively captioning a whole image to actively cropping/zooming/re-examining to gather evidence (popularized by OpenAI o3, studied as "Thinking with Images"). The difficulty is spelled out well: decisive evidence often sits in objects under 1% of a high-res image, so the agent must find a needle in a haystack and reason about it within a bounded turn budget. This one sentence justifies both halves of PixelEyes: precise localization (find the needle) and short search (bounded budget).

11 & 12

The single-model critique — note.

This is the crux of the motivation, and it is a strong argument. Existing methods ask one model to do both fine-grained region-level perception and general reasoning. But that same model is:

  • consistently weaker at grounding than perception-oriented specialists [19, 44], and
  • consistently weaker at reasoning than strong general-purpose VLMs [4, 37, 28, 27, 15].

In other words, joint training buys you a jack-of-both-trades that is master of neither — which is exactly why decoupling (let each specialist operate at its native granularity) is the proposed fix.

13

The two failure modes — and a table of the compared methods.

Two failure modes the paper names:

  • Blind crops from weak perception: when the agent can't localize the correct region, it produces long trajectories of wrong crops.
  • Inattentional blindness from degraded reasoning: even when the correct region is cropped, weak reasoning means the agent sees the target but fails to recognize it — the gap between visiting and answering, quantified in the benchmark.

The compared / contrasted methods, organized:

Comparison methods — single-model / active-perception agents PixelEyes contrasts with
MethodCore ideaGrounding sourceSearch unitWeakness PixelEyes targets
Mini-o3 [18]Pushes the active-perception strategy hardest; scales interaction turns aggressivelyBase VLM's native groundingCoarse bbox crops (dozens per trajectory)Very long trajectories; noisy crops accumulate in context
ZwZ [39]Opposite route: distills zooming into a single forward passBase VLM's native groundingCoarse rectangular cropStill limited by weak base grounding on tiny targets
Early active-
perception [41,49,24,30,21]
Iteratively crop / zoom / re-observe using heuristic or tree-based searchBase VLMRectangular cropsHeuristic search → redundant loops
RL-trained
agents [50,32,45,47,46]
Train the crop/zoom behavior with reinforcement learningBase VLMRectangular cropsLonger trajectories; grounding still coarse
PixelEyes
(this paper)
Decouples reasoning from perception via a tool-call interfaceExternal SAMTok (pixel-precise mask)Semantic regions, BFS-ordered—

Named methods (Mini-o3, ZwZ) are the only ones the excerpts name explicitly; the bracketed clusters are cited as groups, so I list them by role rather than inventing names. Two shared properties PixelEyes drops: (1) reliance on the base VLM's grounding, and (2) search over coarse rectangular crops.

The original related-work text is embedded as screenshot D above, as you asked.

15

What is a “low-IoU region”?

IoU = Intersection over Union, the overlap ratio between two regions (intersection area ÷ union area; 1 = identical, 0 = no overlap). A low-IoU region is a newly proposed region that overlaps little with the regions already tried.

It is the novelty constraint that drives the breadth-first behavior: when grounding fails, jump to a genuinely different area rather than re-cropping near the same wrong spot. That's what prevents the redundant loops.

E1

What does “training the structure of search, not scaling turns” buy you?

Scaling turns is brute force with real costs: more compute, longer trajectories, and a context window that fills with noisy crops that degrade later decisions — with diminishing returns. Training the structure of search (mask-precise grounding + BFS coverage) instead makes each turn count. Benefits:

  • Shorter trajectories for the same or better accuracy.
  • Less context pollution — fewer wrong crops carried forward.
  • Better sample efficiency and generalization — the agent learns how to search rather than memorizing longer rollouts.

The paper calls this a "substantially stronger lever" than adding turns.

E2

Why augment Gemini-3-Flash with mask_based_crop and roll out closed-loop interactions? Why Gemini-3-Flash specifically?

What they're doing: they need expert demonstrations of the decoupled search behavior to train on, but hand-labeling trajectories is expensive. So they bootstrap: take a strong general VLM as a teacher, give it the same mask_based_crop tool, and let it interact in a closed loop on existing image–question pairs. They keep only trajectories that reach the correct answer, yielding PixelEyes-6K (5.8K clean trajectories). Those become supervised fine-tuning data for the student, Qwen-3-VL, and vanilla GRPO then sharpens the policy.

Why Gemini-3-Flash: the "Flash" tier is a fast, cheap, yet strong general VLM — you need to run thousands of multi-turn rollouts, so cost and speed matter, while quality must stay high enough that filtered trajectories are genuinely expert. It is a practical teacher choice, not a claim that it is the single best model.

The VLMs referenced across the paper:

Vision-language models referenced in the paper
ModelOriginRole in this paper
Flamingo [1]DeepMindCited as an early general-purpose VLM
LLaVA [20]Open-sourceCited general VLM lineage
GPT-4o [15]OpenAICited as a strong general-purpose VLM
Gemini / Gemini-3-Flash [27]GoogleTeacher used to synthesize PixelEyes-6K trajectories (Flash = fast/cheap tier → affordable large-scale rollout)
InternVL series [7,6,53,37]Shanghai AI LabCited strong open VLM family
Qwen-VL / Qwen-3-VL [3,34,5,4]AlibabaStudent / base model that is fine-tuned into PixelEyes
SAMTok [52]—Not a VLM — the referring-segmentation perception tool that returns masks

This lists the models the paper cites, which is what matters for reading it. The absolute "current SOTA" leaderboard shifts month to month — treat Gemini-3-Flash / Qwen-3-VL as the paper's generation, and verify against a live leaderboard if you need today's ranking.

E3

What is “vanilla GRPO”?

GRPO = Group Relative Policy Optimization, a reinforcement-learning method for LLMs (introduced with DeepSeekMath and used in DeepSeek-R1). Instead of training a separate value/critic network like PPO, it samples a group of responses per prompt and normalizes each response's reward against the group's mean and standard deviation to get its advantage. Dropping the critic makes it cheaper and simpler, which is why it's popular for reasoning RL.

"Vanilla" just means the standard, unmodified formulation — no custom reward shaping or algorithmic tweaks. The point they're making is that even off-the-shelf GRPO is enough to sharpen the policy once the search structure is already taught by SFT.

E4

Organize the existing benchmarks the paper mentions.

The paper argues existing benchmarks make it hard to evaluate visual search cleanly — some are saturating, some reward reasoning over search, and one lacks spatial labels. Hence Pinpoint-Bench.

Existing visual-search benchmarks named in the paper
BenchmarkWhat it targetsLimitation the paper notesSpatial annotation?
V* [41]Visual search on high-res imagesSaturating (headroom nearly gone)—
HR-Bench [36]High-resolution perceptionSaturating—
TreeBench [30]Visual reasoningEmphasizes reasoning over search—
MME-RealWorld [48]Real-world multimodal understandingEmphasizes reasoning over search—
VisualProbe [18]Fine-grained visual probingRight difficulty, but no spatial labels → can't split localization vs reasoning failuresNo
Pinpoint-Bench
(this paper)
Zero-hint visual search on ultra-high-res images (433 samples; targets ~0.07% of image area)Built to fix the above: instance masks isolate failure typesYes — instance masks + boxes

Pinpoint-Bench adds a strict zero-hint protocol (no spatial cue, no macro-anchors), multi-alias answer matching, and reports LSR + Turn-to-Answer alongside accuracy.

E5

Explain Mean ROI area, LSR, TAE, and the LSR–accuracy gap. Are they unique to this paper?

Short version, then the table:

  • Mean ROI / target area — the average share of the image the target occupies (0.07% here). A difficulty descriptor, not a scored metric.
  • LSR (Localization Success Rate) — a trial counts as a success if any crop covers the target. Isolates perception/search from the final answer.
  • TAE (Turn-to-Answer Efficiency) — how few turns it takes to reach the answer. Isolates trajectory efficiency.
  • LSR − Accuracy gap — the model reached the target but still answered wrong. This is the direct measure of inattentional blindness.
Evaluation metrics in Pinpoint-Bench
MetricMeasuresIsolatesNovel here?
AccuracyFinal answer correctnessEnd-to-end task successStandard
Mean ROI / target areaAvg. fraction of image the target occupies (0.07%)Difficulty of the benchmarkDescriptor, not novel
LSR
(Localization Success Rate)
Trial succeeds if any crop covers the targetPerception / search aloneHas precedent; central here
TAE
(Turn-to-Answer Efficiency)
How few turns are needed to reach the answerTrajectory efficiencyHas precedent
LSR − Accuracy gapTarget was reached but not answered correctlyInattentional blindnessThe paper's signature use

The individual metrics are natural and partly have precedent in active-perception work. What is distinctive is the packaging: zero-hint + instance masks let the LSR−accuracy gap read out inattentional blindness cleanly.

E6

How is the decoupling actually implemented — and is it convenient?

How: through a tool-call interface inspired by visual programming. The reasoner emits a tool call (a referring expression); SAMTok returns a mask; the system crops to it and feeds the crop back into the next turn. Perception is a pluggable external model, so grounding is never fine-tuned into the reasoner — which is what avoids the perception–reasoning trade-off. A Switchable Tool Use fallback handles cases where masks are ill-defined (charts, maps, dense text) by reverting to a plain bounding-box crop.

Convenient? Reasonably, with caveats:

  • Upside: modular — you can swap the segmentation tool, each module runs at its native granularity, and you avoid destructive joint training.
  • Cost: you need a strong segmentation tool available at inference; tool calls add orchestration and latency; and the model still has to be taught to use the tool well, which is exactly why PixelEyes-6K + GRPO exist. Decoupling isn't free — it moves the effort from joint training into building the tool interface and the demonstration data.

Follow-up Questions · Round 2

F1

Why fine-tune grounding into a reasoner? Does it mean the reasoner is more powerful than perception?

It is not a ranking of power — grounding and reasoning are different competencies, exactly as you suspect. Prior methods fine-tune grounding into the reasoner purely for convenience: one model, one forward loop, no external tools to orchestrate. But the two objectives then compete for the same weights. Optimizing precise box/mask prediction pulls parameters toward spatial regression; optimizing open-ended VQA pulls them toward broad semantic reasoning — push on one and the other degrades (the perception–reasoning trade-off). So neither is "stronger"; they simply don't co-habit one set of weights gracefully. PixelEyes' fix is to stop fusing them.

F4

Does the reasoner still possess the capability the perception tool provides?

Partly — and this is the subtle bit. "Perception" is not one skill; it splits into recognition (what is in this crop?) and localization (where exactly is it?). The reasoner (a VLM) keeps the recognition half — that's how it reads a crop and produces the final answer. What it lacks is precise localization, especially for tiny targets. So the reasoner does hold some perception, just not the pixel-precise grounding SAMTok provides (other models can do this job too — see F24). PixelEyes outsources only the where and lets the reasoner keep the what — which is why the neat "reasoner = what / tool = where" slogan slightly hides the overlap.

Perception = recognition + localization

F6

Could outsourcing reasoning (not perception) improve performance? Did the authors test it?

In principle yes — you could call a stronger reasoning model as a tool too. But three things make it the wrong lever here: (1) the diagnosed bottleneck was perception, not reasoning — the base VLM already reasons well; (2) outsourcing reasoning adds latency, cost, and a harder coordination problem (you must ship rich visual context to a second model); (3) it reintroduces exactly the complexity the paper is shedding.

Did they test it? The excerpts I have don't report a "reasoning-outsourced" ablation, so I can't confirm it either way — that would need the full experiments section (Sec. 4). Flagging this as not shown in the provided text rather than a definite no.

F7

Anchor-in-original vs relative-to-previous-crop — illustrate, and do these two cover all modes?

These are two coordinate-referencing schemes:

  • Relative to previous crop: each new region is described inside the last crop's frame. A small error in one crop shifts the frame for the next — errors compound (drift) as you zoom.
  • Anchor to original: every region maps back to the full original image, so a bad crop never poisons the next; the agent can jump anywhere.
Relative-to-crop (drift) vs anchor-to-original (stable) (F7)
Relative to previous crop → drift each frame measured inside the last → errors compound Anchor to original → stable O (origin) every region referenced to O → no drift, can jump anywhere

Do these exhaust the options? No. You could also anchor to an intermediate parent region (a hierarchical zoom stack), use absolute vs normalized coordinates, or a hybrid that anchors to the original but remembers a zoom path. PixelEyes just picks the most robust end of that spectrum — always anchor to the original.

F8

Illustrate “low-IoU region” more vividly.

IoU (Intersection over Union) is the overlap ratio between two regions. A low-IoU proposal barely overlaps the regions already tried — so it points the search at fresh territory instead of re-cropping the same wrong spot.

IoU: high overlap (redundant) vs low overlap (new region) (F8)
High IoU — redundant big overlap → re-cropping the same area IoU = ∩ ÷ ∪ Low IoU — new area tiny overlap → genuinely new region (what BFS wants)
F9

What is the structure of the “semantic regions”?

The regions are organized semantically and hierarchically, not as a fixed pixel grid: the whole image splits into meaningful regions (e.g., "left building", "storefront", "street"), each of which can split further. BFS visits siblings at the same level first, then descends — cover breadth before committing to depth.

Semantic-region tree — BFS visits siblings before descending (F9)
Whole image left building storefront street sign window text 1 2 3 4 5 visit 1→2→3 (siblings) before descending to 4→5
F10

Explain “cropping / zooming / re-examining”.

  • Cropping — cutting out a sub-region of the image to focus on it.
  • Zooming — enlarging (up-sampling) a region so fine detail becomes legible; usually crop + upscale.
  • Re-examining — looking again: re-observing a region (or a new one) at higher resolution to gather more evidence.

These are the actions an active-perception agent takes each turn, versus a passive model that captions the whole image once.

F11

What does “a needle in a haystack” mean?

An English idiom for something extremely hard to find because it is tiny within a vast space (literally, a needle lost in a haystack). Here it is near-literal: the decisive object can occupy under 1% — often ~0.07% — of a high-resolution image, so locating it really is a needle-in-a-haystack search.

F12

In Q13, what does “heuristic” mean?

A heuristic is a rule-of-thumb: a practical strategy that is cheap and usually works but carries no guarantee of being optimal (e.g., "always zoom into the center first"). "Heuristic or tree-based search" therefore means hand-crafted search rules, as opposed to a search policy learned from data — which is what the RL-trained agents and PixelEyes use.

F13

Any other constraints — e.g. depth-first, or a BFS/DFS balance? Could that be a new direction?

Yes — the paper's constraints (anchor-to-original + low-IoU novelty + breadth-first ordering + switchable fallback) are one point in a larger design space. Natural additions:

  • a depth budget or confidence threshold that flips from BFS to DFS once a promising region is found;
  • best-first / beam search — rank candidate regions by a learned score and expand the most promising, blending breadth and depth;
  • a revisit penalty or per-scale IoU thresholds.

Your instinct is a real research direction: an adaptive BFS↔DFS (best-first) semantic search that spends breadth when uncertain and depth when confident. This paper's fixed BFS is a strong, simple baseline such work could build on.

F14

What does “general-purpose” mean — is it perception + reasoning? (fuller definition)

General-purpose does not mean "perception + reasoning" exactly. A general-purpose VLM is one trained broadly to handle many multimodal tasks — VQA, captioning, OCR, chart/diagram reading, dialogue, reasoning — across domains, via instruction-following and transfer, without task-specific fine-tuning. Contrast a specialist (e.g., a segmentation-only model good at one thing). Fuller definition: broad, transferable multimodal competence; strong at recognition and reasoning; instruction-following — but typically only coarse at precise localization. That last gap is exactly why PixelEyes pairs it with a perception specialist.

F15

What is PPO? Give a fuller list of RL methods to compare.

PPO (Proximal Policy Optimization) is an actor–critic policy-gradient method: it improves the policy while a clipped objective keeps each update from moving too far, which stabilizes training. In RLHF it needs a separate value/critic network and a learned reward model. GRPO's whole appeal is dropping the critic. The broader landscape:

RL / preference-optimization methods for LLMs — comparison
MethodTypeNeeds critic?Needs reward model?Core mechanism
REINFORCEPolicy gradientNoOptionalScale log-prob of an action by (reward − baseline); simple but high variance
A2C / A3CActor–criticYes (value net)—Use a value network as baseline → lower-variance advantages
TRPOTrust-region PGYes—Constrain each update inside a KL trust region; stable but heavy
PPOActor–critic (clipped)YesYes (RLHF)Clipped surrogate keeps updates small; the RLHF workhorse
RLOOREINFORCE variantNoYesBaseline = mean reward of the other samples in the batch; critic-free
GRPOGroup-relative PGNoYes / ruleAdvantage = reward normalized within a sampled group; no critic (used here)
DPOPreference opt. (not RL)NoNoDirectly optimize on preferred vs rejected pairs; offline, no rollouts
KTO / ORPO / SimPOOffline preferenceNoNoReward-model-free preference tuning variants; simpler pipelines

Trend for reasoning models: drop the critic (PPO → GRPO / RLOO) and often replace the learned reward model with rule-based / verifiable rewards. "Vanilla GRPO" = this table's GRPO row, unmodified.

F16

What does “even off-the-shelf GRPO is enough once the search structure is taught by SFT” mean?

It means the hard part is already handled by supervised fine-tuning on the expert trajectories — SFT teaches the model the structure of the search (how to call the tool, how to run BFS, when to stop). Once that behavior is in place, you don't need a bespoke RL algorithm to improve it; plain, unmodified ("off-the-shelf") GRPO is enough to sharpen the remaining decisions. In short: good demonstration data does most of the work; RL is a light polish, not the main engine.

F17

How does the localization process actually work, end to end?

  1. The reasoner writes a referring expression (e.g., "the red sign on the left building").
  2. That goes as a tool call to SAMTok.
  3. SAMTok runs referring segmentation and returns a pixel mask of the referred object.
  4. The system converts the mask to a crop region (its bounding box, usually padded) and crops/zooms the original image to it.
  5. The crop is fed back to the reasoner to recognize/answer, or to decide the next region.

If SAMTok fails to ground the target, BFS proposes a new low-IoU region; if masks are ill-defined (charts/maps/dense text), Switchable Tool Use falls back to a plain bounding-box crop.

F18

“A trial succeeds if any crop covers the target” — are crops used as a probe?

Exactly — that's the right reading. The agent's crops are its probes into the image. LSR asks: across the whole trajectory of crops the agent made, did any crop actually cover the target region? If yes, localization succeeded — regardless of whether the final answer was right. So LSR scores the search / probing process; the gap between LSR and accuracy is the target being probed but still mis-answered (inattentional blindness).

F19

More detail on “visual programming”.

Visual programming (in this line of work, e.g. VisProg and ViperGPT) means an LLM solves a visual task by writing a short program — a sequence of calls to modular vision tools (detect, crop, count, OCR, segment…) — and composing their outputs, much like calling functions from a library. The LLM is the "programmer"; the vision modules are the "library". PixelEyes borrows this: the reasoner issues tool calls (to SAMTok) as program steps, letting each module run at its native granularity, instead of forcing one end-to-end model to do everything.

F20

SAMTok returns a mask — why can a mask guide the crop? What's the mechanism?

Because a mask is geometry. A mask marks, per pixel, whether that pixel belongs to the target. From the set of "target" pixels you can compute the object's bounding box (the min/max x and y of those pixels) or its centroid — a deterministic step. The system then crops the original image to that box (often padded, then upscaled). So the chain is mask → box → crop, pure geometry with no learning needed. A mask beats a raw predicted box because it outlines even a small or irregular object tightly, giving a cleaner, better-centered crop.

F22

What are the formats of a mask? Give examples.

  • Binary mask — an H×W array, 1 = object, 0 = background.
  • Soft / probability mask — values in [0,1], a per-pixel confidence.
  • RLE (run-length encoding) — a compressed binary mask (COCO uses this).
  • Polygon — an ordered list of boundary vertices.
  • Instance / multi-class mask — integer labels per pixel saying which object or class.

Example: for a 4000×3000 image, the mask is the same size, almost all 0s with a small blob of 1s where a road sign is; its bounding box is read off from that blob.

F23

How is the alignment between reasoner and external perception handled — any architecture changes?

The alignment problem: the reasoner must phrase referring expressions the segmentation tool can act on, and correctly interpret the masks it gets back. How PixelEyes keeps them aligned:

  • Shared interface. SAMTok "unifies mask generation within the language-model interface" — it's designed to be called in the model's token space (special call/return tokens), so there's less representational mismatch than bolting on a foreign API.
  • Demonstration training. PixelEyes-6K teaches the reasoner, by example, how to phrase queries the tool handles well and how to react when a mask fails — alignment learned from expert trajectories.
  • GRPO then tunes the policy to call the tool effectively.
  • Switchable Tool Use is a safety valve: when the mask modality is ill-suited (charts/maps/dense text), fall back to a box crop rather than force a bad mask.

On architecture: the paper deliberately keeps perception external (a clean tool boundary), avoiding fusion of the two representation spaces. A heavier alternative — feeding mask features directly into the reasoner's latent space — would tighten alignment but reintroduce joint-training coupling, which is what they're avoiding. (Some of this is inferred from the tool-call design; the excerpts describe the interface, not every training detail.)

F24

Besides SAMTok, what other models could do this pixel-precise grounding job?

Yes — mapping a text query to a precise mask is an active area, and several families could slot into the same tool interface. Roughly:

  • Promptable segmentation (need a point/box, not language): SAM and SAM 2 (adds video). Best-in-class mask quality, but they expect a spatial prompt rather than a referring expression — so they're usually paired with a text-to-box stage.
  • Language → mask (referring / reasoning segmentation) — the closest match to SAMTok's role: LISA (LLM + SAM, "reasoning segmentation"), GLaMM, PixelLM, SEEM, GSVA, and classic referring-segmentation nets like LAVT, CRIS, ReLA. These take a phrase directly and return a mask — exactly what an agent wants.
  • Open-vocabulary detection / grounding (text → box), then hand the box to SAM for a mask: GroundingDINO, OWL-ViT / OWLv2, GLIP. The popular open-source recipe Grounded-SAM = GroundingDINO → SAM does text → box → mask.
  • Unified VLMs with grounding output: Florence-2, Kosmos-2, Qwen-VL grounding, PaliGemma — emit boxes (some masks) from text, though often coarser than a dedicated segmenter.

Trade-off: SAM-family gives the cleanest masks but needs a prompt; language-native tools (LISA / GLaMM / SEEM) accept a referring expression directly; Grounded-SAM is a strong, widely-used open pipeline. What PixelEyes actually needs is any tool that turns a referring expression into a precise mask — SAMTok is their choice, but a LISA-style or Grounded-SAM-style tool could occupy the same pluggable slot. (The field moves quickly, so newer options keep appearing; I'm naming established ones, not asserting SAMTok's internal design, which the excerpts don't detail.)

02

Morphing into Hybrid Attention Models

★ Key

Disen Lan, Jianbin Zheng, Yuxi Ren, Xin Xia, Xuanda Wang, Xuefeng Xiao, Xipeng Qiu, Yu Cheng

Fudan University · ByteDance Seed · The Chinese University of Hong Kong

Code: github.com/LanDisen/FlashMorph  ·  June 30, 2026  ·  method: FlashMorph

Abstract
Morphing into Hybrid Attention Models - abstract screenshot

Gist · paraphrased from the abstract

Hybrid attention keeps only a subset of full-attention layers and swaps the rest for linear attention to make long context cheaper — but which layers to keep is usually decided by heuristics (fixed placement, or scoring each layer on its own), which ignores how the choices interact globally.

They cast layer selection as a budget-constrained subset-optimization problem and propose FlashMorph: attach a converted linear-attention branch to every full-attention layer, freeze the weights, and jointly learn per-layer gates on synthetic long-context retrieval data with a regularizer that nudges toward linear attention. The gates are then discretized under a full-attention budget to instantiate the hybrid, followed by logits distillation and long-context finetuning. Reported to find better hybrid configurations, preserve long-context recall and general benchmark scores, and cut selection cost — with claims of efficiency and scalability.

Summary in my own words from the abstract screenshot above — not a verbatim quote; verify against the source before citing.

Questions & Answers — none yet. The abstract is on record; send questions to start this paper's log.

03

When Search Agents Should Ask: DiscoBench for Clarification-Aware Deep Search

★ Key

Yiling Tao, Shihan Deng, Meiling Tao, Pengzhi Wei, Zhichao Hu, Zhihao Zhu

Hunyuan, Tencent · Shenzhen International Graduate School, Tsinghua University

Benchmark: DiscoBench

Abstract
When Search Agents Should Ask: DiscoBench for Clarification-Aware Deep Search - abstract screenshot

Gist · paraphrased from the abstract

Deep-search agents usually assume the user's query is complete and explicit, but real requests are often vague, underspecified, or factually wrong — and that ambiguity compounds along multi-step retrieval, steering agents into wrong trajectories.

DiscoBench is a benchmark for clarification-aware deep search: 211 samples and 463 ambiguity instances across 11 real-world domains and 4 ambiguity types, plus a user simulator for multi-turn interaction and a four-axis evaluation (task utility, ambiguity detection, interaction strategy, cost efficiency). Key finding: detecting ambiguity and asking effective clarifying questions are distinct capabilities, and agents that keep searching instead of asking often do worse than direct guessing — a gap between retrieval skill and interactive problem-solving.

Summary in my own words from the abstract screenshot above — not a verbatim quote; verify against the source before citing.

Questions & Answers — none yet. The abstract is on record; send questions to start this paper's log.

04

Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots

Ling Xu, Chuyu Han, Borui Li, Hao Wu, Shiqi Jiang, Ting Cao, Chuanyou Li, Sheng Zhong, Shuai Wang

Southeast University · Nanjing University · Microsoft Research · AIR, Tsinghua University

Project: github.com/SEU-PAISys/Embodied.cpp  ·  2026-07-01

Abstract
Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots - abstract screenshot

Gist · paraphrased from the abstract

Embodied models now span vision-language-action (VLA) and world-action models (WAMs), but deployment is fragmented across model-specific Python stacks and robot glue code, and existing runtimes are built for request-response serving — not the closed-loop, multi-rate, latency-first, batch-1 reality of robots on heterogeneous edge hardware.

Embodied.cpp is a portable C++ runtime that factors a shared execution path into five layers (input adapters, sequence builders, backbone execution, head plugins, deployment adapters), giving modular multi-rate execution, latency-first fused inference, and extensible operators/I-O behind one backend abstraction. On two VLA models (HY-VLA, pi0.5) it reaches 100.0% and 91.0% closed-loop task success; on a preliminary WAM benchmark it cuts block memory from 312.2 MiB to 88.1 MiB — better deployment efficiency without losing accuracy.

Summary in my own words from the abstract screenshot above — not a verbatim quote; verify against the source before citing.

Questions & Answers — none yet. The abstract is on record; send questions to start this paper's log.

05

ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving

★ Key

Sangjin Choi, Sukmin Cho, Yifan Xiong, Ziyue Yang, Youngjin Kwon, Peng Cheng

KAIST · Microsoft Research · Shanghai Xingyunzhili AI Institute

arXiv 2607.00406  ·  2 Jul 2026  ·  method: ELDR

Abstract
ELDR: Expert-Locality-Aware Decode Routing for PD-Disaggregated MoE Serving - abstract screenshot

Gist · paraphrased from the abstract

In prefill-decode (PD) disaggregated serving, a request is handed to a decode worker after prefill. Existing decode routers balance only load — but for MoE models equally-loaded workers can differ in latency, because each decode step must load the weights of every distinct expert its batch activates.

ELDR is an expert-locality-aware decode router. Offline it builds an expert signature predicting which experts a request will fire, and partitions signature space across workers with balanced K-means; online it routes each request to the least-loaded worker best matching its signature, with a signature cache co-indexed to the KV cache (KV-block granularity) to stay exact under prefix caching. In vLLM on up to 40 GPUs, ELDR cuts median TPOT by 5.9-13.9% over the strongest of four load-balancing baselines across three MoE models and two workloads, leaving outputs unchanged.

Summary in my own words from the abstract screenshot above — not a verbatim quote; verify against the source before citing.

Questions & Answers — none yet. The abstract is on record; send questions to start this paper's log.

06

Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning

★ Key

Shijie Li, Yilin Gao, Siyuan Yang, Tieyuan Chen, Chaofan Gan, Zhihao He, Zicheng Zhao, Yuyu Guo, Weiyao Lin, Hang Yu

Shanghai Jiao Tong University · Ant Group

method: AMVL

Abstract
Multimodal Continuous Reasoning via Asymmetric Mutual Variational Learning - abstract screenshot

Gist · paraphrased from the abstract

MLLMs squeeze visual reasoning into discrete tokens (a language-space bottleneck) that lose perceptual nuance. Continuous latent reasoning is a promising alternative, but it suffers a train-inference mismatch: a training-time posterior conditioned on the ground-truth answer can exploit answer-dependent shortcuts the test-time prior can't access — “answer leakage.”

Asymmetric Mutual Variational Learning (AMVL) resolves this with a bidirectional calibration objective: a forward KL trains the answer-agnostic prior to match the posterior, while a novel reverse KL regularizes the posterior so it doesn't collapse into inference-incompatible regions. They formalize leakage as “prior contamination” and prove the dual-KL objective reduces it. Instantiated in a latent-integrated MLLM, AMVL beats strong discrete and latent-reasoning baselines, improving average BLINK by +10.83 (up to +32.00 on individual tasks) with a more stable latent space.

Summary in my own words from the abstract screenshot above — not a verbatim quote; verify against the source before citing.

Questions & Answers — none yet. The abstract is on record; send questions to start this paper's log.

07

Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination

Subhadeep Pal, Shashwat Sourav, Tirthankar Ghosal, Markus J. Buehler

MIT · Washington University in St. Louis · Oak Ridge National Laboratory · Lawrence Berkeley National Laboratory

method: Graph-PRefLexOR  ·  corresp. mbuehler@mit.edu

Abstract
Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination - abstract screenshot

Gist · paraphrased from the abstract

Materials discovery needs AI that reasons in multi-step, domain-grounded ways, but standard LLMs give fluent yet weakly traceable answers — it's hard to tell whether a conclusion rests on coherent intermediate reasoning.

Graph-PRefLexOR is a family of graph-native reasoning models fine-tuned with GRPO to structure reasoning into explicit phases (mechanism exploration, graph construction, pattern extraction, hypothesis synthesis), tying language generation to a symbolic relational graph so causal links can be built, inspected, and reused. On 100 open-ended materials-science/mechanics questions it improves 40-65% over base models (largest gains in traceability), shows roughly 2×-3× more semantic diversity, and test-time graph expansion mainly increases long-range conceptual recombination within a bounded semantic space rather than just widening coverage — positioning graph-native RL as a route to interpretable scientific-hypothesis AI.

Summary in my own words from the abstract screenshot above — not a verbatim quote; verify against the source before citing.

Questions & Answers — none yet. The abstract is on record; send questions to start this paper's log.

08

DemoPSD: Disagreement-Modulated Policy Self-Distillation

★ Key

Yunhe Li, Hao Shi, Wenhao Liu, Mengzhe Ruan, Hanxu Hou, Zhongxiang Dai, Shuang Qiu, Linqi Song

City University of Hong Kong · Tsinghua University · Shenzhen University of Advanced Technology · CUHK-Shenzhen

method: DemoPSD

Abstract
DemoPSD: Disagreement-Modulated Policy Self-Distillation - abstract screenshot

Gist · paraphrased from the abstract

On-policy self-distillation (one model as both teacher and student, with different information access) trains LLMs to reason — but dense token-level supervision from a teacher conditioned on privileged information can overfit, suppress exploration, and cause “privileged information leakage” (the student learns answer-dependent shortcuts unavailable at test time).

DemoPSD selectively adopts teacher guidance: the student follows the teacher where their distributions agree and leans on its own reasoning where they diverge (a sign the teacher is over-influenced by privileged info). Instead of fitting the full teacher, it steers toward a reverse-KL barycenter — a weighted geometric blend of teacher and student — with the per-token discrepancy controlling the blend. They prove leakage attenuation and exploration preservation; on SciKnowEval it beats GRPO and SDPO while keeping higher training entropy and generalizing to out-of-distribution GPQA.

Summary in my own words from the abstract screenshot above — not a verbatim quote; verify against the source before citing.

Questions & Answers — none yet. The abstract is on record; send questions to start this paper's log.

09

TurboServe: Serving Streaming Video Generation Efficiently and Economically

Youhe Jiang, Haoxu Wang, Haotong Bao, Kai Jiang, Jianfei Chen, Jun Zhu, Fangcheng Fu, Jintao Zhang

Shanghai Jiao Tong University · Shengshu Technology · Tsinghua University

Code: github.com/shengshu-ai/TurboServe  ·  method: TurboServe

Abstract
TurboServe: Serving Streaming Video Generation Efficiently and Economically - abstract screenshot

Gist · paraphrased from the abstract

Streaming video generation is a new serving workload: long-lived sessions produce video chunk-by-chunk, must keep session state across active and idle periods, and have to hit a tight per-chunk latency target. That creates session-duration heterogeneity (placements go stale over time) and temporal demand heterogeneity (active sessions spike and ebb).

TurboServe is the first serving system built for this workload, framing serving as online scheduling that jointly coordinates session placement and GPU provisioning. Its closed-loop scheduler pairs a migration-aware placement controller (rebalances sessions to cut worst-case per-chunk latency) with a load-driven autoscaler (sizes the GPU budget to demand), backed by coalesced chunk batching, GPU-CPU offloading for suspend/resume, and NCCL GPU-GPU migration. On real production traces from Shengshu Technology across model sizes and up to 64 NVIDIA B300 GPUs, it reduces worst-case per-chunk latency by 37.5% and total GPU operating cost by 37.2% on average.

Summary in my own words from the abstract screenshot above — not a verbatim quote; verify against the source before citing.

Questions & Answers — none yet. The abstract is on record; send questions to start this paper's log.

10

ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning

Yanjun Zhao, Ruizhong Qiu, Tianxin Wei, Yuanchen Bei, Zhining Liu, Lingjie Chen, Ismini Lourentzou, Hanghang Tong, Jingrui He

University of Illinois Urbana-Champaign

Code: github.com/Yanjun-Zhao/ReContext  ·  method: ReContext

Abstract
ReContext: Recursive Evidence Replay as LLM Harness for Long-Context Reasoning - abstract screenshot

Gist · paraphrased from the abstract

Long-context LLMs can hold huge windows but frequently fail to use evidence that's actually present — a gap between context access and context utilization.

ReContext (Recursive Evidence Replay) is a training-free inference harness that uses the model's own relevance signals to build a query-conditioned evidence pool and “replays” it before final generation — no training, external memory, or context pruning — separating evidence organization from answer generation. They frame context as a memory store (question = retrieval cue, attention = cue-trace association, replay = trace reactivation). Across eight 128K-token datasets it consistently improves context utilization and earns the best average rank on all three backbones (Qwen3-4B, Qwen3-8B, Llama3-8B).

Summary in my own words from the abstract screenshot above — not a verbatim quote; verify against the source before citing.

Questions & Answers — none yet. The abstract is on record; send questions to start this paper's log.