Core idea
Real-world T2I requests are underspecified — the model lacks the user’s intent, domain knowledge, and up-to-date facts. Qwen-Image-Agent treats the query as partial context and progressively builds generation-ready context: Context-Aware Planning spots what’s missing, Context Grounding gathers it across modalities (search + memory + feedback), and only then is the T2I model invoked. Shifts T2I from monolithic generation to modular agent orchestration; reports gains on IA-Bench, Mindbench, and WISE-Verified.