Core idea
Agentic Self-Instruct: an inner loop where agents build datasets by sampling, reasoning, and self-feedback; an outer meta-optimizer improves the agent’s data-generation policy over time. Turns extra inference compute into higher training-data quality; gains on CS research, legal reasoning, and math vs. classical synthetic-data methods.