LayerCraft: Enhancing Text-to-Image Generation with CoT Reasoning and Layered Object Integration
batch collage editingchain-of-thought reasoningchainarchitectcoordinator agentlayercraftlayered object integrationmulti-step editingnarrative scene generationobject consistencyobject integration networkspatial compositionstructured generationtext-to-image generationvisual content refinement
Text-to-image (T2I) generation has made remarkable progress, yet existing systems still lack intuitive control over spatial composition, object consistency, and multi-step editing. We present **LayerCraft**, a modular framework that uses large language models (LLMs) as autonomous agents to orchestrate structured, layered image generation and editing. LayerCraft supports two key capabilities: (1) *structured generation* from simple prompts via chain-of-thought (CoT) reasoning, enabling it to decompose scenes, reason about object placement, and guide composition in a controllable, interpretable manner; and (2) *layered object integration*, allowing users to insert and customize objects