Large language model agents can produce outputs that span the entire quality spectrum, from insightful, well-reasoned content to repetitive, superficial text often termed “AI slop.” This paper examines why the same underlying architectures yield such divergent outcomes. We identify key factors that determine output quality: model alignment and capability, prompt engineering and instruction quality, agentic workflow design, self-reflection and iterative refinement, and test-time compute allocation. We further analyze the mechanisms that lead to low-quality outputs, including hallucination, sycophancy, insufficient grounding, and mode collapse. Conversely, we examine how chain-of-thought reasoning, tool use, verification loops, and multi-agent collaboration enable high-quality generation. Our analysis suggests that content quality is not an intrinsic property of models but emerges from the interaction between model capabilities, task specification, and inferencetime procedures. These findings have implications for deploying AI agents in applications where output quality is paramount.
Rachel So. Agents Can Write Both Top-Quality Content and Bad AI Slop. Why?. Project Rachel, December 2025. https://doi.org/10.71775/kth.dm8xy-5r459