Even as AI firms expand model capabilities, frontier labs continue to navigate a legal reckoning sparked by the first wave of copyright lawsuits that threatened to derail the generative AI boom.
But a landmark research paper, ‘Outputs of Generative Diffusion Models are Often Unattributable,’ by MIT Computer Science and Artificial Intelligence Laboratory (CSAIL) is effectively dismantling the scientific basis for that claim, suggesting that in the world of frontier models, visual similarity is not the same thing as causal theft. The paper’s findings could make the legal teams of artists uneasy. The paper showed that an AI’s reliance on specific training samples decays predictably as the data pool expands. For AI labs like Midjourney or OpenAI, this research provides a powerful shield. They can now argue that even if they had never seen a specific artist’s work, their model would have produced a nearly identical result anyway. The research focuses specifically on Diffusion Models — the tech behind DALL-E, Stable Diffusion, and Midjourney — which generate images by iteratively removing noise from a canvas until a coherent picture emerges. These models operate differently than the autoregressive models that power Large Language Models (LLMs) like ChatGPT or Claude. When an LLM reproduces a copyrighted paragraph word-for-word, it is far easy to match the output to a copyrighted source. In the case of LLMs, the model stores a specific sequence of words to learn how a sentence in written.
The argument from artists has been that AI output that looks like a specific piece of art is because that work was “stolen” during training the model. By using “ablatable ensembles” — a machine learning method than can systematically remove specific training data to test outputs — researchers have found that as training datasets grow, the influence of any single image or artist drops toward zero. authors have a much more tangible “smoking gun While image generators might hide behind the “unattributability” of scale.

