<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/mit-researchers-find-ai-generated-images-grow-harder-to-trace-back-to-training-data-as-models-scale-wrd3yvkh5" -->

---
title: MIT researchers find AI-generated images grow harder to...
description: MIT CSAIL researchers Zheng Dai and David Gifford introduce the concept of &#x27;attribution decay,&#x27; showing that as diffusion models scale up in training data...
canonical: https://daily.dev/posts/mit-researchers-find-ai-generated-images-grow-harder-to-trace-back-to-training-data-as-models-scale-wrd3yvkh5
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: MIT researchers find AI-generated images grow harder to trace back to training data as models scale | daily.dev
og:description: MIT CSAIL researchers Zheng Dai and David Gifford introduce the concept of &#x27;attribution decay,&#x27; showing that as diffusion models scale up in training data...
og:url: https://daily.dev/posts/mit-researchers-find-ai-generated-images-grow-harder-to-trace-back-to-training-data-as-models-scale-wrd3yvkh5
og:image: https://api.daily.dev/og/posts/WRd3YvkH5.png
og:image:alt: MIT researchers find AI-generated images grow harder to trace back to training data as models scale
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# MIT researchers find AI-generated images grow harder to trace back to training data as models scale

**[Collections](https://daily.dev/sources/collections)** · 2 min read · 0 upvotes · 0 comments

## Summary

MIT CSAIL researchers Zheng Dai and David Gifford introduce the concept of 'attribution decay,' showing that as diffusion models scale up in training data size, the influence of any single training image on generated outputs shrinks dramatically. Using a novel diffusion ensemble architecture that avoids costly full retraining, they tested 24 ensembles across datasets from 256 to over 160,000 images and found an inverse power law relationship between dataset size and per-image influence. The findings challenge a key argument in copyright lawsuits against Stability AI and Midjourney, which hinge on whether outputs can be traced back to specific training images.

## Content

There's a paper out of MIT CSAIL that complicates one of the central arguments in AI copyright lawsuits: the idea that you can trace a generated image back to the training data that produced it.

Zheng Dai and David Gifford call the phenomenon "attribution decay." The basic finding: as diffusion models get bigger and train on more images, any single training image matters less and less to what comes out the other end. Pull one image out of the training set, or even every image by a specific artist, and at scale the model's outputs barely change. Not "change in a way nobody notices." Change in a way that's measurably, statistically almost nothing.

**How they tested it**

The obvious problem with studying this is that removing training data and retraining a diffusion model from scratch is expensive - you'd need to do it over and over to get any real signal. So Dai and Gifford built something they call a diffusion ensemble architecture, which lets them exactly remove the influence of specific training data without retraining the whole model each time. That's the part that makes this study possible at all.

They ran this across 24 diffusion ensembles trained on datasets ranging from 256 images up to more than 160,000, pulling from multiple public datasets. The pattern held consistently: the relationship between dataset size and any single image's influence follows an inverse power law. Bigger dataset, less influence, and the drop-off isn't gentle.

**Why this matters for the lawsuits**

This lands right in the middle of ongoing litigation against Stability AI and Midjourney, where a lot of the legal argument hinges on whether generated images are derivative works of specific training images. If an artist's entire body of work can be stripped from the training set and the model's outputs don't budge, what exactly is the

## Questions this post answers

### What is attribution decay in diffusion models?

Attribution decay is the phenomenon where a single training image's influence on a diffusion model's outputs shrinks as the training dataset grows larger. MIT CSAIL researchers Zheng Dai and David Gifford found this follows an inverse power law across 24 diffusion ensembles trained on datasets ranging from 256 to over 160,000 images, meaning removing one image, or even an artist's entire body of work, barely changes outputs at scale.

_Follow daily.dev for research shaping how AI copyright disputes over training data get argued._

### How did researchers study the effect of removing training data from diffusion models without retraining from scratch?

They built a diffusion ensemble architecture that lets researchers exactly remove the influence of specific training images without retraining the entire model each time. This method made it feasible to test 24 different diffusion ensembles across datasets of varying sizes, from 256 images to over 160,000, and measure how single-image influence changes with scale.

_daily.dev surfaces techniques like this for developers tracking how generative models are evaluated._

---

Tags: [#machine-learning](https://daily.dev/tags/machine-learning), [#genai](https://daily.dev/tags/genai), [#mit](https://daily.dev/tags/mit), [#diffusion-models](https://daily.dev/tags/diffusion-models)

[View this post on daily.dev](https://daily.dev/posts/mit-researchers-find-ai-generated-images-grow-harder-to-trace-back-to-training-data-as-models-scale-wrd3yvkh5)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"MIT researchers find AI-generated images grow harder to trace back to training data as models scale","url":"https://daily.dev/posts/mit-researchers-find-ai-generated-images-grow-harder-to-trace-back-to-training-data-as-models-scale-wrd3yvkh5","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/mit-researchers-find-ai-generated-images-grow-harder-to-trace-back-to-training-data-as-models-scale-wrd3yvkh5"},"datePublished":"2026-08-19T01:49:31.188Z","dateModified":"2026-08-19T01:50:08.951Z","description":"MIT CSAIL researchers Zheng Dai and David Gifford introduce the concept of 'attribution decay,' showing that as diffusion models scale up in training data...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/a121cf1ff5c9007d30366618fb361a24?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/a121cf1ff5c9007d30366618fb361a24?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Collections","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Collections","logo":"https://media.daily.dev/image/upload/s--fk_6ycEi--/f_auto,q_auto/v1780996001/logos/collections?_a=BAMAMiWQ0","url":"https://daily.dev/sources/collections"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/mit-researchers-find-ai-generated-images-grow-harder-to-trace-back-to-training-data-as-models-scale-wrd3yvkh5","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"machine-learning,genai,mit,diffusion-models","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Collections","item":"https://daily.dev/sources/collections"},{"@type":"ListItem","position":3,"name":"MIT researchers find AI-generated images grow harder to trace back to training data as models scale"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/mit-researchers-find-ai-generated-images-grow-harder-to-trace-back-to-training-data-as-models-scale-wrd3yvkh5#faq","mainEntity":[{"@type":"Question","name":"What is attribution decay in diffusion models?","acceptedAnswer":{"@type":"Answer","text":"Attribution decay is the phenomenon where a single training image's influence on a diffusion model's outputs shrinks as the training dataset grows larger. MIT CSAIL researchers Zheng Dai and David Gifford found this follows an inverse power law across 24 diffusion ensembles trained on datasets ranging from 256 to over 160,000 images, meaning removing one image, or even an artist's entire body of work, barely changes outputs at scale. Follow daily.dev for research shaping how AI copyright disputes over training data get argued."}},{"@type":"Question","name":"How did researchers study the effect of removing training data from diffusion models without retraining from scratch?","acceptedAnswer":{"@type":"Answer","text":"They built a diffusion ensemble architecture that lets researchers exactly remove the influence of specific training images without retraining the entire model each time. This method made it feasible to test 24 different diffusion ensembles across datasets of varying sizes, from 256 images to over 160,000, and measure how single-image influence changes with scale. daily.dev surfaces techniques like this for developers tracking how generative models are evaluated."}}]}
```

