OpenAI's opt-out model for its GPTBot web crawler is criticized as fundamentally unfair to content creators. By default, OpenAI scrapes web content for LLM training unless site owners explicitly block it via robots.txt. The author argues this should be opt-in and ideally compensated, drawing a parallel to theft. The piece contends that the opt-out approach reflects a disregard for the value of writing and writers, and that the only truly fair LLMs would be those trained exclusively on content the company owns or has licensed.

3m read timeFrom hidde.blog
Post cover image
25 Impressions