<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/news-sites-are-blocking-internet-archive-over-ai-scraping-fears-2mqqvfze1" -->

---
title: News Sites Are Blocking Internet Archive Over AI...
description: More than 340 local news outlets are blocking the Internet Archive&#x27;s Wayback Machine crawlers, citing fears that their content will be scraped for AI/LLM...
canonical: https://daily.dev/posts/news-sites-are-blocking-internet-archive-over-ai-scraping-fears-2mqqvfze1
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: News Sites Are Blocking Internet Archive Over AI Scraping Fears | daily.dev
og:description: More than 340 local news outlets are blocking the Internet Archive&#x27;s Wayback Machine crawlers, citing fears that their content will be scraped for AI/LLM...
og:url: https://daily.dev/posts/news-sites-are-blocking-internet-archive-over-ai-scraping-fears-2mqqvfze1
og:image: https://api.daily.dev/og/posts/2MQqvfzE1.png
og:image:alt: News Sites Are Blocking Internet Archive Over AI Scraping Fears
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# News Sites Are Blocking Internet Archive Over AI Scraping Fears

**[Hackaday](https://daily.dev/sources/hackaday)** · 2 min read · 0 upvotes · 0 comments

## Summary

More than 340 local news outlets are blocking the Internet Archive's Wayback Machine crawlers, citing fears that their content will be scraped for AI/LLM training. Outlets like The Baltimore Banner claim concern over improper AI citations, while others like The Atlantic have blanket anti-scraping policies. Notably, these same outlets allow paid commercial archiving services like ProQuest and LexisNexis to index their content, suggesting financial motivations. The practical consequence is that researchers lose free access to archived news content and the Wayback Machine's coverage of news becomes increasingly incomplete.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://hackaday.com/2026/06/08/news-sites-are-blocking-internet-archive-over-ai-scraping-fears>

## Similar posts on daily.dev

- [News publishers limit Internet Archive access due to AI scraping concerns](https://daily.dev/posts/news-publishers-limit-internet-archive-access-due-to-ai-scraping-concerns-pfgxshyte) · Hacker News · 2 upvotes · 0 comments
- [More than 340 local news outlets are limiting the Internet Archive’s access to their journalism](https://daily.dev/posts/more-than-340-local-news-outlets-are-limiting-the-internet-archive-s-access-to-their-journalism-owopohxwc) · Hacker News · 1 upvotes · 0 comments

---

Tags: [#crawling](https://daily.dev/tags/crawling)

[View this post on daily.dev](https://daily.dev/posts/news-sites-are-blocking-internet-archive-over-ai-scraping-fears-2mqqvfze1)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"News Sites Are Blocking Internet Archive Over AI Scraping Fears","url":"https://daily.dev/posts/news-sites-are-blocking-internet-archive-over-ai-scraping-fears-2mqqvfze1","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/news-sites-are-blocking-internet-archive-over-ai-scraping-fears-2mqqvfze1"},"datePublished":"2026-06-08T20:27:45.541Z","dateModified":"2026-06-08T20:44:25.449Z","description":"More than 340 local news outlets are blocking the Internet Archive's Wayback Machine crawlers, citing fears that their content will be scraped for AI/LLM...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/1baf5384542ba6208addb1be6b994868?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/1baf5384542ba6208addb1be6b994868?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Hackaday","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Hackaday","logo":"https://media.daily.dev/image/upload/s--JhDfhy70--/f_auto,q_auto/v1767541814/logos/hackaday","url":"https://daily.dev/sources/hackaday"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/news-sites-are-blocking-internet-archive-over-ai-scraping-fears-2mqqvfze1","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"crawling","timeRequired":"PT2M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Hackaday","item":"https://daily.dev/sources/hackaday"},{"@type":"ListItem","position":3,"name":"News Sites Are Blocking Internet Archive Over AI Scraping Fears"}]}
```

