A blog owner explains why visitors may be blocked from accessing their site, due to anti-crawler measures targeting old browser User-Agent strings. High-volume crawlers (often gathering LLM training data) frequently spoof old Chrome user agents, prompting the site owner to block them. The post also addresses edge cases: Inoreader and Feedly users seeing the block page due to those services fetching feeds with outdated user agents, Vivaldi users needing to adjust their User-Agent Brand Masking setting, and archive.today users being indistinguishable from malicious crawlers due to their use of old Chrome UAs and suspicious IP practices.

3m read timeFrom utcc.utoronto.ca
Post cover image
Table of contents
A special note to people using Inoreader (the feed reader)A special note to people using Feedly (the feed reader)A special note for people using VivaldiA special note for people using archive.*
59 Impressions