Is This Slop? Detecting AI-Generated Content Without a Model
This title could be clearer and more informative.Try out Clickbait Shieldfor free (5 uses left this month).
Human detection of AI-generated text is nearly as accurate as a coin flip, yet overconfident moderation tools flag legitimate human writing as AI slop. This piece catalogs research-backed linguistic tells that LLMs exhibit — excess vocabulary words like 'delve,' contrastive phrasing ('It's not X — it's Y'), linguistic hedging, compulsive rule-of-three lists, and em-dash overuse — and explains the mathematical reasons these patterns emerge. SFT trains models on a small, non-representative annotator pool, embedding idiosyncratic word preferences and even recurring fictional name priors (e.g., Marcus Chen, Elena Vasquez). RLHF then reinforces hedging because cautious answers are harder for time-pressured raters to penalize than confident ones. The conclusion: these are weak statistical signals, not fingerprints, and as models improve, even these tells will fade.