Hacker News
Read post

[2405.05417] Fishing for Magikarp: Automatically Detecting Under-trained Tokens in Large Language Models

A comprehensive analysis of Large Language Model (LLM) tokenizers is presented, focusing on the detection of untrained and under-trained tokens. The prevalence of such tokens across various models is demonstrated, along with insights for improving the efficiency and safety of language models.

    #computer-science
May 12, 2024•1m read time•From arxiv.org
Post cover image
35 Impressions
Hacker News's image
Hacker News

Hacker News is a community-driven platform for sharing and discussing technology news, startups, and...

17.4K Followers

•

141.8K Upvotes

Would you recommend this post?

Copy link
WhatsApp
Facebook
X
New Squad
  • © 2026 Daily Dev Ltd.
  • Guidelines
  • Explore
  • Tags
  • Sources
  • Squads
  • Leaderboard