<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/how-to-search-for-long-hashes-and-ids-with-dict-keywords-5cna8pqrd" -->

---
title: How to search for long hashes and IDs with dict=&#x27;keywords
description: Manticore Search&#x27;s default keywords dictionary truncates tokens to 42 bytes after normalization, which breaks searching long values like SHA-256 hashes,...
canonical: https://daily.dev/posts/how-to-search-for-long-hashes-and-ids-with-dict-keywords-5cna8pqrd
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: How to search for long hashes and IDs with dict=&#x27;keywords | daily.dev
og:description: Manticore Search&#x27;s default keywords dictionary truncates tokens to 42 bytes after normalization, which breaks searching long values like SHA-256 hashes,...
og:url: https://daily.dev/posts/how-to-search-for-long-hashes-and-ids-with-dict-keywords-5cna8pqrd
og:image: https://api.daily.dev/og/posts/5CNa8pQrd.png
og:image:alt: How to search for long hashes and IDs with dict=&#x27;keywords
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# How to search for long hashes and IDs with dict='keywords

**[Manticore](https://daily.dev/sources/manticoresearch)** · 10 min read · 0 upvotes · 0 comments

## Summary

Manticore Search's default keywords dictionary truncates tokens to 42 bytes after normalization, which breaks searching long values like SHA-256 hashes, message IDs, and long email addresses since different IDs sharing the same first 42 bytes become indistinguishable. The dict='keywords_32k' option, available since Manticore Search 27.1.1 (27.1.5+ recommended for migrations), raises the maximum token length to 32768 bytes, skipping oversized tokens with a warning instead of truncating them. The guide covers schema choices (string attribute vs string attribute indexed), exact vs full-text matching semantics, prefix/infix search with min_infix_len, blend_chars for emails, CALL KEYWORDS for inspecting tokenization, migration steps for RT and plain tables, why dict='crc' doesn't help, current limitations (no CALL SUGGEST, no percolate tables, no snippet highlighting, no full-text REGEX), and a warning against indexing secrets like API keys and tokens.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://manticoresearch.com/blog/dict_keywords_32k>

## Questions this post answers

### What is the maximum token length for Manticore Search's default keywords dictionary?

The default dict='keywords' dictionary in Manticore Search truncates tokens to 42 bytes after normalization, measured in bytes rather than characters, so UTF-8 characters can consume the limit faster. Tokens exceeding 42 bytes are truncated both when indexing documents and when processing search queries, which can cause two different long IDs sharing the same first 42 bytes to become indistinguishable.

_Anyone indexing hashes or IDs in Manticore Search can track dictionary limits and workarounds on daily.dev._

### How do I search full SHA-256 hashes or long IDs in Manticore Search without truncation?

Use dict='keywords_32k', available starting with Manticore Search 27.1.1 (version 27.1.5+ recommended for converting existing tables), which raises the maximum normalized token length from 42 bytes to 32768 bytes. Tokens exceeding the new limit are skipped with a warning instead of truncated, and morphology is not applied to tokens over 42 bytes since they typically have no useful word stem.

_Developers building log or ID search on Manticore Search can follow updates like this on daily.dev._

### Does dict='crc' in Manticore Search support long token search like keywords_32k does?

No, dict='crc' stores keyword checksums instead of original text but does not increase the allowed token length beyond the regular 42-byte limit. Only dict='keywords_32k' implements the exception to that limit; keywords and keywords_32k also store term text, which allows Manticore to expand prefix and infix wildcard queries against the dictionary, something crc cannot do.

_Teams choosing between Manticore dictionary types can compare trade-offs like this on daily.dev._

## Similar posts on daily.dev

- [Manticore Search 27.1.5: Authentication, sharded tables, conversational search and faster vector search](https://daily.dev/posts/manticore-search-27-1-5-authentication-sharded-tables-conversational-search-and-faster-vector-sea-vaeuj65j7) · Manticore · 2 upvotes · 0 comments
- [Inline Stopwords, Exceptions, and Wordforms](https://daily.dev/posts/inline-stopwords-exceptions-and-wordforms-4n5sirdee) · Manticore · 0 upvotes · 0 comments
- [Building high-performance full-text search for object storage](https://daily.dev/posts/building-high-performance-full-text-search-for-object-storage-l0orx1vdy) · ClickHouse · 1 upvotes · 0 comments
- [Accelerate search queries with full-text search indexes on Databricks](https://daily.dev/posts/accelerate-search-queries-with-full-text-search-indexes-on-databricks-qva3uojdd) · databricks · 0 upvotes · 0 comments
- [Search — The Evolution of the Karpathy LLM Wiki · Los Techies](https://daily.dev/posts/search-the-evolution-of-the-karpathy-llm-wiki-los-techies-ydht2z37z) · LosTechies · 0 upvotes · 0 comments

---

Tags: [#backend](https://daily.dev/tags/backend), [#nlp](https://daily.dev/tags/nlp), [#full-text-search](https://daily.dev/tags/full-text-search), [#manticore-search](https://daily.dev/tags/manticore-search)

[View this post on daily.dev](https://daily.dev/posts/how-to-search-for-long-hashes-and-ids-with-dict-keywords-5cna8pqrd)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"How to search for long hashes and IDs with dict='keywords","url":"https://daily.dev/posts/how-to-search-for-long-hashes-and-ids-with-dict-keywords-5cna8pqrd","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/how-to-search-for-long-hashes-and-ids-with-dict-keywords-5cna8pqrd"},"datePublished":"2026-08-30T05:16:47.556Z","dateModified":"2026-09-13T19:22:15.245Z","description":"Manticore Search's default keywords dictionary truncates tokens to 42 bytes after normalization, which breaks searching long values like SHA-256 hashes,...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/455e94d421b4544e2ed6d75de68b2c23?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/455e94d421b4544e2ed6d75de68b2c23?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Manticore","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Manticore","logo":"https://media.daily.dev/image/upload/s--FP5Ea08h--/f_auto/v1754225081/logos/manticoresearch","url":"https://daily.dev/sources/manticoresearch"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/how-to-search-for-long-hashes-and-ids-with-dict-keywords-5cna8pqrd","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"backend,nlp,full-text-search,manticore-search","timeRequired":"PT10M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Manticore","item":"https://daily.dev/sources/manticoresearch"},{"@type":"ListItem","position":3,"name":"How to search for long hashes and IDs with dict='keywords"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/how-to-search-for-long-hashes-and-ids-with-dict-keywords-5cna8pqrd#faq","mainEntity":[{"@type":"Question","name":"What is the maximum token length for Manticore Search's default keywords dictionary?","acceptedAnswer":{"@type":"Answer","text":"The default dict='keywords' dictionary in Manticore Search truncates tokens to 42 bytes after normalization, measured in bytes rather than characters, so UTF-8 characters can consume the limit faster. Tokens exceeding 42 bytes are truncated both when indexing documents and when processing search queries, which can cause two different long IDs sharing the same first 42 bytes to become indistinguishable. Anyone indexing hashes or IDs in Manticore Search can track dictionary limits and workarounds on daily.dev."}},{"@type":"Question","name":"How do I search full SHA-256 hashes or long IDs in Manticore Search without truncation?","acceptedAnswer":{"@type":"Answer","text":"Use dict='keywords_32k', available starting with Manticore Search 27.1.1 (version 27.1.5+ recommended for converting existing tables), which raises the maximum normalized token length from 42 bytes to 32768 bytes. Tokens exceeding the new limit are skipped with a warning instead of truncated, and morphology is not applied to tokens over 42 bytes since they typically have no useful word stem. Developers building log or ID search on Manticore Search can follow updates like this on daily.dev."}},{"@type":"Question","name":"Does dict='crc' in Manticore Search support long token search like keywords_32k does?","acceptedAnswer":{"@type":"Answer","text":"No, dict='crc' stores keyword checksums instead of original text but does not increase the allowed token length beyond the regular 42-byte limit. Only dict='keywords_32k' implements the exception to that limit; keywords and keywords_32k also store term text, which allows Manticore to expand prefix and infix wildcard queries against the dictionary, something crc cannot do. Teams choosing between Manticore dictionary types can compare trade-offs like this on daily.dev."}}]}
```

