<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/lfm2-5-encoders-for-fast-long-context-inference-on-cpu-x49ilhnc8" -->

---
title: LFM2.5-Encoders for Fast Long-Context Inference on CPU
description: Liquid AI has released two open-weight encoder models — LFM2.5-Encoder-230M and LFM2.5-Encoder-350M — designed for fast, long-context NLP inference on CPU....
canonical: https://daily.dev/posts/lfm2-5-encoders-for-fast-long-context-inference-on-cpu-x49ilhnc8
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: LFM2.5-Encoders for Fast Long-Context Inference on CPU | daily.dev
og:description: Liquid AI has released two open-weight encoder models — LFM2.5-Encoder-230M and LFM2.5-Encoder-350M — designed for fast, long-context NLP inference on CPU....
og:url: https://daily.dev/posts/lfm2-5-encoders-for-fast-long-context-inference-on-cpu-x49ilhnc8
og:image: https://api.daily.dev/og/posts/x49IlHNC8.png
og:image:alt: LFM2.5-Encoders for Fast Long-Context Inference on CPU
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# LFM2.5-Encoders for Fast Long-Context Inference on CPU

**[Hugging Face](https://daily.dev/sources/huggingface)** · 6 min read · 2 upvotes · 1 comments

## Summary

Liquid AI has released two open-weight encoder models — LFM2.5-Encoder-230M and LFM2.5-Encoder-350M — designed for fast, long-context NLP inference on CPU. Built on the LFM2 architecture, they support 8,192-token contexts and are approximately 3.7× faster than ModernBERT-base at long context on CPU. The models are initialized from LFM2 decoder backbones and converted to bidirectional encoders via masked language modeling training. They match or outperform larger models on GLUE, SuperGLUE, and multilingual benchmarks. Use cases include intent routing, policy linting, PII detection, and text classification. Both models are available on Hugging Face and can be fine-tuned with the provided tutorial.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://huggingface.co/blog/LiquidAI/lfm2-5-encoders>

## Community discussion

Top comments from developers on daily.dev.

**@kartiknvj** · 0 upvotes

> Converting decoder backbones into bidirectional encoders through masked language modeling is a smart way to reuse pretraining, and the CPU speedup at 8k context is what makes this practical for the boring high-volume jobs. The listed use cases (intent routing, PII detection, policy linting) are exactly the places where paying for a generative model per call never made sense. I would want to see the long-context latency curve past 4k tokens, since that is usually where encoder throughput on CPU starts to bend.

## Similar posts on daily.dev

- [Up to 3.2x Faster Inference with LFM2.5-DSpark](https://daily.dev/posts/up-to-3-2x-faster-inference-with-lfm2-5-dspark-nwilr6rly) · Hugging Face · 0 upvotes · 0 comments

---

Tags: [#machine-learning](https://daily.dev/tags/machine-learning), [#llm](https://daily.dev/tags/llm), [#nlp](https://daily.dev/tags/nlp), [#transformers](https://daily.dev/tags/transformers)

[View this post on daily.dev](https://daily.dev/posts/lfm2-5-encoders-for-fast-long-context-inference-on-cpu-x49ilhnc8)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"LFM2.5-Encoders for Fast Long-Context Inference on CPU","url":"https://daily.dev/posts/lfm2-5-encoders-for-fast-long-context-inference-on-cpu-x49ilhnc8","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/lfm2-5-encoders-for-fast-long-context-inference-on-cpu-x49ilhnc8"},"datePublished":"2026-07-28T15:02:02.553Z","dateModified":"2026-07-28T20:18:13.819Z","description":"Liquid AI has released two open-weight encoder models — LFM2.5-Encoder-230M and LFM2.5-Encoder-350M — designed for fast, long-context NLP inference on CPU....","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/d7a9bcb7ddb2053eed86b99a0b7cd419?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/d7a9bcb7ddb2053eed86b99a0b7cd419?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Hugging Face","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Hugging Face","logo":"https://media.daily.dev/image/upload/t_logo,f_auto/v1/logos/f1f55c67d81a4330acf5b90b26b0c8e1","url":"https://daily.dev/sources/huggingface"},"commentCount":1,"discussionUrl":"https://daily.dev/posts/lfm2-5-encoders-for-fast-long-context-inference-on-cpu-x49ilhnc8","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":2},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":1}],"keywords":"machine-learning,llm,nlp,transformers","timeRequired":"PT6M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Hugging Face","item":"https://daily.dev/sources/huggingface"},{"@type":"ListItem","position":3,"name":"LFM2.5-Encoders for Fast Long-Context Inference on CPU"}]}
{"@context":"https://schema.org","@type":"WebPage","@id":"https://daily.dev/posts/lfm2-5-encoders-for-fast-long-context-inference-on-cpu-x49ilhnc8","comment":[{"@type":"Comment","text":"Converting decoder backbones into bidirectional encoders through masked language modeling is a smart way to reuse pretraining, and the CPU speedup at 8k context is what makes this practical for the boring high-volume jobs. The listed use cases (intent routing, PII detection, policy linting) are exactly the places where paying for a generative model per call never made sense. I would want to see the long-context latency curve past 4k tokens, since that is usually where encoder throughput on CPU starts to bend.","datePublished":"2026-07-28T18:48:20.423Z","url":"https://daily.dev/posts/x49IlHNC8#c-QaePqIeaE","author":{"@type":"Person","name":"kartik-nvjk","url":"https://daily.dev/kartiknvj","image":"https://media.daily.dev/image/upload/s--3gGgsVCw--/f_auto/v1781456774/avatars/avatar_TvTVeiMdkRCqWUDullFmy?_a=BAMAMiWQ0"}}]}
```

