<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/fai8s2eda" -->

---
title: GitHub - firecrawl/pdf-inspector: Fast Rust library for...
description: pdf-inspector is a fast Rust library for PDF classification and text extraction, available as a crate and with bindings for Python, Node.js, and browser...
canonical: https://daily.dev/posts/fai8s2eda
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: GitHub - firecrawl/pdf-inspector: Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions. | daily.dev
og:description: pdf-inspector is a fast Rust library for PDF classification and text extraction, available as a crate and with bindings for Python, Node.js, and browser...
og:url: https://daily.dev/posts/fai8s2eda
og:image: https://api.daily.dev/og/posts/faI8S2EDa.png
og:image:alt: Post cover image
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# GitHub - firecrawl/pdf-inspector: Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.

**[Alexey Zerkalenkov](https://daily.dev/sources/tzhsbevyajhcmr0fmoxfj)** · [@alexeyzerkalenkov](https://daily.dev/alexeyzerkalenkov) · 108 upvotes · 5 comments

## Summary

pdf-inspector is a fast Rust library for PDF classification and text extraction, available as a crate and with bindings for Python, Node.js, and browser WebAssembly. It detects whether a PDF is text-based, scanned, image-based, or mixed in 10–50ms by sampling content streams, then extracts text with position awareness and converts it to clean Markdown — all without OCR or ML models. Key features include multi-column layout detection, dual-mode table detection (rectangle-based and heuristic), CID font support, and per-page OCR routing. Benchmarked against 200 PDFs, it outperforms pymupdf4llm, markitdown, and opendataloader on overall score, reading order, and table structure, while being the fastest at 0.47s for the full corpus. Built by Firecrawl to skip expensive OCR for the ~54% of PDFs that are already text-based.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://github.com/firecrawl/pdf-inspector>

## Community discussion

Top comments from developers on daily.dev.

**@petecapecod** · 1 upvotes

> sweet I'll be using this later on a pdf project

**@heracles2756** · 1 upvotes

> This looks like a very practical and well-designed PDF processing solution. The combination of smart classification, accurate text extraction, table detection, and browser-based WebAssembly support makes it especially useful. Great work keeping it lightweight without relying on ML models or external services. 👍

**@agustinbarrientos** · 1 upvotes

> I would still send low-confidence tables through a second parser before the Markdown reaches an ingestion pipeline.

## Similar posts on daily.dev

- [Don’t just attend KubeCon \+ CloudNativeCon, Merge Forward your experience\!](https://daily.dev/posts/don-t-just-attend-kubecon-cloudnativecon-merge-forward-your-experience--l0rpp73x8) · CNCF · 1 upvotes · 0 comments
- [Announcing H2 2026 KCDs](https://daily.dev/posts/announcing-h2-2026-kcds-m96goajm1) · CNCF · 1 upvotes · 0 comments
- [Two months of Open Community Groups](https://daily.dev/posts/two-months-of-open-community-groups-asf52zhbs) · CNCF · 0 upvotes · 0 comments

---

Tags: [#nodejs](https://daily.dev/tags/nodejs), [#rust](https://daily.dev/tags/rust), [#markdown](https://daily.dev/tags/markdown)

[View this post on daily.dev](https://daily.dev/posts/fai8s2eda)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"DiscussionForumPosting","mainEntityOfPage":"https://daily.dev/posts/fai8s2eda","headline":"GitHub - firecrawl/pdf-inspector: Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.","text":"Shared: GitHub - firecrawl/pdf-inspector: Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions.","url":"https://daily.dev/posts/fai8s2eda","datePublished":"2026-08-07T14:53:31.054Z","dateModified":"2026-08-07T14:53:31.057Z","author":{"@type":"Person","name":"Alexey Zerkalenkov","url":"https://daily.dev/alexeyzerkalenkov","image":"https://lh3.googleusercontent.com/a/ACg8ocKQBwPwaHC4XrPy0rniDybogCxlwDp7QCLkbwsqNRQSFhqDIlIf=s96-c","description":"Fullstack AI Engineer","interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"EndorseAction"},"userInteractionCount":8460}},"image":"https://media.daily.dev/image/upload/s--0_ODbtD2--/f_auto/v1722860399/public/Placeholder%2008","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":108},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":5}],"sharedContent":{"@type":"WebPage","url":"https://api.daily.dev/r/hQ46jSyRr"},"comment":[{"@type":"Comment","text":"sweet I’ll be using this later on a pdf project","datePublished":"2026-08-09T12:48:33.144Z","url":"https://daily.dev/posts/faI8S2EDa#c-5nKGyvVxh","author":{"@type":"Person","name":"Peter Cruckshank","url":"https://daily.dev/petecapecod","image":"https://media.daily.dev/image/upload/s--ZJhQyKws--/f_auto/v1721235024/avatars/avatar_A9xh33q0QoxtkGoJRCosp"},"interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1}},{"@type":"Comment","text":"This looks like a very practical and well-designed PDF processing solution. The combination of smart classification, accurate text extraction, table detection, and browser-based WebAssembly support makes it especially useful. Great work keeping it lightweight without relying on ML models or external services. 👍","datePublished":"2026-08-09T09:11:56.748Z","url":"https://daily.dev/posts/faI8S2EDa#c-ppUFguXdp","author":{"@type":"Person","name":"heracles 2756","url":"https://daily.dev/heracles2756","image":"https://media.daily.dev/image/upload/s--NoOPHSFv--/f_auto/v1785506917/avatars/avatar_iYIZwiJAMuCGXY2OiXW0x?_a=BAMAMicg0"},"interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1}},{"@type":"Comment","text":"I would still send low-confidence tables through a second parser before the Markdown reaches an ingestion pipeline.","datePublished":"2026-08-08T18:21:24.835Z","url":"https://daily.dev/posts/faI8S2EDa#c-07qpZ40NO","author":{"@type":"Person","name":"Agustin Barrientos","url":"https://daily.dev/agustinbarrientos","image":"https://media.daily.dev/image/upload/s--5ayxQnqn--/f_auto/v1788281802/avatars/avatar_wQYYVe5Tbj0NJ7C7qPoa8?_a=BAMAMicg0"},"interactionStatistic":{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":1}}],"isPartOf":{"@type":"WebPage","url":"https://daily.dev/sources/tzhsbevyajhcmr0fmoxfj","name":"Alexey Zerkalenkov"}}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Alexey Zerkalenkov","item":"https://daily.dev/sources/tzhsbevyajhcmr0fmoxfj"},{"@type":"ListItem","position":3,"name":"GitHub - firecrawl/pdf-inspector: Fast Rust library for PDF inspection, classification, and text extraction. Intelligently detects scanned vs text-based PDFs to enable smart routing decisions."}]}
```

