<!-- mobian-agent-page publisher="dailydev" canonical="https://daily.dev/posts/hands-on-turn-scientific-figures-into-structured-data-with-mistral-ocr-wrwpit8ip" -->

---
title: [Hands-on] Turn Scientific Figures Into Structured Data...
description: A hands-on pipeline shows how to use Mistral OCRv4 combined with a CrewAI-orchestrated multimodal agent (Mistral Small 4) to extract structured data from...
canonical: https://daily.dev/posts/hands-on-turn-scientific-figures-into-structured-data-with-mistral-ocr-wrwpit8ip
twitter:card: summary_large_image
twitter:site: @dailydotdev
og:type: website
og:site_name: daily.dev
og:title: [Hands-on] Turn Scientific Figures Into Structured Data with Mistral OCR | daily.dev
og:description: A hands-on pipeline shows how to use Mistral OCRv4 combined with a CrewAI-orchestrated multimodal agent (Mistral Small 4) to extract structured data from...
og:url: https://daily.dev/posts/hands-on-turn-scientific-figures-into-structured-data-with-mistral-ocr-wrwpit8ip
og:image: https://api.daily.dev/og/posts/wRwPiT8Ip.png
og:image:alt: [Hands-on] Turn Scientific Figures Into Structured Data with Mistral OCR
og:image:width: 1200
og:image:height: 630
og:locale: en
---

> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# [Hands-on] Turn Scientific Figures Into Structured Data with Mistral OCR

**[Daily Dose of Data Science \| Avi Chawla \| Substack](https://daily.dev/sources/dailydoseofds)** · 14 min read · 0 upvotes · 0 comments

## Summary

A hands-on pipeline shows how to use Mistral OCRv4 combined with a CrewAI-orchestrated multimodal agent (Mistral Small 4) to extract structured data from scientific figures in PDFs, rather than relying on surrounding text. The workflow issues a single OCR request that returns page text, layout, embedded images, and schema-defined figure metadata, then matches figures to images, scores extraction confidence, and passes clean figure images to a multimodal agent for chart-type validation and interpretation. Includes a code walkthrough and a note that a newer OCR 4.1 model has since shipped, improving bounding-box accuracy and multi-column handling, with the mistral-ocr-latest alias now pointing to it.

## Full article

daily.dev links to this article rather than hosting it. Read it at the original source: <https://blog.dailydoseofds.com/p/hands-on-turn-scientific-figures>

## Questions this post answers

### What is the difference between mistral-ocr-4-0 and OCR 4.1, and do I need to change my code to use the new version?

OCR 4.1 keeps every capability of mistral-ocr-4-0 but reads busy, marked-up pages more precisely: bounding boxes align to each element instead of drifting, callouts on dense technical diagrams stay as separate regions instead of merging, and multi-column pages return each column separately. The mistral-ocr-latest alias now points to 4.1, so switching only requires changing the model string.

_Track model string changes like this on daily.dev before they break a pinned OCR pipeline._

### Why do standard PDF parsers fail to extract data from scientific figures like charts and graphs?

Standard PDF parsers extract the text layer correctly but save figures only as image references (like PNGs), losing the actual numbers, axis values, and data points inside charts. They break pages into isolated blocks, so a legend gets separated from its chart and a caption loses its connection to the figure it describes, even though the parser itself is working as designed.

_Developers building document pipelines can find OCR and figure-extraction approaches like this on daily.dev._

### How can I automatically match extracted figure captions to the correct embedded images from an OCR response?

Break the extracted caption into words and compare them against the OCR text to find the strongest overlap on a page, then assign the next unused image on that matched page to the figure using a small cursor so multiple figures on one page each get a distinct image. If a caption is too short or ambiguous to match confidently, fall back to the page number reported by the OCR model.

_Engineers wiring up figure-to-image matching logic can follow similar OCR workflows on daily.dev._

## Similar posts on daily.dev

- [Mistral OCR 4: cheap, self-hosted document AI](https://daily.dev/posts/mistral-ocr-4-cheap-self-hosted-document-ai-1wizdbrck) · The Next Web · 1 upvotes · 0 comments
- [Mistral OCR 4.1](https://daily.dev/posts/mistral-ocr-4-1-vodhfccvq) · Hacker News · 0 upvotes · 0 comments
- [Mistral Releases OCR 3 With Improved Accuracy on Handwritten and Structured Documents](https://daily.dev/posts/mistral-releases-ocr-3-with-improved-accuracy-on-handwritten-and-structured-documents-ruteqrhno) · InfoQ · 0 upvotes · 0 comments
- [Introducing Mistral OCR 3](https://daily.dev/posts/introducing-mistral-ocr-3-kdrjfdidf) · Hacker News · 1 upvotes · 0 comments
- [Unlocking Document Understanding with Mistral Document AI in Microsoft Foundry](https://daily.dev/posts/unlocking-document-understanding-with-mistral-document-ai-in-microsoft-foundry-yno9mitm9) · Microsoft Azure · 0 upvotes · 0 comments

---

Tags: [#machine-learning](https://daily.dev/tags/machine-learning), [#multimodal](https://daily.dev/tags/multimodal)

[View this post on daily.dev](https://daily.dev/posts/hands-on-turn-scientific-figures-into-structured-data-with-mistral-ocr-wrwpit8ip)

```json
{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://daily.dev/#organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180},"sameAs":["https://twitter.com/dailydotdev","https://github.com/dailydotdev","https://www.linkedin.com/company/daily-dev-ltd"]},{"@type":"WebSite","@id":"https://daily.dev/#website","url":"https://daily.dev","name":"daily.dev","publisher":{"@id":"https://daily.dev/#organization"},"potentialAction":{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https://daily.dev/search?q={search_term_string}"},"query-input":"required name=search_term_string"}}]}
{"@context":"https://schema.org","@type":"TechArticle","headline":"[Hands-on] Turn Scientific Figures Into Structured Data with Mistral OCR","url":"https://daily.dev/posts/hands-on-turn-scientific-figures-into-structured-data-with-mistral-ocr-wrwpit8ip","mainEntityOfPage":{"@type":"WebPage","@id":"https://daily.dev/posts/hands-on-turn-scientific-figures-into-structured-data-with-mistral-ocr-wrwpit8ip"},"datePublished":"2026-08-26T21:10:00.492Z","dateModified":"2026-09-14T08:51:16.496Z","description":"A hands-on pipeline shows how to use Mistral OCRv4 combined with a CrewAI-orchestrated multimodal agent (Mistral Small 4) to extract structured data from...","image":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/5c54807e18133b3afaf1f63e75d567b8?_a=AQAEuop","thumbnailUrl":"https://media.daily.dev/image/upload/f_auto,q_auto/v1/posts/5c54807e18133b3afaf1f63e75d567b8?_a=AQAEuop","isAccessibleForFree":true,"articleSection":"Daily Dose of Data Science | Avi Chawla | Substack","inLanguage":"en","publisher":{"@type":"Organization","name":"daily.dev","url":"https://daily.dev","logo":{"@type":"ImageObject","url":"https://daily.dev/apple-touch-icon.png","width":180,"height":180}},"author":{"@type":"Organization","name":"Daily Dose of Data Science | Avi Chawla | Substack","logo":"https://media.daily.dev/image/upload/s--4IHQgTOw--/f_auto/v1710503712/logos/dailydoseofds","url":"https://daily.dev/sources/dailydoseofds"},"commentCount":0,"discussionUrl":"https://daily.dev/posts/hands-on-turn-scientific-figures-into-structured-data-with-mistral-ocr-wrwpit8ip","interactionStatistic":[{"@type":"InteractionCounter","interactionType":{"@type":"LikeAction"},"userInteractionCount":0},{"@type":"InteractionCounter","interactionType":{"@type":"CommentAction"},"userInteractionCount":0}],"keywords":"machine-learning,multimodal","timeRequired":"PT14M"}
{"@context":"https://schema.org","@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://daily.dev"},{"@type":"ListItem","position":2,"name":"Daily Dose of Data Science | Avi Chawla | Substack","item":"https://daily.dev/sources/dailydoseofds"},{"@type":"ListItem","position":3,"name":"[Hands-on] Turn Scientific Figures Into Structured Data with Mistral OCR"}]}
{"@context":"https://schema.org","@type":"FAQPage","@id":"https://daily.dev/posts/hands-on-turn-scientific-figures-into-structured-data-with-mistral-ocr-wrwpit8ip#faq","mainEntity":[{"@type":"Question","name":"What is the difference between mistral-ocr-4-0 and OCR 4.1, and do I need to change my code to use the new version?","acceptedAnswer":{"@type":"Answer","text":"OCR 4.1 keeps every capability of mistral-ocr-4-0 but reads busy, marked-up pages more precisely: bounding boxes align to each element instead of drifting, callouts on dense technical diagrams stay as separate regions instead of merging, and multi-column pages return each column separately. The mistral-ocr-latest alias now points to 4.1, so switching only requires changing the model string. Track model string changes like this on daily.dev before they break a pinned OCR pipeline."}},{"@type":"Question","name":"Why do standard PDF parsers fail to extract data from scientific figures like charts and graphs?","acceptedAnswer":{"@type":"Answer","text":"Standard PDF parsers extract the text layer correctly but save figures only as image references (like PNGs), losing the actual numbers, axis values, and data points inside charts. They break pages into isolated blocks, so a legend gets separated from its chart and a caption loses its connection to the figure it describes, even though the parser itself is working as designed. Developers building document pipelines can find OCR and figure-extraction approaches like this on daily.dev."}},{"@type":"Question","name":"How can I automatically match extracted figure captions to the correct embedded images from an OCR response?","acceptedAnswer":{"@type":"Answer","text":"Break the extracted caption into words and compare them against the OCR text to find the strongest overlap on a page, then assign the next unused image on that matched page to the figure using a small cursor so multiple figures on one page each get a distinct image. If a caption is too short or ambiguous to match confidently, fall back to the page number reported by the OCR model. Engineers wiring up figure-to-image matching logic can follow similar OCR workflows on daily.dev."}}]}
```

