
Mistral OCR: cut your bill 10x with open weights
Mistral OCR costs $4 per 1,000 pages. An open-weight OCR on a European GPU drops to $0.15 and stays sovereign. Benchmark, break-even point and pitfalls.
In March 2025, Mistral OCR cost $1 per 1,000 pages. In December, $2. Since June 2026, $4. The price has quadrupled in fifteen months.
Take a company that wants to make its document base searchable by an AI: 10,000 documents of 100 pages on average, or one million pages. At today's rate, OCR alone costs $4,000, before anyone has asked the first question.
We ran the same workload on PaddleOCR-VL, an open-weight model, deployed on a GPU rented in Europe. At list price, with the GPU busy the whole time: $151. Twenty-six times less.
And with no loss of sovereignty: the machines belong to a Finnish operator, and no page is sent to the model's publisher.
A GPU is never 100% busy, which is why our title says ten rather than twenty-six.
Bar chart comparing the cost of OCR for one million pages: $4,000 with Mistral OCR 4.1, $2,000 with Mistral OCR 3 or OCR 4 in batch, $151 with PaddleOCR-VL on an on-demand RTX PRO 6000 GPU at Verda, $76 on the same GPU at spot pricing
Assumptions: RTX PRO 6000 GPU at Verda, 100% busy, at 4 pages per second (our measurement on 173 pages, one machine).
Why a PDF parser is not enough
PDF is a print format. It describes where to put each character on the page. The file has no idea what a heading or a table is: it only holds glyphs and coordinates.
A PDF parser (PyMuPDF, pdfplumber, pypdf) reads that text layer and hands it back to you. It is instant and free. When the layer exists, it is even accurate to the character. The trouble comes from everything the layer does not contain.
The document has no text layer. A contract signed and then scanned, a fax, a photo of a delivery note: the PDF holds nothing but an image. The parser returns an empty page.
The text is inside the images. In our test set, a 25-slide presentation holds 148 images for 470 words of extractable text. On a 34-page technical guide full of screenshots, Mistral OCR found 546 words that appear nowhere in the text layer.
Tables do not exist. In a PDF, a table is a set of lines and numbers placed at coordinates. The parser gives you the numbers one after the other, without the columns. A figure cut off from its column header is useless: the language model that receives it has to guess what it refers to.
Reading order and noise. Two columns, a callout box, a footnote: the parser follows the order of the file, which is not necessarily the order you read in. It also keeps everything that repeats. On a 28-page specification in our test set, 669 words were nothing but header and footer lines repeated from page to page.
For a RAG system, all of this is paid for further down the chain. When an assistant answers beside the point, the cause is often a document that was badly read on the way in. It is one of the mistakes that show up after the POC.
What an AI OCR does
Classic OCR, the Tesseract kind, recognizes characters line by line. It reads a clean scan and stops there: no structure, tables in a jumble.
The current generation works in two steps. A first model splits the page into blocks (heading, paragraph, table, figure, header) and puts them back in reading order. A second one, a vision-language model, transcribes each block. The result is a Markdown file with its headings, tables and images, ready for an LLM to use. This is what people are after when they search for "pdf to markdown".
There are two ways to get it. The first is a managed API, and Mistral OCR is the best-known example in France: you send a PDF, you get Markdown back, you pay per page.
The second is an open-weight model that you host yourself. PaddleOCR-VL, published by Baidu under the Apache 2.0 license, has fewer than one billion parameters and fits on a single GPU. In that case you rent the machine and pay by the hour.
The model comes from China. That takes nothing away from the sovereignty of the setup, and we come back to it after the numbers.
What Mistral OCR costs
| Version | Release | Price per 1,000 pages |
|---|---|---|
| Mistral OCR | March 2025 | $1 |
| Mistral OCR 3 | December 2025 | $2 |
| Mistral OCR 4, then 4.1 | June and July 2026 | $4 ($2 in batch mode) |
Sources: March 2025 announcement, OCR 4.1 model page and Silicon.fr (in French) for the history.
To be fair, each version does more than the one before. OCR 4 reads 170 languages, returns the position of every paragraph and takes on handwriting. But if your documents are typed reports in French or English, you pay for that progress without using it. OCR 3 is still available at $2, and nothing says for how long.
So our million pages cost $4,000 at list price, or $2,000 if you accept batch processing or the older model.
And that amount is only the first invoice. New documents arrive, old ones get updated. Above all, every time you improve your ingestion pipeline (a new model, a different setting for tables), you run the whole stock again. At $4,000 a pass, you hesitate to test an improvement. At $150, you test it.
PaddleOCR-VL on a rented GPU: the same million pages
Our deployment runs at Verda, a Finnish operator whose data centers are in Finland and Iceland. The machine: an RTX PRO 6000 GPU, rented at $2.18 per hour on demand and $1.09 at spot pricing (spare capacity, interruptible at any time).
On September 30, 2026, we measured its throughput on 9 real PDFs totaling 173 pages. A single machine tops out at around 4 pages per second, or 14,400 pages per hour.
One million pages at 14,400 pages per hour is 69 hours of GPU. At $2.18 per hour, $151.
| Cost per 1,000 pages | One million pages | |
|---|---|---|
| Mistral OCR 4.1 | $4 | $4,000 |
| Mistral OCR 3, or OCR 4 in batch | $2 | $2,000 |
| PaddleOCR-VL, on-demand GPU | $0.15 | $151 |
| PaddleOCR-VL, spot GPU | $0.08 | $76 |
Why we say ten, and not twenty-six
That $0.15 assumes a GPU that works during every second you are billed for. But a rented GPU is paid by the hour, whether it is converting pages or waiting.
So everything depends on how busy it is:
| Pages processed per billed hour | GPU occupancy | Result versus Mistral OCR 4.1 |
|---|---|---|
| 545 | 4% | Same cost |
| 1,090 | 8% | 2 times cheaper (the price of Mistral OCR 3) |
| 5,450 | 38% | 10 times cheaper |
| 14,400 | 100% | 26 times cheaper |
A stock to process is the ideal case: you fill the queue, the machine saturates, you switch it off at the end. A trickle of three documents an hour is the worst: the GPU waits, and you pay for the wait.
A GPU that stays on all the time makes no sense
The easiest mistake to make: rent a machine, leave it on and discover the invoice at the end of the month. A GPU at $2.18 per hour running day and night costs $1,591 a month, whether it converts a million pages or none. For that price, Mistral OCR 4.1 processes close to 400,000 pages for you. Below that monthly volume, a machine left on around the clock costs more than the API.
Take 100,000 pages a month:
| Setup | Monthly cost |
|---|---|
| Mistral OCR 4.1 | $400 |
| Rented GPU, always on | $1,591 |
| Rented GPU, on for 7 hours then paused | $15, plus the startup minutes |
So you need a system that starts the container when documents arrive and pauses it as soon as the queue is empty. Without it, open weights cost more than what they replace.
Verda makes this possible. Its Serverless Containers scale down to zero machines when there is nothing to do and hold requests in a queue while a machine starts. Only the minutes it actually ran are billed, startup and shutdown included.
Two settings remain yours to decide. First, the startup threshold: waking a GPU for three documents costs more than sending them to an API. On our side, the machine starts once 50 documents have come in within fifteen minutes and pauses after ten minutes of inactivity. Second, what happens to the documents that arrive during startup, which takes a few minutes. On our side, they go to the fallback engine without waiting.
Cheaper, and just as sovereign
Moving from a French vendor's API to a model published by Baidu takes nothing away from the sovereignty of the setup. A sovereign AI is judged on three layers: the law that applies to the provider, the infrastructure that processes your data and the models that read it.
The provider. Verda is a Finnish company under European law, just as Mistral is a French company.
The infrastructure. Its data centers are in Finland and Iceland, within a European perimeter covered by the GDPR. That is where your documents are converted.
The model. PaddleOCR-VL is published under the Apache 2.0 license: we downloaded its weights once, and it runs in our own container. No page goes back to its publisher. The path your data takes depends on where the model runs and on who operates the machine. The country it was trained in changes nothing.
On one point, open weights even do better than an API, French or not: nobody can change their price. Weights under an Apache 2.0 license do not double in price in December, and nobody removes them from the catalog. Being sovereign also means not depending on someone else's price list.
We keep Mistral OCR in our pipeline, as a fallback. That choice is a matter of cost and risk, not of flags.
Beyond the hourly rate
The price gap is real. It comes with trade-offs.
Mistral is faster
On a single document, Mistral OCR 4.1 processes a page in 0.20 seconds, against 0.43 to 0.56 for our deployment. The gap widens under load: with eight documents sent at once, Mistral reaches 18 to 20 pages per second while our machine stays at around 4. And on a saturated machine, a one-page document that takes 2 seconds on its own waits more than 20 behind the big ones.
That ceiling is the ceiling of a single machine. If throughput matters, rent several: each machine adds its 4 pages per second, and the cost per page does not move as long as they are busy. Our million pages take 69 hours on one GPU, close to three days. On four GPUs, they go through in 17 hours, for the same $151.
Several machines do not shorten the conversion of one document taken alone. For a user who drops a PDF into a conversation and waits for the answer, the API remains faster.
Spot or API, you need fallbacks
At spot pricing, we regularly see the machine taken back from us mid-run. Very rarely at night, though, and night is when most of our processing happens.
The API is no safer. Over two months of production, 4.8% of our calls to Mistral OCR failed with a 404 error: 402 calls out of 8,320, close to one in twenty.
Either way, you need fallbacks. Whether a GPU disappears or an API returns an error, ingestion must not stop: the document goes to another engine, with no manual step. No single link holds on its own, which is true of OCR and of the rest of the AI chain.
Quality: equivalent on human review
We reviewed the output of both engines by eye. PaddleOCR-VL and Mistral OCR give very similar results. Both do better than what we have seen with LightOnOCR and with Docling.
An automatic measurement backs up that review. For each engine, we counted the share of words in the PDFs' text layer that it recovers: 93.6% for PaddleOCR-VL, 97.7% for Mistral OCR 3, 98.3% for Mistral OCR 4.1. Most of that gap is intentional, since our configuration strips headers, footers and footnotes.
Two real differences remain. Mistral transcribes the text in screenshots, while PaddleOCR-VL mostly leaves them as images. And PaddleOCR-VL writes its tables in HTML, with entities in place of apostrophes, which calls for some cleanup.
The automatic measurement covers 9 documents, none of which is a scan. Before deciding, run both engines on fifty of your own documents.
It has to be operated
A container image of more than 15 GB, around five minutes of cold start, billing in ten-minute slices, the controller that starts and pauses the machine, a fallback when it does not respond. None of this is hard taken separately. Together it adds up to a few weeks of engineering work, then monitoring over time.
When to switch, when to stay on the API
| Your situation | The right choice | Why |
|---|---|---|
| Low volume, as documents come in (a few thousand pages a month) | The API | The stakes are a few tens of dollars, a GPU is not worth the effort |
| Large stock or regular reprocessing (hundreds of thousands of pages) | Open weights on a rented GPU | The GPU runs full: 10 to 26 times cheaper |
| A user is waiting for the answer in front of the screen | The API | A single document goes through faster, with no startup to wait for |
| Full control of the chain is required | Open weights on a rented GPU | The model runs on your side, and nobody can change its price |
| Both volume and interactive use | Both, chained | The GPU handles the volume, the API takes over when it is stopped or unavailable |
The last case is the most common. It is also the only one that requires building something.
At Ask This Guy, this work is already done
Our ingestion pipeline chains three engines. PaddleOCR-VL on our GPUs at Verda goes first. If it does not respond, the document goes to Mistral OCR, without waiting for a machine to start. Docling is the last resort. A controller switches the GPU on when the volume justifies it and pauses it after ten minutes of inactivity.
The first link runs on our GPUs in Finland and Iceland, and the fallback is provided by a French vendor. We brought the cost down without touching the level of sovereignty, which we detail in our approach to sovereign AI.
Our customers see none of this. They upload their documents or connect their sources, then query them in their document assistant. We describe the other ways to run models, from purchased GPUs to a turnkey platform, in our guide to enterprise AI inference.
Frequently asked questions about Mistral OCR and open-source OCR
How much does Mistral OCR cost?
As of September 30, 2026, Mistral OCR 4.1 costs $4 per 1,000 pages and $2 in batch mode. The older model, Mistral OCR 3, is still available at $2 per 1,000 pages. At launch in March 2025, the rate was $1 per 1,000 pages. For one million pages, expect $2,000 to $4,000 per pass.
What is the difference between a PDF parser and an OCR?
A PDF parser reads the text already stored in the file: it is instant and free, but it sees neither scanned documents nor the text inside images, and it does not rebuild tables. An OCR analyzes the image of the page. AI-based OCRs also recognize the structure of the document (headings, tables, reading order) and produce structured text, usually Markdown.
Is there an open-source alternative to Mistral OCR?
Yes, several. PaddleOCR-VL, published by Baidu under the Apache 2.0 license, is the one we chose: a model of fewer than one billion parameters that fits on a single GPU. LightOnOCR, Docling, MinerU and DeepSeek-OCR are other options. All of them require renting or owning a GPU and operating the deployment yourself.
Is an open-weight OCR published by a Chinese company sovereign?
Yes, if it runs on infrastructure you control. An open-weight model such as PaddleOCR-VL, published by Baidu under the Apache 2.0 license, is downloaded once and then runs in your own container: no document is sent to its publisher. Sovereignty then depends on who operates the machine and where it is located. Deployed with an operator under European law, in data centers located in Europe, such an OCR is as sovereign as a French vendor's API.
At what volume does self-hosted OCR become cost-effective?
On a GPU rented at $2.18 per hour, self-hosted OCR costs less than Mistral OCR 4.1 from 545 pages processed per billed hour, and less than Mistral OCR 3 from 1,090 pages. To cut the bill by ten, you need around 5,450 pages per hour, which means a GPU that is 38% busy. This calculation assumes the machine is paused as soon as it has nothing left to process: left on around the clock, it costs $1,591 a month and only pays off above 400,000 pages a month.
Is PaddleOCR-VL as accurate as Mistral OCR?
On human review, the results of PaddleOCR-VL and Mistral OCR are very similar, and better than those we have seen with LightOnOCR and Docling. In an automatic measurement over 173 pages, PaddleOCR-VL recovers 93.6% of the words in the PDFs' text layer, against 97.7% for Mistral OCR 3 and 98.3% for Mistral OCR 4.1. Most of that gap comes from headers, footers and footnotes, which our configuration strips on purpose.
Conclusion
OCR looks like a cost line with no story to it. One million pages and a price that doubles twice are enough to give it one.
An open-weight model on a GPU rented in Europe brings that line down from $4,000 to a few hundred, with no loss of sovereignty or quality. On one condition: the machine must only run when it has work to do.
If you have a document stock to make searchable and would rather not build all of this yourself: book a demo.
Measurements taken on September 30, 2026: 9 PDFs, 173 pages, one machine per deployment, one to two runs per engine. Public prices from Mistral and Verda recorded the same day.


