Benchmark · Finance
FinanceBench: ATG analyses financial reports fast and right
The test: 150 demanding questions whose answers sit in long financial reports.
Key takeaways
- On the 150 public FinanceBench questions, ATG answers 138 correctly (92%), searching on its own through 368 financial filings, about 53,900 pages.
- Under the same conditions, Mistral reports 86% correct answers with its Agentic Search (August 2026), at a mean latency of 71 s, against 11 s for ATG.
The test
What is FinanceBench?
FinanceBench is a public question set released in late 2023 by Patronus AI to measure how well an AI answers financial analyst questions from real documents: annual reports (10-K), quarterly reports (10-Q), current reports (8-K) and earnings call transcripts.
Each question has a single, verifiable answer: an amount, a ratio, an explanation taken from the filing. The difficulty comes from the documents: filings of about 150 pages on average, figures in tables, calculations across several lines, accounting vocabulary.
- 150
- public questions
- 368
- SEC filings
- ~53,900
- pages to search
Three families of questions
Metrics extraction
Find or compute an indicator: capital expenditure, operating margin, quick ratio.
Analyst questions
Generic questions an analyst asks about any company: is the company capital-intensive?
Company-specific questions
Questions tied to one company and one filing, which require reasoning over its content.
Results and comparison
Shared corpus
Every filing in one knowledge base, without saying which one holds the answer. Real-world conditions.
ATG, EU-only providers, Fast modeAccuracy93%Mean latency31 s
ATG, Worldwide providers, Fast modeAccuracy92%Mean latency11 s
Mistral Agentic Search, GLM-5.2, search + navigationAccuracy86%Mean latency71 s
Dewey, Claude Opus 4.6, agentic retrievalAccuracy84%Mean latencynot published
Mistral Agentic Search, Mistral Medium 3.5, search + navigation *Accuracy83%Mean latency71 s
Mistral Agentic Search, Mistral Medium 3.5, search only *Accuracy74%Mean latency108 s
Dewey, GPT-5.4, agentic retrievalAccuracy63%Mean latencynot published- FinSage, GPT-4o, multi-aspect RAG149 questions, 83 filingsAccuracy57%Mean latencynot published
Balyasny, GPT-4o, improved RAG, BAM embeddingsAccuracy55%Mean latencynot published
Unstructured, GPT-4, element-based chunking141 questions, 80 filingsAccuracy53%Mean latencynot published
Balyasny, GPT-4o, improved RAG, ada-002 embeddingsAccuracy47%Mean latencynot published
Ragie, shared store (top-k 8)Accuracy27%Mean latencynot published
Mistral, Mistral Medium 3.5, one-shot RAGAccuracy27%Mean latencynot publishedGPT-4 Turbo, shared vector storeAccuracy19%Mean latencynot published
Every published result
The full table, unfiltered: each score with its test conditions, scope, grading method and source.
| System | Conditions | Accuracy | Mean latency | Scope | Grading | Source |
|---|---|---|---|---|---|---|
| No document | 9% | not published | Full benchmark | Human review | Islam et al., FinanceBench (Patronus AI), arXiv:2311.119442023-11-20 | |
| Shared corpus | 93% | 31 s | Full benchmark | LLM judge + human review | This run (2026-10-02) | |
| Shared corpus | 92% | 11 s | Full benchmark | LLM judge + human review | This run (2026-10-02) | |
| Shared corpus | 86% | 71 s | Full benchmark | LLM judge | Mistral AI, Agentic Search2026-08-20 | |
| Shared corpus | 84% | not published | Full benchmark | Numeric match + LLM judge | Dewey, financebench-eval (GitHub)2026-04-08 | |
| Shared corpus | 83% * | 71 s | Full benchmark | LLM judge | Mistral AI, Agentic Search2026-08-20 | |
| Shared corpus | 74% * | 108 s | Full benchmark | LLM judge | Mistral AI, Agentic Search2026-08-20 | |
| Shared corpus | 63% | not published | Full benchmark | Numeric match + LLM judge | Dewey, financebench-eval (GitHub)2026-04-08 | |
| FinSageGPT-4o, multi-aspect RAG | Shared corpus | 57% | not published | 149 questions, 83 filings | Human review | Wang et al., FinSage (CIKM 2025), arXiv:2504.144932025-04-20 |
| Shared corpus | 55% | not published | Full benchmark | Human review | Anderson et al. (Balyasny Asset Management), arXiv:2411.071422024-11-11 | |
| Shared corpus | 53% | not published | 141 questions, 80 filings | Human review | Jimeno Yepes et al. (Unstructured), arXiv:2402.051312024-02-05 | |
| Shared corpus | 47% | not published | Full benchmark | Human review | Anderson et al. (Balyasny Asset Management), arXiv:2411.071422024-11-11 | |
| Shared corpus | 27% | not published | Full benchmark | Not stated | Ragie, How Ragie Outperformed the FinanceBench Test2024-10-22 | |
| Shared corpus | 27% | not published | Full benchmark | LLM judge | Mistral AI, Agentic Search2026-08-20 | |
| Shared corpus | 19% | not published | Full benchmark | Human review | Islam et al., FinanceBench (Patronus AI), arXiv:2311.119442023-11-20 | |
| FinanceBench_RAGhybrid page-level retrieval | One filing, indexed | 76% | not published | Full benchmarkCompany and year filters applied before the search. | Not stated | FinanceBench_RAG (GitHub, aquib8112)2026-02-28 |
| One filing, indexed | 51% | not published | Full benchmark | Not stated | Ragie, How Ragie Outperformed the FinanceBench Test2024-10-22 | |
| One filing, indexed | 50% | not published | Full benchmark | Human review | Islam et al., FinanceBench (Patronus AI), arXiv:2311.119442023-11-20 | |
| Whole filing provided | 79% | not published | Full benchmark | Human review | Islam et al., FinanceBench (Patronus AI), arXiv:2311.119442023-11-20 | |
| Exact page provided | 85% | not published | Full benchmark | Human review | Islam et al., FinanceBench (Patronus AI), arXiv:2311.119442023-11-20 | |
| Pathwaymulti-agent pipeline, with a human answering clarifications | Conditions not stated | 56% | not published | Number of questions not statedA human answers clarification requests: not an autonomous score. Model and grading not stated. | Not stated | Pathway, AI for SEC filings analysis2025-07-09 |
| Pathwaymulti-agent pipeline | Conditions not stated | 42% | not published | Number of questions not statedModel and grading not stated. | Not stated | Pathway, AI for SEC filings analysis2025-07-09 |
| Pathwaysimple RAG | Conditions not stated | 24% | not published | Number of questions not statedModel and grading not stated. | Not stated | Pathway, AI for SEC filings analysis2025-07-09 |
Third-party figures quoted from their publications, not re-run by us. A reduced scope (fewer questions or filings) makes the test easier. * Computed from the deltas published by the source.
Protocol
How we tested
- 1
One corpus
The 368 filings are imported as is into a standard ATG knowledge base, as a customer would.
- 2
No file indicated
Each question is asked on its own, in a new conversation. ATG has to find the right document among the 368 without being told. For 11 questions whose original wording names neither the company nor the period, a fixed sentence gives them ("Added note by ATG for disambiguation"), with the rest of the question unchanged.
- 3
Grading
An independent LLM judge (Mistral Medium 3.5) compares each answer with the reference answer. In each test, 4 verdicts were overturned by a human review, with its justification in the report. A question without an answer counts as an error.
- 4
End-to-end latency
Time measured from sending the question to the end of the answer, search included, with questions asked one at a time. The reports give the mean, the median and the p90.
Transparency
Each downloadable report details our methodology then, for each of the 150 questions, the expected answer, the answer ATG gave and the verdict. Anyone can check a grade and challenge it.
Results
How ATG performs, AI policy by AI policy
The configurations tested
One row per AI policy, both in Fast mode. Each row is a full run over the 150 questions.
| AI policy | Fast |
|---|---|
| Worldwide providers | Highlighted 92% 138 / 150 correct answers
|
| EU-only providers | 93% 139 / 150 correct answers
|
Accuracy and response time
The higher and further left a point sits, the more accurate and faster the answer.
Comparison
Against Mistral Agentic Search
On 20 August 2026, Mistral published its FinanceBench results: the same 150 questions, the same 368 filings, with an agentic search that chains queries and navigates within documents. It is the published result closest to our conditions today.
Two differences with our protocol: Mistral grades answers with an LLM judge calibrated against human labels, and reports latency across the whole benchmark (mean and p90), not per model.
×6.8
In Fast mode, ATG answers in 11 s on average, against 71 s for the fastest configuration published by Mistral AI: answers 6.8 times faster.
Accuracy
ATG, EU-only providers, Fast mode
93%ATG, Worldwide providers, Fast mode
92%
Mistral Agentic Search, GLM-5.2, search + navigation86%
Mistral Agentic Search, Mistral Medium 3.5, search + navigation *83%
Mistral Agentic Search, Mistral Medium 3.5, search only *74%
Mistral, Mistral Medium 3.5, one-shot RAG27%
Mean latency
ATG, EU-only providers, Fast mode
31 sATG, Worldwide providers, Fast mode
11 s
Mistral Agentic Search, GLM-5.2, search + navigation71 s
Mistral Agentic Search, Mistral Medium 3.5, search + navigation *71 s
Mistral Agentic Search, Mistral Medium 3.5, search only *108 s
Mistral, Mistral Medium 3.5, one-shot RAGnot published
Mistral figures quoted from their post, not reproduced by us. * Computed from the deltas published by the source.
Example
Example: one question and ATG's answer
Question (from the test set)
What is the FY2018 capital expenditure amount (in USD millions) for 3M? Give a response to the question by relying on the details shown in the cash flow statement.
Expected answer
$1577.00
ATG's answer (unrestricted policy, Fast mode)
$1,577 million. Per 3M's Consolidated Statement of Cash Flows (fiscal year ended December 31, 2018), the line "Purchases of property, plant and equipment (PP&E)" under Cash Flows from Investing Activities is $(1,577) million, which is the capital expenditure figure.
Document cited by ATG
3M_2018_10K.pdf
Verdict: Correct, answered in 7 seconds
Download the results
One report per configuration, with the protocol, the results and, for every question, ATG's answer and its verdict. Free access, no form.
| AI policy | Fast |
|---|---|
| Worldwide providers | PDF, Worldwide providers, Fast |
| EU-only providers | PDF, EU-only providers, Fast |
Frequently asked questions
Why FinanceBench rather than another benchmark?
Researchers and vendors publish their scores on FinanceBench. That is what lets us compare ourselves with results we did not produce, on real financial documents.
Why not run the other solutions yourselves?
To avoid comparing a solution we master with solutions we might configure poorly. We quote the figures each vendor or researcher published, with their test conditions, and group them by difficulty level.
What is the difference between worldwide and EU-only providers?
It is the AI policy chosen by the administrator. EU-only restricts processing to European providers (Mistral, Scaleway, Nebius); worldwide (the unrestricted policy) adds OpenAI and Google among others. We ran the test with both so everyone knows the cost of their choice, in accuracy and speed.
What does the Fast mode mean?
It is the ATG mode that favours response time. Both tests were run in this mode. The Auto and Thinking modes, which leave more time for reasoning, are not part of this benchmark.
Do these results apply to my documents?
They give an order of magnitude on long, table-heavy financial documents. The best test remains your own documents: we can measure ATG's accuracy on a sample of your questions during a demo.
What about your own filings?
Import your financial reports, ask your teams' questions and measure the accuracy of the answers yourself.