Benchmark · Finance

FinanceBench: ATG analyses financial reports fast and right

The test: 150 demanding questions whose answers sit in long financial reports.

By Jean-Christophe BudinUpdated on

Key takeaways

  • On the 150 public FinanceBench questions, ATG answers 138 correctly (92%), searching on its own through 368 financial filings, about 53,900 pages.
  • Under the same conditions, Mistral reports 86% correct answers with its Agentic Search (August 2026), at a mean latency of 71 s, against 11 s for ATG.

The test

What is FinanceBench?

FinanceBench is a public question set released in late 2023 by Patronus AI to measure how well an AI answers financial analyst questions from real documents: annual reports (10-K), quarterly reports (10-Q), current reports (8-K) and earnings call transcripts.

Each question has a single, verifiable answer: an amount, a ratio, an explanation taken from the filing. The difficulty comes from the documents: filings of about 150 pages on average, figures in tables, calculations across several lines, accounting vocabulary.

150
public questions
368
SEC filings
~53,900
pages to search

Three families of questions

Metrics extraction

Find or compute an indicator: capital expenditure, operating margin, quick ratio.

Analyst questions

Generic questions an analyst asks about any company: is the company capital-intensive?

Company-specific questions

Questions tied to one company and one filing, which require reasoning over its content.

Results and comparison

Who
Test conditions

Shared corpus

Every filing in one knowledge base, without saying which one holds the answer. Real-world conditions.

  • ATG, EU-only providers, Fast modeAccuracy93%Mean latency31 s
  • ATG, Worldwide providers, Fast modeAccuracy92%Mean latency11 s
  • Mistral Agentic Search, GLM-5.2, search + navigationAccuracy86%Mean latency71 s
  • Dewey, Claude Opus 4.6, agentic retrievalAccuracy84%Mean latencynot published
  • Mistral Agentic Search, Mistral Medium 3.5, search + navigation *Accuracy83%Mean latency71 s
  • Mistral Agentic Search, Mistral Medium 3.5, search only *Accuracy74%Mean latency108 s
  • Dewey, GPT-5.4, agentic retrievalAccuracy63%Mean latencynot published
  • FinSage, GPT-4o, multi-aspect RAG149 questions, 83 filingsAccuracy57%Mean latencynot published
  • Balyasny, GPT-4o, improved RAG, BAM embeddingsAccuracy55%Mean latencynot published
  • Unstructured, GPT-4, element-based chunking141 questions, 80 filingsAccuracy53%Mean latencynot published
  • Balyasny, GPT-4o, improved RAG, ada-002 embeddingsAccuracy47%Mean latencynot published
  • Ragie, shared store (top-k 8)Accuracy27%Mean latencynot published
  • Mistral, Mistral Medium 3.5, one-shot RAGAccuracy27%Mean latencynot published
  • GPT-4 Turbo, shared vector storeAccuracy19%Mean latencynot published

Every published result

The full table, unfiltered: each score with its test conditions, scope, grading method and source.

Every published result
SystemConditionsAccuracyMean latencyScopeGradingSource
GPT-4 TurboNo document9%not publishedFull benchmarkHuman reviewIslam et al., FinanceBench (Patronus AI), arXiv:2311.119442023-11-20
ATGEU-only providers, Fast modeShared corpus93%31 sFull benchmarkLLM judge + human reviewThis run (2026-10-02)
ATGWorldwide providers, Fast modeShared corpus92%11 sFull benchmarkLLM judge + human reviewThis run (2026-10-02)
Mistral Agentic SearchGLM-5.2, search + navigationShared corpus86%71 sFull benchmarkLLM judgeMistral AI, Agentic Search2026-08-20
DeweyClaude Opus 4.6, agentic retrievalShared corpus84%not publishedFull benchmarkNumeric match + LLM judgeDewey, financebench-eval (GitHub)2026-04-08
Mistral Agentic SearchMistral Medium 3.5, search + navigationShared corpus83% *71 sFull benchmarkLLM judgeMistral AI, Agentic Search2026-08-20
Mistral Agentic SearchMistral Medium 3.5, search onlyShared corpus74% *108 sFull benchmarkLLM judgeMistral AI, Agentic Search2026-08-20
DeweyGPT-5.4, agentic retrievalShared corpus63%not publishedFull benchmarkNumeric match + LLM judgeDewey, financebench-eval (GitHub)2026-04-08
FinSageGPT-4o, multi-aspect RAGShared corpus57%not published149 questions, 83 filingsHuman reviewWang et al., FinSage (CIKM 2025), arXiv:2504.144932025-04-20
BalyasnyGPT-4o, improved RAG, BAM embeddingsShared corpus55%not publishedFull benchmarkHuman reviewAnderson et al. (Balyasny Asset Management), arXiv:2411.071422024-11-11
UnstructuredGPT-4, element-based chunkingShared corpus53%not published141 questions, 80 filingsHuman reviewJimeno Yepes et al. (Unstructured), arXiv:2402.051312024-02-05
BalyasnyGPT-4o, improved RAG, ada-002 embeddingsShared corpus47%not publishedFull benchmarkHuman reviewAnderson et al. (Balyasny Asset Management), arXiv:2411.071422024-11-11
Ragieshared store (top-k 8)Shared corpus27%not publishedFull benchmarkNot statedRagie, How Ragie Outperformed the FinanceBench Test2024-10-22
MistralMistral Medium 3.5, one-shot RAGShared corpus27%not publishedFull benchmarkLLM judgeMistral AI, Agentic Search2026-08-20
GPT-4 Turboshared vector storeShared corpus19%not publishedFull benchmarkHuman reviewIslam et al., FinanceBench (Patronus AI), arXiv:2311.119442023-11-20
FinanceBench_RAGhybrid page-level retrievalOne filing, indexed76%not publishedFull benchmarkCompany and year filters applied before the search.Not statedFinanceBench_RAG (GitHub, aquib8112)2026-02-28
Ragieper-document store (top-k 32)One filing, indexed51%not publishedFull benchmarkNot statedRagie, How Ragie Outperformed the FinanceBench Test2024-10-22
GPT-4 Turboper-document vector storeOne filing, indexed50%not publishedFull benchmarkHuman reviewIslam et al., FinanceBench (Patronus AI), arXiv:2311.119442023-11-20
GPT-4 TurboWhole filing provided79%not publishedFull benchmarkHuman reviewIslam et al., FinanceBench (Patronus AI), arXiv:2311.119442023-11-20
GPT-4 TurboExact page provided85%not publishedFull benchmarkHuman reviewIslam et al., FinanceBench (Patronus AI), arXiv:2311.119442023-11-20
Pathwaymulti-agent pipeline, with a human answering clarificationsConditions not stated56%not publishedNumber of questions not statedA human answers clarification requests: not an autonomous score. Model and grading not stated.Not statedPathway, AI for SEC filings analysis2025-07-09
Pathwaymulti-agent pipelineConditions not stated42%not publishedNumber of questions not statedModel and grading not stated.Not statedPathway, AI for SEC filings analysis2025-07-09
Pathwaysimple RAGConditions not stated24%not publishedNumber of questions not statedModel and grading not stated.Not statedPathway, AI for SEC filings analysis2025-07-09

Third-party figures quoted from their publications, not re-run by us. A reduced scope (fewer questions or filings) makes the test easier. * Computed from the deltas published by the source.

Protocol

How we tested

  1. 1

    One corpus

    The 368 filings are imported as is into a standard ATG knowledge base, as a customer would.

  2. 2

    No file indicated

    Each question is asked on its own, in a new conversation. ATG has to find the right document among the 368 without being told. For 11 questions whose original wording names neither the company nor the period, a fixed sentence gives them ("Added note by ATG for disambiguation"), with the rest of the question unchanged.

  3. 3

    Grading

    An independent LLM judge (Mistral Medium 3.5) compares each answer with the reference answer. In each test, 4 verdicts were overturned by a human review, with its justification in the report. A question without an answer counts as an error.

  4. 4

    End-to-end latency

    Time measured from sending the question to the end of the answer, search included, with questions asked one at a time. The reports give the mean, the median and the p90.

Transparency

Each downloadable report details our methodology then, for each of the 150 questions, the expected answer, the answer ATG gave and the verdict. Anyone can check a grade and challenge it.

Results

How ATG performs, AI policy by AI policy

The configurations tested

One row per AI policy, both in Fast mode. Each row is a full run over the 150 questions.

The configurations tested
AI policyFast
Worldwide providers
Highlighted

92%

138 / 150 correct answers

Mean
11 s
EU-only providers

93%

139 / 150 correct answers

Mean
31 s

Accuracy and response time

The higher and further left a point sits, the more accurate and faster the answer.

60%70%80%90%100%050100150Mean latency (seconds)Mistral Agentic Search, Mistral Medium 3.5 *Mistral Agentic Search, Mistral Medium 3.5 *Mistral Agentic Search, GLM-5.2FastFast
ATG, Worldwide providersATG, EU-only providersMistral AI, Agentic Search

Comparison

Against Mistral Agentic Search

On 20 August 2026, Mistral published its FinanceBench results: the same 150 questions, the same 368 filings, with an agentic search that chains queries and navigates within documents. It is the published result closest to our conditions today.

Two differences with our protocol: Mistral grades answers with an LLM judge calibrated against human labels, and reports latency across the whole benchmark (mean and p90), not per model.

×6.8

In Fast mode, ATG answers in 11 s on average, against 71 s for the fastest configuration published by Mistral AI: answers 6.8 times faster.

Accuracy

  • ATG, EU-only providers, Fast mode

    93%
  • ATG, Worldwide providers, Fast mode

    92%
  • Mistral Agentic Search, GLM-5.2, search + navigation

    86%
  • Mistral Agentic Search, Mistral Medium 3.5, search + navigation *

    83%
  • Mistral Agentic Search, Mistral Medium 3.5, search only *

    74%
  • Mistral, Mistral Medium 3.5, one-shot RAG

    27%

Mean latency

  • ATG, EU-only providers, Fast mode

    31 s
  • ATG, Worldwide providers, Fast mode

    11 s
  • Mistral Agentic Search, GLM-5.2, search + navigation

    71 s
  • Mistral Agentic Search, Mistral Medium 3.5, search + navigation *

    71 s
  • Mistral Agentic Search, Mistral Medium 3.5, search only *

    108 s
  • Mistral, Mistral Medium 3.5, one-shot RAG

    not published

Mistral figures quoted from their post, not reproduced by us. * Computed from the deltas published by the source.

Example

Example: one question and ATG's answer

  1. Question (from the test set)

    What is the FY2018 capital expenditure amount (in USD millions) for 3M? Give a response to the question by relying on the details shown in the cash flow statement.

  2. Expected answer

    $1577.00

  3. ATG's answer (unrestricted policy, Fast mode)

    $1,577 million. Per 3M's Consolidated Statement of Cash Flows (fiscal year ended December 31, 2018), the line "Purchases of property, plant and equipment (PP&E)" under Cash Flows from Investing Activities is $(1,577) million, which is the capital expenditure figure.

  4. Document cited by ATG

    3M_2018_10K.pdf

Verdict: Correct, answered in 7 seconds

Download the results

One report per configuration, with the protocol, the results and, for every question, ATG's answer and its verdict. Free access, no form.

Download the results
AI policyFast
Worldwide providersPDF, Worldwide providers, Fast
EU-only providersPDF, EU-only providers, Fast

Frequently asked questions

Why FinanceBench rather than another benchmark?

Researchers and vendors publish their scores on FinanceBench. That is what lets us compare ourselves with results we did not produce, on real financial documents.

Why not run the other solutions yourselves?

To avoid comparing a solution we master with solutions we might configure poorly. We quote the figures each vendor or researcher published, with their test conditions, and group them by difficulty level.

What is the difference between worldwide and EU-only providers?

It is the AI policy chosen by the administrator. EU-only restricts processing to European providers (Mistral, Scaleway, Nebius); worldwide (the unrestricted policy) adds OpenAI and Google among others. We ran the test with both so everyone knows the cost of their choice, in accuracy and speed.

What does the Fast mode mean?

It is the ATG mode that favours response time. Both tests were run in this mode. The Auto and Thinking modes, which leave more time for reasoning, are not part of this benchmark.

Do these results apply to my documents?

They give an order of magnitude on long, table-heavy financial documents. The best test remains your own documents: we can measure ATG's accuracy on a sample of your questions during a demo.

What about your own filings?

Import your financial reports, ask your teams' questions and measure the accuracy of the answers yourself.