Benchmarks

Our assistants' accuracy, measured on public tests

An AI assistant that answers fast but wrong costs more than it saves. We run ATG on public, sector-specific benchmarks, under real deployment conditions, and publish everything: protocol, results and ATG's answer to every question.

By Jean-Christophe BudinUpdated on

Our rules for every benchmark

Real-world conditions

Every document goes into a single knowledge base, and ATG is not told which one holds the answer. This matches real-world situations.

Every configuration published

Every configuration tested, EU-only providers as well as worldwide providers, is published, including when it does worse.

Sourced comparisons

We do not re-run other solutions: we quote their published figures, with source, date and grading method, grouped by difficulty level.

Published verdicts

For every question, we publish ATG's answer and the verdict. Anyone can check a grading and challenge it.