Benchmarks
Our assistants' accuracy, measured on public tests
An AI assistant that answers fast but wrong costs more than it saves. We run ATG on public, sector-specific benchmarks, under real deployment conditions, and publish everything: protocol, results and ATG's answer to every question.
Published benchmarks
More sector benchmarks are in preparation.
Our rules for every benchmark
Real-world conditions
Every document goes into a single knowledge base, and ATG is not told which one holds the answer. This matches real-world situations.
Every configuration published
Every configuration tested, EU-only providers as well as worldwide providers, is published, including when it does worse.
Sourced comparisons
We do not re-run other solutions: we quote their published figures, with source, date and grading method, grouped by difficulty level.
Published verdicts
For every question, we publish ATG's answer and the verdict. Anyone can check a grading and challenge it.