Benchmark

AI & Retrieval
Benchmark
Also: benchmarks

A benchmark is a fixed collection of questions with a known way to score answers. Running two systems against the same benchmark turns 'this one feels better' into a number you can compare. It is how the retrieval results on our research pages are grounded.

Benchmarks have limits. A score on one dataset in one domain is strong evidence, not settled law, and a system tuned to ace a benchmark can still stumble on your real data. Read them as current evidence and confirm on your own material.

Where it comes up