~/benchmarks

Benchmarks

corpora we build, maintain and re-run

These are evaluation suites we built and maintain, each with its own corpus, grading method and published metrics. They are distinct from our research studies, which measure models against existing or one-off datasets. Each page here explains what the benchmark tests and how to get the data.