Benchmarks
Original tests of AI models. Each one runs every model once on a fixed question set, checks the answers against sourced keys, and publishes what the run really cost.
Multilingual LLM Bench
Which AI speaks your language best. 150 questions written from scratch in 30 languages, 17 models, run 2026-10-02. Leader: GPT-6 Astra at 70.7.