Provides a repository of community‑submitted benchmarks for local AI models.
Enables users to compare performance of their models against peer results.
Facilitates transparent evaluation and continuous improvement of local AI deployments.
526 community benchmarks across 37 models and 16 chips, from 9 contributors. Decode and prefill tok/s by model, quantisation and chip, measured with BaseRT or llama.cpp.