BenchGecko Labs

Our Own AI Tests · Raw Answers Included

Designed and run by BenchGecko · data under CC BY 4.0

BenchGecko Labs runs Gecko Tests: automated tests we designed and run ourselves on the newest models of every tracked lab. New releases are usually tested within a day of appearing in the catalog. Scores, test items and raw answers are public.

Models tested

Tests with data

Latest run

Benchmarks measure what a model can solve. Gecko Tests measure how it behaves: whether it knows who made it, how it draws the world, what it refuses, where its knowledge stops, what other languages cost, and whether it changes without notice.

Profile tests run once per model version, as soon as a new model appears. Monitoring tests repeat on the same models: the Model Drift Index every week, Same Model Different Host every month. Every run is dated and every answer is stored.

Notable results are published as findings on the blog. Scores, test items and raw answers are free to reuse under CC BY 4.0: as JSON on each test page and as daily exports in the BenchGecko datasets repository on GitHub.

All findings on the blog

Gecko Tests results, test items and raw answers are free to reuse under CC BY 4.0. Download JSON from any test page or get the daily exports on GitHub.

BenchGecko/datasets on GitHub
BenchGecko Labs is where BenchGecko runs its own automated tests on AI models: identity (Who Are You), geography (World Map), refusals (Censorship Index), knowledge cutoff (Knowledge Horizon), language cost (Tokenizer Tax), silent changes (Model Drift Index) and provider quality (Same Model Different Host). The Gecko Scorecard sums them up.