Our Own AI Tests · Raw Answers Included
Designed and run by BenchGecko · data under CC BY 4.0
BenchGecko Labs runs Gecko Tests: automated tests we designed and run ourselves on the newest models of every tracked lab. New releases are usually tested within a day of appearing in the catalog. Scores, test items and raw answers are public.
57
Models tested
6/7
Tests with data
Oct 5, 2026
Latest run
What is BenchGecko Labs?
Benchmarks measure what a model can solve. Gecko Tests measure how it behaves: whether it knows who made it, how it draws the world, what it refuses, where its knowledge stops, what other languages cost, and whether it changes without notice.
Profile tests run once per model version, as soon as a new model appears. Monitoring tests repeat on the same models: the Model Drift Index every week, Same Model Different Host every month. Every run is dated and every answer is stored.
Notable results are published as findings on the blog. Scores, test items and raw answers are free to reuse under CC BY 4.0: as JSON on each test page and as daily exports in the BenchGecko datasets repository on GitHub.
Current Gecko Tests
Gecko Scorecard
Which AI models pass our tests? Grades A to E and a top 5
View testWho Are You
Which AI models say another lab made them?
View testWorld Map
How does each AI imagine the world map?
View testCensorship Index
Which AI refuses the most?
View testKnowledge Horizon
Where does each model's knowledge really stop?
View testTokenizer Tax
What the same text costs in 9 languages
View testModel Drift Index
Which models changed behavior the most this week?
View testSame Model, Different Host
Which providers serve a weaker version of the same model?
View testLatest findings
Open data
Gecko Tests results, test items and raw answers are free to reuse under CC BY 4.0. Download JSON from any test page or get the daily exports on GitHub.
BenchGecko/datasets on GitHub