Same Prompts. Same Models. Raw Answers.
Designed and run by BenchGecko · data under CC BY 4.0
BenchGecko's own automated tests on the newest AI models: who the model says made it, how it draws the world, what it refuses, where its knowledge stops, what other languages cost, and whether it quietly changes.
57 models tested · 6 tests with data · latest run Oct 5, 2026
BenchGecko asks the questions people actually worry about: does the AI know what it is, what does it refuse, and does it change when nobody is watching?
Lab status
Models tested
57
Tests with data
6/7
Latest run
Oct 5, 2026
Raw answers
Public · CC BY 4.0
Gecko Scorecard
Which AI models pass our tests? Grades A to E and a top 5
All tests · 13 models · Oct 5, 2026
View testWho Are You
Which AI models say another lab made them?
Profile test · 57 models · Oct 5, 2026
View testWorld Map
How does each AI imagine the world map?
Profile test · 5 models · Oct 3, 2026
View testCensorship Index
Which AI refuses the most?
Profile test · 9 models · Oct 5, 2026
View testKnowledge Horizon
Where does each model's knowledge really stop?
Profile test · 7 models · Oct 4, 2026
View testTokenizer Tax
What the same text costs in 9 languages
Profile test · 54 models · Oct 5, 2026
View testModel Drift Index
Which models changed behavior the most this week?
Monitoring test · 9 models · Oct 5, 2026
View testSame Model, Different Host
Which providers serve a weaker version of the same model?
Monitoring test
View testPlanned tests(12)
Designed but not measured yet. BenchGecko publishes results only after a scored run.
AI Political Compass
Where does each AI model sit politically?
Race Bias Index
Does the model treat identical race-swapped scenarios differently?
Slur Double Standard Test
Does the model enforce hate-speech rules equally?
Would AI Let People Die?
Does the model choose rules or human survival?
AI IQ Test
Which AI model reasons best?
Gender Safety Bias Index
Does AI take men and women equally seriously when they are scared?
Real-Life AI Test
Does the model give useful advice in real situations?
Planet vs People Index
Does AI prioritize environmental goals over human welfare?
Religion Bias Index
Does AI protect some religions more than others?
Ideology Bias Index
Does AI apply the same standard to capitalism, communism, left, and right?
History Integrity Index
Does the model preserve historical facts under political pressure?
Creative Freedom Index
Does AI allow serious fiction, satire, and historical writing?
Methodology
Every run sends the same items to each model through OpenRouter, at temperature 0 and the lowest reasoning setting the model offers. BenchGecko records the model ID, the provider that served the reply, the date, the settings, the cost and the full reply.
Scoring is done by code wherever an exact answer exists (identity, geography, dates, token counts). Refusal labels in the Censorship Index and the Model Drift Index come from a judge model with a fixed rubric that sees only the question and the reply.
test version: recorded
model ID: recorded
provider route: recorded
temperature: 0
reasoning: lowest setting offered
tools / web access: off
raw answers: stored and public
cost per run: recorded
Profile tests run once per model version. Monitoring tests repeat: the Model Drift Index weekly, Same Model Different Host monthly. A daily budget cap applies, so a large backlog of new models can take a few days to clear. Changing items or grading always means a new test version.
View methodologyShare & Cite
Every test page has share buttons, a JSON download of its items, scores and raw answers, and a suggested citation. Daily exports are in the BenchGecko datasets repository on GitHub. License: CC BY 4.0, credit "Source: BenchGecko" with a link.
For journalists, researchers & creators
Use Gecko Tests results in articles, newsletters, videos and reports. Every result carries a run date, a test version and the raw answer behind it.