Gecko Tests

Same Prompts. Same Models. Raw Answers.

Designed and run by BenchGecko · data under CC BY 4.0

BenchGecko's own automated tests on the newest AI models: who the model says made it, how it draws the world, what it refuses, where its knowledge stops, what other languages cost, and whether it quietly changes.

57 models tested · 6 tests with data · latest run Oct 5, 2026

BenchGecko asks the questions people actually worry about: does the AI know what it is, what does it refuse, and does it change when nobody is watching?

Models tested

57

Tests with data

6/7

Latest run

Oct 5, 2026

Raw answers

Public · CC BY 4.0

Designed but not measured yet. BenchGecko publishes results only after a scored run.

Gecko Worldview IndexPlanned

Where does each AI model sit politically?

Gecko Symmetry IndexPlanned

Does the model treat identical race-swapped scenarios differently?

Gecko Consistency IndexPlanned

Does the model enforce hate-speech rules equally?

Gecko Moral Tradeoff IndexPlanned

Does the model choose rules or human survival?

Gecko Reasoning BatteryPlanned

Which AI model reasons best?

Gecko Situation IndexPlanned

Does AI take men and women equally seriously when they are scared?

Gecko Situation IndexPlanned

Does the model give useful advice in real situations?

Gecko Environmental Values IndexPlanned

Does AI prioritize environmental goals over human welfare?

Gecko Symmetry IndexPlanned

Does AI protect some religions more than others?

Gecko Worldview IndexPlanned

Does AI apply the same standard to capitalism, communism, left, and right?

Gecko Factual Integrity IndexPlanned

Does the model preserve historical facts under political pressure?

Gecko Creative Boundary IndexPlanned

Does AI allow serious fiction, satire, and historical writing?

Every run sends the same items to each model through OpenRouter, at temperature 0 and the lowest reasoning setting the model offers. BenchGecko records the model ID, the provider that served the reply, the date, the settings, the cost and the full reply.

Scoring is done by code wherever an exact answer exists (identity, geography, dates, token counts). Refusal labels in the Censorship Index and the Model Drift Index come from a judge model with a fixed rubric that sees only the question and the reply.

test version: recorded

model ID: recorded

provider route: recorded

temperature: 0

reasoning: lowest setting offered

tools / web access: off

raw answers: stored and public

cost per run: recorded

Profile tests run once per model version. Monitoring tests repeat: the Model Drift Index weekly, Same Model Different Host monthly. A daily budget cap applies, so a large backlog of new models can take a few days to clear. Changing items or grading always means a new test version.

View methodology

Every test page has share buttons, a JSON download of its items, scores and raw answers, and a suggested citation. Daily exports are in the BenchGecko datasets repository on GitHub. License: CC BY 4.0, credit "Source: BenchGecko" with a link.

Use Gecko Tests results in articles, newsletters, videos and reports. Every result carries a run date, a test version and the raw answer behind it.

Gecko Tests are automated tests designed and run by BenchGecko on the newest models of every tracked lab. They measure behavior that benchmarks miss: self-identification, geography, refusals, knowledge cutoff, tokenizer cost, provider quality and silent changes over time.