Gecko Tests · Model Drift IndexLatest run 2026-10-05 · 9 models watched · 0 changed in their latest run
Model Drift Index · Did the Model Quietly Change?
Labs update models behind the same name. Every week we send the same 22 probes to the same flagship models and compare each run with the previous one: exact-answer accuracy, refusals and who the model says it is.
In plain wordsAI companies sometimes change a model without telling anyone. Every week we ask the same questions to the same models and compare. A pink card = the model changed since last week.
Llama 4 Scoutbaseline
| Week | Exact | Refusals | Names maker | Flips |
|---|---|---|---|---|
| 2026-10-05 | 20% | 0/8 | 3/4 | n/a |
Mistral Medium 3.5baseline
| Week | Exact | Refusals | Names maker | Flips |
|---|---|---|---|---|
| 2026-10-05 | 20% | 0/8 | 2/4 | n/a |
Kimi K3baseline
| Week | Exact | Refusals | Names maker | Flips |
|---|---|---|---|---|
| 2026-10-05 | 100% | 0/8 | 4/4 | n/a |
DeepSeek V4 Pro 0813baseline
| Week | Exact | Refusals | Names maker | Flips |
|---|---|---|---|---|
| 2026-10-05 | 100% | 0/7 | 2/4 | n/a |
Grok 4.7baseline
| Week | Exact | Refusals | Names maker | Flips |
|---|---|---|---|---|
| 2026-10-05 | 100% | 0/8 | 4/4 | n/a |
GLM 5.3 Primebaseline
| Week | Exact | Refusals | Names maker | Flips |
|---|---|---|---|---|
| 2026-10-05 | 80% | 0/8 | 4/4 | n/a |
Qwen3.8 Max Primebaseline
| Week | Exact | Refusals | Names maker | Flips |
|---|---|---|---|---|
| 2026-10-05 | 100% | 0/8 | 4/4 | n/a |
Claude Sonnet 5.5baseline
| Week | Exact | Refusals | Names maker | Flips |
|---|---|---|---|---|
| 2026-10-05 | 100% | 0/8 | 4/4 | n/a |
GPT-6.1 Sol Probaseline
| Week | Exact | Refusals | Names maker | Flips |
|---|---|---|---|---|
| 2026-10-05 | 100% | 0/8 | 4/4 | n/a |
How it works
- 22 fixed probes: 10 exact-answer tasks graded by code, 4 identity questions graded by lab name, and 8 legitimate questions (one per Censorship Index category) labelled by a judge model.
- Same prompts, temperature 0, lowest reasoning setting, every week, for the newest flagship model of each tracked lab. Once a model is watched, it stays watched for a year.
- Each run is compared with the model's previous run. Drift is flagged when exact accuracy moves 30 points or more, three or more refusal labels flip, or the model starts or stops naming another lab as its maker.
- Profile tests (Who Are You, World Map, Censorship Index) run once per model version; this index is the one that repeats.
Open dataset · v1
22 items, every model answer and every score are public under CC BY 4.0.
Download JSONHow to cite · data as of 2026-10-05
BenchGecko Gecko Tests, Model Drift Index v1. BenchGecko, data as of 2026-10-05. https://benchgecko.ai/gecko-tests/model-drift-index
BenchGecko measures this data itself: free to reuse under CC BY 4.0 with the credit "Source: BenchGecko" and a link. JSON · llms.txt · MCP
Frequently Asked Questions
Providers update weights, system prompts, safety filters and serving setups behind stable API names. Users only notice when answers change.