WeirdML
WeirdML is a coding benchmark tracked on BenchGecko. Unusual and adversarial machine learning challenges. Tests robustness of reasoning about edge cases in ML systems.
Built from BenchGecko data as of October 2, 2026 · updates when the data changes
WeirdML is a coding benchmark tracked on BenchGecko. Unusual and adversarial machine learning challenges. Tests robustness of reasoning about edge cases in ML systems.
Basic
Unusual and adversarial machine learning challenges. Tests robustness of reasoning about edge cases in ML systems.
Deep
Scores are reported (%, maximum 100); higher is better. BenchGecko collects WeirdML scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models.
Expert
Compare models only on the same benchmark version and settings. A benchmark separates models well while scores are spread out; when top models bunch together near the maximum, it stops telling them apart.
Depending on why you're here
- ·coding benchmark (%, maximum 100)
- ·Source: Epoch AI
- ·Useful if your workload is coding
- ·Check the live leaderboard on /benchmark/weirdml before choosing a model
- ·Labs cite benchmark results in launch announcements; check the source and version
- ·WeirdML is a test of how good an AI is at coding
- ·A higher score means better results on that kind of task