BenchmarksReading · ~3 min · 31 words deep

SWE-Bench verified

SWE-Bench verified is a coding benchmark tracked on BenchGecko. Real-world software engineering tasks from GitHub issues. Models must diagnose bugs and write patches that pass test suites. Human-verified subset of SWE-bench.

Built from BenchGecko data as of October 2, 2026 · updates when the data changes

TL;DR

SWE-Bench verified is a coding benchmark tracked on BenchGecko. Real-world software engineering tasks from GitHub issues. Models must diagnose bugs and write patches that pass test suites. Human-verified subset of SWE-bench.

Live data · updated daily
See full leaderboard
Level 1

Real-world software engineering tasks from GitHub issues. Models must diagnose bugs and write patches that pass test suites. Human-verified subset of SWE-bench.

Level 2

Scores are reported (%, maximum 100); higher is better. BenchGecko collects SWE-Bench verified scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models.

Level 3

Compare models only on the same benchmark version and settings. A benchmark separates models well while scores are spread out; when top models bunch together near the maximum, it stops telling them apart.

The takeaway for you
If you are a
Researcher
  • ·coding benchmark (%, maximum 100)
  • ·Source: Epoch AI
If you are a
Builder
  • ·Useful if your workload is coding
  • ·Check the live leaderboard on /benchmark/swe-bench-verified before choosing a model
If you are a
Investor
  • ·Labs cite benchmark results in launch announcements; check the source and version
If you are a
Curious · Normie
  • ·SWE-Bench verified is a test of how good an AI is at coding
  • ·A higher score means better results on that kind of task
Real-world software engineering tasks from GitHub issues. Models must diagnose bugs and write patches that pass test suites. Human-verified subset of SWE-bench.