Tested on 29 benchmarks · BenchGecko score 39.8. Top scores: Chatbot Arena Elo — Overall (1374.0%), HELM — IFEval (85.6%), Aider — Code Editing (84.2%).
Code editing benchmark from the Aider project. Measures ability to apply targeted code changes while maintaining correctness and style.
Multi-language code editing from Aider. Tests editing ability across Python, JavaScript, TypeScript, Java, C++, Go, Rust, and more.
Computer-aided design evaluation. Tests understanding of CAD concepts, 3D modeling, and engineering design principles.
Stanford HELM WildBench evaluation. Tests reasoning on challenging real-world tasks.
Deceptively simple questions that humans find easy but AI models often get wrong. Tests common sense and reasoning gaps.
Competition-level math from AMC, AIME, and olympiad problems. Level 5 is the hardest tier, requiring creative problem-solving.
Stanford HELM evaluation of mathematical reasoning across diverse problem types.
Mock AIME (American Invitational Mathematics Exam) problems from OTIS. Tests mathematical competition performance.
- Typetext
- ContextN/A
- ReleasedJan 2024
- LicenseProprietary
- Statusbenchmark-only
Frequently Asked Questions
Key facts · as of 2026-03-27
- Claude 3.5 Sonnet by Anthropic. BenchGecko score 39.8, rank 198 of 312 scored models (normalized average of public benchmark scores).
- List price n/a input · n/a output per 1M tokens (as of 2026-03-27).
How to cite · data as of 2026-03-27
Claude 3.5 Sonnet · benchmarks, pricing and providers. BenchGecko, data as of 2026-03-27. https://benchgecko.ai/model/claude-3-5-sonnet
Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP