Home/Models/Claude 3.5 Sonnet
Anthropic logo

Claude 3.5 Sonnet

by Anthropic · Released Jan 2024

39.8
avg score
Rank #199
Compare
Better than 36% of all models
Context
N/A
Input $/1M
n/a
Output $/1M
n/a
Type
text
License
Proprietary
Benchmarks
29 tested
Data as of
About

Tested on 29 benchmarks · BenchGecko score 39.8. Top scores: Chatbot Arena Elo — Overall (1374.0%), HELM — IFEval (85.6%), Aider — Code Editing (84.2%).

Capabilities
coding
39.5
#136 globally
reasoning
46.1
#86 globally
math
17.5
#239 globally
knowledge
47.8
#147 globally
agentic
24.0
#44 globally
general
43.2
#57 globally
multimodal
46.7
#8 globally
language
85.6
#35 globally
safety
13.0
#6 globally
Benchmark Scores
Compare All
Tested on 29 benchmarks · Ranked across 10 categories
Score Distribution (all 312 models)
0255075100
▲ You are here
Aider — Code Editing

Code editing benchmark from the Aider project. Measures ability to apply targeted code changes while maintaining correctness and style.

84.2·
Aider polyglot

Multi-language code editing from Aider. Tests editing ability across Python, JavaScript, TypeScript, Java, C++, Go, Rust, and more.

51.6·
CadEval

Computer-aided design evaluation. Tests understanding of CAD concepts, 3D modeling, and engineering design principles.

48.0·
HELM — WildBench

Stanford HELM WildBench evaluation. Tests reasoning on challenging real-world tasks.

79.2·
SimpleBench

Deceptively simple questions that humans find easy but AI models often get wrong. Tests common sense and reasoning gaps.

13.0·
MATH level 5

Competition-level math from AMC, AIME, and olympiad problems. Level 5 is the hardest tier, requiring creative problem-solving.

51.7·
HELM — Omni-MATH

Stanford HELM evaluation of mathematical reasoning across diverse problem types.

27.6·
OTIS Mock AIME 2024-2025

Mock AIME (American Invitational Mathematics Exam) problems from OTIS. Tests mathematical competition performance.

6.4·
Excellent (85+) Good (70-85) Average (50-70) Below (<50)
Links
Documentation
Community
BenchGecko API
claude-3-5-sonnet
Specifications
  • Typetext
  • ContextN/A
  • ReleasedJan 2024
  • LicenseProprietary
  • Statusbenchmark-only
Available On
Anthropic logoAnthropicn/a
Share & Export
Tweet
Claude 3.5 Sonnet is a proprietary text AI model by Anthropic, released in January 2024. It has an average benchmark score of 39.8.

Key facts · as of 2026-03-27

  • Claude 3.5 Sonnet by Anthropic. BenchGecko score 39.8, rank 198 of 312 scored models (normalized average of public benchmark scores).
  • List price n/a input · n/a output per 1M tokens (as of 2026-03-27).

How to cite · data as of 2026-03-27

Claude 3.5 Sonnet · benchmarks, pricing and providers. BenchGecko, data as of 2026-03-27. https://benchgecko.ai/model/claude-3-5-sonnet

Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP