Home/Models/Llama 3.1 405B
Meta logo

Llama 3.1 405B

by Meta · Released Jul 2024

Open Source
43.8
avg score
Rank #180
Compare
Better than 42% of all models
Context
N/A
Input $/1M
n/a
Output $/1M
n/a
Type
text-generation
License
Open Source
Benchmarks
22 tested
Data as of
About

Meta-llama text generation model. 383K downloads on HuggingFace.

Tested on 22 benchmarks · BenchGecko score 43.8. Top scores: ARC AI2 (93.7%), HellaSwag (85.6%), TriviaQA (82.7%).

Capabilities
coding
14.4
#182 globally
reasoning
29.0
#122 globally
math
19.8
#226 globally
knowledge
59.0
#64 globally
agentic
7.4
#59 globally
general
21.7
#158 globally
language
18.1
#169 globally
Benchmark Scores
Compare All
Tested on 22 benchmarks · Ranked across 7 categories
Score Distribution (all 312 models)
0255075100
▲ You are here
WeirdML

Unusual and adversarial machine learning challenges. Tests robustness of reasoning about edge cases in ML systems.

21.4·
Cybench

Capture-the-flag cybersecurity challenges. Tests vulnerability analysis, reverse engineering, cryptography, and exploitation skills.

7.5·
BBH

BIG-Bench Hard. 23 challenging tasks from BIG-Bench where prior language models fell below average human performance.

77.2·
SimpleBench

Deceptively simple questions that humans find easy but AI models often get wrong. Tests common sense and reasoning gaps.

7.6·
MUSR

HuggingFace MuSR (Multi-Step Reasoning). Tests multi-hop reasoning requiring chaining multiple facts together.

2.2·
MATH level 5

Competition-level math from AMC, AIME, and olympiad problems. Level 5 is the hardest tier, requiring creative problem-solving.

49.8·
OTIS Mock AIME 2024-2025

Mock AIME (American Invitational Mathematics Exam) problems from OTIS. Tests mathematical competition performance.

9.6·
MATH Level 5

HuggingFace evaluation of MATH Level 5 problems. Competition math requiring advanced reasoning and proof construction.

0.0·
Excellent (85+) Good (70-85) Average (50-70) Below (<50)
Links
Documentation
Community
BenchGecko API
meta-llama-llama-31-405b
Specifications
  • Typetext-generation
  • ContextN/A
  • ReleasedJul 2024
  • LicenseOpen Source
  • StatusActive
Available On
Meta logoMetan/a
Share & Export
Tweet
Llama 3.1 405B is an open-source text-generation AI model by Meta, released in July 2024. It has an average benchmark score of 43.8.

Key facts · as of 2026-04-09

  • Llama 3.1 405B by Meta. BenchGecko score 43.8, rank 180 of 312 scored models (normalized average of public benchmark scores).
  • List price n/a input · n/a output per 1M tokens (as of 2026-04-09).

How to cite · data as of 2026-04-09

Llama 3.1 405B · benchmarks, pricing and providers. BenchGecko, data as of 2026-04-09. https://benchgecko.ai/model/meta-llama-llama-31-405b

Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP