Tested on 18 benchmarks · BenchGecko score 41.6. Top scores: ARC AI2 (77.1%), OpenBookQA (76.8%), MMLU (72.4%).
Capture-the-flag cybersecurity challenges. Tests vulnerability analysis, reverse engineering, cryptography, and exploitation skills.
HuggingFace MuSR (Multi-Step Reasoning). Tests multi-hop reasoning requiring chaining multiple facts together.
Competition-level math from AMC, AIME, and olympiad problems. Level 5 is the hardest tier, requiring creative problem-solving.
HuggingFace evaluation of MATH Level 5 problems. Competition math requiring advanced reasoning and proof construction.
Mock AIME (American Invitational Mathematics Exam) problems from OTIS. Tests mathematical competition performance.
- Typetext
- ContextN/A
- ReleasedJan 2024
- LicenseOpen Source
- Statusbenchmark-only
Frequently Asked Questions
Key facts · as of 2026-06-22
- Llama 3-70B by Meta. BenchGecko score 41.6, rank 192 of 312 scored models (normalized average of public benchmark scores).
- List price n/a input · n/a output per 1M tokens (as of 2026-06-22).
How to cite · data as of 2026-06-22
Llama 3-70B · benchmarks, pricing and providers. BenchGecko, data as of 2026-06-22. https://benchgecko.ai/model/llama-3-70b
Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP