Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 70B instruct-tuned version is optimized for high quality dialogue usecases. It has demonstrated strong...
Tested on 18 benchmarks · BenchGecko score 49.7. Top scores: Chatbot Arena Elo — Overall (1293.3%), IFEval (86.7%), MMLU (73.5%).
Qwen3 Next 80B A3B Instruct scores 49.2 (99% as good) at $0.09/1M input · 78% cheaper
Code editing benchmark from the Aider project. Measures ability to apply targeted code changes while maintaining correctness and style.
Unusual and adversarial machine learning challenges. Tests robustness of reasoning about edge cases in ML systems.
HuggingFace MuSR (Multi-Step Reasoning). Tests multi-hop reasoning requiring chaining multiple facts together.
HuggingFace evaluation of MATH Level 5 problems. Competition math requiring advanced reasoning and proof construction.
Competition-level math from AMC, AIME, and olympiad problems. Level 5 is the hardest tier, requiring creative problem-solving.
Mock AIME (American Invitational Mathematics Exam) problems from OTIS. Tests mathematical competition performance.
- Typetext
- Context131K tokens (~66 books)
- ReleasedJul 2024
- LicenseOpen Source
- StatusActive
- Cost / Message~$0.001
Frequently Asked Questions
Key facts · as of 2026-10-05
- Llama 3.1 70B Instruct by Meta. BenchGecko score 49.7, rank 147 of 312 scored models (normalized average of public benchmark scores).
- List price $0.40 input · $0.40 output per 1M tokens (as of 2026-10-05).
- Sold by 2 providers (as of 2026-10-05): DeepInfra (fp8) $0.40 in / $0.40 out · Amazon Bedrock $0.72 in / $0.72 out. Every provider
How to cite · data as of 2026-10-05
Llama 3.1 70B Instruct · benchmarks, pricing and providers. BenchGecko, data as of 2026-10-05. https://benchgecko.ai/model/llama-3-1-70b-instruct
Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP