Compare · ModelsLive · 2 picked · head to head

Grok-2 (Dec 2024) vs Llama 3.3 70B Instruct (free)

Side by side · benchmarks, pricing, and signals you can act on.

Winner summary

Grok-2 (Dec 2024) wins 5 of 5 shared benchmarks. Leads in knowledge · math · reasoning.

Category leads
knowledge·Grok-2 (Dec 2024)math·Grok-2 (Dec 2024)reasoning·Grok-2 (Dec 2024)coding·Grok-2 (Dec 2024)
Hype vs Reality
Grok-2 (Dec 2024)
#230 by perf·no signal
QUIET
Llama 3.3 70B Instruct (free)
#256 by perf·#18 by attention
QUIET
Best value
Grok-2 (Dec 2024)
n/a
no price
Llama 3.3 70B Instruct (free)
n/a
$0.00/M
Vendor risk
xAI logo
xAI
$250.0B·Tier 1
Medium risk
Meta logo
Meta AI
$1.87T·Tier 1
Low risk
Head to head
Grok-2 (Dec 2024)Llama 3.3 70B Instruct (free)
GPQA diamond
Grok-2 (Dec 2024) leads by +8.5
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Grok-2 (Dec 2024)
38.4
Llama 3.3 70B Instruct (free)
29.9
MATH level 5
Grok-2 (Dec 2024) leads by +21.9
MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics.
Grok-2 (Dec 2024)
63.5
Llama 3.3 70B Instruct (free)
41.6
OTIS Mock AIME 2024-2025
Grok-2 (Dec 2024) leads by +6.4
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Grok-2 (Dec 2024)
11.4
Llama 3.3 70B Instruct (free)
5.0
SimpleBench
Grok-2 (Dec 2024) leads by +3.4
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking.
Grok-2 (Dec 2024)
7.2
Llama 3.3 70B Instruct (free)
3.9
WeirdML
Grok-2 (Dec 2024) leads by +7.8
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
Grok-2 (Dec 2024)
22.2
Llama 3.3 70B Instruct (free)
14.4
Full benchmark table
BenchmarkGrok-2 (Dec 2024)Llama 3.3 70B Instruct (free)
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
38.429.9
MATH level 5
MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics.
63.541.6
OTIS Mock AIME 2024-2025
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
11.45.0
SimpleBench
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking.
7.23.9
WeirdML
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
22.214.4
Pricing · per 1M tokens · projected $/mo at 10M tokens
ModelInputOutputContextProjected $/mo
xAI logoGrok-2 (Dec 2024)————
Meta logoLlama 3.3 70B Instruct (free)$0.00$0.00131K tokens (~66 books)—