Llama-3.1-Nemotron-Ultra-253B-v1 is a large language model (LLM) optimized for advanced reasoning, human-interactive chat, retrieval-augmented generation (RAG), and tool-calling tasks. Derived from Meta’s Llama-3.1-405B-Instruct, it has been significantly customized using Neural...
Tested on 4 benchmarks. Top scores: Chatbot Arena Elo — Overall (1346.8%), Artificial Analysis — Quality Index (15.0%), Artificial Analysis — Coding Index (13.1%).
Chatbot Arena overall Elo rating. Crowdsourced human preference ranking from blind head-to-head comparisons across all topics.
Artificial Analysis Quality Index. Composite quality score combining multiple benchmark results into a single metric.
Artificial Analysis Coding Index. Composite coding quality score from multiple code benchmarks.
Artificial Analysis Agentic Index. Composite score measuring agent capability across tool use and planning tasks.
- Typetext
- Context131K tokens (~66 books)
- ReleasedApr 2025
- LicenseOpen Source
- StatusActive
- Cost / Message~$0.003
Frequently Asked Questions
Related Models
Qwen3 Coder Next FP8 · AlibabaQwen2.5 Coder 32B Instruct AWQ · AlibabaSpeaker Diarization Community 1 · PyannoteGemma 3 12B (free) · Google DeepMindQwen3 Next 80B A3B Instruct (free) · Alibaba QwenKey facts · as of 2026-04-14
- Llama 3.1 Nemotron Ultra 253B v1 by NVIDIA. Not enough public benchmark scores to rank yet.
- List price $0.60 input · $1.80 output per 1M tokens (as of 2026-04-14).
How to cite · data as of 2026-04-14
Llama 3.1 Nemotron Ultra 253B v1 · benchmarks, pricing and providers. BenchGecko, data as of 2026-04-14. https://benchgecko.ai/model/llama-3-1-nemotron-ultra-253b-v1
Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP