Home/Models/GPT-5.3-Codex
OpenAI logo

GPT-5.3-Codex

by OpenAI · Released Feb 2026

Multimodal
68.1
avg score
Rank #46
Compare
Better than 85% of all models
Context
400K tokens (~200 books)
Input $/1M
$1.75
Output $/1M
$14.00
Type
multimodal
License
Proprietary
Benchmarks
18 tested
Data as of
About

GPT-5.3-Codex is OpenAI’s most advanced agentic coding model, combining the frontier software engineering performance of GPT-5.2-Codex with the broader reasoning and professional knowledge capabilities of GPT-5.2. It achieves state-of-the-art results...

Tested on 18 benchmarks · BenchGecko score 68.1. Top scores: Artificial Analysis · GPQA Diamond (91.5%), Artificial Analysis · tau2-Bench Telecom (86.0%), Artificial Analysis · Long Context Reasoning (83.3%).

Looking for similar performance at lower cost?
DeepSeek V4 Flash scores 68.6 (101% as good) at $0.03/1M input · 99% cheaper
Capabilities
coding
77.1
#8 globally
knowledge
17.8
#274 globally
agentic
32.1
#37 globally
speed
61.4
#6 globally
general
74.5
#3 globally
Benchmark Scores
Compare All
Tested on 18 benchmarks · Ranked across 5 categories
Score Distribution (all 312 models)
0255075100
▲ You are here
WeirdML

Unusual and adversarial machine learning challenges. Tests robustness of reasoning about edge cases in ML systems.

79.3·
Terminal Bench

Complex terminal-based engineering tasks. Models must use command-line tools, navigate filesystems, and debug systems through shell interaction.

77.3·
SWE-Bench verified

Real-world software engineering tasks from GitHub issues. Models must diagnose bugs and write patches that pass test suites. Human-verified subset of SWE-bench.

74.8·
PostTrainBench

Evaluates post-training behaviors including instruction following, safety, and helpfulness balance.

17.8·
SWE Atlas — Codebase QnA

SEAL SWE Atlas Codebase Q&A. Tests understanding of large codebases through question answering.

32.6·
APEX-Agents

Agent performance evaluation testing multi-step tool use, planning, and execution in realistic environments.

31.7·
Excellent (85+) Good (70-85) Average (50-70) Below (<50)
Links
Documentation
Community
BenchGecko API
gpt-5-3-codex
Specifications
  • Typemultimodal
  • Context400K tokens (~200 books)
  • ReleasedFeb 2026
  • LicenseProprietary
  • StatusActive
  • Cost / Message~$0.018
Available On
OpenAI logoOpenAI$1.75
Share & Export
Tweet
GPT-5.3-Codex is a proprietary multimodal AI model by OpenAI, released in February 2026. It has an average benchmark score of 68.1. Context window: 400K tokens.

Key facts · as of 2026-10-05

  • GPT-5.3-Codex by OpenAI. BenchGecko score 68.1, rank 45 of 312 scored models (normalized average of public benchmark scores).
  • List price $1.75 input · $14.00 output per 1M tokens (as of 2026-10-05).
  • Sold by 3 providers (as of 2026-10-05): Azure $1.75 in / $14.00 out · OpenAI $1.75 in / $14.00 out · OpenAI $3.50 in / $28.00 out. Every provider

How to cite · data as of 2026-10-05

GPT-5.3-Codex · benchmarks, pricing and providers. BenchGecko, data as of 2026-10-05. https://benchgecko.ai/model/gpt-5-3-codex

Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP