ChipsReading · ~3 min · 55 words deep

Microsoft Maia 200

Maia 200 is Microsoft's successor to Maia 100, its custom AI accelerator for Azure. BenchGecko publishes its figures once they are in the hardware dataset.

Text reviewed October 5, 2026

Maia 200 on hardware map
TL;DR

Maia 200 is Microsoft's successor to Maia 100, its custom AI accelerator for Azure. BenchGecko publishes its figures once they are in the hardware dataset.

Level 1

Microsoft Maia is Azure's custom AI silicon. Maia 100 shipped late 2024 serving select Azure OpenAI workloads. Not publicly benchmarked but Azure claims per-token cost below H100 Bedrock equivalents.

Level 2

Maia architecture: custom cores designed with OpenAI feedback, focused on large context (100K+) serving and reasoning workloads. Maia 200 specs leaked: HBM3e, 5nm, rack-scale via custom Cobalt-like interconnect. Software stack is an evolution of Microsoft's internal AI compiler (likely Triton-based). Primarily deployed inside Azure OpenAI · customers see it via lower prices on specific endpoints.

Level 3

Microsoft's Maia + Cobalt strategy: Maia for AI, Cobalt for general-purpose ARM compute. Both reduce NVIDIA dependency for Azure. Maia 200 reportedly gets priority silicon at TSMC · evidence of Microsoft commitment. OpenAI influence on design means Maia is optimized for transformer inference patterns from GPT-series: large context + long generation + reasoning. Not a generalist AI chip like H200; a specialized LLM-inference chip.

The takeaway for you
If you are a
Curious · Normie
  • ·Microsoft's own AI chip for Azure
  • ·Makes OpenAI models cheaper to run on Azure
  • ·Not sold to customers directly
If you are a
Builder
  • ·Access via Azure OpenAI only
  • ·Lower cost on specific endpoints
  • ·Not directly programmable
If you are a
Investor
  • ·Microsoft's vertical AI play · reduces NVIDIA exposure
  • ·Strategic lock with OpenAI via hardware co-design
  • ·Azure margin protection
If you are a
Researcher
  • ·Microsoft's 2nd-gen AI ASIC · HBM3e · 5nm
  • ·Designed with OpenAI input · transformer-specialized
  • ·Rack-scale interconnect (Cobalt-adjacent)
Gecko's take

Maia 200 is Microsoft's hardware bet on OpenAI · architecture co-design with GPT series makes it hard for competitors to replicate.

Not yet fully GA. Select Azure OpenAI workloads use it; broader rollout through 2026.