ChipsReading · ~3 min · 29 words deep

Inferentia 3

AWS Inferentia is AWS's line of custom chips for AI inference · this entry covers the next generation after Inferentia2.

Text reviewed October 5, 2026

Inferentia 3 on hardware map
TL;DR

AWS Inferentia is AWS's line of custom chips for AI inference · this entry covers the next generation after Inferentia2.

Level 1

AWS designs Inferentia chips to serve trained models at lower cost than general-purpose GPUs on its own cloud. The first Inferentia launched in 2019 and Inferentia2 followed. BenchGecko has no confirmed specifications for a third generation in its dataset; this entry will be updated from the dataset when they exist.

Level 2

Inferentia chips are used through AWS services such as EC2 Inf instances and Amazon Bedrock, with the AWS Neuron software stack. For training, AWS offers the separate Trainium line.

Level 3

Check the hardware pages for chips with confirmed specs. Until a chip is in the dataset, BenchGecko does not publish performance or price figures for it.

The takeaway for you
If you are a
Curious · Normie
  • ·AWS-made chips for running AI models
If you are a
Builder
  • ·Available only on AWS
If you are a
Investor
  • ·Custom silicon lowers AWS dependence on NVIDIA
If you are a
Researcher
  • ·Inference-focused AWS silicon with the Neuron SDK
BenchGecko has no confirmed specifications for a third-generation Inferentia in its dataset yet.