Inferentia 3
AWS Inferentia is AWS's line of custom chips for AI inference · this entry covers the next generation after Inferentia2.
Text reviewed October 5, 2026
AWS Inferentia is AWS's line of custom chips for AI inference · this entry covers the next generation after Inferentia2.
Basic
AWS designs Inferentia chips to serve trained models at lower cost than general-purpose GPUs on its own cloud. The first Inferentia launched in 2019 and Inferentia2 followed. BenchGecko has no confirmed specifications for a third generation in its dataset; this entry will be updated from the dataset when they exist.
Deep
Inferentia chips are used through AWS services such as EC2 Inf instances and Amazon Bedrock, with the AWS Neuron software stack. For training, AWS offers the separate Trainium line.
Expert
Check the hardware pages for chips with confirmed specs. Until a chip is in the dataset, BenchGecko does not publish performance or price figures for it.
Depending on why you're here
- ·AWS-made chips for running AI models
- ·Available only on AWS
- ·Custom silicon lowers AWS dependence on NVIDIA
- ·Inference-focused AWS silicon with the Neuron SDK