Skip to main content
ultraData Center

A100 SXM 40GB GPU

The 40 GB version of the Ampere A100 in the SXM4 form factor, with HBM2 memory at 1,555 GB/s and NVLink for multi-GPU nodes. It is the A100 in AWS p4d, Azure ND96asr_v4 and Google a2-highgpu instances.

VRAM 40GB
CUDA Cores 6,912
Tensor Cores 432
TDP 400W
Process 7nm
From
$0.644/hr
across 8 providers
A100 SXM 40GB GPU

Cloud Pricing

Cheapest on Verda — 63% below avg
ProviderPrice / hr
$0.644/hr
1×8×
Runpod logo
RunpodCommunity Cloud
$1.00/hr
1×
$1.22/hr36mo
16×
$1.29/hr36mo
2×4×8×
$1.29/hr
8×
$1.29/hr
1×
Shadeform logo
ShadeformDenvr Dataworks
$1.40/hr
8×
Omega Gradient logo
Omega GradientGlobal · flexible
$1.47/hr
8×
$1.54/hr
1×
$1.99/hr
1×8×
$2.09/hr
1×
$2.09/hr
16×
$2.19/hr12mo
16×
$2.31/hr12mo
2×4×8×
Amazon AWS logo
Amazon AWSus-east-1
$2.74/hr
8×
$3.48/hr
16×

Prices updated daily. Last check: Sep 30, 2026

A100 SXM 40GB pricing by provider

Every configuration, price rank, and alternative for one provider at a time.

Performance

FP16
312 TFLOPS
FP32
19.5 TFLOPS
BF16
312 TFLOPS
INT8
624 TOPS
Bandwidth
1555 GB/s

Common Use Cases

Multi-GPU training and fine-tuning of models that fit in 40 GB per GPU, batch inference, HPC workloads using FP64, and MIG partitioning into up to seven instances

Full Specifications

Hardware

Manufacturer
NVIDIA
Architecture
Ampere
CUDA Cores
6,912
Tensor Cores
432
RT Cores
0
Process Node
7nm
TDP
400W

Memory & Performance

VRAM
40GB
Memory Interface
5120-bit
Memory Bandwidth
1555 GB/s
FP32
19.5 TFLOPS
FP16
312 TFLOPS
BF16
312 TFLOPS
FP64
9.7 TFLOPS
INT8
624 TOPS
Release
2020