NVIDIA A40 dedicated server for large models and rendering
The NVIDIA A40 is an Ampere data center GPU with 48 GB of ECC memory. It handles both AI and graphics: large language models, image generation, 3D rendering with RT cores and virtual workstations. At IPHOST the A40 runs in a dedicated server, so the whole card and all server resources belong to you, with no virtualization layer.
NVIDIA A40 specs
| Architecture | Ampere |
|---|---|
| GPU memory | 48 GB GDDR6 with ECC |
| Memory bandwidth | 696 GB/s |
| CUDA cores | 10,752 |
| RT and Tensor cores | 2nd generation RT cores, 3rd generation Tensor cores |
| Power draw (TDP) | 300 W |
| Interface | PCIe Gen4 x16, dual slot, passive cooling |
Server configuration
The A40 plan runs on an HPE DL380 Gen10 with two Intel Xeon Gold 6230R processors (52 cores, 104 threads), 192 GB of ECC RAM and two 1.6 TB enterprise SAS-SSDs. It includes a 1 Gbps network port with 30 TB of monthly traffic, one IPv4 address, basic DDoS protection and iLO remote management.
What you can run on an NVIDIA A40
- 70B language models – Llama 3.3 70B or Qwen 72B quantized to 4-bit fit in 48 GB, and models up to about 20B run in FP16.
- Image and video generation – Stable Diffusion XL, Flux and ComfyUI workflows at high resolution.
- Fine-tuning – LoRA and QLoRA on 7–13B models.
- 3D rendering – Blender Cycles, V-Ray, Octane and Unreal Engine with hardware ray tracing.
- Virtual workstations – remote desktops for design, CAD and video editing.
NVIDIA A40 price and billing
The price in the table above is the monthly price with quarterly billing, 3 months in advance. If you pay for 12 months, the monthly price is up to 14% lower. GPU servers are delivered on pre-order within 14 days; if we miss the date, you get a full refund.
A40 or another GPU?
The A40 gives the most video memory for the price in our range. For smaller models, video transcoding and lower cost, see the NVIDIA L4 dedicated server. For training and fine-tuning larger models, see the 2x NVIDIA A100 dedicated server. All plans are listed on the GPU server hosting page.