141GB HBM3e memory and 4.8TB/s bandwidth on a Hopper-generation PCIe accelerator for AI and HPC workloads in qualified server systems. Passive cooling requires server-provided airflow.
Pre-order. UK delivery: £19.95. Estimated supplier availability: 1 January 2027, subject to NVIDIA supply; not a dispatch or delivery date. Delivery by 31 January 2027 unless you expressly agree an extension; otherwise cancel and refund in full. You may cancel before dispatch for a full refund.
Warranty
24-month Percepta return-to-base hardware warranty from delivery for items sold as New. Subject to the Hardware Warranty Policy. Your statutory rights are unaffected.
Pre-order. Full payment is taken at checkout. You can cancel before dispatch for a full refund.
The NVIDIA H200 NVL pairs 141GB of HBM3e with 4.8TB/s memory bandwidth for LLM inference, model training and memory-intensive research. Its Hopper architecture combines Tensor Core acceleration for AI with FP64 compute for scientific applications, in a dual-slot PCIe server card.
This new H200 NVL PCIe accelerator (900-21010-0040-000) requires a compatible server with suitable power and forced airflow.
141GBHBM3e memory per GPU
4.8TB/sGPU memory bandwidth
PCIe 5.0x16 host interface
Up to 600WConfigurable GPU TDP
AI inference, training and scientific computing
LLM inference and model training
High-capacity HBM3e provides space for model weights and runtime data, while Hopper Tensor Cores support mixed-precision inference and training. Inference memory requirements include context and concurrent requests, not just the model itself. Supported NVLink configurations allow applications designed for multiple GPUs to scale across cards.
Simulation and numerical research
FP64 and FP64 Tensor Core capabilities serve GPU-accelerated numerical applications that need double precision. The 4.8TB/s memory bandwidth is particularly relevant to simulations and scientific analysis that repeatedly move large working datasets through GPU memory.
Separate workloads on one GPU
MIG can partition the H200 NVL into up to seven isolated instances, letting smaller inference services or research jobs share the accelerator. Available profiles depend on the driver, software stack and host platform.
Technical specification
Specifications per GPU unless stated
01 / Architecture and memory
GPU
NVIDIA H200 NVL
Architecture
NVIDIA Hopper
Memory
141GB HBM3e
Memory bandwidth
4.8TB/s
Multi-Instance GPU
Up to 7 instances; supported profiles depend on the software and platform
Decode engines
7 NVDEC and 7 JPEG engines
02 / Peak compute performance
FP64
30 TFLOPS
FP64 Tensor Core
60 TFLOPS
FP32
60 TFLOPS
TF32 Tensor Core
835 TFLOPS with sparsity
BF16 / FP16 Tensor Core
1,671 TFLOPS with sparsity
FP8 Tensor Core
3,341 TFLOPS with sparsity
INT8 Tensor Core
3,341 TOPS with sparsity
03 / Connectivity and multi-GPU support
Host interface
PCIe 5.0 x16
PCIe bandwidth
128GB/s total bidirectional
GPU interconnect
NVIDIA NVLink, up to 900GB/s per GPU total bidirectional
NVLink configurations
Supported 2- or 4-GPU bridge configurations; compatible server layout and bridge hardware required
04 / Power, cooling and physical format
Maximum TDP
Up to 600W, configurable within supported platform limits
4.4 inches high x 10.5 inches long; dual-slot width
Compute figures are peak values; results depend on the application and server configuration. Sparse performance requires supported sparse operations. Bidirectional bandwidth is the total across both directions, and TDP applies to the GPU alone. Allow additional installation space for cabling and airflow. NVIDIA lists the H200 specifications as preliminary and subject to change; confirm the board specifications before purchase.
GPU supply
H200 NVL accelerator
The listing is for the H200 NVL accelerator. A server and additional GPUs are not included.
The card is new and includes a 24-month Percepta return-to-base hardware warranty from delivery. A server, NVLink bridges, additional power cables and installation services are not included. NVIDIA AI Enterprise entitlement is subject to NVIDIA activation and system requirements; contact us about licensing for your deployment.
Plan your installation
What your server needs.
Explicit H200 NVL support from the system manufacturer, with appropriate firmware and a supported PCIe slot or riser.
Chassis airflow and power delivery rated for the intended GPU configuration, including approved auxiliary power cabling.
Mechanical clearance for the card, cabling and any NVLink bridges.
A supported operating system, NVIDIA driver and application stack; additional software licences where required.
Will H200 NVL work in my server?
Check explicit H200 NVL support with the server manufacturer, including firmware, PCIe slot or riser, auxiliary power cabling, cooling and mechanical clearance. The passive heatsink requires forced chassis airflow, and GPU TDP is configurable up to 600W within supported platform limits.
H200 NVL is a PCIe card and is not interchangeable with an H200 SXM module. Send us the server make and model, riser and power details, and intended software environment to discuss the installation.
How do I check whether an AI model will fit?
Memory use depends on the model, precision, context length, batch size, concurrency and serving software. Include working buffers and runtime overhead when assessing the 141GB capacity. Send us these workload details to discuss a suitable configuration.
Can the GPU be divided between workloads?
MIG supports up to seven GPU instances. The available profiles and setup requirements depend on the driver, software release and host platform. Check the supported profile table for your intended configuration.
Can I connect several H200 NVL cards?
Suitable systems support two- or four-GPU NVLink bridge configurations. Check card spacing, bridge hardware, power, cooling and software support with the system manufacturer.
Each GPU has its own memory; applications must support distributing work across cards to use them together. Confirm any additional cards and bridges separately from this listing.
Is NVIDIA AI Enterprise included?
NVIDIA advertises a five-year AI Enterprise subscription for eligible H200 NVL products. Contact us to confirm entitlement, activation eligibility and system requirements for the card being supplied. The software subscription is separate from the hardware warranty.
What comes with the accelerator?
The listing covers one new H200 NVL card, part number 900-21010-0040-000, rather than a complete server or multi-GPU package. It includes a 24-month Percepta return-to-base hardware warranty from delivery.
Discuss any NVLink bridges, auxiliary power cabling or installation service required for your server so these can be agreed as part of the supply.
When will my pre-order arrive, and can I cancel?
The supplier currently estimates availability on 1 January 2027, subject to NVIDIA supply. This is not a dispatch or delivery date. The agreed delivery deadline is 31 January 2027 unless you expressly agree an extension; otherwise we will cancel and refund in full.
Full payment is taken at checkout. Both business and retail customers can cancel before dispatch for a full refund. Read the pre-order terms for delivery, cancellation and refund details.