Hardware
A list of the currently available hardware.
If you are looking for information on maximum resource requests, see Job Limits.
Compute Nodes¶
Your jobs will land on appropriately sized nodes automatically based on your CPU to memory ratio. For example in the Genoa partition:
- A job requesting ≤ 2 GB/core will run on a 2 GB/core node, or if full, a 4 GB/core node.
- A job requesting ≤ 4 GB/core will run on a 4 GB/core node, or if full, a 8 GB/core node.
And so on. You will always get the amount of memory you requested, even if running on a node with a higher ratio.
| Architecture | Cores | Memory | GPU | Nodes | |
| 2 x AMD Milan 7713 CPU └ 8 x Chiplets └ 8 x Cores |
128 | 512GB | (4GB / Core) | - | 55 |
| 1024GB | (8GB / Core) | - | 8 | ||
| 1 x AMD Milan 7713P CPU └ 8 x Chiplets └ 8 x Cores |
64 | 512GB | (8GB / Core) | 4 x NVIDIA HGX A100 | 4 |
| 2 x AMD Genoa 9634 CPU └ 12 x Chiplets └ 7 x Cores |
168 | 384GB | (2GB / Core) | - | 44 |
| 768GB | (4GB / Core) | 2 x NVIDIA RTX PRO 6000 | 4 | ||
| 1536GB | (8GB / Core) | - | 8 | ||
| 2 x NVIDIA H100 NVL | 4 | ||||
| 4 x NVIDIA L4 | 4 | ||||
| 2 x Intel Xeon Gold 6230 CPU (Cascade Lake) |
40 | 1.5TB | (38GB / Core) | - | 2 |
| 4 x Intel Xeon Gold 6238M CPU (Cascade Lake) |
88 | 6TB | (69GB / Core) | - | 1 |
Memory figures
Memory shown is the amount physically installed. A small amount is reserved for the operating system,
so the memory actually available to jobs is a few percent lower — for example a 512GB Milan node offers
480GB to Slurm. A job requesting exactly the full per-core ratio across every core of a node will
therefore not fit. Run sinfo -o '%n %m' for the exact schedulable figures.
hugemem
Jobs will not automatically land on the Intel 'hugemem' nodes. You must specifically request --partition hugemem.
The CPU architecture is different enough from the milan and genoa nodes, you will probably have to recompile your software.
GPUs¶
REANNZ HPC has a range of Graphical Processing Units (GPUs) to accelerate compute-intensive research and support more analysis at scale.
Depending on the type of GPU, you can access them in different ways, such as via batch scheduler (Slurm), or Virtual Machines (VMs).
For information about how to request these GPUs in a Slurm job, see Using GPUs.
| Architecture | Purpose/Note | VRAM | GPUs on Node | Nodes | |
| NVIDIA A100 SXM4 | 80GB | 4 | Milan | 4 | |
| NVIDIA RTX PRO 6000 | Avoid for double precision floating point (fp64) as it is slow | 96GB | 2 | Genoa | 4 |
| NVIDIA H100 NVL | 94GB | 2 | Genoa | 4 | |
| NVIDIA L4 | Avoid for double precision floating point (fp64) as it is slow | 24GB | 4 | Genoa | 4 |
Choosing a GPU¶
A rough guide to which GPU suits which kind of work. See Using GPUs for how to request a specific type.
| Workload | Best fit | Avoid |
| fp64 HPC (VASP, Quantum ESPRESSO, CP2K, OpenFOAM, Gaussian) | H100 NVL, then A100 | L4, RTX PRO 6000 |
| Molecular dynamics (GROMACS, AMBER, OpenMM) | RTX PRO 6000 ≳ H100 NVL > A100 | L4 |
| 4-GPU tightly-coupled | A100 (only option — see Milan) | everything else |
| 2-GPU communication-bound | H100 NVL (600GB/s NVLink between the pair) | RTX PRO 6000, L4 |
| Working set larger than 24GB | A100, H100 NVL, RTX PRO 6000 | L4 |
| Single-GPU fine-tuning, large-model inference | RTX PRO 6000 or H100 NVL | L4 |
| Small-model inference, video encoding, teaching | L4 (best performance per watt) | - |
Multi-node GPU jobs
GPUDirect RDMA is not currently enabled, and the GPU nodes are split across two InfiniBand switches. Keep GPU jobs within a single node.
GPU specifications¶
Peak dense throughput per GPU.
Dense figures, datasheets will show higher speeds
Datasheets often quote tensor figures twice as large as the ones below, marked with an asterisk for "with sparsity". Those speeds rely on structured sparsity: Tensor cores can skip zeros in a weight matrix, provided exactly two of every four consecutive values are zero. Half the multiplies, so twice the rate.
| NVIDIA A100 SXM4 | NVIDIA RTX PRO 6000 | NVIDIA H100 NVL | NVIDIA L4 | |
| Architecture | Ampere GA100 | Blackwell GB202 | Hopper GH100 | Ada AD104 |
| FP64 | 9.7 TFLOPS | ~1.9 TFLOPS | 30 TFLOPS | ~0.5 TFLOPS |
| FP64 tensor | 19.5 TFLOPS | - | 60 TFLOPS | - |
| FP32 | 19.5 TFLOPS | ~120 TFLOPS | 60 TFLOPS | 30.3 TFLOPS |
| TF32 tensor | 156 TFLOPS | ~126 TFLOPS | 418 TFLOPS | 60 TFLOPS |
| FP16 / BF16 tensor | 312 TFLOPS | ~250 TFLOPS | 836 TFLOPS | 121 TFLOPS |
| FP8 tensor | - | ~500 TFLOPS | 1,671 TFLOPS | 242 TFLOPS |
| FP4 tensor | - | ~2,000 TFLOPS | - | - |
| INT8 tensor | 624 TOPS | ~1,000 TOPS | 1,671 TOPS | 242 TOPS |
| VRAM | 80GB HBM2e | 96GB GDDR7 ECC | 94GB HBM3 | 24GB GDDR6 |
| Memory bandwidth | 2.04TB/s | ~1.6TB/s | 3.94TB/s | 300GB/s |
| TDP | 400W | 600W | 350-400W | 72W |
TF32 on the RTX PRO 6000
Unlike the A100 and H100, the RTX PRO 6000 gets no TF32 tensor speedup — TF32 runs at roughly FP32 rate. PyTorch and cuBLAS defaults that rely on TF32 see no speedup on these cards; the gain has to come from BF16, FP8 or FP4.
If you have any questions about hardware or the status of anything listed in the table, Contact our Support Team.