NVIDIA GPU benchmark suite — automate CUDA workloads, run HPL & HPCG, profile custom kernels, compare hardware, and generate rich reports.
Memory bandwidth, latency, matrix multiply, attention — run standard CUDA workloads with zero setup.
Self-contained Linpack and CG benchmarks via NVIDIA HPC Benchmarks binaries — downloads and runs automatically.
Run MLPerf inference benchmarks with configurable mode, batch size, and precision.
Generate interactive HTML reports with charts comparing GPUs, precisions, and problem sizes.
No system CUDA toolkit required — use cupy-cuda13x[ctk] to bundle CUDA libs with pip.
Submit, monitor, and collect results from HPC clusters — all from a single CLI.
Memory read/write, copy, stride, pointer-chase latency
CUDA CMatMul, tiled MatMul, fused attention, reduction
CuPyHigh Performance Linpack (FP64)
MPI Datacenter GPUsHigh Performance Conjugate Gradients
MPI Datacenter GPUsMLPerf Inference (test / find_performance / full)
TFLite