QuickMark
QuickMark is the default choice in Create New Job and runs on NVIDIA and AMD. The portal summarizes it as FP32, FP16, and memory bandwidth. It reports per-GPU and combined numbers for FP32 and FP16 compute, memory bandwidth, and temperature, which you can compare with manufacturer specs. On a multi nodes job it also measures how fast the machines talk to each other. It is meant to give a short, standardized view of GPU and network performance. The test stresses compute and memory on each GPU, then on all GPUs together if there is more than one. After the speed pass it runs a heat-and-cool check. If that thermal check fails, the speed numbers still stand. QuickMark measures both vector cores and matrix cores (CUDA cores and Tensor cores on NVIDIA). FP32 uses the vector cores; FP16, BF16, and FP8 use the matrix cores when the GPU has them. Teraflops (TF) is processing speed, so higher means faster math. FP32 is full precision. FP16, BF16, and FP8 are the smaller formats used in much of AI. FP8 needs an NVIDIA Ada Lovelace or Hopper GPU (or newer); Ampere and earlier do not support it. L2 cache speed is how quickly data moves through the GPU’s on-chip cache. Memory bandwidth is how fast the GPU can feed its own memory; when it is low, the cores sit idle waiting for data. Host to device and device to host are the traffic between CPU memory and the GPU. With two or more GPUs on the same machine, QuickMark also measures how well the cards share data with each other, including GPU-to-GPU bandwidth and NCCL collective bandwidth for all-reduce and broadcast. Power, temperature, and energy used during the run tell you about cooling and electricity cost. The SiliconMark Score (SM Score) is a letter on the job for many enterprise GPUs. It blends memory bandwidth with a compute mix, and compares the result with other units of the same model in SiliconMark’s fleet. A is the top 75%, B the top 50%, C the top 25%, and D the bottom 25%. Consumer cards usually show N/A. A failed precision test shows ERR. QuickMark needs no extra API keys. Each machine’s QuickMark is independent, and the network test between machines is a separate measurement (see Cluster network below).Llama 3 inference
This test runs on NVIDIA and reports tokens per second and time to first token. It uses NVIDIA DGX recipes for Llama 3 inference and measures how quickly the system starts responding and how many tokens it generates per second, across different numbers of simultaneous users and prompt lengths. SiliconMark starts an NVIDIA NIM container serving a fixed Nemotron Llama3 model, then runs a load generator against it. Pulling the images and waiting for the server can take tens of minutes the first time. The settings you usually change are how many requests hit the GPU at once (concurrency) and how long the prompt and the reply are. Tokens per second is how much text the system produces under that load; higher means more capacity. Time to first token is how long a user waits before the answer starts, and lower is better for interactive chat. Time between tokens shows how smoothly the rest of the answer streams. Request latency is the time from send to fully done. The machine needsNGC_API_KEY. Podman is used if present, otherwise Docker.
Llama 3 fine-tuning
This test runs on NVIDIA and reports tokens per second and step time. It uses NVIDIA DGX recipes for Llama 3 fine-tuning and measures training speed, how consistent each training step is, and a rough time to train on a trillion tokens. Model size is configurable: 8B, 70B, or 405B. Under the hood it runs NVIDIA NeMo on one server. The defaults are an 8B model, FP8, LoRA, and a short run of 50 steps. You can choose 8B, 70B, or 405B, FP8 or BF16, and LoRA or a full fine-tune. SiliconMark decides how to split the work across the GPUs you have. In practice 8B fits on one GPU (about 16 GB for LoRA, 24 GB for a full fine-tune), 70B needs two or more GPUs, and 405B needs eight or more GPUs with FP8 and LoRA only. Tokens per second is training throughput, so higher means a dataset finishes sooner. Step time is seconds per training step, and lower and steadier is better. Time to 1T tokens is a rough estimate of how many days it would take to process a trillion tokens at this speed, useful for capacity planning. The machine needsHF_TOKEN and Docker, usually with sudo. An NGC key lets it pull the NeMo container. Disk use is about 75 GB for 8B, 205 GB for 70B, and 905 GB for 405B, counting the container and the weights.
Inference Benchmark
Inference Benchmark runs on NVIDIA and AMD and reports tokens per second and time to first token. It measures how well the system serves AI requests: how fast it responds, how many requests it handles at once, and how quickly it generates tokens, across different parallelism settings and prompt sizes. It works like the Llama 3 inference test but serves an open model with vLLM on the host instead of a NIM container. The default model is a large open checkpoint unless you change it. No NGC key is needed, and a Hugging Face token is only needed if the model you pick is gated. NVIDIA machines need CUDA 12.6 or newer. AMD machines need ROCm 7.0 or newer; the current agent works with ROCm 7.0 through 7.2. You can set concurrency, prompt length, reply length, how many GPUs to split the model across, and how many prompts to send. The result screens use a few labels. Input sequence length is the number of tokens in the prompt, and output sequence length is the number in the reply. Time to first token is the wait until the first word appears. Time per output token and inter-token latency are the gap between later tokens. End-to-end is the time from the start of the request to the last token. Use Inference Benchmark on mixed NVIDIA and AMD fleets, or when you do not want a NIM setup.SiliconAudit
SiliconAudit runs on NVIDIA only and covers memory health, DCGM diagnostics, and power telemetry. The machine needs cuda-memtest and NVIDIA DCGM (dcgmi); the agent sets those up when that benchmark is selected. Run the agent with sudo (raised system privileges); SiliconAudit needs that access.
This is a health check, not a speed test. It sweeps GPU memory for stuck or flaky cells, records page retirement (memory the driver has already mapped out), and runs NVIDIA DCGM diagnostics across software, hardware, and PCIe. It also reports NVIDIA XID messages from the kernel log, which are driver codes that flag past GPU, driver, or workload problems. Each part comes back as pass or fail with error counts, and the job produces a PDF.
Run it when a card looks slow, when you take delivery of a used or refurbished GPU, or as a periodic fleet health scan. A full scan can take hours. A failing audit is a reason to pull the card.