Abstract
The Xeon Max 9470C is a specialized CPU platform: its clearest wins come from HBM bandwidth and Intel-optimized AI kernels, while general integer, compile, video encoding, and several HPC workloads are more mixed. The strongest application-level results are PyTorch CPU inference, TensorFlow CPU inference, oneDNN BF16/INT8, STREAM bandwidth, and OpenVINO GenAI throughput. Results that depend on high clock speed, broad generic integer throughput, or modern dual-socket scale are less favorable.
Excellent / platform strengths
- STREAM / HBM memory bandwidth
- oneDNN BF16 and INT8 convolution kernels
- PyTorch 2.11 CPU inference
- TensorFlow 2.21 CPU inference
- OpenVINO GenAI sustained throughput
Good or mixed
- OpenVINO classic INT8 inference
- y-cruncher 1B
- FFTW 1D 4096
- spaCy transformer pipeline
- vLLM CPU latency
Ordinary or weak for class
- 7-Zip compression/decompression
- llama.cpp CPU BLAS
- GROMACS / LAMMPS / CP2K versus modern dual-socket servers
- QuantLib
- Linux kernel compile
- SVT-AV1 heavy 4K and 10-bit encoding
- spaCy en_core_web_lg
Not a platform-wide conclusion
- NumPy 1.2.1 because the PTS profile is forced single-threaded
STREAM 2013 Memory Bandwidth
Higher is better. Mixed platform comparison; labels show socket class and memory type/channel where the public source identifies it.
7-Zip CPU Throughput
Higher is better. MIPS comparison for compression and decompression; labels include core/thread counts where known.
oneDNN Convolution Kernels
Higher is better. GFLOPS comparison for oneDNN BF16 and INT8 convolution kernels; labels include CPU/socket context where known.
oneDNN FP16 / FP32 Limited Comparison
Higher is better. Separate horizontal view because public comparison coverage is sparse for these oneDNN modes.
OpenVINO 2026.0 CPU Inference
Higher FPS is better. Labels retain measured latency in ms where available; complex throughput comparisons use README/OpenBenchmarking values.
OpenVINO GenAI CPU Throughput
Higher is better. Tokens/sec for CPU OpenVINO GenAI profiles; use the model tabs to compare one profile at a time.
OpenVINO GenAI CPU Time To First Token
Lower original ms is better. Bars use first-token speed relative to milliseconds so taller bars mean lower TTFT; labels keep original ms.
llama.cpp b4154 CPU BLAS
Higher is better. granite-3.0-3b-a800m-instruct-Q8_0 prompt processing 2048 tokens.
vLLM CPU Hermes 3B
Latency and throughput are shown as separate horizontal views because they use different units. Latency is converted to requests/s equivalent.
Molecular Dynamics Throughput
Higher is better. GROMACS and LAMMPS are shown separately; each page now includes server, HEDT/workstation, and consumer/unified-memory comparison points where OpenBenchmarking has public data.
CP2K Molecular Dynamics
Original metric is seconds, lower is better. Bars show speed relative to 9470C; labels show original seconds plus relative speed.
QuantLib 1.39
Higher is better. QuantLib benchmark index, size S, tasks/s.
Compile / FFT / y-cruncher
Three separate horizontal views. Bars are normalized to 9470C so taller is better; labels keep the original seconds or Mflops.
NumPy 1.2.1 Single-threaded Score
Higher is better, but this PTS profile forces OMP_NUM_THREADS=1; use as single-thread context only.
PyTorch 2.11 CPU Inference
Higher is better. Three model families with batch-size tabs; values are batches/s from the PyTorch CPU benchmark.
TensorFlow 2.21 CPU Inference
Higher is better. Four TensorFlow CPU model families with batch-size tabs; values are images/s.
SVT-AV1 4.0 CPU Encoding
Higher is better. Source tabs and preset tabs cover all nine local SVT-AV1 subtests from README.
spaCy NLP Throughput
Higher is better. Tokens/sec for spaCy CPU NLP pipelines; en_core_web_lg is single-threaded statistical NLP, en_core_web_trf is transformer NLP.