Skip To Main Content

Intel Software Optimizations Boost AI Inference in MLPerf v6.1

MLPerf Inference v6.1 results showcase how software optimization can improve AI inference performance, alongside broader model coverage and expanded ecosystem participation

What’s New: In the latest MLPerf Inference v6.1 benchmark results from MLCommons, Intel highlighted AI inference performance gains achieved through software optimization on Intel Xeon 6 processors and Intel Arc Pro B-series GPUs. The results also reflect growing customer and partner participation, with more systems validating Intel platforms across more models, workloads and deployment scenarios.

 

On the same silicon and socket count used as MLPerf v6.0, Intel Xeon 6980P processors delivered a 2.4x increase in Llama 3.1 8B Server throughput and a 56% increase in Offline throughput¹. Those gains were achieved through software improvements alone, showing how continued optimization can deliver more performance from hardware customers have already deployed. Intel partners demonstrated similar momentum.

“MLPerf Inference v6.1 demonstrates how sustained software optimization can unlock significant new performance from hardware that customers already have deployed. Across Xeon and Arc Pro, we are continuing to enhance the full AI software stack while expanding the range of models, workloads, and systems our platforms can support.”

- Bill Pearson, Intel vice president, Data Center Software, Data Center Group

Why It Matters: As AI inference moves into production, customers need platforms that can support larger and more diverse models while realizing performance improvements, flexibility and value over time. Intel’s MLPerf Inference v6.1 results prove how software optimization can extend the useful performance of deployed infrastructure, helping customers get more from systems they already use while scaling across CPU and GPU compute.

 

Customer and partner participation on Intel platforms also increased from 29 results in v6.0 to 39 in v6.1, expanding third-party validation beyond Intel’s own submissions². Oracle contributed its first Intel-based submission; Red Hat delivered its first Xeon CPU inference submission, and Quanta Cloud Technology and Supermicro provided the first partner submissions using Intel Arc Pro B70.

 

About Intel Xeon 6 Results: Intel broadened its Xeon participation in MLPerf v6.1 from two benchmarked Xeon 6 SKUs in v6.0 to five, increasing the number of CPU inference results from 24 to 35. Intel Xeon remains the only standalone server CPU represented in MLPerf Inference submissions.

 

About Intel Arc Pro B70 Results: Intel Arc Pro B70 submissions expand on how software improvements can extend performance across a range of AI models. A single node configured with four Intel Arc Pro B70 GPUs provides 128GB of VRAM and supported submissions across Llama 3.1 8B, Llama 2 70B, gpt-oss-120B, Whisper and end-to-end retrieval-augmented generation (E2E-RAG). On the same four-GPU system used in v6.0, gpt-oss-120B Server performance improved 36%, while Offline performance increased 27%, reflecting continued maturity of Intel’s software stack³.

 

About the New E2E-RAG Benchmark: Intel also co-developed and submitted results for the new MLPerf Inference v6.1 end-to-end retrieval-augmented generation benchmark, one of the round’s most complex new workloads. In a single measured run on a system combining an Intel Xeon 6787P processor with four Intel Arc Pro B70 GPUs, the workload was split across the CPU and GPUs, with each handling different stages of the AI pipeline: Xeon handles embedding, reranking, vector search and Small Language Model (SLM), while Arc Pro B70 GPUs perform Large Language Model (LLM) generation. The result demonstrates how optimized CPU and GPU compute can work together across a complete AI workflow.

 

Intel’s ongoing optimization work extends beyond benchmark results. Xeon improvements are upstreamed into widely used AI frameworks so customers can benefit from software advances on deployed infrastructure, while optimization work on Intel Arc Pro B70 helps advance the kernels, frameworks, and serving stack for future Intel GPU products.

 

More Context: MLPerf Inference v6.1 results

Notices & Disclaimers

 

Performance varies by use, configuration and other factors. Learn more at www.Intel.com/PerformanceIndex.

 

Performance results are based on testing as of dates shown in configurations and may not reflect all publicly available updates. Visit MLCommons for more details. No product or component can be absolutely secure.

 

¹ Based on MLPerf Inference v6.1 (ID: mint-lark-1a28) and MLPerf Inference v6.0 (ID: 6.0-0057) results, submitted by Intel. Configuration: 1-node, 2x Intel® Xeon® 6980P processors (128 cores), 24x 96GB DDR5 MRDIMM 8800 MT/s (2304GB total), no discrete accelerator. Server scenario: 1,139.51 vs. 471.54 tokens/s (2.4x). Offline scenario: 1,916.74 vs. 1,229.56 tokens/s (1.56x). Results may not reflect all publicly available updates. MLPerf name and logo are trademarks of MLCommons. See mlcommons.org for more information.

² Based on MLPerf Inference v6.1 and v6.0 closed division results published at mlcommons.org. "Intel platforms" defined as submissions using Intel® Xeon® processors as primary compute (CPU-only inference) or Intel® Arc™ Pro GPUs as accelerator. Partner/customer count excludes Intel's own submissions. MLPerf name and logo are trademarks of MLCommons. See mlcommons.org for more information.

³ Based on MLPerf Inference v6.1 and MLPerf Inference v6.0 (ID: 6.0-0101) closed division results, submitted by Intel. Configuration: 1-node, 1x Intel® Xeon® 698X processor (86 cores), 4x Intel® Arc™ Pro B70 GPUs (32GB GDDR6 each, PCIe Gen5 x16), 8x 16GB DDR5 6400 MT/s (128GB total). gpt-oss-120B Server: 1,296.87 vs. 951.67 tokens/s (+36%). gpt-oss-120B Offline: 1,956.40 vs. 1,536.90 tokens/s (+27%). Software: PyTorch vLLM on Ubuntu 24.04. Results may not reflect all publicly available updates. MLPerf name and logo are trademarks of MLCommons. See mlcommons.org for more information.

 

Performance comparisons between MLPerf Inference v6.0 and v6.1 are based on comparable configurations using the same processor or GPU hardware and socket or GPU count, as applicable. See MLCommons results and Intel performance documentation for full system configurations and benchmark details.

 

Your costs and results may vary.

 

Intel technologies may require enabled hardware, software, or service activation.