By: Nithya Rao, System and Software Optimization Engineer, Intel Corporation
Contributors
Saumya Buliya, System and Software Optimization Engineer, Intel Corporation
For years, AI infrastructure discussions have focused on model quality, accelerator performance, and inference latency. That approach made sense when most AI applications consisted of relatively simple interactions: a prompt was submitted, a model generated a response, and the transaction ended.
The rise of agentic AI is changing that paradigm. Modern AI systems are increasingly expected to reason, plan, retrieve information, invoke tools, execute actions, validate results, and iteratively refine their behavior. What appears to be a single user interaction may trigger dozens of independent tasks spanning multiple software components before a result is delivered.
Rather than focusing exclusively on latency, organizations are increasingly asking a more practical question: How much useful work can the system complete over time? In many agentic deployments, throughput is becoming the metric that ultimately determines productivity.
Agentic AI Creates New Infrastructure Challenges
Traditional AI applications typically follow a straightforward execution model: prompt to model to response. Agentic AI introduces planning, tool selection, execution, validation, and iteration.
As a result, modern agentic AI workloads generate significantly more activity across the system stack than traditional inference workloads. Organizations deploying coding agents, enterprise copilots, research assistants, operational automation systems, and AI-powered business workflows often discover that the limiting factor is not model generation speed. Instead, bottlenecks emerge across task orchestration, tool execution, memory access, storage operations, and runtime management.
Why Intel® Xeon® CPUs Are Well-Suited for Agentic AI
A common misconception is that agentic AI performance can be measured solely by model inference speed. While strong CPU performance remains foundational for reasoning, tool execution, orchestration, and security operations, agentic workloads also place significant demands on memory, storage, and I/O resources.
Unlike traditional AI applications, agents maintain conversation history, execution context, retrieved knowledge, intermediate outputs, and workflow state across multiple stages of execution. They continuously invoke tools, access data sources, launch subprocesses, and coordinate dependent tasks. As a result, overall performance depends not only on compute capability, but also on how efficiently the platform moves work through each stage of the workflow.
Enterprise-scale agent deployments are characterized by high concurrency, significant memory utilization, intensive orchestration requirements, and continuous execution of heterogeneous tasks. Intel Xeon processors with built-in acceleration combine strong compute performance with large memory capacity, substantial memory bandwidth, and robust I/O capabilities, enabling efficient execution across diverse workload requirements.
As organizations scale from individual agents to thousands of concurrent workflows, platform efficiency increasingly determines overall productivity. The platforms that deliver the highest throughput are those that balance compute, memory, storage, and orchestration resources to minimize bottlenecks and sustain work under continuous load. This makes Intel Xeon CPUs well suited to the demands of modern agentic AI deployments.
Building a Repeatable Agentic AI Benchmark
One of the challenges associated with benchmarking agentic AI is that real-world agent behavior is inherently dynamic. Different prompts can produce different execution paths and different tool selections.
To address this challenge, a record-and-replay methodology was adopted. During the recording phase, representative agent workflows were executed while capturing workflow composition and execution characteristics. During replay, each platform processes the exact same workload trace, ensuring observed differences are attributable to infrastructure behavior rather than application-level randomness.
Evaluating Agentic AI Using a Real-World Workflow
To understand how platform capabilities translate into real-world agent performance, we evaluated the same deterministic Terminal-Bench workload across Intel Xeon 6 processor-based, 5th Gen AMD EPYC processor-based, and Arm v9.2A-based instances. The benchmark consisted of a representative set of agent tasks executed at 24-way concurrency over a 60-minute evaluation window, with identical code, inputs, and agent trajectories across all platforms
To assess performance across different real-world deployment scenarios, we evaluated three representative industry verticals: healthcare, banking and financial services industry (FSI) and manufacturing. While the specific application workflows differ, these workflows reflect common agentic AI characteristics, including reasoning, retrieval, analytics, document processing, and orchestration.
Agentic AI for Healthcare
The healthcare benchmark includes representative agent workflows spanning clinical documentation, evidence retrieval, diagnostic analysis, fraud detection, risk assessment, molecular research, genomic analysis, and laboratory diagnostics.
| Representative Task | Description | Category |
|---|---|---|
| Clinical transcription | Converts patient-provider interactions into structured clinical documentation. | Analytical and reasoning |
| Evidence retrieval | Retrieves relevant clinical information to support diagnosis and care decisions. | Analytical and reasoning |
| Radiology triage | Triages imaging studies and ranks cases by abnormality severity. | Analytical and reasoning |
| Radiology analysis | Generates, reviews, and validates radiology reports and diagnostic findings. | Analytical and reasoning |
| Claims fraud detection | Identifies anomalous claims and potentially fraudulent activity. | Analytical and reasoning |
| Readmission risk analysis | Analyzes patient data to predict the likelihood of hospital readmission. | Analytical and reasoning |
| Demand forecasting | Forecasts service utilization and demand and flags potential supply shortages. | Analytical and reasoning |
| Molecular candidate prioritization | Evaluates and ranks potential therapeutic or drug candidates. | Analytical and reasoning |
| Genomic variant interpretation | Interprets genetic variants and correlates findings with clinical evidence. | Analytical and reasoning |
| Laboratory diagnostics | Investigates laboratory results and identifies potential causes of diagnostic anomalies. | Analytical and reasoning |
| Audit logging | Compresses and encrypts clinical audit logs for secure, tamper-evident retention. | Secure data processing |
| Secure data exchange | Encrypts, verifies, and transfers clinical data packages between systems. | Secure data processing |
| Backup and restore validation | Verifies the integrity and recoverability of clinical data backups. | Control |
Figure 1. Representative of healthcare agent tasks
As shown in Figure 2, the Intel Xeon 6 processor-based instance achieved the highest normalized agent-task throughput, delivering 1.66 times the throughput of the 5th Gen AMD EPYC processor-based instance and 4.5 times the throughput of the Arm v9.2A-based instance during the 60-minute evaluation. Higher throughput enables organizations to support more concurrent agent workflows on the same infrastructure, improving overall productivity and accelerating completion of large-scale agent-driven workloads.
Figure 2. The Intel Xeon 6 processor-based instance delivers up to 1.66 times higher agent throughput than the 5th Gen AMD EPYC processor-based instance and delivers up to 4.5 times higher agent throughput than the Arm v9.2A-based instance on healthcare tasks.
Intel Xeon 6 CPU’s throughput advantage is not driven by a single workload type. Instead, it delivers consistently higher task completion rates across reasoning, security, and control operations, enabling faster end-to-end execution of real-world healthcare agent workflows.
As shown in Figure 3, the Intel Xeon 6 processor-based instance delivered the highest normalized throughput across all workload categories evaluated. Relative to the 5th Gen AMD EPYC processor-based instance, Intel Xeon 6 achieved 1.67 times higher throughput for analytical and reasoning workloads, 1.63 times higher throughput for secure data processing workloads, and 1.61 times higher throughput for control workloads. Relative to the Arm v9.2A based instance, Intel Xeon 6 achieved 4.67 times, 4 times, and 4.14 times higher throughput, respectively. This consistent performance advantage enables organizations to support more concurrent AI agent workflows, improve infrastructure utilization, and complete agent-driven work more quickly.
Figure 3. In the same 60-minute run, Intel Xeon 6 processor-based instance completed more healthcare tasks in reasoning, security, and control operations than 5th Gen AMD EPYC processor-based instance and Arm v9.2A-based instance.
Agentic AI for Banking and Financial Services
The banking and financial services benchmark includes representative agent workflows spanning lending, risk analysis, investment research, trading analytics, fraud detection, compliance, and regulatory reporting.
| Representative Task | Description | Category |
|---|---|---|
| Commercial loan underwriting | Evaluates applicants' credit and financial data to support lending decisions. | Analytical and reasoning |
| Behavioral risk scoring | Assesses customer behavior and transaction patterns to identify financial risk. | Analytical and reasoning |
| Investment recommendation ranking | Generates and prioritizes personalized investment insights and recommendations. | Analytical and reasoning |
| Liquidity stress testing | Evaluates portfolios and balance sheets under simulated market conditions. | Analytical and reasoning |
| Covenant extraction | Extracts key terms, obligations, and conditions from financial documents. | Analytical and reasoning |
| Trading analytics | Analyzes trading activity, execution quality, and performance metrics. | Analytical and reasoning |
| KYC verification | Processes and validates customer identity and onboarding documentation. | Analytical and reasoning |
| AML investigation | Identifies potentially suspicious transactions and money-laundering activity for review. | Analytical and reasoning |
| General ledger reconciliation | Matches and reconciles financial records and end-of-day ledger entries. | Analytical and reasoning |
| Regulatory reporting | Supports compliance, audit, and regulatory reporting workflows. | Secure data processing |
| Trade message attestation | Parses, validates, and cryptographically attests trade settlement messages. | Secure data processing |
| Portfolio optimization | Evaluates asset allocation strategies and investment trade-offs to improve portfolio performance. | Control |
Figure 4. Representative banking and financial services agent tasks
As shown in Figure 5, the Intel Xeon 6 processor-based instance delivered the highest normalized agent-task throughput, achieving 1.64 times the throughput of the 5th Gen AMD EPYC processor-based instance and 4.57 times the throughput of the Arm v9.2A-based instance. This higher throughput enables financial institutions to support more concurrent agent-driven workloads, including risk analysis, document processing, compliance operations, and customer-facing services, within the same infrastructure footprint.
Figure 5. Intel Xeon 6 processor-based instance delivers up to 1.64 times higher agent throughput than 5th Gen AMD EPYC processor-based instance and up to 4.57 times higher agent throughput than Arm v9.2A-based instance on banking tasks.
As shown in Figure 6, the Intel Xeon 6 processor-based instance delivered the highest normalized throughput across all workload categories evaluated. Relative to the 5th Gen AMD EPYC processor-based instance, Intel Xeon 6 achieved 1.66 times higher throughput for analytical and reasoning workloads, 1.61 times higher throughput for secure data processing workloads, and 1.59 times higher throughput for control workloads. Relative to the Arm v9.2A-based instance, Intel Xeon 6 achieved 4.84 times, 3.95 times, and 3.91 times higher throughput, respectively. This consistent performance advantage helps prevent bottlenecks across key stages of AI agent workflows, enabling financial institutions to process more agent-driven work within a given infrastructure footprint.
Figure 6. In the same 60-minute run, an Intel Xeon 6 processor-based instance completed more banking/FSI tasks in reasoning, security and control operations than a 5th Gen AMD EPYC processor-based instance and an Arm v9.2A-based instance.
Agentic AI for Manufacturing
The manufacturing benchmark includes representative agent workflows spanning predictive maintenance, quality inspection, supply chain optimization, production planning, factory operations, industrial security, and process optimization.
| Representative Task | Description | Category |
|---|---|---|
| Predictive Maintenance Anomaly Detection | Analyzes vibration and acoustic sensor data to identify equipment anomalies and potential failures. | Analytical and reasoning |
| Semiconductor Yield Prediction | Evaluates manufacturing process data to predict wafer and production yield outcomes. | Analytical and reasoning |
| Demand Forecasting | Forecasts future product demand across multiple supply-chain tiers and inventory locations. | Analytical and reasoning |
| Energy Load Forecasting | Predicts plant energy consumption to improve operational efficiency and capacity planning. | Analytical and reasoning |
| Visual Defect Detection | Identifies surface defects and quality issues from manufacturing inspection images. | Analytical and reasoning |
| Inspection Report Analysis | Processes inspection data and findings to generate actionable quality insights. | Analytical and reasoning |
| Failure Signature Similarity Search | Compares historical equipment failure patterns to accelerate troubleshooting and root-cause determination. | Analytical and reasoning |
| Supplier Compliance Verification | Validates supplier BOMs, component requirements, and compliance documentation. | Analytical and reasoning |
| Supplier Risk Assessment | Evaluates supplier performance, reliability, and operational risk indicators. | Analytical and reasoning |
| Scrap and Yield Root-Cause Analysis | Investigates production losses and identifies factors impacting yield and manufacturing efficiency. | Analytical and reasoning |
| Industrial Telemetry Protection | Processes and secures OT/IoT telemetry data transmitted across factory environments. | Secure data processing |
| Firmware Distribution Validation | Validates and cryptographically signs firmware updates for secure deployment across industrial devices. | Secure data processing |
| Quality Archive Compression | Compresses and manages high-volume inspection and quality-control data. | Secure data processing |
| Spare Parts Cross-Reference | Matches parts, components, and maintenance records across manufacturing systems. | Secure data processing |
| Process Historian Archiving | Compresses and retains long-term process-historian and sensor time-series data. | Secure data processing |
| Secure Site-to-Site Transfer | Encrypts and compresses telemetry and production data transferred between facilities. | Secure data processing |
| Digital Twin Simulation Calibration | Calibrates engineering and simulation models to align with physical production systems. | Control |
| Robotic Motion Planning | Optimizes robotic movement paths for manufacturing operations. | Control |
| CNC Toolpath Optimization | Improves machining efficiency through optimized toolpath generation. | Control |
| Production Changeover Scheduling | Evaluates production schedules to minimize downtime and improve line utilization. | Control |
Figure 7. Representative manufacturing agent tasks
As shown in Figure 8, the Intel Xeon 6 processor-based instance delivered the highest normalized agent-task throughput, achieving 1.58 times the throughput of the 5th Gen AMD EPYC processor-based instance and 4.24 times the throughput of the Arm v9.2A-based instance. This throughput advantage enables manufacturers to support more concurrent agent-driven workflows per server, accelerating tasks such as predictive maintenance, quality inspection, yield analysis, supplier management, and production optimization while maintaining the same infrastructure footprint.
Figure 8. Intel Xeon 6 processor-based instance delivers up to 1.58 times higher than 5th Gen AMD EPYC processor- based instance and up to 4.24 times higher agent throughput than Arm v9.2A- based instance on manufacturing tasks.
As shown in Figure 9, the Intel Xeon 6 processor-based instance delivered the highest normalized throughput across all workload categories evaluated. Relative to the 5th Gen AMD EPYC processor-based instance, Intel Xeon 6 achieved 1.59 times higher throughput for analytical and reasoning workloads, 1.55 times higher throughput for secure data processing workloads, and 1.55 times higher throughput for control workloads. Relative to the Arm v9.2A-based instance, Intel Xeon 6 achieved 4.48 times, 3.78 times, and 3.78 times higher throughput, respectively. This consistent performance advantage helps maintain throughput and responsiveness across end-to-end manufacturing agent workflows, from production analysis and quality inspection to factory operations and process control.
Figure 9. In the same 60-minute run, the Intel Xeon 6 processor-based instance completed more manufacturing tasks in reasoning, security and control operations than a 5th Gen AMD EPYC processor-based instance and Arm v9.2A-based instance.
Why Throughput Matters
Agentic AI shifts performance measurement away from isolated operations and toward complete workflow execution. What ultimately matters is how efficiently the platform can execute tasks, invoke tools, process data, manage state, validate outputs, coordinate dependent actions, and sustain productivity under continuous load.
The most valuable platform is the one that can continuously move the greatest amount of work through the system while maintaining efficiency, responsiveness, and scalability.
Final Takeaway
As organizations deploy larger numbers of agents and increasingly complex workflows, sustaining high throughput becomes a key determinant of operational efficiency. These results demonstrate that an Intel Xeon 6 processor-based instance provides a strong foundation for agentic AI, delivering higher throughput across real-world healthcare, manufacturing and financial services workflows and enabling organizations to process more agent-driven work on the same infrastructure.
Product and Performance Information:
Terminal Bench:
Intel Xeon 6: 1-instance m8id.12xlarge: 48 vCPU, 185 GB total memory, Ubuntu 26.04, 7.0.0-1006-aws, Terminal-Bench 2.0. Tested by Intel as of September 2026. Results may vary.
5th Gen AMD EPYC: 1-instance m8a.12xlarge: 48 vCPU, 185 GB total memory, Ubuntu 26.04, 7.0.0-1006-aws, Terminal-Bench 2.0. Tested by Intel as of September 2026. Results may vary.
Arm v9.2A: 1-instance m9gd.12xlarge: 48 vCPU, 185 GB total memory, Ubuntu 26.04, 7.0.0-1006-aws, Terminal-Bench 2.0. Tested by Intel as of September 2026. Results may vary.
Test-Methodology:
Benchmark
Terminal-Bench + Harbor 0.16.1, terminus-2 agent. Deterministic fixture replays - A recorded Claude trajectory is replayed through a local proxy, so runs do no model inference and no network (identical work on every system); one canonical terminus on all systems; verifier enabled, all three instances ran all the tasks without any failures.
Load
One sandbox per physical core, pinned 1 core/task via cpuset_slot_pinner, memory local to the socket. 24 concurrent sandboxes. 60-minute run, refill setting -k 500 (slots stay full); each task cycles many times.
Images
Task images pre-built and staged; runs make no network access (LLM replaced by the replay proxy on :4001). Per-container thread cap = 1.
Primary metric
whole-box system throughput calculated as tasks completed per 60 minutes (primary).
1
Performance results are based on testing as of the dates shown in configurations and may not reflect all publicly available updates. See backup for configuration details. No product or component can be absolutely secure.
Your costs and results may vary.
1 Performance varies by use, configuration, and other factors. Learn more on the Performance Index site.