TECH
AMD report says Zen 6 EPYC Venice smokes NVIDIA Vera in revised agentic AI benchmarks
AMD has updated its performance estimates for its upcoming 6th Gen EPYC 9006 "Venice" server CPUs, and the latest numbers make an aggressive case against Intel and NVIDIA in high-density data center workloads. A new AMD white paper, titled "AMD EPYC Server CPU Architecture And Performance Overview" revises the company's modeling for Venice and compares the 256-core EPYC 9996 against Intel's 128-core Xeon 6980P and NVIDIA's 88-core Vera CPU.
The updated analysis places particular priority on agentic AI infrastructure, where CPUs handle tasks such as orchestration, databases, web services, caching, APIs, retrieval, and other work surrounding accelerator-based AI inference. This work is now comprising a larger and larger percentage of the actual workload of Agentic AI, reinforcing the role of CPUs in the datacenter, once thought to merely be orchestration for GPUs that do all the 'real' work.
AMD expects that the EPYC 9996 will deliver 2.4× the performance of Xeon 6980P in server-side Java, 2.5× in OpenSSL, 3.5× in MongoDB with YCSB, 2.9× in Redis Benchmark, 3.7× in NGINX with WRK, and 2.6× in transaction processing based on TPC-C. Those results are normalized to a 128-core Xeon 6980P system and represent the CPU-heavy enterprise and cloud-native portions of an agentic AI infrastructure stack. Notably, Vera was not included in this modeling, but not because AMD didn't have numbers to compare against.
Besides, the more eye-catching update comes at the rack level. AMD has revisited its earlier 100-kW rack-capacity model from June, which estimated how much aggregate workload throughput each platform could deliver within a fixed rack-level power envelope. The company says the previous analysis put 5th Gen EPYC 9965 at 2.37× the rack-level throughput of its NVIDIA Vera baseline. The revised study now models the 6th Gen EPYC 9996 at 3.4× Vera's rack-level performance.
This is partially because the underlying methodology has changed. AMD's earlier rack estimate was based on a six-workload geometric mean that included SPECrate 2017 Int, while the updated model incorporates SPECrate 2026 Int, server-side Java, NGINX, Redis, Memcached, and TPROC-C. That makes the latest figure a refreshed model rather than simply adding Venice to the old chart. There's also a small but interesting change for AMD's previous-generation part. The earlier 9965 estimate was 2.37× Vera, while the new chart shows the 9965 at about 2.3× in the revised model. In other words, AMD appears to have rerun the calculations rather than simply carrying its previous numbers forward.
However, it's critical to keep in mind with these rack-scale comparisons that AMD is modeling performance rather than publishing results from physical 100-kW racks containing both Venice and Vera hardware. The company describes the figures as estimates based on a combination of benchmark data and system-level assumptions, so they absolutely should not be treated as apples-to-apples measurements of shipping systems.
But AMD's also making a broader argument about the role of CPUs in agentic AI. Rather than treating the CPU as merely a host for an accelerator, the company says increasingly complex AI workflows require substantial CPU resources for orchestration, retrieval, database access, tool execution, networking, and response generation. The latest EPYC 9996 figures reinforce AMD's pitch for Venice as a high-core-count CPU designed to maximize useful work under real-world rack constraints rather than simply chasing peak processor performance.
Still, Vera is a more specialized CPU than Venice, and it has a very different architecture at the platform level. While these simulations put NVIDIA's new chip behind AMD's finest, NVIDIA claims the win in its own testing using different benchmarks and constraints. It's entirely possible that Vera may end up being a better choice for some workloads. We'll just have to wait for independent benchmarks on real hardware to know for sure.
AMD EPYC Venice is expected to outsell NVIDIA Vera by 2027... To make things even more exciting, a Morgan Stanley report predicts that AMD's EPYC Venice CPUs will reach 6.75 million units sold by next year—17% more than NVIDIA's Vera (and 5.4 times the volume compared to 2026).
According to reports, NVIDIA will remain TSMC's primary customer for CoWoS packaging capacity, with the Taiwanese company expected to reach a capacity of 200,000 wafers per month by 2027.
"Team Green" utilizes TSMC's CoWoS packaging solution for two main products: CoWoS-L for AI GPUs (such as Blackwell and Rubin) and CoWoS-R for Vera CPUs.
CoWoS-L production capacity is expected to reach approximately 910,000 units—a 40% year-over-year increase—while Vera shipments are projected to double. This would drive a 52% increase in revenue for NVIDIA compared to the previous year.
Finally, the report projects that NVIDIA's Vera CPUs will reach 5.75 million units by 2027. This is a significant figure for a new CPU launch, especially given NVIDIA's stated goal of becoming the leading CPU supplier by 2026.
Great news (for AMD)...The key takeaway is that NVIDIA is currently facing stiff competition. While its Vera processors are already in mass production at TSMC, the same applies to AMD's next-generation EPYC platform, codenamed Venice.
It is worth noting that Venice is based on the upcoming Zen 6 architecture, which is expected to deliver significant gains in performance and efficiency. As previously highlighted, the report projects that EPYC Venice CPUs will reach a volume of 6.75 million units—17% more than NVIDIA’s Vera (and 5.4 times the volume compared to 2026).
AMD is also utilizing TSMC’s advanced 2nm manufacturing process, whereas Vera is based on 3nm process technology. Furthermore, Vera is designed for agentic AI, while AMD’s EPYC Venice addresses both AI and HPC workloads.
The challenge at hand is not simply AMD versus NVIDIA or NVIDIA versus AMD, but rather the rise of custom silicon, as many AI companies are now venturing into that field.
Just this week, we reported that Google and MediaTek are collaborating on a chip that integrates CPU and AI capabilities into a single package for the next generation of intelligent agents. OpenAI and Amazon are also either in talks to produce or are already manufacturing custom chips, a trend that will intensify the debate between in-house development and external sourcing.
In short, with the growing popularity of custom chip manufacturing, NVIDIA, AMD, and other manufacturers may be facing a critical situation. While the demand for computing power remains high, AI companies producing their own chips will further exacerbate the supply-demand imbalance.
mundophone
No comments:
Post a Comment