All News
amdepychardwarememory-bandwidthdata-centerai-news

AMD's EPYC 9006 'Venice' gets a price list: $700 to $14,904

AMD's EPYC 9006 'Venice' gets full specs and prices from $700 to $14,904. We check whether 16-channel DDR5-12800 really matches an RTX 5090's bandwidth.

Vlad MakarovVlad Makarovreviewed and published
6 min read
AMD's EPYC 9006 'Venice' gets a price list: $700 to $14,904

AMD's EPYC 9006 server line, the Zen 6 family AMD still calls Venice, finally has a public price list. The disclosure arrived in coverage dated September 29, 2026, after StorageReview obtained the numbers: SKUs from eight to 256 cores, priced from $700 to $14,904. The pricing is the news, not the chip. AMD sketched the platform in a July 23 blog post and The Register walked the architecture on September 23; money was the last piece left to publish, as Tom's Hardware reported.

The lineup, and what it costs

Two sockets split the range. The flagship SP7 platform carries up to 16 channels of DDR5, RDIMM or MRDIMM, with MRDIMMs rated to 12,800 MT/s. It adds PCIe Gen 6 and CXL 3.1, up to 256 cores and 512 threads per socket, and 400W to 600W TDPs. A trimmed SP8 line keeps eight memory channels for storage, database and virtualization boxes. The bandwidth argument only concerns SP7, so here is that price list:

ProcessorCores / Threads1Ku priceL3 cacheTDP
EPYC 9996256 / 512$14,9041,024 MB600 W
EPYC 9966192 / 384$14,079768 MB600 W
EPYC 9846168 / 336$13,114768 MB500 W
EPYC 9756128 / 256$12,498512 MB500 W
EPYC 9686F96 / 192$11,434384 MB500 W
EPYC 965696 / 192$9,713512 MB400 W
EPYC 955664 / 128$8,008384 MB300 W

Per-core math favors the top. The 256-core EPYC 9996 lands near $58 a core; the entry SP7 part, the 64-core EPYC 9556, near $125. The "F" SKUs trade cores for clock speed: the 96-core 9686F boosts to 5 GHz and costs more than the denser 9656. Pair two 9996s in a 2P board and you get 512 Zen 6 cores for $29,808 in silicon before a single DIMM.

Why 16 channels beat 256 cores

For language-model work, the number to watch is not core count but memory bandwidth. Server DDR5 moves data in 64-bit, 8-byte chunks, so 16 channels running at 12.8 gigatransfers per second give 16 x 12.8 x 8 = 1.6384 TB/s per socket. The Register puts AMD's own claim near 1.6 TB/s, with roughly 1.3 TB/s in the STREAM Triad benchmark, so the arithmetic holds.

Now compare it with the GPU many people run locally. NVIDIA's RTX 5090 has a 512-bit GDDR7 bus rated at 1,792 GB/s, which is 1.792 TB/s. Divide one by the other: 1.6384 / 1.792 = 0.914. One socket of EPYC 9006 memory sits within roughly a tenth of a flagship consumer GPU's bandwidth. That is the claim a r/LocalLLaMA thread made when the prices surfaced, and the arithmetic checks out.

The point of the number is that decode on a CPU is a memory-bound loop. Every token streams the model's active weights and the key-value cache through the memory controllers, and a wider bus spreads that traffic across more channels at once. Adding channels is the cheapest way to move a bandwidth-limited workload without reworking the compute dies. It is also why the SP8 parts, which halve the count to eight channels, are pitched at database and virtualization work rather than inference.

Bandwidth parity is not inference parity

The framing invites a wrong conclusion. Equal bandwidth does not mean equal tokens per second. A GPU's GDDR7 is soldered to the card and priced into it; server RDIMM and MRDIMM kits are bought separately, and the thread's own running joke, that a 2TB DDR5-12800 RDIMM kit "is going to cost only 2 kidneys and a small micronation's GDP," is the community pointing at exactly that bill. Capacity of this class is expensive and, right now, scarce.

There is also no independent benchmark of tokens per second on an EPYC 9006 system. Nobody outside AMD has published decode throughput on the platform, and the bandwidth figure is a ceiling, not a result. In an agentic pipeline the CPU's job is orchestration, retrieval and tool execution, while token generation still runs on GPUs. That division is AMD's own: the company writes that "agentic AI is fundamentally a systems-level workload," which is the honest version of the claim, and a narrower one than "GPU bandwidth in a CPU socket."

The vendor's benchmark, unaudited

AMD's July blog offers a CPU-centric comparison and, to its credit, labels its own limits. It pits the EPYC 9996 (256 cores / 512 threads) and EPYC 9965 (192 / 384) against Intel's Xeon 6980P (128 / 256) and AWS Graviton5 (192 / 192), then reports a geomean across gateway, context-assembly, retrieval and tool-execution stages. It excludes GPU inference outright, noting the blog "concentrates on the CPU infrastructure... and does not cover GPU inference performance." The numbers are AMD's, run by AMD's engineers, and unaudited. Treat the geomean as a marketing artifact until someone reproduces it. Note what the comparison leaves out, too: Graviton5 is an Arm part without hyperthreading, so the thread counts are not apples to apples, and folding nine dissimilar stages into one geomean hides whichever stage a buyer actually cares about. A single leaderboard figure says more about which stages AMD chose to weight than about how the chip behaves in your pipeline.

What would settle it

Two things, and neither exists yet. First, third-party tokens-per-second figures for a quantized model running on a 16-channel SP7 socket, measured at the same prompt and batch size as a GPU run. Second, street prices for DDR5-12800 MRDIMM kits, because the platform's bandwidth advantage evaporates if the memory costs more than the accelerator it is meant to feed. AMD's chip is impressive on paper; the skepticism in the thread, at roughly a thousand points and more than two hundred comments, is about throughput and total cost, not the spec sheet. Given the memory crunch already pushing prices up, and AMD's announced 10% Q4 price increase across its server line, the cost question is the live one, not the core count.

Related Articles

Scroll down

to load the next article