All News
amdradeonrdna-5gpu-leaksnvidiartx-5090hardwarelocal-ai

Radeon RX 10800 XT TFLOPS: the leak's FP32 number is 79-87 — and it is not the number that matters

AMD has published no Radeon RX 10800 XT specs, so the 79-87 TFLOPS FP32 figure is arithmetic on a leak. Here is the math behind the number, and its caveats.

Vlad MakarovVlad Makarovreviewed and published
5 min read
Radeon RX 10800 XT TFLOPS: the leak's FP32 number is 79-87 — and it is not the number that matters

AMD has published no specifications and no performance figures for any RDNA 5 card, which means there is no official Radeon RX 10800 XT TFLOPS number to look up — no product page, no launch date, no price and no first-party benchmark. The figure that circulates instead is arithmetic performed on a leak. Take the leaked 200 Compute Units at 64 stream processors each and you get 12,800 shader cores; at two floating-point operations per clock, the range between 3.1 and 3.4 GHz works out to 79-87 TFLOPS FP32. That number is derived, not measured — and it is not the one that will decide how fast the card runs a language model.

The 79-87 figure is arithmetic on a leak, not an AMD number

The only numbers in circulation come from a single leak stream, attributed by the hardware press to Moore's Law is Dead and Kepler_L2 and relayed by GameGPU and Sportskeeda: RDNA 5, a 3 nm process, up to 200 Compute Units, clocks of 3.1-3.4 GHz, a 512-bit memory bus (384-bit has also been floated), up to 48 GB of GDDR7 (24-36 GB also floated), a 350-450 W board, a rumored $2,000-$2,500 price and a rumored mid-to-late 2027 launch. None of it is AMD's.

The TFLOPS figure is one multiplication on top of that. Each Compute Unit carries 64 stream processors, so 200 CUs is 12,800 ALUs. A shader core that issues one fused multiply-add per clock does two floating-point operations per clock. Multiply 12,800 by two by 3.1 GHz and you get 79.4 TFLOPS; run the same sum at 3.4 GHz and you get 87.0. The range is the clock range, nothing more. Every step here is arithmetic on a leaked spec sheet, which is worth saying plainly: this is an estimate for a chip nobody has benchmarked, not a measurement.

The other precisions need a multiplier you also have to state

FP32 is only the first row. Lower precisions run faster on matrix hardware, and the multiplier depends on whether the architecture keeps a packed path. For the RTX 5090, TechPowerUp's GPU database lists FP16 at 104.8 TFLOPS — exactly the same as its FP32 figure, a 1:1 ratio. If RDNA 5 keeps the classic packed-FP16 path, the leaked part would sit near double its FP32 rate, and INT8 or FP8 higher again. That "if" is doing all the work.

PrecisionDerived rateAssumption
FP32~79-87 TFLOPS12,800 ALUs at 2 FLOPs per clock, 3.1-3.4 GHz
FP16~158-174 TFLOPS2x FP32, only if RDNA keeps its packed-FP16 path
INT8 / FP8~316-348 TFLOPS2x FP16, only if the classic rate holds

Every row above FP32 is conditional, so a cross-vendor "TFLOPS" comparison built on it is not like-for-like. The 1:1 FP16 line on the 5090 datasheet is the proof: the same word means different things depending on whose shader core is doing the multiply.

Against the RTX 5090 the core-count gap survives

Set the two side by side and the derived numbers land where the leak's own framing does. The 5090's GB202 runs 21,760 FP32 cores at a 2.41 GHz boost for 104.8 TFLOPS; the leaked Radeon part, at 12,800 shader cores, comes in below it on both counts. Fewer cores, higher clocks, more cache — the case for a generational leap rests on everything except the raw ALU count.

Leaked RX 10800 XTRTX 5090 (shipping)
Shader cores12,800 (leaked)21,760
FP32~79-87 TFLOPS (arithmetic)104.8 TFLOPS
FP16not stated104.8 TFLOPS (1:1)
Memory bus512-bit (leaked)512-bit
Memory bandwidthunknown (no speed in leak)1,792 GB/s

For scale, the shipping RX 9070 XT carries 64 Compute Units, a 2.97 GHz boost, 16 GB of GDDR6 and 640 GB/s. That chip derives to roughly 24 TFLOPS of FP32. The leaked flagship is a different class, but it is the same kind of arithmetic.

Bandwidth, not TFLOPS, decides tokens per second

Here is where the TFLOPS question stops being the useful one. Local inference on a quantized model is memory-bound: the weights have to move from VRAM through the memory bus for every token the card produces, and peak FP32 throughput barely enters that loop. The leak quotes a bus width — 512 bits — and stops there. It names no memory speed. Put 28 Gbps modules on a 512-bit bus, the speed the RTX 5090 uses to reach 1,792 GB/s, and the Radeon part would land in the same place; nothing in the leak says it will run at that speed. A capacity figure in the leak — up to 48 GB — changes what model fits, not how fast tokens come out. Our RDNA 5 architecture explainer works through the rest of the design questions; for inference, bandwidth is the line to watch.

Every peak number assumes every ALU fires every clock

The derived 79-87 TFLOPS is a ceiling, and a loose one. It assumes all 12,800 ALUs issue a fused multiply-add on every cycle at the top of the clock range, with no stalls and no thermal throttle. Real kernels are memory-stalled, and the leak's own 350-450 W envelope is where any sustained figure would be set. Review-site numbers are measured on shipping silicon, which is why they are the only ones worth arguing about. Our leak-versus-5090 comparison keeps the two apart for the same reason.

What would settle it

Three things would replace the arithmetic with a reading: a retail launch with board-partner cards and a real price; third-party 4K reviews measuring raster and ray tracing on shipping silicon; and published tokens-per-second numbers from llama.cpp or vLLM, cited next to the ROCm build they ran on. The leaked release-date and price hub tracks the launch window. Until then, the honest reading of "Radeon RX 10800 XT TFLOPS" is a derived 79-87, conditional on a multiplier nobody has confirmed, and secondary to a bandwidth number the leak never gave.

Related Articles

Scroll down

to load the next article