Four R9700s and 128 GB of VRAM: one builder's numbers, unattested
A 10-person startup's 128 GB R9700 box serves 16 sessions on its builder's own numbers. We run the arithmetic on when such a rig beats a hosted API bill.

On 11 October a user posting as sayamss put a photograph of a four-card rig on r/LocalLLaMA under the title "Building a 4x R9700 setup for a 10 person startup." The machine is for a client, he writes: four AMD Radeon R9700 AI Pro cards on a Threadripper 9970X, 128 GB of VRAM in one tower, serving Qwen 3.8 and DeepSeek models to 16 simultaneous sessions. It is a build report, not a benchmark, and that gap is the story — because the case for a box like this is that it replaces somebody's API bill.
What the post actually specifies
| Component | As posted |
|---|---|
| CPU | Threadripper 9970X |
| System memory | 128 GB DDR5 ECC 5600 |
| GPUs | 4x AMD Radeon R9700 AI Pro 32 GB (128 GB total) |
| Power | 1600 W PSU; GPUs undervolted, under 210 W each |
| Serving engine | "a fork of Radiance" |
| Models | Qwen 3.8 Next Flash, Qwen 3.8 27B, DSV4 Flash |
AMD's own product page lists the R9700 at 300 W board power, so four cards at stock would pull 1.2 kW before the CPU. Undervolted and under 210 W each, the poster's four cards draw roughly 840 W, which is why a 1600 W supply is not overkill. The card itself is 32 GB of GDDR6 on a 256-bit bus at 640 GB/s over PCIe 5.0 x16. The other half of the build is the platform: 128 GB of ECC system memory can hold weights or a KV cache that will not fit in the cards at the context lengths being quoted.
Whose performance numbers these are
| Workload (16 sessions) | Aggregate prefill | Aggregate decode |
|---|---|---|
| Qwen 3.8 27B MXFP4 | 6.3–6.8k tok/s | 900 tok/s at 4k/user, 80 tok/s at 128k/user |
| Qwen 3.8 Next Flash, 48K/user | 6,880 tok/s | 538.5 tok/s |
| Qwen 3.8 Next Flash, 128K/user | 3,050 tok/s | 365 tok/s |
BF16 KV cache throughout, per the post. These are the builder's own measurements from one machine, with no prompt set, no time-to-first-token distribution and no second opinion. The engine is the harder problem for anyone trying to repeat it: "a fork of Radiance" is not upstream vLLM or llama.cpp, so reproducing these figures needs that fork and the same quantizations rather than a standard install. The decode spread is the honest part of the table. At 128k per user the box returns roughly 80 tok/s across all 16 sessions, about an eleventh of its 4k figure, and the post never says what prompt lengths the sessions actually ran at.
The parts the post leaves out
No idle draw, no TTFT, no published prompt set, no monthly cost, and no comparison against a hosted API on the same workload. The workload itself is described only as work for a client. A commenter, dethswatch, asked the obvious question — "nice but what's the workload?" — and the thread does not answer it. Three models are named, Qwen 3.8 Next Flash, Qwen 3.8 27B and DSV4 Flash, but throughput is published for only the first two, and never per session. Sixteen is the concurrency the box was configured for, not a measured ceiling: nothing says whether the sessions ran at once or in sequence, whether the figures are averages or peaks, or how long each run lasted. Without that, the throughput numbers can be read almost any way; a rig that saturates on long context and chats lightly is a different purchase from one running agents all day.
The memory bill behind the price
The R9700 is a GDDR6 card, which makes the memory-supply story relevant to it. TechPowerUp reported on 10 October that AMD has raised the price of the GDDR6 it supplies to board partners, effective 1 October, after a roughly 10% increase on Radeon kits in July. AMD's own site still lists the R9700 at an MSRP of $1,299 as of 1 October 2025. Today Micro Center lists a Sapphire single-fan R9700 at $1,799.99, down from a $2,100 sticker, in-store pickup only. That is about 38% above the listed MSRP: $56 per gigabyte of VRAM, against $40.6 at the launch figure. None of this proves the memory increase caused the shelf price, but the number the builder paid is not the number in AMD's footnote — and the same pressure is reshaping the Nvidia side, where the RTX 5090 is ending production with GB202 reserved for professional cards.
When a box beats a meter
Here the arithmetic is doable, if every assumption is stated. Take the EIA's US residential average of 18.31 cents per kWh for July 2026, the current figure on its monthly update, as a stand-in for the client's rate. At roughly 1,050 W at the wall — 840 W of GPUs plus CPU, memory and supply losses — that is 8.4 kWh a day at eight hours of load, about $47 a month, and 25.2 kWh a day at 24/7, about $140 a month. Idle draw is unreported, and a warm machine on standby still bills.
On the other side, our provider records list Qwen 3.8 Flash Next on Alibaba's cloud at $0.16 per million input tokens and $0.47 per million output. A light office pattern — ten seats at 0.4M input and 0.05M output each per day — costs about $27 a month to serve hosted. Against $7,200 of GPU capex, four cards at today's $1,799.99, that is 270 months, and the electricity alone is larger than the API bill. Push the same box harder, to 16 seats at 200 output tokens a second for two hours each, and hosted serving costs roughly $890 a month: an eight-month payback. Run them six hours a day and it is about $2,670 a month, repaid in under three. The same $7,200 buys roughly 45 billion hosted input tokens, or about 197 days of continuous decode at the builder's 4k aggregate figure.
That spread, from three weeks to twenty-two years, is the point. Nothing about the hardware decides it; the workload does, and the post does not state one. The earlier 128 GB server built on second-hand parts had the same shape: real throughput, real wattage, and an economics section that only closes once you know what the machine is for. What would settle this build is a workload description, a prompt set, TTFT numbers and a month of real traffic priced against the same tokens through an API. Until then it is one person's tower, and the numbers are his.


