All News
nvidiadgx-sparkgb10local-llmlocal-llamahome-labmemory-pricingpower

From one RTX 3090 to 20 DGX Sparks: the home lab that blew its own fuses

One r/LocalLLaMA builder scaled from a single RTX 3090 to 20 GB10 boxes and blew his fuses, days after NVIDIA halved DGX Spark memory and raised the price.

Vlad MakarovVlad Makarovreviewed and published
5 min read
From one RTX 3090 to 20 DGX Sparks: the home lab that blew its own fuses

A build log posted to r/LocalLLaMA on 4 October follows one hobbyist from a single RTX 3090 to a twenty-unit GB10 cluster, and records that the first thing to give way was not the software stack but the house wiring. The poster, u/ciprianveg, was sitting at roughly 810 upvotes and about 440 comments as of 5 October, a snapshot, since both counts keep moving. His thread arrived two days after NVIDIA announced a smaller-memory, higher-priced DGX Spark, and that timing is the interesting part: a single buyer's escalation and a vendor's product decision are now governed by the same scarcity. What breaks a home-AI build in 2026 is memory and mains power, not compute.

The climb, in the builder's own words

His account runs in steps. One RTX 3090 for LLaMA 33B, then a second card when LLaMA 65B appeared. A Threadripper with 512GB of DDR4 followed, running DeepSeek 671B MoE at 8 t/s with experts offloaded to RAM, or Qwen 235B at 10-12 t/s when he wanted speed, used for real coding at his job through OpenWebUI. When agentic coding raised the token bill, he moved to sixteen RTX 3090s spread across P620-based nodes on a 100Gbit network, running MiniMax M2, Qwen 235B and Qwen 397B. Every step is a plausible answer to the previous one being too slow.

Then the circuit breaker intervened

That is where the electrical story starts. "The house circuit said no: the fuses blew whenever the rig and the electric oven ran together. Heat and stability were issues too," he writes. The fix was not more GPUs but a different kind of computer. Four ASUS GB10 units ran a 397B model at 30 t/s on 400W, against 50-60 t/s at 6kW for the older rig, and he calls the new boxes "rock solid and almost silent." Eight GB10s followed for MiMo 2.5 Pro and Kimi 2.6, then two eight-unit clusters were joined across two houses (his brother lives a five-minute walk away) into a sixteen-unit pool, at which point Kimi K3 went "from an unusable 7 t/s at 100k context to a fairly usable 20 t/s at 300k." He is now adding four more units so a smaller model, GLM 5.3 Flash, can run around the clock while the big cluster handles GLM 5.3, MiMo 2.6 Pro, Kimi K3 or Qwen 3.8 2.4T. He adds that he has never paid for a commercial model subscription, "not because of the cost, but, because of my strong confidence in local models future."

NVIDIA halved the memory and raised the price

NVIDIA's 2 October announcement, under the newsroom headline "NVIDIA DGX Spark 64GB Gives Developers More Ways to Build and Scale Local AI," introduced a 64GB configuration of the same GB10 Grace Blackwell machine, rated by the company for models up to 100 billion parameters against up to 200 billion on the 128GB model. It starts at $4,999 from Acer, ASUS, Dell, Gigabyte, HP and MSI on 23 October. The 128GB Founders Edition had its own MSRP raised to $4,699 in February 2026 from $3,999, on worldwide memory-supply constraints, so the new half-memory OEM systems list $300 above the current 128GB price. VideoCardz, which flagged the move, notes the 128GB street price has crossed $6,000.

What twenty units would cost and draw

QuantityBasisEstimate
Twenty units at the current 128GB MSRP$4,699 eachabout $93,980
Twenty units at the new 64GB OEM price$4,999 eachabout $99,980
Twenty units at full load240W each4.8 kW
One typical US household circuit120V times 15Aabout 1.8 kW

Every row is arithmetic, not a quote, and every price is an estimate: he bought across several price eras, so no single list price describes what he paid. The load figure is the sharper one. Twenty units at full tilt would draw about 4.8 kW, roughly 2.7 times what a single 120-volt, 15-amp household circuit is rated to carry, which is a tidy restatement of the fuse problem he hit with the 3090 racks. The post's title also counts four units that had not arrived when he wrote — "They arrive October 2" — inside a pool that is otherwise ASUS's GB10 variant rather than NVIDIA's Founders Edition.

One post, one self-report

The throughput numbers, 8 t/s, 30 t/s, 20 t/s at 300k context, are all his own, and none carries a reproducible configuration or an independent run. Read them as one enthusiast's log rather than a benchmark: decode speed depends on quantization, context length, batch size and offload settings that the post does not fully specify. The community's reaction split along that line. One reply questioned the power budget directly ("16x 3090 isn't something like 32 KW range for spikes? That is 135A at 240V"), and another made the accessibility case bluntly, writing that the knowledge "is not going to be applicable to 99.999% of the community, because most people out there can't afford and can't even dream about owning a cluster of 20 DGX Sparks." Trending summaries of the thread render the mood as money envy, along the lines of asking where the poster's money tree grows; treat that phrasing as a summary of tone, not a verified quote.

The bill is memory, not compute

The Spark's price move is not really about the Spark. It is a memory story wearing a workstation's casing. On 30 September, Micron's chief executive said 75% of the company's 2027 output is already committed and conditions will be "much tighter" in 2027 and 2028, the same premium that just moved the Spark's price. A buyer paying $4,699 to $6,000 for a 128GB unit is paying the AI-datacenter memory tax, and a new 64GB model that costs more than the larger one it replaces is the clearest sign that memory, not the GB10 die, sets the price.

What would settle the throughput claims

Nothing in the thread is unfalsifiable, which is the good news. A published configuration, covering model, quantization, context length, offload layout and network topology, would let a second party reproduce the t/s figures on identical hardware. An MLPerf-style run at a stated context length and batch would put them on a scale that survives comparison. A measured wall-plug draw for one unit, rather than the 240W rated figure, would test the power arithmetic directly. Until one of those exists, the honest summary is that a determined individual built something remarkable and told us it works, and we have his word, his photos and his receipts, but not a number anyone else can rerun.

Related Articles

Scroll down

to load the next article