All News
local-llmapple-siliconiphonellama-cppinference

Backburner makes an iPhone a second GPU for a 24 GB Mac, on reads only

StayLameBro's open-source Backburner turns an iPhone into a second GPU for a 24 GB M4 Pro MacBook. Its own bench scripts show prefill up 29-44%, decode flat.

Vlad MakarovVlad Makarovreviewed and published
6 min read
Backburner makes an iPhone a second GPU for a 24 GB Mac, on reads only

The top post in r/LocalLLaMA on October 2 was a phone doing part of a 27B model's work beside a laptop. The project behind it is Backburner, MIT-licensed, by GitHub user StayLameBro: an iPhone or M-series iPad plugged into an Apple Silicon Mac over a 10 Gb/s USB-C cable, helping run Qwen3.8-27B locally through a llama.cpp fork with its own Mac kernels. The Reddit thread drew about 1,370 upvotes and 239 comments. The video is the hook, but the README is the story: measured prefill and context deltas from one machine, published alongside the bench scripts, and a limits section that is unusually frank about where the phone stops helping.

How the split works

For every batch of 256 prompt tokens, the Mac runs layers 1 through 40 and streams the residual to the phone, which runs layers 41 through 64 on its GPU while the Mac starts the next batch. On the A19 Pro inside an iPhone 17 Pro Max, those phone-side layers use the GPU's matrix units through Metal 4 tensor operations and run 2.4x faster than the same phone with those units switched off. The engine is a llama.cpp fork with three additions: SME2 and Metal kernel fusions, plus a speculative draft model the project calls DFlash2. The author also tested an iPhone 16 Pro Max with an A18 Pro. Every measurement is dated October 1, 2026, on a MacBook Pro M4 Pro with 24 GB running macOS 26.2, with the model at IQ4_XS.

Prefill: what the phone buys on reads

The headline claim is time saved on the reads an agent does constantly, such as a two-thousand-token file or tool result pulled into a saved session. The comparison is the same build with and without the phone attached.

ContextMac aloneMac + iPhoneChange
16k109 tok/s157 tok/s+44%
32k101 tok/s130 tok/s+29%
48k87 tok/s113 tok/s+30%

A cold 26,849-token agent session tells the same story end to end: 245 seconds on stock llama.cpp, 228 seconds on the fork with the Mac alone, and 168 seconds with the phone attached. The author asks readers to quote these end-to-end figures rather than the per-layer tokens-per-second the phone app displays, because that on-screen number covers only the layers the phone holds.

Context: what 24 GB holds

A 24 GB Mac fits 64k tokens of 8-bit context beside the model. Attach the phone and the server sizes the total from the phone's free memory at startup, reaching 196k to 229k tokens at 8-bit on an iPhone 17 Pro Max. The author tested end to end to 128k at 8-bit and 140k at 4-bit. At 128k, three planted facts at positions 1,500, 40,000 and 100,000 tokens were all recalled. Without the phone, the same machine has to drop to 4-bit context past 64k, trading precision for room.

Past 64k the phone switches jobs

Past 64k the phone stops running layers 41 through 64 and does a different task: computing attention over the oldest KV pages it holds, 4,096 keys per page, while the Mac runs all 64 layers. The phone's Neural Engine compiles each 16,384-key page into an ANE model with the keys and values as its weights. At 140k that changed the cost of writing a token from 279 to 176 milliseconds compared with the phone's GPU doing the same work. The phone is doing one job here, not two.

Decode: the phone is not the lever

Below 64k the phone does not speed up writing at all. What does is the fork's kernels and draft model: stock llama.cpp decodes at 11.3 tokens per second, the fork with the Mac alone at 25.0, and the fork plus the iPhone at 25.1, all at 27k to 33k context. At 128k with the phone, decode is 12.6 tokens per second. The reason to attach a phone is prefill and context length, not generation speed. The author also reports that greedy output is token-identical with and without the phone, 256 of 256 tokens at 8k and 32k and 32 of 32 at 140k, so this is not a cheaper approximation. One negative result stands out: the Mac's own Neural Engine shares the Mac's memory bandwidth and slowed decode by 26% when running, so the project leaves it off.

The limits the author states

The README is candid about where Backburner does not apply. The phone only joins reads longer than about 512 tokens; in a real 36-request agent session, only 7 requests were big enough, though those 7 carried about 83% of the tokens read. The phone handles one request at a time. Past 64k the app must stay open and in front, or the server stops after 15 seconds, and the author recommends Guided Access because iOS will not let an app use the GPU in the background. If the phone fails during a read, the phone is switched off for 60 seconds and the batch reruns on the Mac. Saving a session while the phone holds keys needs app version 0.0.4. On security, the app answers only the paired Mac over the cable, or a paired Mac over an encrypted tunnel, and versions before 0.0.3 accepted Wi-Fi connections, which the author says to update past. This is pre-release software.

What would settle it

Every number here is the project author's own measurement on one machine, published with the bench scripts so others can reproduce it, and none of it has been reviewed by a third party. The model is a Qwen3.8-27B that arrived without an API, and the fork's speculative decoding sits in the same family of cost-cutting work as recent prompt-lookup speedups in a llama.cpp fork. Independent runs on other Macs and other iPhones are the missing piece. The install is a one-command curl script, and the phone app goes through AltStore with a free Apple ID and no developer account, friction low enough that reproduction should be quick.

Related Articles

Scroll down

to load the next article