Google open-sourced ML Drift, whose loudest number sits in the README rather than the blog
Google open-sourced ML Drift, a cross-platform on-device GPU inference engine, but the order-of-magnitude claim in the README never appears in the blog.

Google's AI Edge team open-sourced ML Drift on October 8 under Apache 2.0, a cross-platform GPU compute engine for on-device inference that abstracts OpenGL ES, OpenCL, Metal and WebGPU behind a single runtime. It is the acceleration layer inside LiteRT and now also ships as a standalone C++ library. The repository's first feature bullet is the loud one, an "order-of-magnitude performance improvement relative to existing open-source GPU inference engines". It names no hardware, and the announcement post never repeats the multiplier. What the post does quantify reads more modestly, and all of it comes from vendors measuring their own software.
A delegate replacement, not a new category
ML Drift is the stated successor to the GPU delegate in TensorFlow Lite, which is a narrower thing than a new inference engine. Chintan Parikh and Juhyun Lee, product manager and software engineer on the Google AI Edge team, frame the old delegate as foundational but built for an ecosystem that has moved on: on-device workloads now span real-time vision, audio and depth processing as well as high-parameter generative models, and the legacy runtime was not designed for those bottlenecks. The delegate "will no longer receive new feature updates"; the migration target is the LiteRT ML Drift GPU accelerator, which Google says keeps backwards compatibility for existing models. Android developers running unbundled runtimes get acceleration in standalone LiteRT packages today, with the Play Services build promised "soon".
The engine is not a fresh bet. Chrome, YouTube Shorts, Google Photos and Meet already ship it, alongside Adobe Lightroom, Adobe Photoshop and Snapchat lenses, and it sits under the same AI Edge stack that produced EmbeddingGemma 2 two days earlier. Those adoption figures are the most concrete numbers in the release: YouTube Shorts reports up to 40% lower average frame latency, Adobe up to 30% faster on-device photo editing, Snap 30% lower model latency in lenses. All three are partner-reported and unaudited.
Tensor virtualization, and what it replaces
Historically, keeping optimized shaders for the TFLite GPU delegate meant hardcoding logical tensor mappings directly to physical GPU objects, textures and buffers, separately across the OpenGL, OpenCL and Metal backends. ML Drift replaces that with tensor virtualization: a tensor's logical representation is decoupled from its physical allocation, and dynamic shader templates resolve coordinates during the compiler's initialization phase rather than at inference time. One unified shader model stands in for three backend-specific codebases. Google's phrase for the trade-off is "minimal runtime overhead", and that phrase is doing a lot of work here, since the release ships no profile of that overhead and none of the published charts isolate it.
Where the shaders get written by an agent
The genuinely unusual part of this release is not in any benchmark chart. ML Drift's modernized custom-op framework exposes direct registration APIs with low-level shading-language access, and the repository ships an agentic SKILL.md guide so that, in Google's words, coding agents can "author, register, and verify performant custom shaders in minutes". Notice which verb comes third. When an agent writes a GPU kernel, correctness and speed are precisely the properties a human reviewer used to supply, and a markdown instruction file does not supply them. That does not make the feature useless. It makes it a developer-experience change rather than a performance one, and it is the only part of the announcement with no prior equivalent in the open-source stack.
Coverage the old delegate could not offer
The GPU delegate was structurally hardcoded to 4D tensors, which forced layout hacks on anyone whose model needed five. ML Drift supports 5D out of the box in the LiteRT GPU accelerator, so 3D convolutional networks for volumetric and spatiotemporal work, and architectures such as YOLO 11n, MobileViT v2 and Swin Transformer v2, run directly on edge GPUs. This is coverage rather than acceleration: models that previously could not run at all now run. Both kinds of progress are real, but only one of them can be expressed as a percentage of an existing baseline, and this is the one that cannot. It is the same edge-silicon argument that the iPhone's second GPU is chasing from the other direction.
The multiplier, and where it stops
The edge-LLM work is the most concrete. Autoregressive models have two stages with opposite bottlenecks, a compute-bound KV cache prefill and a memory-bandwidth-bound decode that emits one token at a time, and ML Drift switches kernels and layout configuration between them, using a convolution-aligned KV cache layout and in-kernel activation quantization to avoid redundant memory round trips. Google publishes the configuration behind its mobile chart, which matters more than the bars do.
| Setting | Mobile | Desktop preview |
|---|---|---|
| Prefill | 512 | 8192 |
| Decode | 128 | 1024 |
| Context | 640 | 9216 |
| Weights | 4-bit, block 32 | 4-bit, block 32 |
| KV cache | fp16 | fp16 |
Desktop remains a preview: Google compiled the same WebGPU backend natively outside the browser with Dawn to reach Windows and Linux, and labels those figures "an early snapshot" while stating its own priority plainly, "delivering lightweight, production-grade inference where resource constraints are tightest: mobile and edge devices". Up to 12% lower memory across Gemma benchmarks completes the number set.
Commenters on r/LocalLLaMA on October 9 made the same distinction the blog's own numbers do: GPU acceleration on mobile is not new, and the multiple, not the acceleration, is the thing worth arguing about. Nothing in Google's material contradicts that reading. Its mobile charts are phone-class silicon, its desktop figures are a snapshot, and the third-party corroboration, from Arm, describes co-designed kernels, texture cache locality and workgroup tuning for Mali and Immortalis GPUs, real engineering in a post that makes no speedup claim of its own. What would settle the headline is a third party running the same Gemma models on the same phones against another open-source GPU engine, publishing kernel variants and device names. Three days after launch, the repository showed about 140 stars. That is not a measure of quality; it is a measure of how few developers have looked yet.


