Gemma 4 on Arm marks a significant step in moving AI processing from cloud servers directly onto Android smartphones, giving developers a path to faster, more private app experiences for billions of users. Google launched the model on April 2, 2026, in collaboration with Arm, targeting the Armv9 CPU architecture that runs across the Android device ecosystem.
What Gemma 4 Adds for Mobile Developers
The model expands support for multimodal tasks, covering text, audio, and image inputs. It handles reasoning, agentic workflows, and vision-and-audio use cases without increasing memory footprint on the device. Broader language support is also included, making the model more accessible across global markets.
Gemma 4 on Arm: Performance Numbers from Early Tests
Arm engineering tests using the Scalable Matrix Extension 2 (SME2) instruction set showed an average 5.5x speedup in prefill, which is the phase where the device processes user input. Decode speed, the phase that generates responses, improved by up to 1.6x. These results came from tests on the Gemma 4 E2B (Effective 2 Billion) model and include upcoming patches to Google XNNPACK and Arm KleidiAI.
Arm KleidiAI is a software acceleration layer built into runtime libraries such as Google XNNPACK, Google LiteRT, and MediaPipe. It delivers SME2 performance gains to developers without requiring changes to existing code, models, or deployment pipelines.
Accessibility App Envision Tests Local Inference
Envision, an app designed for blind and low-vision users, tested an on-device version of its scene interpretation feature using Gemma 4 running on SME2-enabled Arm CPUs. Previously, the feature depended on cloud connectivity. In the prototype, users could photograph a scene and receive a detailed description directly on the device, with no network connection required and no data sent off-device.
“Envision is excited to work with Arm and Google to bring powerful accessibility experiences directly onto smartphones. Running visual understanding models like Gemma 4 on-device on SME2-enabled Arm CPUs opens the door to reliable, low-latency scene description and visual Q&A for blind and low-vision users. For our community, the ability to access these capabilities offline is incredibly meaningful because it ensures the technology works wherever they are, while also improving privacy by keeping more processing on the device itself.”
Karthik Mahadevan, CEO, Envision
Why the Armv9 Architecture Matters Here
SME2 is built into Arm C1 CPUs, which are already integrated into the latest Android smartphone devices. The instruction set accelerates matrix-heavy AI workloads within the power limits of a smartphone, sustaining higher performance without draining the battery or raising device temperature. Developers targeting Arm-based Android devices automatically receive these optimizations with no code changes needed.
“Delivering Gemma 4 efficiently across the Android ecosystem requires deep collaboration across hardware and software. Our work with Arm reflects a shared commitment to advancing on-device AI, combining the benefits of the Armv9 architecture and built-in acceleration technologies, like SME2, with the Android operating system to unlock greater performance and efficiency at scale. Together, we’re making it easier for developers to bring fast, responsive, and privacy-preserving AI experiences to our users, without needing to modify their existing applications.”
Sandeep Patil, Engineering Director, Android, Google
Broader Implications for Mobile AI
Moving AI inference from the cloud to the device reduces latency, lowers infrastructure costs for developers, and keeps user data on the device. These factors matter for real-time applications that need consistent performance regardless of network conditions. Alex Spinelli, SVP of AI and Developer Platforms and Services at Arm, said the collaboration aims to make on-device AI the default architecture across the Android ecosystem rather than an optional add-on.
Audio support in Gemma 4 applies only to the E2B (Effective 2 Billion) and E4B (Effective 4 Billion) model variants. As adoption grows, Arm and Google said they will continue providing developers with performance guidance and optimization tools for all Arm-based mobile devices.





