Google releases Gemma 4 12B open multimodal model

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Google released Gemma 4 12B, an 11.95B-parameter unified, encoder-free open multimodal model that can run locally on devices with 16GB of VRAM or unified memory, contrasting with the industry's broader push toward ever-larger models.
- Gemma 4 12B processes text, image, and audio inputs natively within a single architecture, and Google researcher Michael Tschannen framed it as aligned with years of research into unifying models and training paradigms across modalities.
- Google released Gemma 4 12B under an Apache 2.0 license, with the model already available on Hugging Face, on Ollama (including via MLX for Apple hardware), and discussed in dedicated guides from Google Developers Blog and The Keyword.
- Sundar Pichai said the model hits a "sweet spot between size + performance," enabling multi-step reasoning and agentic workflows while running locally on a laptop, and Google claims its performance nears the company's larger Gemma models at a far smaller memory footprint.
- Coverage converged across VentureBeat, Tech Times, Android Authority, The Decoder, MarkTechPost, Digit, and WinBuzzer, with community discussion on r/LocalLLaMA, r/technology, r/GeminiAI, and r/Bard — the cross-outlet framing consistently highlighted local deployment on 16GB laptops and the encoder-free architecture, while the Apache 2.0 license angle (developer-friendly, no usage restrictions) was less emphasized as a standalone story.
- Ollama demonstrated agent-style integrations, showing Gemma 4 12B running as a Hermes Agent and behind Claude Code via simple launch commands, signaling that Google is targeting not just chat use cases but local agentic workflows.
Why it matters: By shipping a 12B multimodal model under Apache 2.0 that runs locally on 16GB laptops, Google gives individual developers and small teams a free, license-permissive alternative to cloud-only API models — the encoder-free single-architecture design is the technical hook, but the combination of zero cost, local execution, and native audio/vision is what actually changes the build-vs-buy calculus for on-device AI agents.
