Google Releases Gemma 4 12B, Open Multimodal AI for 16GB Laptops

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Google released Gemma 4 12B, an 11.95B-parameter unified multimodal model that processes text, image, and audio inputs and runs locally on devices with 16GB of VRAM or unified memory.
- The model uses a unified, encoder-free architecture that handles raw inputs across modalities directly, which researcher Michael Tschannen said reflects his years-long focus on unifying models and training paradigms.
- Gemma 4 12B is released under an Apache 2.0 license, making it freely available for commercial and enterprise use.
- Sundar Pichai called it a "sweet spot between size + performance," saying it enables multi-step reasoning and agentic workflows while remaining small enough to run on a laptop.
- Ollama added same-day support for Gemma 4 12B via MLX, including integrations for Hermes Agent and Claude Code workflows.
Why it matters: For developers, Gemma 4 12B runs text, image, and audio on a 16GB laptop, bringing agentic multimodal workflows to local hardware. Google releasing this under Apache 2.0 — while the VentureBeat framing notes most open-source providers are chasing larger models — gives enterprises a free, self-hosted alternative to cloud multimodal APIs.
