Google Releases Gemma 4 12B Local Multimodal AI

Get the Tech newsletter
Daily tech — startups, AI labs, chips, the launches that shape the next decade. Free.
- Google released Gemma 4 12B, an 11.95B-parameter open multimodal model with a unified encoder-free architecture that processes text, image, and audio inputs natively
- The model runs locally on devices with 16GB of VRAM or unified memory, bringing agentic reasoning, vision, and audio capabilities to standard laptops without cloud infrastructure
- Google DeepMind released Gemma 4 12B under an Apache 2.0 license, making the model freely available for developers to use, modify, and deploy commercially
- Google stated the model delivers "performance nearing our larger Gemma models" while maintaining a much smaller memory footprint, according to the company's announcement
- The release positions Google in the local/on-device AI market, explicitly contrasting with the broader industry trend toward increasingly larger AI models, as VentureBeat noted
- Gemma 4 12B is available through Hugging Face, Ollama, and LM Studio, with Google also publishing a developer guide and a visual architecture guide from researcher Michael Tschannen
Why it matters: Google is making advanced multimodal AI — combining text, vision, and audio processing — accessible on standard consumer hardware without cloud connections or expensive GPUs. The Apache 2.0 license means developers can build commercial applications on top of it for free, undercutting the assumption that capable AI requires massive infrastructure.
