Gemma 4 12B
Gemma 4 12B is an open-weight language model from Google DeepMind, released on June 3, 2026. It is a dense model with 11.95B parameters. Google DeepMind lists a context window of 256K tokens. The model accepts text, image and audio input.
- Developer
- Google DeepMind
- Released
- June 3, 2026
- Total parameters
- 11.95B
- Architecture
- Dense
- Context window
- 256K
- Input
- Text, image, audio
- License
- Apache 2.0
- Commercial use
- Yes
- Official weights, GB
- 23.9 GB
- Hugging Face repository
- huggingface.co
- Checked on
The weights are released under Apache 2.0, which allows commercial use.
The official weights on Hugging Face take about 23.9 GB in BF16. Running the model needs at least that much memory across GPUs and system RAM, plus room for the context cache; quantized versions need less.
- Total parameters, billions
- 11.95B
- Context window, tokens
- 256K
- License type
- Apache 2.0
- License conditions
- Apache 2.0: keep the license and notices; no limits on field of use or number of users. Gemma 4 is the first Gemma generation under Apache 2.0; earlier versions used the Gemma Terms of Use.
- Weights precision
- BF16
- Official quantized or alternative versions
- google/gemma-4-12B-it-qat-q4_0-gguf; google/gemma-4-12B-it-qat-q4_0-unquantized; google/gemma-4-12B-it-qat-w4a16-ct
- License text
- ai.google.dev
- Sources
Page URL Hugging Face model card https://huggingface.co/google/gemma-4-31B-it/raw/main/README.md blog.google announcement https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/ opensource.googleblog.com announcement https://opensource.googleblog.com/2026/03/gemma-4-expanding-the-gemmaverse-with-apache-20.html Hugging Face API https://huggingface.co/api/models?author=google&search=gemma-4 blog.google announcement https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12b/