Gemma 4 E4B
Gemma 4 E4B is an open-weight language model from Google DeepMind, released on April 2, 2026. It is a dense model with 8B parameters. Google describes it as having 4.5B effective parameters. Google DeepMind lists a context window of 128K tokens. The model accepts text, image and audio input.
- Developer
- Google DeepMind
- Released
- April 2, 2026
- Total parameters
- 8B
- Architecture
- Dense
- Context window
- 128K
- Input
- Text, image, audio
- License
- Apache 2.0
- Commercial use
- Yes
- Official weights, GB
- 16 GB
- Hugging Face repository
- huggingface.co
- Checked on
The weights are released under Apache 2.0, which allows commercial use.
The official weights on Hugging Face take about 16.0 GB in BF16. Running the model needs at least that much memory across GPUs and system RAM, plus room for the context cache; quantized versions need less.
- Total parameters, billions
- 8B
- Context window, tokens
- 128K
- License type
- Apache 2.0
- License conditions
- Apache 2.0: keep the license and notices; no limits on field of use or number of users. Gemma 4 is the first Gemma generation under Apache 2.0; earlier versions used the Gemma Terms of Use.
- Weights precision
- BF16
- Official quantized or alternative versions
- google/gemma-4-E4B-it-qat-q4_0-gguf; google/gemma-4-E4B-it-qat-q4_0-unquantized; google/gemma-4-E4B-it-qat-w4a16-ct; google/gemma-4-E4B-it-qat-mobile-ct; google/gemma-4-E4B-it-qat-mobile-transformers
- License text
- ai.google.dev
- Sources
Page URL Hugging Face model card https://huggingface.co/google/gemma-4-31B-it/raw/main/README.md blog.google announcement https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/ opensource.googleblog.com announcement https://opensource.googleblog.com/2026/03/gemma-4-expanding-the-gemmaverse-with-apache-20.html Hugging Face API https://huggingface.co/api/models?author=google&search=gemma-4