Gemma 4 26B A4B
Gemma 4 26B A4B is an open-weight language model from Google DeepMind, released on April 2, 2026. It is a mixture-of-experts model with 25.2B total parameters, of which 3.8B are active for each token. Google DeepMind lists a context window of 256K tokens. The model accepts text and image input.
- Developer
- Google DeepMind
- Released
- April 2, 2026
- Total parameters
- 25.2B
- Active parameters per token
- 3.8B
- Architecture
- MoE
- Context window
- 256K
- Input
- Text, image
- License
- Apache 2.0
- Commercial use
- Yes
- Official weights, GB
- 51.6 GB
- Hugging Face repository
- huggingface.co
- Checked on
The weights are released under Apache 2.0, which allows commercial use.
The official weights on Hugging Face take about 51.6 GB in BF16. Running the model needs at least that much memory across GPUs and system RAM, plus room for the context cache; quantized versions need less.
- Total parameters, billions
- 25.2B
- Active parameters, billions
- 3.8B
- Context window, tokens
- 256K
- License type
- Apache 2.0
- License conditions
- Apache 2.0: keep the license and notices; no limits on field of use or number of users. Gemma 4 is the first Gemma generation under Apache 2.0; earlier versions used the Gemma Terms of Use.
- Weights precision
- BF16
- Official quantized or alternative versions
- google/gemma-4-26B-A4B-it-qat-q4_0-gguf; google/gemma-4-26B-A4B-it-qat-q4_0-unquantized
- License text
- ai.google.dev
- Sources
Page URL Hugging Face model card https://huggingface.co/google/gemma-4-31B-it/raw/main/README.md blog.google announcement https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/ opensource.googleblog.com announcement https://opensource.googleblog.com/2026/03/gemma-4-expanding-the-gemmaverse-with-apache-20.html Hugging Face API https://huggingface.co/api/models?author=google&search=gemma-4 Hugging Face weights index https://huggingface.co/google/gemma-4-26B-A4B-it/raw/main/model.safetensors.index.json