How to Run Gemma 4 Locally
Google released Gemma 4 on April 2, 2026, and this time the license is plain Apache 2.0. The family has four sizes you will actually run on your own hardware. E4B fits a phone or an 8 GB laptop. The 12B, added on June 3, reads images and audio without separate encoders. The 26B-A4B is a mixture of experts with 3.8B active parameters. The 31B is the dense flagship. This guide explains the three design choices that set each model's memory bill, gives every quant size, and tells you which build your machine runs. Everything here was checked against the repositories on October 2, 2026. My local speed run was blocked by download bandwidth, so the speed section gives computed ceilings and credits published runs.
The family and its dates
Google's launch post went up on April 2, 2026, with four models: E2B, E4B, the 26B mixture of experts, and the 31B dense. Ollama 0.20.0, llama.cpp, and the ggml-org GGUF repositories had support on the same day. The 12B followed on June 3, aimed at laptops with 16 GB of memory. Each size ships as a pre-trained base and an instruction-tuned -it checkpoint. This guide covers the -it models.
One date trap is worth knowing. The Hugging Face API lists the E4B repository as created on March 2 and the 31B on March 11. The 12B repository dates to May 23. Google created those repositories before the launches, so the API date marks repository creation, not public release. The public dates are April 2 for the first four sizes and June 3 for the 12B.
Later releases added companion checkpoints. Google announced small "assistant" drafter models for speculative decoding on May 5. It announced quantization-aware training (QAT) checkpoints on June 5, as Q4_0 GGUFs and as w4a16 weights for vLLM. The previous generation is covered in the Gemma 3 guide. The jump is large: Google's card shows the 31B at 85.2% on MMLU Pro, against 67.6% for Gemma 3 27B.
The license changed
Gemma 1, 2, and 3 shipped under the Gemma Terms of Use, a custom license with a prohibited-use policy attached. Gemma 4 does not. Every model card points to ai.google.dev/gemma/docs/gemma_4_license, which redirects to the unmodified text of the Apache License 2.0, last updated April 1, 2026. I found no separate use policy linked from the cards or the license page.
In practice, this means you can use, modify, fine-tune, and redistribute the weights commercially. You must keep the license text and attribution notices, and mark files you change. The patent grant ends for anyone who sues over the work. Apache 2.0 does not grant trademark rights, so you cannot brand your product as Gemma. The QAT, drafter, and GGUF repositories from Google and ggml-org carry the same license tag.
Architecture from config.json
All four are Gemma4ForConditionalGeneration models, except the 12B, which uses a gemma4_unified_text decoder. The values below are from the raw config.json files and the Hugging Face safetensors metadata.
| Field | E4B | 12B | 26B-A4B | 31B |
|---|---|---|---|---|
| Parameters stored | 8.00B | 11.96B | 25.81B | 31.27B |
| Card label | 4.5B effective | 11.95B | 3.8B active | 30.7B |
| Layers | 42 | 48 | 30 | 60 |
| Global layers | 7 (every 6th) | 8 | 5 | 10 |
| Sliding window | 512 | 1,024 | 1,024 | 1,024 |
| Hidden size | 2,560 | 3,840 | 2,816 | 5,376 |
| Query heads | 8 | 16 | 16 | 32 |
| KV heads, local / global | 2 / 2 | 8 / 1 | 8 / 2 | 16 / 4 |
| Head dim, local / global | 256 / 512 | 256 / 512 | 256 / 512 | 256 / 512 |
| Keys equal values (global) | no | yes | yes | yes |
| KV-shared tail layers | 18 | 0 | 0 | 0 |
| Per-layer embedding width | 256 | none | none | none |
| Experts | 128 routed, top 8, plus a dense MLP | |||
| Context | 131,072 | 262,144 | 262,144 | 262,144 |
| Vocabulary | 262,144 | 262,144 | 262,144 | 262,144 |
| Inputs | text, image, audio | text, image, audio | text, image | text, image |
Three shared traits matter for memory. Every model alternates five sliding-window layers with one global layer, and the last layer is always global. Global layers use a 512-wide head and rotate only a quarter of it, which the card calls proportional RoPE. In the 12B, 26B, and 31B, the global layers also compute keys and values from one projection. If you want the background on any of these parts, the transformer internals guide covers them from scratch.
Per-layer embeddings
The "E" in E4B means effective. The model stores 8.0B parameters, but the card counts 4.5B. The difference is two lookup tables. The first is the normal token embedding, 262,144 rows by 2,560 values, or 0.67B parameters. The second is new: a per-layer embedding (PLE) table, which the GGUF names per_layer_token_embd. It holds 262,144 rows of 10,752 values each, which is 2.82B parameters.
Each row belongs to one token. The 10,752 values split into 42 slices of 256, one slice per decoder layer. In layer n, a small gate projects the hidden state down to 256 values and multiplies it by the token's slice. Another projection takes the result back up. So each layer gets its own learned signal about which token it is processing, without growing the hidden size.
The table is large but it is never multiplied. A new token reads one row: 10,752 values, about 6 KB at Q4_0. That changes where it should live. The llama.cpp loader places input-layer tensors on the CPU, and per_layer_token_embd is one of them. So the table stays in host memory even with every layer on the GPU. Pick a token and a placement to see what it costs:
The token ids are real entries from the Gemma 4 vocabulary. Each token reads one row of the table, and the row splits into one 256-value slice per layer.
model
GGUF type
where the table lives
per_layer_token_embd
Table sizes: 262,144 rows × layers × 256 values, at 2 bytes (BF16), 34 bytes per 32 values (Q8_0), or 18 bytes per 32 values (Q4_0). The GGUF type of the table matches the file type in the ggml-org builds, which I checked in each file header. "Rest of model" is the ggml-org file size minus the table.
The arithmetic explains the "effective" label. On a phone or a small GPU, the 2.82B table can sit in slower memory because each token touches 0.0004% of it. At 30 tokens per second, E4B pulls about 0.18 MB per second from the Q4_0 table. Even a slow storage path keeps up. The rest of the model, about 3.0 GB at Q4_0, is what the accelerator must hold and read for every token.
Two other details complete the E-series design. The tied token embedding doubles as the output head, so all 0.67B of its parameters are read on every token, unlike the PLE table. And the last 18 of E4B's 42 layers compute no keys or values at all. They reuse the cache of layer 22, the last local layer before them, or layer 23, the last global one. The GGUF confirms it: blk.24 has attn_q but no attn_k or attn_v. E2B does the same with 20 of its 35 layers.
Local and global attention
Five of every six layers are sliding-window layers. Each attends only to the most recent 512 tokens in E4B, or 1,024 in the larger models. The sixth layer is global and attends to the whole prompt. A local layer's cache is capped at its window, so only the global layers keep a cache that grows with the conversation. The KV cache arithmetic, at 2 bytes per value:
- 31B: 10 global layers × 4 KV heads × 512 × 2 (keys and values) × 2 bytes = 81,920 bytes per token. The 50 local layers hold a fixed 50 × 16 × 256 × 2 × 2 × 1,024 = 838.9 MB.
- 26B-A4B: 5 × 2 × 512 × 4 = 20,480 bytes per token, plus 209.7 MB of local cache.
- 12B: 8 × 1 × 512 × 4 = 16,384 bytes per token, plus 335.5 MB of local cache.
- E4B: only the first 24 layers own a cache. Its 4 global layers give 4 × 2 × 512 × 4 = 16,384 bytes per token, and its 20 local layers hold 21.0 MB.
The "keys equal values" flag does not halve these numbers. In the Transformers code, the value tensor is the same projection as the key, but it skips RoPE and gets its own normalization. So the cache still stores two different tensors. The llama.cpp cache also keeps a little more than one window per local layer, so treat the local figures as minimums.
| Model | Per token | Local, fixed | At 32K | At 128K | At 256K |
|---|---|---|---|---|---|
| E4B | 16 KB | 21 MB | 0.56 GB | 2.17 GB | above max |
| 12B | 16 KB | 336 MB | 0.87 GB | 2.48 GB | 4.63 GB |
| 26B-A4B | 20 KB | 210 MB | 0.88 GB | 2.89 GB | 5.58 GB |
| 31B | 80 KB | 839 MB | 3.52 GB | 11.58 GB | 22.31 GB |
The memory saving has a cost in reach. A local layer cannot read a token that is more than one window back. Long-range facts reach the newest token only through the global layers, which then carry them forward in the hidden state for the local layers above. Drag the prompt length and tap a part of the prompt to see which layers can read it directly:
One coding-agent prompt, laid out on a token ruler. The bracket at the right end is the local window, measured back from the newest token. The sizes of the named parts are an example layout.
model
layers, first to last
Layer types come from layer_types in each config.json. Cache = per-token global bytes × prompt length + the fixed local cache, at 2 bytes per value. E4B caps at 131,072 tokens. Distance is counted from the end of the selected part to the newest token.
The ruler shows why Gemma 4's long-context scores trail its short-context ones. On Google's 8-needle MRCR test at 128K, the 31B scores 66.4%, but the 26B-A4B with 5 global layers scores 44.1% and E4B scores 25.4%. Fewer global layers means fewer direct paths to a far fact. If your workload retrieves details from deep in a long prompt, the 31B is the safer pick. The cache table above tells you what that costs.
The 26B-A4B experts
Each of the 30 layers holds 128 experts, and a router picks 8 per token. Each expert is a small gated MLP with an intermediate size of 704. Beside the experts, every layer also runs a dense MLP with an intermediate size of 2,112, which the card counts as the shared expert. The arithmetic from config.json:
- Routed experts: 3 matrices × 2,816 × 704 = 5.95M per expert. Across 128 experts and 30 layers, that is 22.84B, or 88% of the model.
- Read per token: 8 of 128 experts is 1.43B. Add the dense MLPs (0.54B), attention (1.11B), the router, and the tied output head (0.74B). The total is about 3.8B, which matches the card.
The consequence is that the 26B-A4B reads about as many bytes per token as E4B, but it must keep all 25.8B parameters resident. Decode speed follows bytes read, so the two should decode at similar rates on one machine. The 31B reads eight times as much per token. The speed ceilings below show the pattern. On a GPU that cannot hold all 17 GB, the llama.cpp flag --n-cpu-moe keeps the experts of some layers in system RAM. During decode, the CPU computes the selected experts for those layers, so their weights stay in RAM.
One thing the MoE does not do well is speculative decoding. The author of the llama.cpp drafter pull request measured more than 2x for the dense models on a DGX Spark. The same author saw no speedup on the 26B-A4B. A drafter helps when the target is slow, and this target is already cheap per token.
The encoder-free 12B
The other Gemma 4 models put images through a separate vision tower: about 150M parameters for E4B and 550M for the 26B and 31B. E4B also runs audio through a 12-layer, 300M-parameter conformer. The 12B removes both. Per Google's developer guide, one 35M-parameter embedder projects raw 48×48 pixel patches into the decoder. Audio at 16 kHz is cut into 40 ms frames, and a linear layer projects each frame directly. Its decoder has the same structure as the 31B.
For local use this has two effects. The projector file is tiny: 0.18 GB in BF16, against 1.20 GB for the 31B's. And images cost decoder compute, so the visual token budget matters more. Gemma 4 accepts budgets of 70, 140, 280, 560, or 1,120 tokens per image. Google's own video demo used 313 frames at a budget of 70, which is 21,910 tokens before the prompt and audio. At 16 KB per token that is 0.36 GB of cache, plus the 0.34 GB local cache.
Every quant, with sizes
Sizes are sums of the files from the Hugging Face tree API on October 2, 2026, in decimal gigabytes. They exclude the vision and audio projector, which you add if you want image or audio input. The Google QAT builds are Q4_0 files trained to be quantized, so they lose less quality than a plain Q4_0 conversion. Google's card says they preserve "similar quality to bfloat16".
E4B
| Build | Size | Notes |
|---|---|---|
unsloth UD-Q2_K_XL | 3.76 GB | smallest sane option, for 4 to 6 GB devices |
ggml-org Q4_0 | 4.59 GB | official build, the repo lists the QAT checkpoint as a source |
unsloth Q4_K_M | 4.98 GB | common default |
google qat-q4_0-gguf | 5.15 GB | Google's own QAT file |
unsloth Q6_K | 7.07 GB | |
ggml-org Q8_0 | 8.03 GB | near lossless |
ggml-org BF16 | 15.05 GB | reference |
| MLX 4-bit / 8-bit (lmstudio-community) | 6.83 / 8.94 GB | MLX keeps more tensors at higher precision |
mmproj BF16 / Q8_0 | 0.99 / 0.56 GB | vision and audio encoders |
12B
| Build | Size | Notes |
|---|---|---|
unsloth UD-Q3_K_XL | 6.02 GB | for 8 GB GPUs |
google qat-q4_0 | 6.98 GB | the best 4-bit pick |
unsloth Q4_K_M | 7.12 GB | |
ggml-org Q4_0 | 7.22 GB | |
unsloth Q6_K | 9.79 GB | |
ggml-org Q8_0 | 12.67 GB | |
ggml-org BF16 | 23.83 GB | reference |
google qat-w4a16-ct | 10.26 GB | vLLM, compressed-tensors |
| RedHatAI FP8-Dynamic / cyankiwi AWQ-INT4 | 15.04 / 11.22 GB | vLLM |
| MLX 4-bit / 8-bit (lmstudio-community) | 6.74 / 12.72 GB |
26B-A4B
| Build | Size | Notes |
|---|---|---|
unsloth UD-IQ2_M | 10.01 GB | 12 GB cards, with a quality cost |
unsloth UD-Q3_K_XL | 12.91 GB | 16 GB cards |
google qat-q4_0 | 14.44 GB | the best 4-bit pick |
ggml-org Q4_0 | 14.62 GB | |
unsloth MXFP4_MOE | 16.55 GB | experts in MXFP4 |
unsloth UD-Q4_K_M | 16.95 GB | common default |
unsloth UD-Q6_K | 23.17 GB | |
ggml-org Q8_0 | 26.86 GB | |
ggml-org BF16 | 50.51 GB | one 80 GB GPU, per Google |
| RedHatAI FP8-dynamic / cyankiwi AWQ-4bit | 28.64 / 17.19 GB | vLLM |
| unsloth NVFP4 | 16.91 GB | Blackwell GPUs |
| MLX 4-bit / 8-bit (mlx-community) | 15.34 / 27.95 GB |
31B
| Build | Size | Notes |
|---|---|---|
unsloth UD-Q2_K_XL | 11.77 GB | 16 GB cards, with a quality cost |
unsloth UD-Q3_K_XL | 15.38 GB | 20 GB cards |
google qat-q4_0 | 17.65 GB | the best 4-bit pick |
ggml-org Q4_0 | 17.99 GB | |
unsloth Q4_K_M | 18.32 GB | |
unsloth Q6_K | 25.20 GB | |
ggml-org Q8_0 | 32.64 GB | |
ggml-org BF16 | 61.41 GB | one 80 GB GPU, per Google |
google qat-w4a16-ct | 23.27 GB | vLLM |
| RedHatAI FP8-block / cyankiwi AWQ-4bit | 33.26 / 20.90 GB | vLLM |
| MLX 4-bit / 8-bit (lmstudio-community) | 18.41 / 33.76 GB |
What fits your machine
Each verdict below adds the weights, the KV cache at 32,768 tokens from the table above, and 1 GB for buffers. Usable memory is 75% of unified memory on a Mac, at the default wired limit. On a GPU it is 90 to 96% of the card, with small cards losing the larger share to the driver and display. Add the projector if you need images.
| Machine | Usable | Best fit at 32K | Need |
|---|---|---|---|
| 8 GB laptop GPU | 7.2 GB | E4B Q4_K_M | 6.5 GB |
| 12 GB GPU, 16 GB Mac | 11 to 12 GB | 12B Q4_K_M, or E4B Q8_0 | 9.0 GB |
| 16 GB GPU | 15 GB | 12B Q8_0 | 14.5 GB |
| 24 GB GPU | 22.5 GB | 26B-A4B UD-Q4_K_M | 18.8 GB |
| 32 GB Mac | 24 GB | 31B Q4_K_M, or 26B-A4B for speed | 22.8 GB |
| 32 GB GPU | 30.5 GB | 26B-A4B Q8_0, or 31B Q4_K_M | 28.7 GB |
| 48 GB Mac | 36 GB | 26B-A4B Q8_0 with long context | 28.7 GB |
| 64 GB Mac, 48 GB GPU | 46 to 48 GB | 31B Q8_0 | 37.2 GB |
Two edge cases deserve a note. A 24 GB GPU misses the 31B Q4_K_M by 0.3 GB at 32K, but it fits at 16K, or at 32K with a q8_0 cache. And the 26B-A4B runs on an 8 GB or 12 GB GPU with 32 GB of system RAM if you offload experts. The non-expert part of the Q4 model is roughly 2 GB. The GPU holds that part and the cache, and RAM holds the experts you offload.
Ollama
Ollama added Gemma 4 in v0.20.0, released the same day as the models. The library tags carry their own sizes, and the default gemma4 tag is E4B.
# E4B, the default tag, 6.6 GB
ollama run gemma4:e4b
# 12B, 26B-A4B, and 31B at Q4_K_M: 8.0, 18, and 20 GB
ollama run gemma4:12b
ollama run gemma4:26b
ollama run gemma4:31b
# Google's QAT builds: 6.1, 7.2, 16, and 19 GB
ollama run gemma4:26b-a4b-it-qat
The library also has -mlx, -mxfp8, and -nvfp4 tags for Apple Silicon, and a 26b-a4b-it-mtp-q4_K_M tag that bundles the drafter. Set the context with /set parameter num_ctx 32768 in a session, or with OLLAMA_CONTEXT_LENGTH for the server. The default is smaller than the model maximum.
llama.cpp
Support landed in PR #21309, "model: support gemma 4 (vision + moe, no audio)", merged on April 2 and tagged b8637. Audio input followed on April 12. The 12B needs build b9493 or later, from June 3, and drafter support needs b9549. Many fixes followed, such as a Gemma 4 tool-call parser on April 4 and a September 19 fix for required tool calls. Use a recent build.
# E4B on any 8 GB machine
llama-server -hf ggml-org/gemma-4-E4B-it-GGUF:Q4_0 \
--jinja -ngl 99 -fa on -c 32768 \
--temp 1.0 --top-p 0.95 --top-k 64
# 26B-A4B on a 24 GB GPU or a 32 GB Mac
llama-server -hf unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4_K_M \
--jinja -ngl 99 -fa on -c 65536 \
--temp 1.0 --top-p 0.95 --top-k 64
# 26B-A4B on a 12 GB GPU: experts of the first 20 layers stay in RAM
llama-server -hf unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4_K_M \
--jinja -ngl 99 -fa on -c 32768 --n-cpu-moe 20
# 31B on a 24 GB GPU: 16K context, or a q8_0 cache
llama-server -hf unsloth/gemma-4-31B-it-GGUF:Q4_K_M \
--jinja -ngl 99 -fa on -c 32768 -ctk q8_0 -ctv q8_0
Lower --n-cpu-moe until the model no longer fits, then step back up by one. With -hf, llama-server also downloads the mmproj file, which enables image input and, for E4B and the 12B, audio. Pass --no-mmproj to save that memory for text-only work. Always pass --jinja, because the chat template handles thinking and tool calls.
vLLM and SGLang
vLLM added Gemma 4 in v0.19.0 on April 3, with a requirement of transformers>=5.5.0. The 12B and drafter support arrived in v0.23.0 on June 15. The vLLM recipe shows the parsers:
vllm serve google/gemma-4-26B-A4B-it \
--max-model-len 32768 \
--gpu-memory-utilization 0.90 \
--enable-auto-tool-choice \
--reasoning-parser gemma4 \
--tool-call-parser gemma4 \
--limit-mm-per-prompt '{"image": 0, "audio": 0}'
Set --max-model-len every time. The default is the full 262,144 tokens, which reserves cache you probably do not need. The last flag skips multimodal profiling for text-only work. The BF16 26B-A4B needs an 80 GB card. On a 48 GB card, use the RedHatAI FP8 build. On 24 to 32 GB, use the AWQ build, or Google's qat-w4a16-ct weights for the 12B and 31B.
SGLang documents Gemma 4 in its cookbook with the same parser names. It selects the Triton attention backend on its own, because image tokens use bidirectional attention during prefill.
sglang serve --model-path google/gemma-4-E4B-it \
--reasoning-parser gemma4 \
--tool-call-parser gemma4 \
--host 0.0.0.0 --port 30000
Once several users or an agent fleet share the server, paged attention and batching change the economics. The vLLM internals guide explains why.
MLX on a Mac
mlx-lm added Gemma 4 text support in v0.31.2 on April 7. The 12B alias (gemma4_unified) merged on September 11 in PR #1386, after the last tagged release, so the 12B needs mlx-lm from the main branch.
pip install -U mlx-lm
mlx_lm.generate --model mlx-community/gemma-4-26b-a4b-it-4bit \
--max-tokens 1024 --temp 1.0 --top-p 0.95 \
--prompt "Explain the failing test in the log below."
# the 12B, until the next mlx-lm release
pip install -U git+https://github.com/ml-explore/mlx-lm.git
For image input on MLX, use mlx-vlm or LM Studio, which ships MLX builds for every size. For the GGUF path on a Mac, use the llama.cpp commands above. Metal runs them unchanged.
Drafter models
Google publishes a small "assistant" model for each size, from 0.16 GB for E4B to 0.94 GB for the 31B in BF16. The assistant shares the target's KV cache and proposes several tokens, and the target checks them in one pass. The E2B and E4B assistants also limit their output to about 4,000 candidate tokens through "centroids masking". The vLLM recipe says this cuts the output-head work by about 45x.
# vLLM
vllm serve google/gemma-4-31B-it --tensor-parallel-size 2 \
--max-model-len 8192 \
--speculative-config '{"model": "google/gemma-4-31B-it-assistant", "num_speculative_tokens": 4}'
# llama.cpp, with the drafter file from the same ggml-org repo
llama-server -m gemma-4-31B-it-Q4_0.gguf \
--model-draft mtp-gemma-4-31B-it-Q4_0.gguf \
--spec-type draft-mtp --spec-draft-n-max 4 \
--jinja -ngl 99 -fa on -c 32768
Match the drafter to the target. Google's QAT card says a QAT target needs a QAT assistant of the same precision. The ggml-org repositories list both the QAT and the plain assistant as sources, so pair files of the same type from the same repo. The vLLM recipe recommends 2 draft tokens for E2B, 4 for E4B and the 26B-A4B, and 4 to 8 for the 12B and 31B.
Thinking, sampling, tools
Sampling. Google gives one setting for every use: temperature=1.0, top_p=0.95, top_k=64. The ggml-org GGUFs store these as defaults in their metadata.
Thinking. Thinking is off by default. The chat template turns it on when you pass enable_thinking: true, which puts a <|think|> token at the start of the system turn. In llama.cpp, add --chat-template-kwargs '{"enable_thinking":true}'. Over an OpenAI-compatible API, send "chat_template_kwargs": {"enable_thinking": true}. In vLLM, --default-chat-template-kwargs sets it for every request. The thought arrives between <|channel>thought and <channel|>, and the gemma4 reasoning parser moves it to reasoning_content. With thinking off, the 12B, 26B, and 31B still emit an empty thought block, but E2B and E4B do not.
History. Do not send old thoughts back. The card says previous turns must contain only the final answer, except for tool-call turns, where the thinking stays.
Tools. Turns use <|turn> and <turn|>, and a call is written as <|tool_call>call:name{...}<tool_call|>. The template embedded in current GGUFs is dated July 9, 2026, and its header says it fixed "tool-calling loops, turn closures, and thinking content-ordering". If an agent loops on tool calls, check that your GGUF or engine carries that template.
Multimodal order. Put images before the text and audio after it. Use a visual budget of 70 or 140 for video frames, and 560 or 1,120 for documents and small print.
Speed: ceilings and published runs
I planned to measure E4B and the 26B-A4B on my M1 Max with 64 GB, using llama.cpp build 10330. Download bandwidth blocked the run. The E4B file reached about 7% and the 26B-A4B file about 2% before the time box ended, so this guide has no local measurements. The partial file headers did confirm two facts: both load as architecture gemma4 in that build, with 42 and 30 blocks. The Ollama E4B file also turned out to be a mixed quant that keeps attention at Q6 and Q8, not a plain Q4_K_M.
What I can give you is a ceiling. Decode reads the weights it uses once per token, so memory bandwidth divided by bytes read per token is the fastest any engine can go. Real engines land well under it, often at half or less. The bytes per token below use the ggml-org Q4_0 files:
- E4B: the file minus the PLE table, about 3.0 GB, because the table supplies only 6 KB per token.
- 12B: all 7.22 GB.
- 26B-A4B: about 2.2 GB. That is the non-expert weights plus 8 of 128 experts per layer, scaled from the 14.62 GB file by parameter share.
- 31B: all 17.99 GB.
| Ceiling, tokens/s | Bandwidth | E4B | 12B | 26B-A4B | 31B |
|---|---|---|---|---|---|
| DGX Spark | 273 GB/s | 91 | 38 | 124 | 15 |
| M1 Max | 400 GB/s | 133 | 55 | 182 | 22 |
| RTX 4090 | 1,008 GB/s | 336 | 140 | 458 | 56 |
These are upper bounds from arithmetic, not measurements. They also ignore the KV cache reads, which grow with context. The pattern is the useful part. The 26B-A4B has a higher ceiling than E4B, and the 31B ceiling is eight times lower than the 26B-A4B ceiling.
Two published runs give real data points. In llama.cpp PR #23398, its author ran the 31B on a DGX Spark at about 6 tokens per second without a drafter. With the drafter at 4 draft tokens, the same nine-prompt run took 120.65 seconds instead of 290.01, with 59% of drafts accepted. The PR does not state the quant. In PR #24282, its author ran the E4B QAT Q4_0 build on a Galaxy S26+ phone at 14.04 tokens per second. The drafter ran on the CPU, with 48% of drafts accepted.
Benchmarks, vendor-reported
These numbers come from Google's model card, for the instruction-tuned models, under Google's settings. No independent run is shown here.
| Benchmark | 31B | 26B-A4B | 12B | E4B | Gemma 3 27B |
|---|---|---|---|---|---|
| MMLU Pro | 85.2% | 82.6% | 77.2% | 69.4% | 67.6% |
| AIME 2026, no tools | 89.2% | 88.3% | 77.5% | 42.5% | 20.8% |
| LiveCodeBench v6 | 80.0% | 77.1% | 72.0% | 52.0% | 29.1% |
| GPQA Diamond | 84.3% | 82.3% | 78.8% | 58.6% | 42.4% |
| Tau2, average of 3 | 76.9% | 68.2% | 69.0% | 42.2% | 16.2% |
| MMMU Pro | 76.9% | 73.8% | 69.1% | 52.6% | 49.7% |
| MRCR v2, 8 needles, 128K | 66.4% | 44.1% | 43.4% | 25.4% | 13.5% |
The 26B-A4B stays within three points of the 31B on knowledge and math. It falls behind on agentic tool use (Tau2) and on long-context retrieval, which matches the reach figure above. The 12B beats the 26B-A4B on Tau2, so for a laptop agent it is the stronger pick per gigabyte. Google's chart also placed the 31B third among open models on the Arena text leaderboard as of April 1.
Failure modes and fixes
- "unknown model architecture: gemma4". Your llama.cpp is older than
b8637, or your 12B load predatesb9493. Update the build. In vLLM, the 12B needs v0.23.0 or later. - Out of memory at startup in vLLM. The default context is 262,144 tokens. Set
--max-model-lenand limit multimodal inputs you do not use. - Agent loops on tool calls. Old GGUFs embed an older template. Download a current GGUF, or pass the July 9 template with
--chat-template-file. - Thinking text in the answer. Start the server with the
gemma4reasoning parser, or use--jinjain llama.cpp. Strip old thoughts from the history you send. - No audio. Only E2B, E4B, and the 12B accept audio. In llama.cpp, the
mmprojfile must be loaded, and audio support needs an April 12 or later build. - The drafter does not speed up the 26B-A4B. This is expected. Use the drafter with the dense models.
- Long-context recall is weak on small models. E4B has 4 cache-owning global layers. Use the 31B for retrieval from long prompts.
- The 31B is slow on a laptop. It reads its full weight set for every token. Use the 26B-A4B unless you need the extra quality.
FAQ
Which Gemma 4 model should I run?
On 8 GB of memory, run E4B. On 12 to 16 GB, run the 12B at Q4. On a 24 GB GPU or a 32 GB Mac, run the 26B-A4B for speed or the 31B for quality. The 26B-A4B reads about as many bytes per token as E4B, so it decodes much faster than the 31B.
Is Gemma 4 open source?
The weights are under the standard Apache License 2.0. Gemma 1 to 3 used the custom Gemma Terms of Use. Gemma 4 has no separate prohibited-use addendum on its license page.
What does the E in E4B mean?
Effective parameters. E4B stores 8.0B parameters, but 2.82B of them are a per-layer embedding table. The model only looks up 10,752 values of it per token. By default, llama.cpp keeps that table in host memory.
When was Gemma 4 released?
Google announced Gemma 4 on April 2, 2026, with E2B, E4B, 26B-A4B, and 31B. The 12B followed on June 3, 2026. The Hugging Face repositories were created earlier, in March and May, before the weights went public.
Does Gemma 4 support speculative decoding?
Yes. Google publishes assistant drafter models for each size. vLLM takes them through --speculative-config, SGLang through --speculative-algorithm NEXTN, and llama.cpp through --spec-type draft-mtp from build b9549.
Keep reading