How to Run Gemma 4 Locally

Google released Gemma 4 on April 2, 2026, and this time the license is plain Apache 2.0. The family has four sizes you will actually run on your own hardware. E4B fits a phone or an 8 GB laptop. The 12B, added on June 3, reads images and audio without separate encoders. The 26B-A4B is a mixture of experts with 3.8B active parameters. The 31B is the dense flagship. This guide explains the three design choices that set each model's memory bill, gives every quant size, and tells you which build your machine runs. Everything here was checked against the repositories on October 2, 2026. My local speed run was blocked by download bandwidth, so the speed section gives computed ceilings and credits published runs.

E4B 8.0B stored · 4.5B effective 12B dense · encoder-free 26B-A4B 128 experts · 3.8B active 31B dense license Apache 2.0 context 128K to 256K

The family and its dates

Google's launch post went up on April 2, 2026, with four models: E2B, E4B, the 26B mixture of experts, and the 31B dense. Ollama 0.20.0, llama.cpp, and the ggml-org GGUF repositories had support on the same day. The 12B followed on June 3, aimed at laptops with 16 GB of memory. Each size ships as a pre-trained base and an instruction-tuned -it checkpoint. This guide covers the -it models.

One date trap is worth knowing. The Hugging Face API lists the E4B repository as created on March 2 and the 31B on March 11. The 12B repository dates to May 23. Google created those repositories before the launches, so the API date marks repository creation, not public release. The public dates are April 2 for the first four sizes and June 3 for the 12B.

Later releases added companion checkpoints. Google announced small "assistant" drafter models for speculative decoding on May 5. It announced quantization-aware training (QAT) checkpoints on June 5, as Q4_0 GGUFs and as w4a16 weights for vLLM. The previous generation is covered in the Gemma 3 guide. The jump is large: Google's card shows the 31B at 85.2% on MMLU Pro, against 67.6% for Gemma 3 27B.

The license changed

Gemma 1, 2, and 3 shipped under the Gemma Terms of Use, a custom license with a prohibited-use policy attached. Gemma 4 does not. Every model card points to ai.google.dev/gemma/docs/gemma_4_license, which redirects to the unmodified text of the Apache License 2.0, last updated April 1, 2026. I found no separate use policy linked from the cards or the license page.

In practice, this means you can use, modify, fine-tune, and redistribute the weights commercially. You must keep the license text and attribution notices, and mark files you change. The patent grant ends for anyone who sues over the work. Apache 2.0 does not grant trademark rights, so you cannot brand your product as Gemma. The QAT, drafter, and GGUF repositories from Google and ggml-org carry the same license tag.

Architecture from config.json

All four are Gemma4ForConditionalGeneration models, except the 12B, which uses a gemma4_unified_text decoder. The values below are from the raw config.json files and the Hugging Face safetensors metadata.

FieldE4B12B26B-A4B31B
Parameters stored8.00B11.96B25.81B31.27B
Card label4.5B effective11.95B3.8B active30.7B
Layers42483060
Global layers7 (every 6th)8510
Sliding window5121,0241,0241,024
Hidden size2,5603,8402,8165,376
Query heads8161632
KV heads, local / global2 / 28 / 18 / 216 / 4
Head dim, local / global256 / 512256 / 512256 / 512256 / 512
Keys equal values (global)noyesyesyes
KV-shared tail layers18000
Per-layer embedding width256nonenonenone
Experts128 routed, top 8, plus a dense MLP
Context131,072262,144262,144262,144
Vocabulary262,144262,144262,144262,144
Inputstext, image, audiotext, image, audiotext, imagetext, image

Three shared traits matter for memory. Every model alternates five sliding-window layers with one global layer, and the last layer is always global. Global layers use a 512-wide head and rotate only a quarter of it, which the card calls proportional RoPE. In the 12B, 26B, and 31B, the global layers also compute keys and values from one projection. If you want the background on any of these parts, the transformer internals guide covers them from scratch.

Per-layer embeddings

The "E" in E4B means effective. The model stores 8.0B parameters, but the card counts 4.5B. The difference is two lookup tables. The first is the normal token embedding, 262,144 rows by 2,560 values, or 0.67B parameters. The second is new: a per-layer embedding (PLE) table, which the GGUF names per_layer_token_embd. It holds 262,144 rows of 10,752 values each, which is 2.82B parameters.

Each row belongs to one token. The 10,752 values split into 42 slices of 256, one slice per decoder layer. In layer n, a small gate projects the hidden state down to 256 values and multiplies it by the token's slice. Another projection takes the result back up. So each layer gets its own learned signal about which token it is processing, without growing the hidden size.

The table is large but it is never multiplied. A new token reads one row: 10,752 values, about 6 KB at Q4_0. That changes where it should live. The llama.cpp loader places input-layer tensors on the CPU, and per_layer_token_embd is one of them. So the table stays in host memory even with every layer on the GPU. Pick a token and a placement to see what it costs:

Fig. 1 · the per-layer embedding slicer

The token ids are real entries from the Gemma 4 vocabulary. Each token reads one row of the table, and the row splits into one 256-value slice per layer.

model

GGUF type

where the table lives

decode speed = 30 tokens/s

per_layer_token_embd

slice for a local layerslice for a global layerlayer that reuses another layer's KV cache
table resident0
read for this token0
read per second0

Table sizes: 262,144 rows × layers × 256 values, at 2 bytes (BF16), 34 bytes per 32 values (Q8_0), or 18 bytes per 32 values (Q4_0). The GGUF type of the table matches the file type in the ggml-org builds, which I checked in each file header. "Rest of model" is the ggml-org file size minus the table.

The arithmetic explains the "effective" label. On a phone or a small GPU, the 2.82B table can sit in slower memory because each token touches 0.0004% of it. At 30 tokens per second, E4B pulls about 0.18 MB per second from the Q4_0 table. Even a slow storage path keeps up. The rest of the model, about 3.0 GB at Q4_0, is what the accelerator must hold and read for every token.

Two other details complete the E-series design. The tied token embedding doubles as the output head, so all 0.67B of its parameters are read on every token, unlike the PLE table. And the last 18 of E4B's 42 layers compute no keys or values at all. They reuse the cache of layer 22, the last local layer before them, or layer 23, the last global one. The GGUF confirms it: blk.24 has attn_q but no attn_k or attn_v. E2B does the same with 20 of its 35 layers.

Local and global attention

Five of every six layers are sliding-window layers. Each attends only to the most recent 512 tokens in E4B, or 1,024 in the larger models. The sixth layer is global and attends to the whole prompt. A local layer's cache is capped at its window, so only the global layers keep a cache that grows with the conversation. The KV cache arithmetic, at 2 bytes per value:

The "keys equal values" flag does not halve these numbers. In the Transformers code, the value tensor is the same projection as the key, but it skips RoPE and gets its own normalization. So the cache still stores two different tensors. The llama.cpp cache also keeps a little more than one window per local layer, so treat the local figures as minimums.

ModelPer tokenLocal, fixedAt 32KAt 128KAt 256K
E4B16 KB21 MB0.56 GB2.17 GBabove max
12B16 KB336 MB0.87 GB2.48 GB4.63 GB
26B-A4B20 KB210 MB0.88 GB2.89 GB5.58 GB
31B80 KB839 MB3.52 GB11.58 GB22.31 GB

The memory saving has a cost in reach. A local layer cannot read a token that is more than one window back. Long-range facts reach the newest token only through the global layers, which then carry them forward in the hidden state for the local layers above. Drag the prompt length and tap a part of the prompt to see which layers can read it directly:

Fig. 2 · the attention reach ruler

One coding-agent prompt, laid out on a token ruler. The bracket at the right end is the local window, measured back from the newest token. The sizes of the named parts are an example layout.

model

prompt length = 131,072 tokens
token 0newest token

layers, first to last

global layer that reads itlocal layer that reads itreuses another layer's cachecannot read it directly
layers that read it directly0
distance to newest token0
KV cache at this length0

Layer types come from layer_types in each config.json. Cache = per-token global bytes × prompt length + the fixed local cache, at 2 bytes per value. E4B caps at 131,072 tokens. Distance is counted from the end of the selected part to the newest token.

The ruler shows why Gemma 4's long-context scores trail its short-context ones. On Google's 8-needle MRCR test at 128K, the 31B scores 66.4%, but the 26B-A4B with 5 global layers scores 44.1% and E4B scores 25.4%. Fewer global layers means fewer direct paths to a far fact. If your workload retrieves details from deep in a long prompt, the 31B is the safer pick. The cache table above tells you what that costs.

The 26B-A4B experts

Each of the 30 layers holds 128 experts, and a router picks 8 per token. Each expert is a small gated MLP with an intermediate size of 704. Beside the experts, every layer also runs a dense MLP with an intermediate size of 2,112, which the card counts as the shared expert. The arithmetic from config.json:

The consequence is that the 26B-A4B reads about as many bytes per token as E4B, but it must keep all 25.8B parameters resident. Decode speed follows bytes read, so the two should decode at similar rates on one machine. The 31B reads eight times as much per token. The speed ceilings below show the pattern. On a GPU that cannot hold all 17 GB, the llama.cpp flag --n-cpu-moe keeps the experts of some layers in system RAM. During decode, the CPU computes the selected experts for those layers, so their weights stay in RAM.

One thing the MoE does not do well is speculative decoding. The author of the llama.cpp drafter pull request measured more than 2x for the dense models on a DGX Spark. The same author saw no speedup on the 26B-A4B. A drafter helps when the target is slow, and this target is already cheap per token.

The encoder-free 12B

The other Gemma 4 models put images through a separate vision tower: about 150M parameters for E4B and 550M for the 26B and 31B. E4B also runs audio through a 12-layer, 300M-parameter conformer. The 12B removes both. Per Google's developer guide, one 35M-parameter embedder projects raw 48×48 pixel patches into the decoder. Audio at 16 kHz is cut into 40 ms frames, and a linear layer projects each frame directly. Its decoder has the same structure as the 31B.

For local use this has two effects. The projector file is tiny: 0.18 GB in BF16, against 1.20 GB for the 31B's. And images cost decoder compute, so the visual token budget matters more. Gemma 4 accepts budgets of 70, 140, 280, 560, or 1,120 tokens per image. Google's own video demo used 313 frames at a budget of 70, which is 21,910 tokens before the prompt and audio. At 16 KB per token that is 0.36 GB of cache, plus the 0.34 GB local cache.

Every quant, with sizes

Sizes are sums of the files from the Hugging Face tree API on October 2, 2026, in decimal gigabytes. They exclude the vision and audio projector, which you add if you want image or audio input. The Google QAT builds are Q4_0 files trained to be quantized, so they lose less quality than a plain Q4_0 conversion. Google's card says they preserve "similar quality to bfloat16".

E4B

BuildSizeNotes
unsloth UD-Q2_K_XL3.76 GBsmallest sane option, for 4 to 6 GB devices
ggml-org Q4_04.59 GBofficial build, the repo lists the QAT checkpoint as a source
unsloth Q4_K_M4.98 GBcommon default
google qat-q4_0-gguf5.15 GBGoogle's own QAT file
unsloth Q6_K7.07 GB
ggml-org Q8_08.03 GBnear lossless
ggml-org BF1615.05 GBreference
MLX 4-bit / 8-bit (lmstudio-community)6.83 / 8.94 GBMLX keeps more tensors at higher precision
mmproj BF16 / Q8_00.99 / 0.56 GBvision and audio encoders

12B

BuildSizeNotes
unsloth UD-Q3_K_XL6.02 GBfor 8 GB GPUs
google qat-q4_06.98 GBthe best 4-bit pick
unsloth Q4_K_M7.12 GB
ggml-org Q4_07.22 GB
unsloth Q6_K9.79 GB
ggml-org Q8_012.67 GB
ggml-org BF1623.83 GBreference
google qat-w4a16-ct10.26 GBvLLM, compressed-tensors
RedHatAI FP8-Dynamic / cyankiwi AWQ-INT415.04 / 11.22 GBvLLM
MLX 4-bit / 8-bit (lmstudio-community)6.74 / 12.72 GB

26B-A4B

BuildSizeNotes
unsloth UD-IQ2_M10.01 GB12 GB cards, with a quality cost
unsloth UD-Q3_K_XL12.91 GB16 GB cards
google qat-q4_014.44 GBthe best 4-bit pick
ggml-org Q4_014.62 GB
unsloth MXFP4_MOE16.55 GBexperts in MXFP4
unsloth UD-Q4_K_M16.95 GBcommon default
unsloth UD-Q6_K23.17 GB
ggml-org Q8_026.86 GB
ggml-org BF1650.51 GBone 80 GB GPU, per Google
RedHatAI FP8-dynamic / cyankiwi AWQ-4bit28.64 / 17.19 GBvLLM
unsloth NVFP416.91 GBBlackwell GPUs
MLX 4-bit / 8-bit (mlx-community)15.34 / 27.95 GB

31B

BuildSizeNotes
unsloth UD-Q2_K_XL11.77 GB16 GB cards, with a quality cost
unsloth UD-Q3_K_XL15.38 GB20 GB cards
google qat-q4_017.65 GBthe best 4-bit pick
ggml-org Q4_017.99 GB
unsloth Q4_K_M18.32 GB
unsloth Q6_K25.20 GB
ggml-org Q8_032.64 GB
ggml-org BF1661.41 GBone 80 GB GPU, per Google
google qat-w4a16-ct23.27 GBvLLM
RedHatAI FP8-block / cyankiwi AWQ-4bit33.26 / 20.90 GBvLLM
MLX 4-bit / 8-bit (lmstudio-community)18.41 / 33.76 GB

What fits your machine

Each verdict below adds the weights, the KV cache at 32,768 tokens from the table above, and 1 GB for buffers. Usable memory is 75% of unified memory on a Mac, at the default wired limit. On a GPU it is 90 to 96% of the card, with small cards losing the larger share to the driver and display. Add the projector if you need images.

MachineUsableBest fit at 32KNeed
8 GB laptop GPU7.2 GBE4B Q4_K_M6.5 GB
12 GB GPU, 16 GB Mac11 to 12 GB12B Q4_K_M, or E4B Q8_09.0 GB
16 GB GPU15 GB12B Q8_014.5 GB
24 GB GPU22.5 GB26B-A4B UD-Q4_K_M18.8 GB
32 GB Mac24 GB31B Q4_K_M, or 26B-A4B for speed22.8 GB
32 GB GPU30.5 GB26B-A4B Q8_0, or 31B Q4_K_M28.7 GB
48 GB Mac36 GB26B-A4B Q8_0 with long context28.7 GB
64 GB Mac, 48 GB GPU46 to 48 GB31B Q8_037.2 GB

Two edge cases deserve a note. A 24 GB GPU misses the 31B Q4_K_M by 0.3 GB at 32K, but it fits at 16K, or at 32K with a q8_0 cache. And the 26B-A4B runs on an 8 GB or 12 GB GPU with 32 GB of system RAM if you offload experts. The non-expert part of the Q4 model is roughly 2 GB. The GPU holds that part and the cache, and RAM holds the experts you offload.

Ollama

Ollama added Gemma 4 in v0.20.0, released the same day as the models. The library tags carry their own sizes, and the default gemma4 tag is E4B.

# E4B, the default tag, 6.6 GB
ollama run gemma4:e4b

# 12B, 26B-A4B, and 31B at Q4_K_M: 8.0, 18, and 20 GB
ollama run gemma4:12b
ollama run gemma4:26b
ollama run gemma4:31b

# Google's QAT builds: 6.1, 7.2, 16, and 19 GB
ollama run gemma4:26b-a4b-it-qat

The library also has -mlx, -mxfp8, and -nvfp4 tags for Apple Silicon, and a 26b-a4b-it-mtp-q4_K_M tag that bundles the drafter. Set the context with /set parameter num_ctx 32768 in a session, or with OLLAMA_CONTEXT_LENGTH for the server. The default is smaller than the model maximum.

llama.cpp

Support landed in PR #21309, "model: support gemma 4 (vision + moe, no audio)", merged on April 2 and tagged b8637. Audio input followed on April 12. The 12B needs build b9493 or later, from June 3, and drafter support needs b9549. Many fixes followed, such as a Gemma 4 tool-call parser on April 4 and a September 19 fix for required tool calls. Use a recent build.

# E4B on any 8 GB machine
llama-server -hf ggml-org/gemma-4-E4B-it-GGUF:Q4_0 \
  --jinja -ngl 99 -fa on -c 32768 \
  --temp 1.0 --top-p 0.95 --top-k 64

# 26B-A4B on a 24 GB GPU or a 32 GB Mac
llama-server -hf unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4_K_M \
  --jinja -ngl 99 -fa on -c 65536 \
  --temp 1.0 --top-p 0.95 --top-k 64

# 26B-A4B on a 12 GB GPU: experts of the first 20 layers stay in RAM
llama-server -hf unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4_K_M \
  --jinja -ngl 99 -fa on -c 32768 --n-cpu-moe 20

# 31B on a 24 GB GPU: 16K context, or a q8_0 cache
llama-server -hf unsloth/gemma-4-31B-it-GGUF:Q4_K_M \
  --jinja -ngl 99 -fa on -c 32768 -ctk q8_0 -ctv q8_0

Lower --n-cpu-moe until the model no longer fits, then step back up by one. With -hf, llama-server also downloads the mmproj file, which enables image input and, for E4B and the 12B, audio. Pass --no-mmproj to save that memory for text-only work. Always pass --jinja, because the chat template handles thinking and tool calls.

vLLM and SGLang

vLLM added Gemma 4 in v0.19.0 on April 3, with a requirement of transformers>=5.5.0. The 12B and drafter support arrived in v0.23.0 on June 15. The vLLM recipe shows the parsers:

vllm serve google/gemma-4-26B-A4B-it \
  --max-model-len 32768 \
  --gpu-memory-utilization 0.90 \
  --enable-auto-tool-choice \
  --reasoning-parser gemma4 \
  --tool-call-parser gemma4 \
  --limit-mm-per-prompt '{"image": 0, "audio": 0}'

Set --max-model-len every time. The default is the full 262,144 tokens, which reserves cache you probably do not need. The last flag skips multimodal profiling for text-only work. The BF16 26B-A4B needs an 80 GB card. On a 48 GB card, use the RedHatAI FP8 build. On 24 to 32 GB, use the AWQ build, or Google's qat-w4a16-ct weights for the 12B and 31B.

SGLang documents Gemma 4 in its cookbook with the same parser names. It selects the Triton attention backend on its own, because image tokens use bidirectional attention during prefill.

sglang serve --model-path google/gemma-4-E4B-it \
  --reasoning-parser gemma4 \
  --tool-call-parser gemma4 \
  --host 0.0.0.0 --port 30000

Once several users or an agent fleet share the server, paged attention and batching change the economics. The vLLM internals guide explains why.

MLX on a Mac

mlx-lm added Gemma 4 text support in v0.31.2 on April 7. The 12B alias (gemma4_unified) merged on September 11 in PR #1386, after the last tagged release, so the 12B needs mlx-lm from the main branch.

pip install -U mlx-lm

mlx_lm.generate --model mlx-community/gemma-4-26b-a4b-it-4bit \
  --max-tokens 1024 --temp 1.0 --top-p 0.95 \
  --prompt "Explain the failing test in the log below."

# the 12B, until the next mlx-lm release
pip install -U git+https://github.com/ml-explore/mlx-lm.git

For image input on MLX, use mlx-vlm or LM Studio, which ships MLX builds for every size. For the GGUF path on a Mac, use the llama.cpp commands above. Metal runs them unchanged.

Drafter models

Google publishes a small "assistant" model for each size, from 0.16 GB for E4B to 0.94 GB for the 31B in BF16. The assistant shares the target's KV cache and proposes several tokens, and the target checks them in one pass. The E2B and E4B assistants also limit their output to about 4,000 candidate tokens through "centroids masking". The vLLM recipe says this cuts the output-head work by about 45x.

# vLLM
vllm serve google/gemma-4-31B-it --tensor-parallel-size 2 \
  --max-model-len 8192 \
  --speculative-config '{"model": "google/gemma-4-31B-it-assistant", "num_speculative_tokens": 4}'

# llama.cpp, with the drafter file from the same ggml-org repo
llama-server -m gemma-4-31B-it-Q4_0.gguf \
  --model-draft mtp-gemma-4-31B-it-Q4_0.gguf \
  --spec-type draft-mtp --spec-draft-n-max 4 \
  --jinja -ngl 99 -fa on -c 32768

Match the drafter to the target. Google's QAT card says a QAT target needs a QAT assistant of the same precision. The ggml-org repositories list both the QAT and the plain assistant as sources, so pair files of the same type from the same repo. The vLLM recipe recommends 2 draft tokens for E2B, 4 for E4B and the 26B-A4B, and 4 to 8 for the 12B and 31B.

Thinking, sampling, tools

Sampling. Google gives one setting for every use: temperature=1.0, top_p=0.95, top_k=64. The ggml-org GGUFs store these as defaults in their metadata.

Thinking. Thinking is off by default. The chat template turns it on when you pass enable_thinking: true, which puts a <|think|> token at the start of the system turn. In llama.cpp, add --chat-template-kwargs '{"enable_thinking":true}'. Over an OpenAI-compatible API, send "chat_template_kwargs": {"enable_thinking": true}. In vLLM, --default-chat-template-kwargs sets it for every request. The thought arrives between <|channel>thought and <channel|>, and the gemma4 reasoning parser moves it to reasoning_content. With thinking off, the 12B, 26B, and 31B still emit an empty thought block, but E2B and E4B do not.

History. Do not send old thoughts back. The card says previous turns must contain only the final answer, except for tool-call turns, where the thinking stays.

Tools. Turns use <|turn> and <turn|>, and a call is written as <|tool_call>call:name{...}<tool_call|>. The template embedded in current GGUFs is dated July 9, 2026, and its header says it fixed "tool-calling loops, turn closures, and thinking content-ordering". If an agent loops on tool calls, check that your GGUF or engine carries that template.

Multimodal order. Put images before the text and audio after it. Use a visual budget of 70 or 140 for video frames, and 560 or 1,120 for documents and small print.

Speed: ceilings and published runs

I planned to measure E4B and the 26B-A4B on my M1 Max with 64 GB, using llama.cpp build 10330. Download bandwidth blocked the run. The E4B file reached about 7% and the 26B-A4B file about 2% before the time box ended, so this guide has no local measurements. The partial file headers did confirm two facts: both load as architecture gemma4 in that build, with 42 and 30 blocks. The Ollama E4B file also turned out to be a mixed quant that keeps attention at Q6 and Q8, not a plain Q4_K_M.

What I can give you is a ceiling. Decode reads the weights it uses once per token, so memory bandwidth divided by bytes read per token is the fastest any engine can go. Real engines land well under it, often at half or less. The bytes per token below use the ggml-org Q4_0 files:

Ceiling, tokens/sBandwidthE4B12B26B-A4B31B
DGX Spark273 GB/s913812415
M1 Max400 GB/s1335518222
RTX 40901,008 GB/s33614045856

These are upper bounds from arithmetic, not measurements. They also ignore the KV cache reads, which grow with context. The pattern is the useful part. The 26B-A4B has a higher ceiling than E4B, and the 31B ceiling is eight times lower than the 26B-A4B ceiling.

Two published runs give real data points. In llama.cpp PR #23398, its author ran the 31B on a DGX Spark at about 6 tokens per second without a drafter. With the drafter at 4 draft tokens, the same nine-prompt run took 120.65 seconds instead of 290.01, with 59% of drafts accepted. The PR does not state the quant. In PR #24282, its author ran the E4B QAT Q4_0 build on a Galaxy S26+ phone at 14.04 tokens per second. The drafter ran on the CPU, with 48% of drafts accepted.

Benchmarks, vendor-reported

These numbers come from Google's model card, for the instruction-tuned models, under Google's settings. No independent run is shown here.

Benchmark31B26B-A4B12BE4BGemma 3 27B
MMLU Pro85.2%82.6%77.2%69.4%67.6%
AIME 2026, no tools89.2%88.3%77.5%42.5%20.8%
LiveCodeBench v680.0%77.1%72.0%52.0%29.1%
GPQA Diamond84.3%82.3%78.8%58.6%42.4%
Tau2, average of 376.9%68.2%69.0%42.2%16.2%
MMMU Pro76.9%73.8%69.1%52.6%49.7%
MRCR v2, 8 needles, 128K66.4%44.1%43.4%25.4%13.5%

The 26B-A4B stays within three points of the 31B on knowledge and math. It falls behind on agentic tool use (Tau2) and on long-context retrieval, which matches the reach figure above. The 12B beats the 26B-A4B on Tau2, so for a laptop agent it is the stronger pick per gigabyte. Google's chart also placed the 31B third among open models on the Arena text leaderboard as of April 1.

Failure modes and fixes

FAQ

Which Gemma 4 model should I run?

On 8 GB of memory, run E4B. On 12 to 16 GB, run the 12B at Q4. On a 24 GB GPU or a 32 GB Mac, run the 26B-A4B for speed or the 31B for quality. The 26B-A4B reads about as many bytes per token as E4B, so it decodes much faster than the 31B.

Is Gemma 4 open source?

The weights are under the standard Apache License 2.0. Gemma 1 to 3 used the custom Gemma Terms of Use. Gemma 4 has no separate prohibited-use addendum on its license page.

What does the E in E4B mean?

Effective parameters. E4B stores 8.0B parameters, but 2.82B of them are a per-layer embedding table. The model only looks up 10,752 values of it per token. By default, llama.cpp keeps that table in host memory.

When was Gemma 4 released?

Google announced Gemma 4 on April 2, 2026, with E2B, E4B, 26B-A4B, and 31B. The 12B followed on June 3, 2026. The Hugging Face repositories were created earlier, in March and May, before the weights went public.

Does Gemma 4 support speculative decoding?

Yes. Google publishes assistant drafter models for each size. vLLM takes them through --speculative-config, SGLang through --speculative-algorithm NEXTN, and llama.cpp through --spec-type draft-mtp from build b9549.

rg
Rohit Ghumare

CNCF Ambassador and Google Developer Expert. I build agent infrastructure and write about the fundamentals underneath the AI stack. Config values, file sizes, and tensor types here come from the Hugging Face repositories of Google and the quant uploaders, and from the GGUF headers. Dates and versions come from Google's posts and model cards and the llama.cpp, vLLM, and mlx-lm release records, read on October 2, 2026.

Gemma 3 · Hardware guide · More guides · X