How to Run Granite 4.2 Locally
IBM released Granite 4.2 on August 25, 2026: three dense reasoning models at 3B, 8B, and 30B, all under Apache 2.0. Each one thinks by default, calls tools, and has a 128K window that the base models stretch to 512K. The family is the easiest of this year's releases to run. The 3B Q4_K_M file is 2.24 GB. The 8B runs on an 8 GB card at short context, and the 30B fits one 24 GB GPU. The hard part is the cache. Every layer is full attention, so long context costs more memory than the weights do. This guide gives the cache arithmetic and every file size. It adds a size grid by machine and task, the commands for each engine, and local measurements.
What IBM shipped
Granite 4.2 is IBM's first Granite generation with built-in reasoning. The three models are post-trained from the Granite 4.1 base models, which IBM trained from scratch on about 15 trillion tokens. Post-training has three stages: supervised fine-tuning on about 7.2 million samples, a chain of reinforcement learning runs, and a final preference and safety stage. The IBM technical blog gives the full recipe.
The sizes do not get the same training. The 8B and 30B go through an agentic block of reinforcement learning. It has a software engineering stage in repository sandboxes, a terminal stage, and a web search stage. The 3B skips that block. The model card prints "NA" for the 3B on every coding agent benchmark. The context story also differs by size. The Granite 4.1 blog lists the 512K extension stage for the 8B and 30B only. So the 3B is a chat and reasoning model with a 128K window. The 8B and 30B are agent models with a window that can go past 128K.
About the date. The Hugging Face repositories have a creation time of August 7, 2026. IBM's model card gives the release date as August 25, and the technical blog and the first community quants appeared that day. This guide uses August 25 as the public release. The license is plain Apache 2.0 on all three models and on IBM's quantized repositories. There is no user cap and no use restriction beyond the license.
The architecture
All three are GraniteForCausalLM, a standard decoder: grouped-query attention, RoPE, SwiGLU, RMSNorm, and separate input and output embeddings. There is no mixture of experts, no sliding window, and no linear attention. The values below come from each repository's config.json. The parameter counts come from the safetensors metadata.
| Field | 3B | 8B | 30B |
|---|---|---|---|
| Parameters (BF16) | 3,659,737,600 | 8,791,592,960 | 29,276,770,304 |
| Layers | 40 | 40 | 64 |
| Hidden size | 2,560 | 4,096 | 4,096 |
| Query heads / KV heads | 40 / 8 | 32 / 8 | 32 / 8 |
| Head dimension | 64 | 128 | 128 |
| MLP hidden size | 8,192 | 12,800 | 32,768 |
| RoPE theta (config) | 10,000,000 | 10,000,000 | 50,000,000 |
| max_position_embeddings | 131,072 | 131,072 | 131,072 |
| Vocabulary | 100,352 | 100,352 | 100,352 |
| BF16 checkpoint | 7.32 GB | 17.58 GB | 58.56 GB |
One mismatch: the 30B model card says RoPE theta is 10,000,000, but its config.json says 50,000,000. Engines read the config, so the config value is the one that runs. The 30B is the odd shape of the three. It has the same width as the 8B, 24 more layers, and an MLP that is 2.56 times wider. Most of its extra parameters sit in the MLP, which is why its cache grows less than its weights.
The cache passes the weights
A dense model with full attention in every layer stores a key and a value for every token in every layer. The cache size per token is the product of five numbers: layers, KV heads, head dimension, two tensors, and bytes per value. In f16 the arithmetic is:
- 3B: 40 layers × 8 heads × 64 × 2 × 2 bytes = 81,920 bytes, or 80 KiB per token.
- 8B: 40 × 8 × 128 × 2 × 2 = 163,840 bytes, or 160 KiB per token.
- 30B: 64 × 8 × 128 × 2 × 2 = 262,144 bytes, or 256 KiB per token.
Multiply by the context length to get the cache. The table is in decimal gigabytes for one session.
| Context | 3B f16 | 8B f16 | 30B f16 | 8B q8_0 | 8B q4_0 |
|---|---|---|---|---|---|
| 8,192 | 0.67 | 1.34 | 2.15 | 0.71 | 0.38 |
| 32,768 | 2.68 | 5.37 | 8.59 | 2.85 | 1.51 |
| 131,072 | 10.74 | 21.47 | 34.36 | 11.41 | 6.04 |
| 524,288 | 42.95 | 85.90 | 137.44 | 45.63 | 24.16 |
Now compare the cache to the weights. The 8B Q4_K_M file is 5.35 GB. The f16 cache reaches that size at 5,348 MB ÷ 163,840 bytes = 32,642 tokens. Past that point the cache is the bigger number. At the full 128K window it is 4.0 times the weights. The 3B crosses at 27,393 tokens and the 30B at 67,600 tokens. A hybrid or sliding-window model keeps most layers out of this sum. Granite 4.2 keeps every layer in it, so plan memory by context first and by model size second.
Two levers cut the cache. A q8_0 cache stores 34 bytes per 32 values, so it is 0.53 of f16. A q4_0 cache stores 18 bytes per 32 values, 0.28 of f16. In llama.cpp a quantized V cache needs flash attention on. The second lever is fewer sessions. Each parallel session at full context holds its own cache, so four agents at 32K cost the same as one session at 128K.
One memory map for the build you pick. Weights stay fixed. The cache grows with the context and the session count. The ruler marks where the cache passes the weights. Machine chips show GB needed over GB usable.
memory map
Weights are IBM's GGUF file sizes. Cache = layers × 8 KV heads × head dim × 2 × bytes per value × tokens × sessions, with q8_0 at 34/64 and q4_0 at 18/64 of f16. Usable memory is assumed at 92% of VRAM on NVIDIA cards and 75% of unified memory on Macs. Buffers are an assumed 1 GB.
Every build and its size
IBM published the quants itself, which is rare on release day. File sizes below are from the Hugging Face tree API on October 2, 2026, in decimal gigabytes. The IBM GGUF repositories carry 14 types from Q2_K to Q8_0 plus BF16. The blog says IBM made them with the standard llama.cpp converter. IBM gives no perplexity table for them.
| GGUF (ibm-granite) | 3B | 8B | 30B |
|---|---|---|---|
Q2_K | 1.46 | 3.41 | 10.86 |
Q3_K_M | 1.84 | 4.35 | 14.13 |
Q4_0 | 2.13 | 5.06 | 16.58 |
Q4_K_M | 2.24 | 5.35 | 17.72 |
Q5_K_M | 2.61 | 6.25 | 20.78 |
Q6_K | 3.01 | 7.22 | 24.02 |
Q8_0 | 3.89 | 9.35 | 31.11 |
BF16 | 7.32 | 17.59 | 58.56 (2 shards) |
For vLLM and SGLang, IBM made three more formats with LLM Compressor. FP8 uses dynamic per-channel weights and per-token activations with no calibration. NVFP4 and MXFP4 use GPTQ, calibrated on 2,000 samples from the fine-tuning set at a 2K context. For Macs, IBM and LM Studio both publish MLX builds at the same sizes.
| Build | 3B | 8B | 30B | Runs on |
|---|---|---|---|---|
-fp8 | 4.18 | 9.62 | 30.11 | vLLM, SGLang on FP8 GPUs |
-nvfp4 | 2.80 | 6.13 | 17.65 | Blackwell GPUs |
-mxfp4 | 2.70 | 5.88 | 16.76 | vLLM |
-q4-mlx | 2.06 | 4.95 | 16.47 | Apple Silicon |
-q8-mlx | 3.89 | 9.34 | 31.11 | Apple Silicon |
LM Studio MLX-6bit | 23.79 | Apple Silicon |
The Ollama tags granite4.2:3b, :8b, and :30b are 2.2, 5.3, and 18 GB. The 8B manifest's model layer is 5,347,917,952 bytes, the same byte count as IBM's Q4_K_M file. Community GGUFs from bartowski, mradermacher, and LM Studio appeared on release day and after. Use them when you want an imatrix build at Q3 or below.
Which size for which job
A dense ladder makes the choice simple to compute. Each size has a fixed cost per token of context, and each task has a typical context and a minimum size. Two minimums come from IBM's own training notes. Coding agents need the 8B or the 30B, because the 3B skipped agentic training. Contexts past 128K need the 8B or the 30B, because the 3B was not trained there. The grid below picks, per machine and task, the biggest size that fits at Q4_K_M or better. Switch the rule to get the smallest size that clears the minimum. That choice is faster, and the agent loop in the bandwidth essay explains why speed matters there.
Six machines, five jobs, thirty computed picks. Each cell tries Q8_0, Q6_K, Q5_K_M, then Q4_K_M per size, and shows GB needed over GB usable. Tap a cell for its arithmetic.
Contexts: 8,192 for quick answers, 32,768 for reasoning, 65,536 for a coding agent session, 131,072 for one long document, 524,288 for the extended window. Memory model as in Fig. 1. Decode ceilings divide published memory bandwidth by the weight file size. Real decode lands below the ceiling: the 3B measured 66.1 tokens/s on the M1 Max, 37% of its 178 ceiling.
Read the grid by column. With an f16 cache, quick answers fit all six machines, and the 30B takes every machine with 22 GB or more usable. At 32K for reasoning, the 24 GB machines drop to the 8B at Q8_0. The coding agent column removes the 8 GB card and the 16 GB laptop, because the 8B needs 17.09 GB at 64K. The long document column is the real filter. The 24 GB machines can hold only the 3B there, at its weakest score. The whole-repository column fits only the 128 GB Mac.
Now switch the cache. With q8_0, the 24 GB machines hold the 8B for a long document and the 8 GB card runs the 8B for quick answers. With q4_0, the 64 GB Mac takes the 8B to 512K and the 128 GB Mac takes the 30B there. The cache type moves more cells than the model size does.
Measured on an M1 Max
I ran the 3B and the 8B on a MacBook Pro with an M1 Max and 64 GB of unified memory, on October 2, 2026. The engine was llama.cpp build 10330 with Metal, all layers on the GPU, and an f16 cache. The files were Ollama's granite4.2:3b and granite4.2:8b blobs. The 8B blob has the same sha256 as IBM's Q4_K_M file. Speeds are the median of three llama-bench runs.
| Measure | 3B Q4_K_M | 8B Q4_K_M | 30B Q4_K_M |
|---|---|---|---|
| Prompt, 512 tokens | 1,049.6 tok/s | 426 tok/s | not measured |
| Prompt, 4,096 tokens | 693.2 tok/s | 272 tok/s | not measured |
| Decode, 128 tokens | 66.1 tok/s | 24.8 tok/s | not measured |
| Time to first token, 1,000-token prompt | 1,409 ms | 4,819 ms | not measured |
| Peak RSS, bench / server at 8K | 2.66 / 3.59 GB | 6.13 / 7.52 GB | not measured |
The 30B was skipped for disk space. Use these rows to calibrate the ceilings in Fig. 2. The ceiling for the 3B on this machine is 400 GB/s ÷ 2.24 GB = 178 tokens per second, and the measured 66.1 is 37% of it. The ceiling for the 8B is 400 GB/s ÷ 5.35 GB = 75 tokens per second, and the measured 24.8 is 33% of it. A decode-only rerun of the 8B gave 20.0 tokens per second, plus or minus 6.0, while the desktop also used the GPU. Treat the 8B as about 20 to 25 tokens per second on this machine.
Prompt speed also falls with length. It drops from 1,049.6 to 693.2 tokens per second between 512 and 4,096 tokens, because every new token attends to every earlier one. At 128K that cost grows further, which is a second reason to keep contexts short on this family.
The thinking switch behaved as the card says, with one lesson. I used the bench prompt "a bat and a ball cost $1.10" plus a one-line Python task, at temperature 0:
- Thinking on: correct answers, $0.05 and
s[::-1], after 1,203 tokens and about 21 seconds. - Thinking on with a 1,500-token cap: one run hit the cap inside the thinking block and gave no answer.
- Thinking off with
enable_thinking: false: 10 tokens in 0.2 seconds, and the wrong answer, $0.10. - The key
thinkingchanged nothing. Onlyenable_thinkingandreasoning_effortare template switches.
So the 3B needs its thinking for anything with a trap in it, and thinking needs a large output budget. Budget at least 2,048 new tokens per answer, as the SGLang cookbook advises.
Ollama
Ollama has an official library entry. The tag without a size is the 8B.
ollama run granite4.2:3b
ollama run granite4.2:8b
ollama run granite4.2:30b
# raise the context for every model the server loads
OLLAMA_CONTEXT_LENGTH=32768 ollama serve
# or per session, inside ollama run
/set parameter num_ctx 32768
The tags list a 128K context window, but Ollama allocates the cache for the context you set, not for the maximum. Set it on purpose, because the cache table above applies here too. The library params set temperature 1 and top_p 0.95, which match IBM's advice. The manifest has no separate template layer, so the template inside the GGUF applies. I did not check how Ollama's thinking switch maps to this template. Use the OpenAI-compatible endpoint with chat_template_kwargs when you need exact control.
llama.cpp and 512K
The granite architecture has been in llama.cpp for a long time, so no new build is needed for 4.2. IBM's GGUF card has a commented-out line that names build b9768. My measurements use build 10330. Always pass --jinja, because the thinking and tool switches live in the chat template.
# 8B on an 8 GB card: Q4_K_M, 8K of context, q8_0 cache
llama-server -hf ibm-granite/granite-4.2-8b-GGUF:Q4_K_M \
--jinja -ngl 99 -fa on -c 8192 \
-ctk q8_0 -ctv q8_0 \
--temp 1.0 --top-p 0.95 --port 8080
# 30B on a 24 GB card: q8_0 cache, 16K of context
llama-server -hf ibm-granite/granite-4.2-30b-GGUF:Q4_K_M \
--jinja -ngl 99 -fa on -c 16384 \
-ctk q8_0 -ctv q8_0 \
--temp 1.0 --top-p 0.95 --port 8080
# 3B on a CPU-only laptop
llama-server -hf ibm-granite/granite-4.2-3b-GGUF:Q4_K_M \
--jinja -c 8192 --temp 1.0 --top-p 0.95
To go past 128K with the 8B or 30B, do these steps:
- Pick the 8B or the 30B. Do not use the 3B past 131,072 tokens.
- Set
-cto the context you need, for example 262144 or 524288. - Add
-ctk q4_0 -ctv q4_0and keep-fa on. - Read the log for a warning about the training context and continue.
- Test recall on your own long input before you trust the output.
The GGUF records 131,072 as the training context, because it copies the config. The 512K figure comes from IBM's base-model training, not from a released serving recipe. IBM reports RULER only to 128K. I have not run either model past 128K. The memory is the certain part: the 8B at 524,288 tokens needs 24.16 GB of q4_0 cache plus 5.35 GB of weights.
vLLM and SGLang
The model card's vLLM command uses a reasoning parser plugin that ships in the repository, granite_thinking_parser.py, and needs vLLM 0.20 or newer. The card says the built-in nemotron_v3 parser also works. Tool calls use the qwen3_coder parser.
huggingface-cli download ibm-granite/granite-4.2-8b granite_thinking_parser.py --local-dir .
vllm serve ibm-granite/granite-4.2-8b \
--served-model-name granite-4.2-8b \
--dtype bfloat16 \
--max-model-len 131072 \
--reasoning-parser granite_thinking_parser \
--reasoning-parser-plugin ./granite_thinking_parser.py \
--tool-call-parser qwen3_coder \
--enable-auto-tool-choice
On a single 24 GB GPU, use ibm-granite/granite-4.2-8b-fp8 at 9.62 GB. Lower --max-model-len to what the cache allows. vLLM refuses a length above the config value. To test the extended window, the usual route is --hf-overrides '{"max_position_embeddings": 524288}'. That is an untested path for this model, so treat it as an experiment.
SGLang needs 0.5.18 or newer, per the card. Both parsers resolve on their own with auto:
sglang serve --model-path ibm-granite/granite-4.2-8b \
--dtype bfloat16 --context-length 131072 \
--reasoning-parser auto --tool-call-parser auto \
--host 127.0.0.1 --port 30000
The SGLang cookbook verified all three BF16 checkpoints on one H200 and one B200 with --tp 1 --mem-fraction-static 0.8. It also notes that a stable image failed during its tests and recommends lmsysorg/sglang:dev. The auto reasoning parser resolves to nemotron_3, which returns thinking in reasoning_content.
MLX on a Mac
IBM's MLX repositories run on stock mlx-lm. The card makes one point that is easy to miss: mlx_lm generate does not read the sampling values in generation_config.json. Without flags it decodes greedily at temperature 0. Pass the values yourself.
pip install mlx-lm
mlx_lm generate --model ibm-granite/granite-4.2-8b-q4-mlx \
--temp 1.0 --top-p 0.95 --max-tokens 4096 \
--prompt "Explain the failing test below."
# thinking off
mlx_lm generate --model ibm-granite/granite-4.2-30b-q4-mlx \
--chat-template-config '{"enable_thinking": false}' \
--temp 1.0 --top-p 0.95 \
--prompt "Summarize this paragraph in one sentence."
A 16 GB Mac runs the 8B q4 build with room for about 32K of f16 cache. A 32 GB Mac runs the 30B q4 build at short context. On a 64 GB Mac, the 30B at 128K with an f16 cache needs about 53 GB. That is above the default GPU limit of about 48 GB. Raise the limit with sudo sysctl iogpu.wired_limit_mb or use a quantized cache in llama.cpp.
Thinking, sampling, tools
Sampling. IBM says to use temperature=1.0 and top_p=0.95 for every task and every engine, including tool calls. The card suggests 8,192 new tokens for thinking mode and 2,048 without it.
Thinking. The template has three modes. Thinking is on by default and the prompt ends inside an open <think> block. Pass enable_thinking: false to close the block before the model writes. Pass low_effort: true for a short trace, which appends {reasoning effort: low} to the last user message. The template also reads reasoning_effort: "low" as the same switch, which is the form the MLX card uses. Over an OpenAI-compatible API, send these in the request body:
{"model": "granite-4.2-8b",
"messages": [{"role": "user", "content": "Plan the migration."}],
"temperature": 1.0, "top_p": 0.95, "max_tokens": 8192,
"chat_template_kwargs": {"enable_thinking": true, "low_effort": true}}
History. By default the template strips thinking from earlier assistant turns (truncate_history_thinking is true). That keeps the cache small in long chats. Set it to false only when a later turn must see the earlier reasoning.
Tools. Tool calls come out in the XML shape <tool_call><function=name><parameter=city>. That is why vLLM and SGLang use the qwen3_coder parser. The card has ready configs for OpenCode, Pi, and OpenHands that point at http://localhost:8000/v1.
Benchmarks, vendor-reported
These are IBM's numbers from the model card, run with a harness based on the NeMo Evaluator SDK. No independent evaluation was available when I checked. The table shows a selection of rows.
| Benchmark | 3B | 8B | 30B |
|---|---|---|---|
| SWE-Bench Verified | NA | 47.67 | 57.00 |
| SWE-Bench Pro | NA | 19.11 | 33.29 |
| Terminal-Bench 2.1 | NA | 20.56 | 29.24 |
| τ³-bench (avg) | 45.78 | 58.06 | 62.00 |
| BFCL v4 | 52.41 | 52.39 | 61.39 |
| AIME25 | 78.33 | 86.67 | 89.17 |
| GPQA | 54.80 | 64.14 | 66.41 |
| LiveCodeBench v6 | 69.71 | 73.24 | 75.77 |
| MMLU-Pro | 67.84 | 74.04 | 77.60 |
| IFBench (prompt) | 74.33 | 79.33 | 77.17 |
| RULER 64K | 67.52 | 80.99 | 89.96 |
| RULER 128K | 55.30 | 71.41 | 81.38 |
Three readings matter for a local choice. First, the gap from 8B to 30B is large on agent tasks and small on reasoning tasks. AIME25 moves 2.5 points, while SWE-Bench Pro moves 14.2. Second, the 3B loses the most at long context, from 67.52 at 64K to 55.30 at 128K. Third, one cell disagrees between sources: the card gives the 8B 52.39 on BFCL v4 and the blog table gives 50.29.
Failure modes and fixes
- Out of memory at load with a long context. The cache, not the model, is too big. Lower
-c, use a q8_0 or q4_0 cache with flash attention, or cut parallel sessions. - Answers cut off inside the thinking block. The output budget is too small. Give thinking mode at least 2,048 new tokens, and 8,192 for hard problems.
- Thinking text appears inside the answer. The server has no reasoning parser. Add the parser flag on vLLM or SGLang, or pass
--jinjaon llama.cpp. - Raw
<tool_call>text in the reply. The tool parser is missing. Useqwen3_coderon vLLM andautoon SGLang. - Flat, repetitive answers on MLX. mlx-lm ran greedy decoding. Pass
--temp 1.0 --top-p 0.95. - The SGLang stable image fails before load. The cookbook reports this. Use
lmsysorg/sglang:dev. - The 3B fails at agent tasks. It had no agentic training. Move to the 8B.
- Recall drops past 128K. No released score covers that range. Test with your own data, and keep critical facts near the end of the prompt.
FAQ
Which Granite 4.2 size should I run?
Run the 8B for most work. It has the agentic training. It runs on an 8 GB card at Q4_K_M with 8K of context and a q8_0 cache. It runs on a 24 GB card with 64K of context. Run the 30B when you have 24 GB or more and short contexts, or 48 GB and more for long ones. Run the 3B on 8 to 16 GB machines for chat and reasoning.
Why does it use so much memory at long context?
Every layer is full attention with 8 KV heads, so every layer stores keys and values for every token. The 8B stores 160 KiB per token in f16. At 131,072 tokens that is 21.47 GB, four times its Q4_K_M weights.
Can I run it at 512K?
The 8B and 30B base models were trained at 512K, but the released config sets 131,072 and IBM publishes no 512K recipe or score. You must raise the limit yourself. The 8B needs about 31 GB at 512K with a q4_0 cache. The 3B was not trained past 128K.
How do I turn off thinking?
Pass enable_thinking: false in chat_template_kwargs. For a short trace, pass low_effort: true. Thinking is on by default.
Is it free for commercial use?
Yes. All three models and IBM's quantized repositories are under Apache 2.0.
Keep reading