How to Run Qwen3.8-27B Locally

On August 5 Qwen published Qwen3.8-27B under Apache-2.0. It is a dense 27B model that reads images and video, thinks by default, and fits one 24 GB card at 4-bit. Three of every four layers are Gated DeltaNet layers. Those layers keep a fixed-size state, so they do not grow a cache as the conversation gets longer. That one design choice decides how much context fits on your card, and why llama.cpp sometimes reads your whole prompt again. This guide covers the memory math from config.json and every quant with its exact size. It shows what fits from 24 to 48 GB, the commands per engine, and what my M1 Max bench run found. I checked everything on October 2, 2026.

27B dense · VLM 64 layers · 48 DeltaNet + 16 attention 262K native · 1M with YaRN license Apache-2.0 4-bit 16 to 18 GB

What it is

Qwen3.8-27B is the small member of the Qwen3.8 release. The card describes it as "a compact, deployment-friendly dense model" and "a native vision-language model that understands images and videos." The weights are one checkpoint with a 27-layer vision tower attached. There is no separate text-only or instruct variant. You switch thinking off per request instead. Thinking is on by default, and you control its depth with a new reasoning_effort field that takes xhigh, medium or low.

The license is plain Apache-2.0, the same file that ships in the repository. The weights were in a public repository on August 5. Qwen's own FP8 checkpoint followed on August 13, and the community GGUF, MLX and NVFP4 builds arrived from August 13 to 14. Ollama added the model in v0.32.12 on August 14.

Three days after the 27B, on August 8, Qwen published Qwen3.8-2.4T-A95B, a 92-layer MoE with 2.4T total and 95B active parameters. It is text only, and it uses the same DeltaNet and attention layout. Its license is different. The "Qwen3.8-Max License" is permissive, with two conditions. A product with more than 100 million monthly users or 20 million US dollars of monthly revenue must display the model name. A model-as-a-service or AI work assistant business above 50 million US dollars of revenue in 12 months needs a separate license. Internal use is exempt. Unsloth's smallest GGUF of it is 397 GB, so it is a workstation project, not a laptop one. The rest of this guide is about the 27B.

The architecture

The values below come from the raw config.json. The architecture is Qwen3_5ForConditionalGeneration with model type qwen3_5. That is the Qwen3.5 code path, so engines that already ran Qwen3.5 and Qwen3.6 run this model without new model code. SGLang's cookbook says the serving-relevant architecture is identical to Qwen3.6-27B.

ComponentValue
Layers64, as 16 blocks of (3 Gated DeltaNet + 1 gated attention), full_attention_interval: 4
Hidden size / FFN5,120 / 17,408, SiLU
Gated attention24 query heads, 4 KV heads, head dim 256, output gate
Rotary25% of each head (64 dims), interleaved M-RoPE, base 10,000,000
Gated DeltaNet16 key heads, 48 value heads, head dim 128, conv kernel 4
DeltaNet state dtypemamba_ssm_dtype: float32
Context262,144 native, YaRN factor 4 for about 1M
Vocabulary248,320, untied embeddings
Vision tower27 layers, width 1,152, patch 16, 2x2 merge, video via 2-frame patches
Speculative head1 MTP layer, trained with multiple steps
BF16 weights on disk55.56 GB in 18 safetensors shards

A Gated DeltaNet layer is a form of linear attention. It does not store a key and value for every past token. It keeps one matrix per value head and updates that matrix with each new token. A learned gate decides how much of the old content to forget. The attention layers every fourth position still see every past token exactly. If the internals are new to you, the transformer internals guide covers the attention half of this design.

Fixed state, growing cache

The two layer types use memory in different ways. Here is the arithmetic for each.

ContextKV, BF16KV, q8_0DeltaNet stateIf all 64 layers were attention
8,1920.54 GB0.29 GB0.154 GB2.15 GB
32,7682.15 GB1.14 GB0.154 GB8.59 GB
131,0728.59 GB4.56 GB0.154 GB34.36 GB
262,14417.18 GB9.13 GB0.154 GB68.72 GB
1,000,000 (YaRN)65.54 GB34.82 GB0.154 GB262.14 GB

The q8_0 column uses llama.cpp's block format of 34 bytes per 32 values. The last column is a counterfactual stack with the same attention heads in all 64 layers. The hybrid cuts the long-context cost by four, and the state it adds is constant. On a 24 GB card the difference decides whether 128K tokens fit at all.

The state has a price that the table does not show. An attention cache can drop its last N tokens by moving a pointer. A recurrent state cannot go back. The matrix at token 20,000 holds no copy of the matrix at token 18,000. The next section shows what the engines do about that.

Why the state needs checkpoints

Chat clients and agent harnesses edit history all the time. A user regenerates a reply, or edits a message. A harness compacts a long tool log. Each of these sends a prompt that matches the cached prompt only up to some token. For the 16 attention layers that is cheap, because the server keeps the matching keys and values and drops the rest. For the 48 DeltaNet layers the server needs the state as it was at that exact token.

llama-server solves this with context checkpoints. Each checkpoint is a full copy of the recurrent state at one position. The defaults are --ctx-checkpoints 32 per slot and --checkpoint-min-step 8192. Per the server source, a prompt gets a checkpoint at the start of its last user message. It gets two more at 516 tokens (4 plus n_ubatch) and 4 tokens before its end. Other user messages get one only if they sit more than the minimum step past the previous checkpoint. On a mismatch, the server restores the latest checkpoint at or before the matching prefix and processes the tokens from there. With no usable checkpoint, it starts again from token zero.

Fig. 1 · the rewind ledger

A 23,230-token agent session that fixes a failing checkout test. Pick what the client sends next. Diamonds are DeltaNet checkpoints. The amber lane is work the attention layers alone would not need. The blue lane is new or changed text, which every architecture must read.

How the session reached the server

--checkpoint-min-step

Next request

systemuserassistanttoolcheckpointrestorederased
token 023,230
checkpoints after the cut0
tokens every layer type reads0
extra tokens for the state0

A model of the llama-server rules in tools/server/server-context.cpp as of build b11342: checkpoint placement, the min-step rule, no checkpoint in a batch that starts with an image, restore from the latest checkpoint at or before the matching prefix, erase checkpoints past it. Tool results count as non-user messages. Each checkpoint is about 154 MB. PR #29463 measured about 3.2 MB per recurrent layer. Token counts are illustrative. The 1280x800 screenshot is 1,000 image tokens at 32 by 32 pixels per token.

Three things stand out in the ledger. A normal next message costs nothing extra, because the state is already at the end of the cached prompt. Regenerating, or sending with preserve_thinking off, costs 4 extra tokens, because the checkpoint 4 tokens before the end exists for that case. The expensive cases are edits that land before the first checkpoint. If a harness reorders its tool list, the prompt diverges at token 3,100. The DeltaNet layers must then read those 3,100 tokens again, on top of everything after them.

The rules translate into practice. Keep the system prompt and tool schemas byte-stable across turns. Leave preserve_thinking on, which the Qwen card says "improves KV cache utilization." If you resume saved chats by pasting them as one prompt, lower --checkpoint-min-step so older user messages get checkpoints. Budget host memory for the copies too, because 32 checkpoints at 154 MB is 4.9 GB per slot at worst.

The serving engines do the same thing with different names. SGLang reserves several state slots per running request for its prefix cache. Its default extra_buffer strategy uses 5 slots. The cookbook says that on a 32 GB card "the state pool bounds concurrency long before KV does."

Every quant

Sizes are sums of the files from the Hugging Face tree API on October 2, in decimal gigabytes. The vision projector is a separate file for GGUF builds. It is 0.93 GB at F16 or BF16, and ggml-org also ships a 0.63 GB Q8_0 one.

GGUF, for llama.cpp, Ollama and LM Studio

BuildSizeUse when
unsloth UD-IQ2_XXS7.27 GB12 GB cards, expect real quality loss
unsloth UD-Q2_K_XL9.83 GBthe smallest Unsloth recommends
unsloth UD-Q3_K_XL13.15 GB16 GB cards
unsloth UD-IQ4_XS14.25 GB16 GB cards with some context
unsloth UD-Q4_K_M16.46 GB24 GB cards, long context
bartowski Q4_K_M17.44 GB24 GB cards, plain K-quant
unsloth UD-Q4_K_XL17.56 GBUnsloth's default pick
ggml-org Q4_K_M18.97 GBthe official conversion
unsloth UD-Q5_K_XL20.88 GB24 GB cards, short context
unsloth UD-Q6_K21.98 GB32 GB cards
Q8_0 (unsloth / bartowski / ggml-org)29.05 / 29.12 / 28.60 GB32 GB cards with short context, 48 GB cards
BF1653.81 to 54.66 GBreference, 64 GB and up

Unsloth ships an MTP drafter as MTP/mtp-Qwen3.8-27B-Q4_0.gguf at 1.37 GB. The ggml-org repo has MTP sidecars at 1.68 GB (Q4_0), 3.16 GB (Q8_0) and 5.95 GB (BF16). It also has DFlash2 drafters from 1.09 to 3.86 GB. LM Studio's own GGUFs are Q4_K_M at 16.81 GB, Q6_K at 22.43 GB and Q8_0 at 29.05 GB.

MLX, FP8, NVFP4, AWQ

BuildSizeEngine
mlx-community/Qwen3.8-27B-4bit16.05 GBmlx-vlm, group size 64
lmstudio-community MLX 5bit / 6bit19.42 / 22.78 GBLM Studio, mlx-vlm
MLX 8bit (mlx-community, lmstudio)29.50 GBmlx-vlm
RedHatAI/Qwen3.8-27B-INT419.45 GBvLLM, W4A16
cyankiwi/Qwen3.8-27B-AWQ-INT421.02 GBvLLM, SGLang
nvidia/Qwen3.8-27B-NVFP421.92 GBvLLM, SGLang, Blackwell only
unsloth/Qwen3.8-27B-NVFP423.42 GBvLLM, Blackwell only, FP8 layers mixed in
Qwen/Qwen3.8-27B-FP830.87 GBvLLM, SGLang, block size 128
Qwen/Qwen3.8-27B (BF16)55.56 GBvLLM, SGLang, Transformers 5.8+

Unsloth publishes KL divergence and top-1 agreement for its quants rather than a perplexity table. For its NVFP4 build it reports top-1 agreement of 92 to 97% against BF16 across its test corpora. I did not find a published perplexity table for the GGUF ladder of this model, so pick by size and step up when you can.

What fits from 24 to 48 GB

A single device in this range runs the model at 4 to 8 bits. The question is how much context fits once the weights load. Each tile in the map below is half a gigabyte of the device. Move the context and only the amber KV tiles grow. The green state tile stays the same size, and it grows only when you add parallel slots.

Fig. 2 · the half-gigabyte tile map

Real file sizes, the KV and state arithmetic above, and a context ceiling solved for your device. Grey tiles are memory the system keeps. Red dashed tiles are memory the setup needs but the device lacks.

Device

GGUF build

KV cache type and parallel slots

context = 65,536 tokens

Usable memory: GPUs at 95% of VRAM (the vLLM recipe found 31.4 of 32 GiB usable on an RTX 5090), Macs at 75% of unified memory (raise it with sysctl iogpu.wired_limit_mb). Runtime and compute buffers are an assumed 1.2 GB. MTP counts 2 GB, the top of Unsloth's 1 to 2 GB guidance. KV per token: f16 65,536 B, q8_0 34,816 B, q4_0 18,432 B. State: 154 MB per slot. With llama.cpp, -c is shared across slots.

The map makes the tiers plain. A 24 GB card runs a 4-bit build with 64K to 128K tokens of q8_0 cache, and the 5-bit build with less. A 32 GB card runs Q6_K with long context, or Q8_0 with a short one. A 48 GB card runs Q8_0 at the full 262,144 tokens. On Macs, the 75% limit means a 24 GB machine is tight at 4-bit, and a 36 GB or larger machine is comfortable.

Throughput is a separate question from fit. Decode reads every weight once per token, so a 17.5 GB file on a 1 TB/s card has a ceiling near 57 tokens per second. The memory bandwidth essay explains the rule. Dense models gain the most from the MTP drafter for that reason.

Ollama and LM Studio

Ollama v0.32.12 added the model. The default tag is a Q4_K_M build of 18 GB with image input and a 256K window.

ollama run qwen3.8:27b

# other tags: 27b-q8_0 (30 GB), 27b-bf16 (56 GB),
# 27b-mtp-q4_K_M (18 GB), 27b-mtp-q8_0 (30 GB),
# 27b-mlx (18 GB), 27b-mxfp8 (32 GB, MLX)
OLLAMA_CONTEXT_LENGTH=65536 ollama serve

Ollama's default context is much smaller than the model's window, so set the length you want. v0.32.13 added support for developer instructions. LM Studio lists both the lmstudio-community GGUFs and MLX builds. On a Mac, the MLX builds are the faster choice in LM Studio.

llama.cpp

The GGUF architecture is qwen35, so any recent llama.cpp build runs it. Per issue #27431, build 10524 ran the plain Q4_K_M. The flags below match the server docs at build b11342.

# 24 GB GPU or 32 GB+ Mac: 4-bit, 64K context, q8_0 cache
llama-server -hf unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_XL \
  --jinja -ngl 99 -fa on -c 65536 \
  -ctk q8_0 -ctv q8_0 \
  --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0 \
  --host 127.0.0.1 --port 8080

# Add the MTP drafter (download the MTP folder first)
llama-server -m Qwen3.8-27B-UD-Q4_K_XL.gguf \
  --mmproj mmproj-F16.gguf \
  --model-draft MTP/mtp-Qwen3.8-27B-Q4_0.gguf \
  --spec-type draft-mtp --spec-draft-n-max 2 \
  --jinja -ngl 99 -fa on -c 65536 -ctk q8_0 -ctv q8_0

# Shorter thinking, or none
--reasoning-effort medium
--chat-template-kwargs '{"enable_thinking":false}'

Follow these steps for a first run.

  1. Pick a build from the tile map that fits your device with the context you need.
  2. Start llama-server with --jinja, because tool calls depend on the chat template.
  3. Add --no-mmproj if you never send images and want the 0.93 GB back.
  4. Point your client at http://127.0.0.1:8080/v1.
  5. Check the log for "forcing full prompt re-processing" after a few turns.

Unsloth suggests --spec-draft-n-max 2 as a starting point and says to try values from 1 to 6. The ggml-org repo also carries DFlash2 drafters that use --spec-type draft-dflash. On Vulkan, both speculative paths have an open bug (#28158), so leave speculation off there.

MLX on a Mac

The MLX builds were made with mlx-vlm, which keeps the vision tower. The mlx-community card gives this command:

pip install -U mlx-vlm

python -m mlx_vlm.generate --model mlx-community/Qwen3.8-27B-4bit \
  --max-tokens 100 --temperature 0.0 \
  --prompt "Describe this image." --image screenshot.png

The 4-bit MLX build is 16.05 GB, so a 32 GB Mac runs it with room for context. The 8-bit build at 29.50 GB wants 48 GB or more. A 24 GB Mac has about 18 GB of GPU-usable memory at the default limit, which leaves almost nothing for cache after the 4-bit weights. Raise the limit with sudo sysctl iogpu.wired_limit_mb, or use a 3-bit GGUF. Ollama's 27b-mlx tag is the low-effort route to the same backend.

vLLM and SGLang

The vLLM recipe lists vLLM 0.17.0 or newer with Transformers 5.8.0 or newer. Its consumer-card runs used a 0.26.1 development build. This is its verified command for one RTX 5090:

vllm serve Inferact/Qwen3.8-27B-NVFP4 \
  --tensor-parallel-size 1 \
  --max-model-len 32768 \
  --kv-cache-dtype fp8 \
  --enforce-eager \
  --reasoning-parser qwen3

# MTP, per the recipe
--speculative-config '{"method":"mtp","num_speculative_tokens":3}'
# tool calls
--enable-auto-tool-choice --tool-call-parser qwen3_xml

Without --enforce-eager, the recipe says startup dies in CUDA graph capture with an out-of-memory error, and changing --gpu-memory-utilization does not help. On two RTX 5090s at 262,144 context, the recipe measured a KV pool of 377,456 tokens with FP8 weights. The Unsloth NVFP4 build gave 920,517 tokens. MTP acceptance was 0.77 to 0.90. DFlash2 drafting needs vLLM 0.28.0 or newer. MXFP4 checkpoints do not load on NVIDIA in vLLM, so use NVFP4 there.

The SGLang cookbook verified 202 configurations on v0.5.19. This is its H200 FP8 command:

sglang serve --trust-remote-code \
  --model-path Qwen/Qwen3.8-27B-FP8 \
  --kv-cache-dtype fp8_e4m3 \
  --mem-fraction-static 0.85 \
  --attention-backend flashinfer \
  --chunked-prefill-size 32768 \
  --max-prefill-tokens 32768 \
  --reasoning-parser qwen3 \
  --tool-call-parser qwen3_coder \
  --mamba-full-memory-ratio 4.59 \
  --mamba-radix-cache-strategy extra_buffer \
  --mamba-ssm-dtype float32 \
  --host 0.0.0.0 --port 30000

The --mamba-full-memory-ratio flag splits memory between the state pool and the KV pool. The cookbook's formula is (S + D) × state bytes divided by (average request length × KV bytes per token). It says the default of 0.9 "silently clamps concurrency." The --mamba-ssm-dtype bfloat16 flag halves the state to 78.4 MB. On an RTX 5090 the cookbook measured 97,280 KV tokens with a BF16 state against 68,588 with float32. It calls BF16 state "an accuracy gate," so validate it on your own tasks. Once several agents share a box, the vLLM internals guide explains why the serving engines pull ahead.

Thinking, sampling, tools, vision

Sampling. The card gives two sets. For thinking: temperature=1.0, top_p=0.95, top_k=20, min_p=0, presence_penalty=0. For non-thinking: temperature=0.7, top_p=0.8, top_k=20, presence_penalty=1.5. The shipped generation_config.json carries the thinking set.

Thinking. Send "chat_template_kwargs": {"enable_thinking": false} to answer directly. Send reasoning_effort as medium or low to think less. The card warns that a lower effort "does not always reduce overall task completion time" in agent loops, because shallow turns fail and retry. With preserve_thinking on, the default, earlier thinking blocks stay in the prompt. That costs tokens but keeps the prefix stable, which matters for the checkpoints above.

Tools. The template emits <tool_call> blocks with <function=...> and <parameter=...> inside. SGLang decodes that with qwen3_coder, and the vLLM recipe uses qwen3_xml. A Hermes parser expects JSON inside the tag, so tool calls never parse with it. The card also allots 262,144 tokens for reasoning and 131,072 for the final answer in long agent tasks.

Vision. Each 32 by 32 pixel area becomes one token after the 2x2 merge, so a 1280x800 screenshot costs about 1,000 tokens. For hour-scale video, the card recommends raising longest_edge in video_preprocessor_config.json to 469,762,048, which is about 224K video tokens.

M1 Max bench status

The bench run on my M1 Max (64 GB, macOS 26.2, llama.cpp Homebrew build 10330) did not reach this model on October 2. Downloads from Hugging Face and the Ollama registry ran at 2 to 3 MB per second. The 16.8 GB file was not complete in the time box. I do not publish speed numbers I did not measure. This section gets the tokens per second, time to first token and memory when the run finishes.

The partial download still settled one question. The GGUF header of qwen3.8:27b-q4_K_M reads general.architecture=qwen35, 65 blocks and file type 15 (Q4_K_M). I read the 65 blocks as 64 layers plus the MTP layer. Build 10330 includes the qwen35 architecture, so a Homebrew llama.cpp from August or later loads the file.

Until then, use a ceiling and two published measurements. Decode reads about 17.4 GB per token at Q4_K_M. The M1 Max's 400 GB/s caps it near 23 tokens per second, and real runs land below that. The vLLM recipe reports 64 tokens per second for one stream on an Ascend 950PR with MTP on. Unsloth reports 133.7 tokens per second for one user on a B200 with its NVFP4 build. Both are their measurements, on hardware far from a laptop.

Benchmarks, vendor-reported

These are Qwen's numbers from the model card, under Qwen's harness choices. Most coding rows ran in the Claude Code harness at a 256K context. Treat them as the vendor's claims until independent runs exist.

BenchmarkQwen3.8-27BQwen3.6-27BMuse Glimmer-30BOpus 4.6 Max (as printed)
Terminal Bench 2.1 (Terminus)73.063.451.778.2
SWE-bench Pro61.753.551.253.4
NL2Repo-Bench42.336.247.6
IFBench79.569.177.062.5
GPQA Diamond89.287.883.591.3
OSWorld-Verified84.363.965.972.7
OmniDocBench 1.591.189.475.886.6
CharXiv (RQ), no code interpreter83.778.478.866.0

The jump over Qwen3.6-27B is largest on agent rows: OSWorld-Verified rises by 20.4 points and SWE-bench Pro by 8.2. The Muse Glimmer guide covers the nearest dense competitor at this size.

Failure modes and fixes

FAQ

Does it fit a 24 GB GPU?

Yes, at 4-bit. A 16.5 to 17.6 GB GGUF leaves room for 64K to 128K tokens of q8_0 cache. Q8_0 at 29 GB needs a 32 or 48 GB device.

Why is the cache so small?

Only 16 of 64 layers keep a per-token cache, at 64 KB per token in BF16. The other 48 keep a fixed state of about 154 MB per sequence.

Why does llama.cpp reprocess my prompt?

The DeltaNet state cannot be rolled back. If your prompt diverges before the earliest usable checkpoint, the server starts again from token zero. Keep the prefix stable.

How do I turn thinking off?

Send enable_thinking: false in chat_template_kwargs, and switch to the non-thinking sampling set. Use reasoning_effort to keep thinking but shorten it.

Should I run the 2.4T instead?

Only on a workstation with several hundred gigabytes of memory, and only after you read its license. For one GPU or one Mac, the 27B is the model.

rg
Rohit Ghumare

CNCF Ambassador and Google Developer Expert. I build agent infrastructure and write about the fundamentals under the AI stack. Config values and file sizes come from the Qwen, unsloth, bartowski, ggml-org, lmstudio-community, mlx-community, nvidia, RedHatAI and cyankiwi repositories on Hugging Face. Engine details come from the vLLM recipe, the SGLang cookbook, Ollama, and the llama.cpp source and issues. I read all of them on October 2, 2026.

Qwen3.8-Flash-Next · Hardware guide · More guides · X