How to Run GLM-5.3 Locally: 744B, 40B Active, and the Sparse-Attention KV Cache
GLM-5.3 is z.ai's flagship open-weights model. It is a text-only mixture of experts with 744 billion parameters in the main stack and about 40 billion active per token. The API opened on August 14, 2026, and the weights followed on August 28 after a safety review. Every one of its 78 layers keeps a compressed attention cache, and a sparse indexer decides which 2,048 cached tokens each step reads. This guide covers that memory arithmetic, every published checkpoint with its real size, the vLLM, SGLang, and llama.cpp commands, and what fits on which machine.
reasoning_effort settings, and max is the default. No GLM-5.3-Max repository exists on Hugging Face. The family has two models: GLM-5.3, covered here, and GLM-5.3-Flash, the smaller multimodal one.Max is a setting
The z.ai documentation defines one model, glm-5.3, with a reasoning_effort field. It accepts low, high, and max. The default is max, and the chat template falls back to max for any other value. Thinking is always on. A request that sets thinking to disabled fails, and the migration notice tells you to send reasoning_effort: "low" instead.
So a leaderboard row named "GLM-5.3 (max)" is the default model at its default effort. Artificial Analysis scores it 60 on its Intelligence Index at max and 34 at low. That gap is the effort setting, not a second set of weights. Set the effort per request:
# z.ai API (OpenAI-compatible path from the docs quick start)
curl https://api.z.ai/api/paas/v4/chat/completions \
-H "Authorization: Bearer $ZAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"glm-5.3",
"messages":[{"role":"user","content":"Plan the refactor."}],
"reasoning_effort":"high"}'
# a local vLLM or SGLang server reads it from the chat template
"chat_template_kwargs": {"reasoning_effort": "low", "clear_thinking": true}
# llama-server
llama-server ... --jinja --reasoning-effort low
The z.ai docs also list a coding-plan path, api.z.ai/api/coding/paas/v4, an OpenAI Responses path, and an Anthropic Messages path. The docs do not say which chat path is canonical, so use the one your plan names.
What shipped, and when
z.ai launched GLM-5.3 on August 14, 2026 as an API and coding-plan model. The launch post says: "We will release the weights in two weeks after launch, once safety evaluation and hardening are complete." The zai-org/GLM-5.3 repository was created empty on August 25. Its first content commit, titled "Initial commit 0828", landed at 17:16 UTC on August 27, which is August 28 in Beijing. Community evaluation pull requests followed on August 28, so this guide dates the open weights to August 28. A chat-template fix for tool-result handling followed on September 4.
The weights are not MIT. The GLM-5.3 License permits use, copying, modification, distribution, and sale. It adds one condition. It applies to model-as-a-service businesses with more than $10 billion of revenue in any twelve months. Such a licensee must pass a z.ai security review before commercial use. For a local run, the condition does not apply. GLM-5.3-Flash, released August 26, is plain MIT.
Counting the parameters
Three totals circulate. z.ai's GitHub README says 744B with 40B active. The vLLM recipe says 743B with 39B active. Artificial Analysis says 753B. All three are consistent once you count the layers from config.json.
- Each MLA attention block holds 165.0M parameters. Each routed expert holds 3 × 6,144 × 2,048 = 37.75M.
- The 78-layer main stack, with embeddings, output head, and 21 indexers, comes to 743.38B.
- The single MTP layer, the built-in draft head, adds 9.95B. The sum is 753.33B.
- The BF16 repository's index lists 1,506,659,919,872 bytes. Divided by 2 bytes, that is 753.33B, an exact match.
So 743B or 744B is the main model, and 753B includes the draft head. The active count works the same way. Attention, 8 routed experts, and the shared expert give 39.35B without embeddings and 41.25B with them. Plan memory around 753B, because the MTP layer ships in the file.
The architecture
The model type is glm_moe_dsa, the same structure as GLM-5.2 and DeepSeek-V3.2. The vLLM recipe says the architecture and serving flags are identical to GLM-5.2. The new part is that the default checkpoint is now FP8.
| Component | Value | Why it matters here |
|---|---|---|
| Layers / hidden | 78 / 6,144 | every layer is attention plus MLP, no linear layers |
| Attention | MLA, 64 heads, q rank 2,048, kv rank 512 | the cache holds one 512-wide latent per token per layer |
| Head dims | qk 192 + rope 64, v 256 | the 64-wide rope key is cached beside the latent |
| Sparse indexer | 32 heads, dim 128, top-k 2,048 | picks the 2,048 cached tokens each step reads |
| Indexer layout | 21 full, 57 shared | shared layers reuse the previous full layer's selection |
| MoE | 256 routed, 8 active, 1 shared, expert dim 2,048 | first 3 layers use dense MLPs |
| MTP | 1 next-token layer | the drafter vLLM and SGLang call MTP or EAGLE |
| Context / vocab | 1,048,576 / 154,880 | rope theta 8,000,000 |
Flash is built differently. It has 45 layers, and 34 of them are linear attention with a fixed state and no per-token cache. GLM-5.3 has no linear layers. Every layer caches every token. Its sparse attention saves the work of reading the cache, not the memory that holds it. That one difference drives most of the sizing below.
The cache, and what sparse attention saves
Each layer caches the 512-value latent plus the 64-value rope key for every token. That is 576 values. In BF16, 576 × 2 = 1,152 bytes per token per layer, and 1,152 × 78 = 89,856 bytes per token. An FP8 cache halves it to 44,928 bytes. The vLLM and SGLang recipes use FP8 cache on Blackwell by default.
| Context | BF16 cache | FP8 cache |
|---|---|---|
| 32,768 | 2.94 GB | 1.47 GB |
| 131,072 | 11.78 GB | 5.89 GB |
| 262,144 | 23.56 GB | 11.78 GB |
| 524,288 | 47.11 GB | 23.56 GB |
| 1,048,576 | 94.22 GB | 47.11 GB |
These numbers exclude the indexer's own key cache. No primary source states its size. If each of the 21 full indexer layers keeps one 128-wide BF16 key per token, that adds 5,376 bytes per token. Treat that as an estimate.
The sparse indexer changes what a decode step reads. Dense attention reads the whole cache on every step. Here, the 21 full layers score every cached token with the small indexer and pick the top 2,048. All 78 layers attend only to those. At 1M tokens the cache is 94 GB, but one step reads the 2,048 selected latents per layer plus the indexer keys. Drive it:
One agent session as 64 slices of its context. Step through one decode step and watch what the cache holds and what the step reads. The byte counts come from config.json. Which slices the indexer picks is illustrative, because the real choice depends on learned scores.
Cache per token = 78 layers × (kv rank 512 + rope 64) × bytes per value. Indexer keys = 21 full layers × 128 values per token, an estimate because no source states their size. Indexer read = every cached token's key on the 21 full layers. Sparse attention read = 78 layers × min(context, 2,048) latents. Dense read = the whole cache. Engines add their own overheads.
The reading for local use: the cache stays modest until you pass 100K tokens, and decode speed stays flat as the context grows. The weights still decide whether you fit. Even the smallest quant is 216 GB.
Pick a checkpoint
Sizes are sums of each repository's files from the Hugging Face API, in decimal gigabytes, checked on October 2, 2026.
| Checkpoint | Size | Runs on | Notes |
|---|---|---|---|
| zai-org/GLM-5.3 (FP8) | 755.66 GB | vLLM, SGLang | the default; one 8x H200 or 8x B200 node |
| zai-org/GLM-5.3-BF16 | 1,506.69 GB | vLLM, SGLang | multi-node, or one 8x B300 node per SGLang |
| Inferact/GLM-5.3-NVFP4 | 464.87 GB | vLLM, Blackwell only | only routed experts drop to 4-bit |
| gpustack/GLM-5.3-W4A8 | 399.75 GB | vLLM, Hopper only | INT4 experts, FP8 elsewhere |
| mlx-community/GLM-5.3-4bit | 418.34 GB | mlx-lm, Apple Silicon | needs a 512 GB Mac |
| unsloth/GLM-5.3-GGUF | 216.72 to 1,507.99 GB | llama.cpp | the ladder below |
| GGUF quant | Files | Use when |
|---|---|---|
UD-IQ1_S | 216.72 GB | the floor; 256 GB machines with offload |
UD-IQ1_M | 228.49 GB | a little more quality than IQ1_S |
UD-IQ2_M | 238.58 GB | the best 2-bit tier for 256 GB of RAM |
UD-Q2_K_XL | 253.88 GB | unsloth's usual 2-bit pick |
UD-IQ3_XXS | 281.69 GB | 384 GB boxes |
UD-Q3_K_XL | 342.97 GB | a 512 GB Mac Studio at the default memory limit |
UD-IQ4_XS | 365.31 GB | the largest tier for a stock 512 GB Mac |
UD-Q4_K_XL | 467.29 GB | 512 GB and up with offload |
UD-Q5_K_XL | 562.47 GB | 768 GB and up |
UD-Q6_K_XL | 684.37 GB | near the FP8 original |
Q8_0 | 801.36 GB | larger than the FP8 safetensors |
BF16 | 1,507.99 GB | reference only |
Unsloth has not published per-quant retention figures for this ladder. I also could not confirm whether these files carry any naming change like the one that broke Flash's GGUFs. Re-check the repository before a 300 GB download. Unsloth re-embedded a fixed chat template on August 29 for llama.cpp tool-call parsing, so download after that date.
What fits where
Four things share memory: the weights, the cache, the engine's buffers, and any experts you push to system RAM. The vLLM and SGLang servers want the whole model in GPU memory. The llama.cpp server can keep attention on a GPU and stream the experts from RAM. A Mac holds everything in one pool. Build your machine and every checkpoint reports whether it fits:
Add hardware to the rig, or start from a preset. Every published checkpoint checks itself against the rig: the engine it needs, the GPU generation it requires, and whether its weights and cache fit. Tap a part in the rig to remove it.
presets
add a part
Sizes are the published repositories. GPU memory is counted at 90% usable, plus 2 GB of buffers per GPU. The vLLM and SGLang rows cache in FP8 at 44,928 B per token. The llama.cpp and MLX rows cache in BF16 at 89,856 B. Offload in llama.cpp keeps about 2.5% of the file on the GPU, the non-expert share from config.json. The rest must fit in RAM. NVFP4 offload follows the vLLM DGX Station recipe: 252 GB of HBM or more, experts moved to host RAM. Mac usable share is reader-set.
| Machine | Realistic setup | Source |
|---|---|---|
| 8x H200 (141 GB each) | FP8 or W4A8 with vLLM or SGLang | vLLM and SGLang recipes |
| 8x B200 (180 GB each) | FP8 at the full 1M context, or NVFP4 | vLLM recipe |
| 4x GB300 | NVFP4 at tensor parallel 4 | SGLang cookbook, experimental |
| One GB300 with 252 GB HBM plus host RAM | NVFP4 with 222 GB of experts offloaded, 262K context, 8 sequences | vLLM recipe |
| Mac Studio 512 GB | UD-IQ4_XS, or the 418 GB MLX build with a raised memory limit | file sizes above |
| One 32 GB GPU plus 384 GB RAM | UD-IQ2_M or UD-Q2_K_XL with expert offload | file sizes above, slow |
vLLM
vLLM serves the glm_moe_dsa architecture in tagged releases. The official recipe lists a minimum of vLLM 0.29.0 and verified B300, H200, Ascend 950PR, and DGX Station GB300 setups. It enables MTP with five draft tokens. The standard command for one 8x H200 node:
uv pip install "vllm>=0.29.0" --torch-backend=auto
vllm serve zai-org/GLM-5.3 \
--tensor-parallel-size 8 \
--kv-cache-dtype fp8 \
--speculative-config.method mtp \
--speculative-config.num_speculative_tokens 5 \
--tool-call-parser glm47 \
--reasoning-parser glm47 \
--enable-auto-tool-choice \
--served-model-name glm-5.3
On B200 the recipe uses --kv-cache-dtype fp8_e4m3 and --max-num-seqs 32 to reach the full 1M window. On Blackwell, the NVFP4 build halves the weights:
vllm serve Inferact/GLM-5.3-NVFP4 \
--tensor-parallel-size 8 \
--enable-expert-parallel \
--kv-cache-dtype fp8_e4m3 \
--reasoning-parser glm47 \
--tool-call-parser glm47 \
--enable-auto-tool-choice
Two optional drafters exist. The recipe lists incoai/GLM-5.3-DFlash2 with 7 draft tokens and RedHatAI/GLM-5.3-speculator.dspark with 8. The DFlash2 drafter is licensed CC BY-NC-ND 4.0 for research and evaluation. Read that before you put it in production. The mechanics of paged KV blocks and speculative decoding are in the vLLM internals guide.
SGLang
The SGLang cookbook picks the sparse attention backends and the cache type for you. It uses an FP8 cache on Blackwell and BF16 on Hopper. The latest SGLang release is v0.5.21, tagged October 2, 2026. A low-latency launch on 8x H200:
uv pip install --prerelease=allow sglang
sglang serve --model-path zai-org/GLM-5.3 \
--tp 8 --mem-fraction-static 0.8 \
--speculative-algorithm EAGLE \
--speculative-num-steps 5 \
--speculative-eagle-topk 1 \
--speculative-num-draft-tokens 6 \
--reasoning-parser glm45 \
--tool-call-parser glm47
The parser names differ between engines, and both are correct for their engine. SGLang's auto reasoning parser resolves to glm45, and vLLM's recipe uses glm47. The tool parser must be glm47 in both. The cookbook warns that the older glm45 tool parser leaves calls as raw text, so an agent loop never sees them. On AMD MI300X, MI325X, and MI355X, the cookbook uses the tilelang sparse backends and disables MTP until the draft kernel is validated.
llama.cpp
llama.cpp has run glm_moe_dsa since GLM-5.2, and MTP drafting for that architecture merged on July 29 in PR #25980. A current build loads the unsloth files. Keep attention on the GPU and send the routed experts to system RAM:
cmake -B build -DCMAKE_BUILD_TYPE=Release -DGGML_CUDA=ON
cmake --build build -j
./build/bin/llama-server \
-hf unsloth/GLM-5.3-GGUF:UD-Q2_K_XL \
--jinja -ngl 99 \
-ot ".ffn_.*_exps.=CPU" \
-c 32768 --temp 1.0 --top-p 0.95 \
--reasoning-effort low \
--host 127.0.0.1 --port 8080
Multi-shard GGUFs load from the first shard. On this repository the first shard holds metadata only, about 9.4 MB. One open report matters for offload: issue #29010 shows CPU-only inference (-ngl 0) printing repeated @ tokens with UD-Q4_K_XL. That run also used a q4_0 V cache. Keep some layers on a GPU and leave the cache type at its default until the issue closes.
On a Mac
Among Macs, only a 512 GB Mac Studio holds this model. Two paths exist. The GGUF path is the llama.cpp build above with -DGGML_METAL=ON. At the usual three-quarters of unified memory, 384 GB is usable, so UD-IQ4_XS at 365.31 GB fits with a short context. The MLX path uses mlx-community/GLM-5.3-4bit, converted with mlx-lm 0.31.3. At 418.34 GB it needs more than 384 GB, so you must raise the GPU wired-memory limit:
pip install -U mlx-lm
# raise the wired limit for this boot (value in MB; leave room for macOS)
sudo sysctl iogpu.wired_limit_mb=460000
mlx_lm.generate --model mlx-community/GLM-5.3-4bit \
--max-tokens 4096 --temp 1.0 --top-p 0.95 \
--prompt "Review this diff for race conditions."
A third-party build, pipenetwork/GLM-5.3-MLX-4bit at 418.62 GB, reports a problem in stock mlx-lm. Its authors say mlx-lm loads the 57 shared-indexer layers leniently and leaves their indexer weights uninitialized, which matters past 2,048 tokens. Their bundled runtime fixes it. I have not verified this claim, so test long prompts before you trust either build.
Effort, sampling, tools
Effort is the cost dial. z.ai reports its in-house Code Bench at 34.5% for max effort with about 75K output tokens per task. High effort scored 31.4% with about 50K tokens. GLM-5.2 needed about 96K tokens for 23.4%. These are vendor numbers. z.ai publishes no token count for low effort. Output tokens are the expensive class on the API at $4.40 per million, so the effort setting moves your bill more than anything else. Price it:
One coding task priced at each effort setting on the z.ai API, and timed on a local box. Output tokens per task are z.ai's own Code Bench averages. Input size, cache share, volume, and local speed are yours to set.
Prices: $1.40 per million input tokens, $0.26 per million cached input tokens, $4.40 per million output tokens (z.ai docs). Output tokens are about 75K at max and 50K at high. These are vendor-reported averages on the private Z.ai Code Bench, which scored 34.5% and 31.4%. z.ai publishes no token count for low. Local time counts decode only, at the speed you set.
Sampling. The checkpoint's generation config and the SGLang cookbook give temperature 1.0 and top-p 0.95. The z.ai benchmark runs vary top-p between 0.95 and 1.0 by task.
Thinking in the transcript. The chat template defaults clear_thinking to false, which keeps earlier reasoning in the conversation. Pass clear_thinking: true for chat. For agents, leaving it false costs context and keeps decisions consistent across turns.
Tool calls. The model emits <tool_call> blocks with <arg_key> and <arg_value> pairs. Use the glm47 tool parser in vLLM and SGLang, and --jinja in llama.cpp. Hand-built prompts break the format.
API price. z.ai charges $1.40 per million input tokens, $0.26 per million cached input tokens, and $4.40 per million output tokens. OpenRouter and SiliconFlow list the same rates.
Benchmarks, labeled
The first table is z.ai's own, from its model documentation, run in z.ai's chosen harness at max effort. The README notes that several agent benchmarks ran inside Claude Code with domain allow-lists against cheating.
| Benchmark (vendor-reported) | GLM-5.2 | GLM-5.3 |
|---|---|---|
| Terminal-Bench 3.0 | 4.6 | 28.3 |
| DeepSWE v1.1 | 46.2 | 66.9 |
| Agents' Last Exam (CLI) | 23.8 | 28.5 |
| CyberGym | 77.2 | 84.5 |
| ExploitBench | 24.4 | 54.4 |
| Z.ai Code Bench (private), max effort | 23.4% | 34.5% |
Independent numbers are thinner. Artificial Analysis gives 60 on its Intelligence Index v4.1.1 at max effort and 34 at low. Two secondary trackers list an LMArena Elo of 1487 for the max setting but disagree on its rank, so I leave the rank out. Treat the vendor table as direction, and run your own tasks before you commit hardware. The evals guide shows how to tell a real gap from noise.
GLM-5.3 or Flash
Pick GLM-5.3-Flash when you need images, when you have less than 384 GB of fast memory, or when you serve many long sessions. Flash is 320B with 18B active, its cache is about 11 KB per token, and it is MIT. Pick GLM-5.3 when the task is long-horizon coding or security work and you have a GPU node or a 512 GB machine. Its cache is about eight times larger per token. Its weights are more than twice the size. It costs more per session in every dimension. The API makes the comparison cheap to test first. Flash costs $0.15 and $0.50 per million input and output tokens, and GLM-5.3 costs $1.40 and $4.40.
Failure modes
- Tool calls arrive as plain text. The server uses the
glm45tool parser. Switch toglm47in vLLM or SGLang. - Answers and reasoning arrive mixed in one field. No reasoning parser is set. Add it, or the reply carries a stray
</think>. - Replies are slow and long. Effort is at max, the default. Send
lowfor tool dispatch and keepmaxfor planning turns. - CPU-only llama.cpp prints
@tokens. This is open issue #29010. Offload attention to a GPU and use the default cache type. - The FP8 checkpoint will not load on 8x H100. 640 GB of HBM is smaller than 755.66 GB of weights. Use W4A8 on Hopper, or a larger node.
- NVFP4 fails on Hopper. The NVFP4 build is Blackwell-only. W4A8 is the Hopper 4-bit path.
- MLX output degrades past 2,048 tokens. A third-party report points at uninitialized shared-indexer weights. Compare with the pipenetwork build before you blame the model.
- Thinking cannot be turned off. A request with thinking disabled fails. Use
reasoning_effort: "low".
FAQ
Is GLM-5.3-Max a separate download?
No. "Max" is the default reasoning_effort of the one GLM-5.3 model. No Max repository exists.
Is it 744B or 753B?
Both. The main stack is 743.38B, which rounds to 743B or 744B. The MTP draft layer adds 9.95B, for 753.33B in the file.
What is the smallest machine that runs it?
A machine with about 256 GB of combined GPU and system memory runs UD-IQ1_S with expert offload, slowly. A 512 GB Mac Studio runs UD-IQ4_XS. For full quality you need an 8-GPU H200 or B200 node.
Why is the cache bigger than Flash's?
All 78 layers cache every token here. Flash caches in only 11 of 45 layers. Sparse attention in GLM-5.3 cuts the reads per step, not the stored bytes.
Is the license open?
It is permissive for local use and nearly all businesses. Only model-as-a-service providers with more than $10 billion in yearly revenue need a z.ai security review first.
Keep reading