How to Run Hy4 Preview Locally
On August 27 Tencent's Hy team put Hy4 preview on Hugging Face under Apache-2.0. It is a mixture-of-experts model with 770B total parameters and 49B active per token. Each token picks 8 of 256 routed experts plus one shared expert. Attention is Gated DeepSeek Sparse Attention, so each token reads only 2,048 cached positions per layer. The context limit is 1M tokens. The name says preview, and Tencent means it. This guide covers what that label implies, the memory arithmetic from config.json, every published build with its exact size, and which machines can hold it. I did not run it on my own hardware. Every fit verdict here is computed, and every speed figure names who measured it. Sources were read on October 2, 2026.
A preview checkpoint
Tencent ships two repositories: the BF16 weights at tencent/Hy4-preview and an MXFP8 copy at tencent/Hy4-preview-FP8. Both appeared at about 08:55 UTC on August 27. The model card is direct about the status. It calls this "an early version of Hy4" with "real headroom left in both pre-training and post-training." It names two known issues: the model spends "longer than necessary reasoning through complex tasks" and has "a tendency to over-verify its own work."
There is a precedent for what happens next. Tencent published Hy3-preview on April 13, 2026, and Hy3 on July 2, 2026, about eleven weeks later. Both repositories report the same parameter count, 298.8B. If Hy4 follows the same path, a final checkpoint will replace this one. For a local setup, that has four practical consequences:
- Budget for a second download. The smallest useful build is about 230 GB. A replacement means you fetch it again.
- Expect long answers. Reasoning defaults to
high. Plan for more output tokens per task than the benchmark scores suggest, and set a token cap. - Treat quants as provisional. Community GGUFs were cut from the preview weights. They will not carry over to a final release.
- Pin your engine version. Engine support landed in the first two weeks of September. Behavior can change as the model code settles.
The license is not provisional. The LICENSE file states that "Tencent Hy4 preview is licensed under the Apache-2.0," followed by the standard Apache 2.0 text. I found no acceptable-use policy and no user-count clause in the repository. That differs from many earlier Hunyuan releases, so check the file yourself if your use case depends on it.
The architecture
The model type is hy_v4 and the class is HYV4ForCausalLM. The values below come from the raw config.json. The parameter split comes from the safetensors headers of all 131 BF16 shards, which I summed per tensor class.
| Component | Value |
|---|---|
| Parameters, backbone | 769.9B total, 49.1B active per token |
| MTP draft layer | 10.05B, about 0.7B active (card) |
| Layers | 78: layer 0 dense FFN, layers 1 to 77 MoE |
| Hidden size | 6,144 |
| Experts per MoE layer | 256 routed + 1 shared, top-8, sigmoid scores |
| Expert FFN width | 2,048 (dense layer 0: 18,432) |
| Attention | MLA, 64 heads, q rank 2,048, kv rank 512 |
| Head dims | 192 no-RoPE + 64 RoPE for keys, 256 for values |
| Output gate | element-wise sigmoid gate, 16,384 x 6,144 per layer |
| Sparse attention | DSA indexer, 32 heads x 128 dims, top-k 2,048 |
| Indexer layers | 21 "full", 57 "shared" (IndexCache) |
| Residual | iHC, 4 residual streams |
| Context | 1,048,576, RoPE base 10,000,000 |
| Vocabulary | 120,832 |
Four parts deserve a plain description. MLA (multi-head latent attention) caches one compressed 512-value latent plus a 64-value RoPE key per token per layer. It does not cache full keys and values per head. The gate multiplies each attention output by a learned sigmoid before the output projection. That gate matrix is as large as the output projection. It adds about 7.9B parameters across the 78 layers and nothing to the cache. DSA adds a small indexer that scores every earlier token and keeps the top 2,048 for the real attention. IndexCache lets most layers reuse the picks of an earlier layer instead of scoring again. The sections below put numbers on each.
The parameter count also shows where the weight is. Routed experts hold 744.1B of the 769.9B backbone, which is 96.6%. Everything else is 25.8B. Attention is 20.7B, the shared experts are 2.9B, and the embeddings and output head are 1.5B. Norms, routers and indexers make up the rest. Per token, the model reads all 25.8B of that, plus 8 of 256 routed experts in each of 77 layers, which is 23.3B. The sum is the 49B active figure on the card.
256 experts, 8 at a time
Each routed expert is a small SwiGLU block of three 6,144 by 2,048 matrices, 37.7M parameters. A router scores all 256 for the current token and keeps the 8 best. The shared expert runs for every token. For one token alone, the model reads 8 experts per layer, 3.1% of the routed weights.
That fraction holds for one token at a time only. A forward pass that carries several tokens reads the union of their picks. Two tokens rarely pick the same 8 experts. With uniform routing, a pass of B tokens touches 256 x (1 − (248/256)B) distinct experts per layer. That is 30.5 experts for 4 tokens, 102 for 16, and nearly all 256 for a 512-token prefill chunk. This matters twice for a local run. Prefill reads almost the full model per chunk. Speculative decoding verifies several drafted tokens in one pass, and that pass reads more experts than a single decode step. Drive the meter to see the union grow.
One MoE layer, blk.40, as 256 routed expert cells. Each token in the pass picks 8 cells. A cell darkens as more tokens pick it. The shared expert runs once per pass for all tokens. The picks are simulated with a fixed seed, and the skewed option is an assumption, not a measured Hy4 router trace. The byte counts use measured file sizes.
Routed bytes = distinct experts per layer, averaged over all 77 simulated layers, x 77 layers x bytes per expert. Bytes per expert are the routed tensor bytes of each GGUF divided by 19,712 experts: IQ4_XS and Q2_K from the qtum shard headers, STQ1_0 from the anemll sidecar (807,665,664 bytes per expert slot across 77 layers). MXFP8 is computed from the format: 8 bits per weight plus one 8-bit scale per 32 weights. The other weights (attention, shared experts, embeddings) are read once per pass. Skewed routing draws picks from a Zipf-like popularity with exponent 0.8, an assumption for illustration.
The meter explains a line in the SGLang cookbook for this model. It offers two recipes: low latency with the MTP drafter on, and high throughput with it off. The cookbook says that "at saturation the draft+verify overhead outweighs the speedup." With one user, a 4-token verify pass reads about 2.0 times the weight bytes of a plain decode step at IQ4_XS. On a machine limited by memory bandwidth, the drafter must land about 2.0 tokens per pass to break even. With many users, each pass already touches a large share of the experts. The drafted tokens then add compute and save little bandwidth.
The cache arithmetic
MLA is why the per-token cache is small for a model of this size. Each layer stores 576 values per token: the 512-value latent and the 64-value RoPE key. In BF16 that is 1,152 bytes per layer.
- Latent cache: 78 layers x 576 x 2 bytes = 89,856 bytes per token.
- Index cache: only the 21 "full" indexer layers keep a 128-dim index key per token. SGLang stores it in FP8, so 21 x 128 = 2,688 bytes per token.
- Total: 92,544 bytes per token, about 92.5 KB.
llama.cpp stores the index keys with the same type as the main cache, F16 by default, as I read its source. That gives 21 x 256 = 5,376 bytes, and 95,232 bytes per token in total. Without IndexCache, all 78 layers would keep index keys: 78 x 128 = 9,984 bytes in FP8. The shared indexers save 7,296 bytes per token.
| Context | SGLang layout | llama.cpp F16 | llama.cpp q8_0 |
|---|---|---|---|
| 8,192 | 0.76 GB | 0.78 GB | 0.41 GB |
| 32,768 | 3.03 GB | 3.12 GB | 1.66 GB |
| 131,072 | 12.13 GB | 12.48 GB | 6.63 GB |
| 262,144 | 24.26 GB | 24.96 GB | 13.26 GB |
| 1,048,576 | 97.04 GB | 99.86 GB | 53.06 GB |
The q8_0 column applies 34 bytes per 32 values to both caches. Next to the weights these numbers are small. At 131,072 tokens the cache is about 3% of the IQ4_XS file. The weights decide what you can run, and the context is a minor term until you pass 256K.
What sparse attention reads
Sparse attention does not shrink the cache. Every token stays cached, as the table above shows. What it changes is how much of the cache each new token reads. In dense MLA, a decode step at 131,072 tokens reads all 131,072 latent rows in all 78 layers, 11.78 GB. With DSA, the indexer first scores every cached token using its small 128-dim keys. Then the real attention reads only the 2,048 best rows per layer, 184 MB in total.
The scoring pass is not free. It reads one index key per cached token in every layer that owns an indexer. This is where IndexCache earns its place. Hy4 owns indexers in layers 0, 1, 5, 9 and every fourth layer after that, up to 77. The 57 layers in between reuse the picks of the nearest owner above them. Tap any layer to see whose picks it uses.
All 78 layers of Hy4 preview, numbered as in config.json. Blue layers own a DSA indexer and pick 2,048 positions. Grey layers reuse the picks of the owner shown under the number. Set a context length and an attention design, then read what one decode step costs on the attention side.
attention designLatent row = 576 BF16 values, 1,152 bytes. Index key = 128 FP8 bytes, the SGLang layout. Reads per decode step: latent rows x 1,152 x 78 layers, plus context x 128 bytes x indexer layers. Owner layers from indexer_types: 0, 1, 5, 9, 13 and every fourth layer to 77. Weight reads and compute are left out.
Two readings follow from the board. Past 2,048 tokens, the latent read per step stops growing, and the index scan becomes the part that scales with context. At 1M tokens, the shipped design scans 2.82 GB of index keys per step, against 10.47 GB if every layer kept its own indexer. Dense MLA would read 94.2 GB per step at that length, which no single machine streams at a usable rate. The second reading is about memory. You still need room for the full cache, so a 1M session needs about 97 GB for the cache alone.
The gated part of "Gated DSA" is the output gate from the architecture table. It changes the math of each attention output, not what is read or stored. For sizing, you can ignore it.
Every build, with its size
Sizes are sums of the files in each Hugging Face repository, in decimal gigabytes, read on October 2. The KL divergence and top-1 agreement figures are measurements by qtum against the BF16 weights on 8x H100. Compare them only within this table.
GGUF
| Build | Files | Quality vs BF16 | Runs on |
|---|---|---|---|
AngelSlim STQ1_0 | 229.41 GB | not published | patched llama.cpp only |
qtum IQ2_XS | 235.11 GB | KLD 0.599, top-1 73.4% | llama.cpp b10813+ |
AngelSlim UD-IQ1_M | 235.35 GB | not published | built with the patch, see below |
avar6 IQ2_XS 2.73 bpw | 262.43 GB | not published | no model card |
qtum Q2_K | 289.54 GB | KLD 0.345, top-1 80.4% | llama.cpp b10813+ |
qtum IQ4_XS | 416.97 GB | KLD 0.079, top-1 90.5% | llama.cpp b10813+ |
AngelSlim Q4_K_M | 467.29 GB | not published | built with the patch, see below |
Three notes on that table. The 6block/Hy4-preview-GGUF repository has the same three files as qtum, at the same sizes. AngelSlim is Tencent's own compression group. Its card says that "neither file runs on stock llama.cpp" because the architecture was not upstream at the time. The upstream support that came later was "developed from" that patch, per its pull request. I have not confirmed that the AngelSlim Q4_K_M and UD-IQ1_M files load on a stock build. The STQ1_0 file cannot, because the STQ1_0 type is in llama.cpp PR #22836, which is still open.
MLX, NVFP4, FP8 and BF16
| Build | Files | Engine |
|---|---|---|
mlx-community 4bit | 433.36 GB | mlx-lm from a branch |
inferencerlabs MLX-Q4i | 456.03 GB | Inferencer app 2.3.7 |
0xTank NVFP4 W4A16 + MTP | 453.01 GB | custom vLLM image, 4x DGX Spark |
Oxmiq NVFP4 W4A16 | 490.36 GB | vLLM hy4-preview image, 8x 96 GB Blackwell |
tencent Hy4-preview-FP8 (MXFP8) | 813.77 GB | vLLM 0.29+, SGLang 0.5.20+ |
tencent Hy4-preview (BF16) | 1,559.98 GB | vLLM, SGLang, multi-node |
The NVFP4 builds quantize only the routed experts and keep everything else at its original precision. The FP8 repository is MXFP8 made with NVIDIA ModelOpt, with one shared 8-bit scale per 32 weights. The vLLM recipe notes that MI325X GPUs dequantize MXFP8 weights to BF16 at load, so plan memory for BF16 on that hardware.
Which machines fit
The rule I use: weights plus cache plus 4 GB of buffers must fit in usable memory. Usable means 92% of GPU memory, 80% of system RAM for offloaded experts, and 75% of a Mac's unified memory at the default wired limit. Cache is the llama.cpp F16 figure at 32,768 tokens, 3.12 GB. None of these are my measurements. Where someone published a measured speed, the last column names them.
| Machine | Usable | Best build that fits, need | Verdict |
|---|---|---|---|
| Laptop, 64 GB unified | 48 GB | none, smallest needs 242.2 GB | no |
| RTX 4090 24 GB + 128 GB RAM | 22.1 + 102.4 GB | none, 124.5 GB short of 242.2 GB | no |
| RTX 4090 24 GB + 256 GB RAM | 22.1 + 204.8 GB | none, 226.9 GB is below 242.2 GB | no |
| 48 GB GPU + 256 GB RAM | 44.2 + 204.8 GB | qtum IQ2_XS, 242.2 GB | fits with expert offload, 6.8 GB spare |
| RTX 4090 24 GB + 384 GB RAM | 22.1 + 307.2 GB | qtum Q2_K, 296.7 GB | fits with expert offload, tight on the GPU |
| Mac Studio 256 GB | 192 GB | none at the default limit | no |
| Mac Studio 512 GB | 384 GB | qtum Q2_K, 296.7 GB | fits; IQ4_XS needs the limit raised to 425 GB |
| 8x H100 80 GB | 588.8 GB | qtum IQ4_XS, 424.1 GB | fits, about 165 GB for more context |
| 8x H20 96 GB | 706.6 GB | AngelSlim STQ1_0 measured by AngelSlim | 20.47 tok/s decode, 204.56 tok/s prefill |
| 8x RTX PRO 6000 96 GB | 706.6 GB | Oxmiq NVFP4, 490.36 GB | fits, per the Oxmiq card |
| 4x DGX Spark | 4 x 128 GB | 0xTank NVFP4 + MTP | 14.71 tok/s decode on prose, measured by 0xTank |
| 8x B300 or 16x B200 | vendor recipe | MXFP8, 813.77 GB | the vLLM recipe target |
For a ceiling on decode speed, divide memory bandwidth by bytes read per token. At IQ4_XS a single decode step reads about 34.0 GB of weights, by the arithmetic behind Fig. 1. An M3 Ultra at 819 GB/s therefore tops out near 24 tokens per second, before any compute or overhead. The MLX-Q4i card from inferencerlabs reports about 10.3 tokens per second on that chip at 1,000 tokens of context. Those two numbers agree, because real engines usually reach a third to a half of the bandwidth ceiling.
llama.cpp and Ollama
Upstream support arrived in PR #28127, "add Tencent Hy 4 (hy_v4) preview architecture support," merged on September 4. The first release that contains it is b10813. The implementation covers iHC, the gated MLA, the learnable sink, the MoE, and the DSA indexer with shared layers. It drops the MTP layer, so speculative decoding with the built-in drafter does not run in llama.cpp.
# 8 GPUs that hold the whole model (qtum IQ4_XS, 417 GB)
llama-server -m Hy4-preview-IQ4_XS-00001-of-00039.gguf \
--jinja -fa on -c 65536 \
--temp 0.9 --top-p 1.0 \
--host 127.0.0.1 --port 8080
# One GPU plus system RAM: attention on the card, routed experts in RAM
llama-server -m Hy4-preview-Q2_K-00001-of-00039.gguf \
--jinja -fa on -c 32768 -ngl 999 \
-ot "blk\..*\.ffn_.*_exps\..*=CPU" \
-ctk q8_0 \
--temp 0.9 --top-p 1.0
Download all 39 shards into one folder and name only the first on the command line. The qtum card advises against passing -ngl for the multi-GPU case, because llama.cpp fits layers to free VRAM by itself. The -ot regex in the second command sends every routed expert tensor to the CPU. What stays on the GPU is the 25.8B non-routed part, 20.18 GB at Q2_K by the shard headers. On a 24 GB card that leaves under 2 GB for cache and compute buffers, so keep the context short. A 48 GB card has room to spare.
The --jinja flag is required. AngelSlim's card notes that the chat template "matches no llama.cpp built-in family." Keep the GGUF on a local disk. The same card reports that loading over NFS ran at about 12 MB/s and turned "a 1-minute load into hours."
Ollama has no official library entry, and ollama.com/library/hy4-preview returns 404. A community upload linked in the Ollama issue thread, frob/hy4-preview, offers q2_k (280 GB), q4_K_M (465 GB) and q8_0 (818 GB) tags. Ollama synced past b10813 on September 10, and the first release with that sync is v0.34.1.
ollama run frob/hy4-preview:770b-a49b-q2_k
vLLM and SGLang
vLLM support merged in PR #54160 on August 29 and first shipped in v0.29.0 on September 9. The vLLM recipe targets 16x B200, 8x B300 and 8x MI355X. Its command with the MTP drafter on:
export VLLM_ENABLE_HPC_OPS=1
vllm serve tencent/Hy4-preview-FP8 \
--tensor-parallel-size 8 \
--speculative-config '{"num_speculative_tokens":3,"method":"mtp"}' \
--attention-backend FLASHMLA_SPARSE \
--tool-call-parser hy_v4 \
--reasoning-parser hy_v4 \
--enable-auto-tool-choice \
--served-model-name hy4-preview --port 8000
The recipe calls 262,144 tokens "the default verified window" on MI355X. It reports that MXFP8 also starts at the full 1,048,576 at a memory utilization of 0.92. For the prebuilt container, use vllm/vllm-openai:hy4-preview.
SGLang support merged in PR #36805 on September 5 and first shipped in v0.5.20 on September 18. The SGLang cookbook covers BF16 on H200, B200, B300 and GB300, and MXFP8 on Blackwell. It suggests a context of 262,144, or 131,072 on H200. The model card's command:
sglang serve --model-path tencent/Hy4-preview-FP8 \
--tp-size 8 \
--reasoning-parser auto --tool-call-parser auto \
--speculative-algorithm NEXTN \
--speculative-num-steps 3 \
--speculative-eagle-topk 1 \
--speculative-num-draft-tokens 4 \
--served-model-name hy4-preview --port 8000
Drop the four speculative flags for batch work, as Fig. 1 explains. The auto parsers resolve Hy4's suffixed special tokens, such as <think:opensource>, from the tokenizer at run time. Once several people or an agent fleet share the server, the vLLM internals guide explains how paged KV and continuous batching change the arithmetic.
On a Mac
Only the 512 GB Mac Studio holds a standard build in memory. Two paths exist. The llama.cpp path is the Metal build of b10813 or later with the qtum Q2_K at the default wired limit. For IQ4_XS, raise the limit first, for example sudo sysctl iogpu.wired_limit_mb=440000, and leave the rest of the machine idle.
The MLX path is ahead of its library. The mlx-community/Hy4-preview-4bit build was converted with mlx-lm 0.32.0. Its card says the model code is in a branch "not yet included in an mlx-lm release":
pip install git+https://github.com/kernelpool/mlx-lm.git@add-hy4-preview
mlx_lm.generate --model mlx-community/Hy4-preview-4bit \
--max-tokens 2048 --temp 0.9 --top-p 1.0 \
--prompt "Summarize the failing test output below."
On a 128 GB Mac, one experimental path exists. anemll/Hy4-preview-FlashMoE-STQ1_0 splits the AngelSlim STQ1_0 build into a 22.65 GB dense file and per-layer expert files. A patched llama.cpp fork streams experts from the SSD into a bank of slots in memory. The uploader measured 4.8 generated tokens per second on an M5 Max with 128 GB and a 96-slot bank, using about 93.3 GiB. That is a measurement on one machine with a 2,048-token context. It needs a fast internal SSD, and it is the only path below 192 GB.
Reasoning, sampling, tools
Sampling. Tencent recommends temperature=0.9 and top_p=1.0, and generation_config.json sets the same values. The SGLang cookbook advises against hardcoding sampling in client code, because the server applies these defaults.
Reasoning. The chat template accepts one variable, reasoning_effort, with two legal values: high (the default) and no_think. Any other value raises a template error, so low or medium from another model will fail. Send it in the request body:
{"model": "hy4-preview",
"messages": [{"role": "user", "content": "Rename this column."}],
"chat_template_kwargs": {"reasoning_effort": "no_think"}}
In llama.cpp, set it for the whole server with --chat-template-kwargs '{"reasoning_effort":"no_think"}'. Given the known over-reasoning issue, no_think is the better default for short edits and lookups.
Tools. Tool calls use suffixed XML-style tokens: <tool_calls:opensource>, <tool_call:opensource>, and an <arg_key> and <arg_value> pair per argument. When tools are present, the template keeps earlier reasoning in the history by default (preserved_thinking). Use --tool-call-parser hy_v4 in vLLM and auto in SGLang. For an agent, serve an OpenAI-compatible endpoint and point the harness at http://127.0.0.1:8000/v1.
Benchmarks, vendor-reported
Tencent publishes its scores as a chart image on the model card. I read the values below from that chart. They use Tencent's harness choices, and I found no independent evaluation of the open weights.
| Benchmark | Hy4 preview | Hy3 | GLM 5.3 | Kimi K3 | Claude Opus 5 |
|---|---|---|---|---|---|
| Terminal Bench 2.1 | 85.4 | 70.8 | 88.3 | 85.7 | 85.4 |
| DeepSWE | 64.3 | 28.0 | 68.1 | 74.0 | 74.7 |
| SWE Atlas Refactoring | 53.3 | 32.9 | 51.9 | 37.4 | 60.0 |
| Toolathlon-Verified | 74.1 | 56.2 | 73.8 | 74.7 | 76.5 |
| HLE, text only, no tools | 43.4 | 34.4 | 42.3 | 46.6 | 53.2 |
The jump over Hy3 is large on every row, and largest on DeepSWE, where the score more than doubles. Against the other open models in the chart, Hy4 preview is level or close, and it trails Claude Opus 5 on most rows. The card also reports a blind side-by-side test: 163 Tencent experts rated outputs on 203 engineering tasks. Hy4 preview averaged 2.99, against 2.92 for GLM 5.3 and 2.94 for Kimi K3.
Failure modes and fixes
- "unknown model architecture: hy_v4." Your llama.cpp build is older than b10813, or your Ollama is older than v0.34.1. Update.
- The STQ1_0 file fails to load. Stock llama.cpp has no STQ1_0 type yet. Use a qtum build, or apply both AngelSlim patches at the commit their card names.
- Template error naming reasoning_effort. You sent a value other than
highorno_think. Change it. - Answers run to the token cap. This is the known over-reasoning issue. Send
no_think, or set a hardmax_tokens. - Raw tool tokens in the output. The tool parser is off or wrong. Pass
--jinjain llama.cpp,hy_v4in vLLM, orautoin SGLang. - A load that takes hours. The files are on a network share. Copy them to a local NVMe disk.
- Out of memory after a successful load. The cache grew with the context. Lower
-c, or use the q8_0 cache. - Image input is rejected. Hy4 preview is text only, and the SGLang cookbook says the endpoint rejects images by design.
FAQ
Can I run Hy4 preview on a laptop or a single 24 GB GPU?
No for any standard build. The smallest GGUF is 229 GB, so a 24 GB card needs about 230 GB of system RAM beside it. On a 128 GB Mac, only the experimental SSD-streaming fork runs it, at 4.8 tokens per second as measured by its uploader.
What does the preview label mean for Hy4?
Tencent calls it an early version with known issues, such as long reasoning and repeated self-verification. Hy3 preview became Hy3 in about eleven weeks with the same parameter count. Expect a replacement checkpoint and plan to download again.
Does sparse attention make the 1M context cheap in memory?
No. DSA cuts the rows each token reads to 2,048 per layer, but every token stays cached. At about 92.5 KB per token, 1M tokens need about 97 GB of cache.
Which build should I download?
IQ4_XS from qtum at 417 GB if it fits, because it has the lowest measured KL divergence of the published GGUFs. Q2_K at 290 GB is the next step down. The STQ1_0 build needs a patched llama.cpp.
Does Hy4 preview have a license restriction?
The weights are under the Apache License 2.0, with a Tencent copyright notice in the LICENSE file. There is no separate use policy or user-count limit in the repository.
Keep reading