How to Run Hy4 Preview Locally

On August 27 Tencent's Hy team put Hy4 preview on Hugging Face under Apache-2.0. It is a mixture-of-experts model with 770B total parameters and 49B active per token. Each token picks 8 of 256 routed experts plus one shared expert. Attention is Gated DeepSeek Sparse Attention, so each token reads only 2,048 cached positions per layer. The context limit is 1M tokens. The name says preview, and Tencent means it. This guide covers what that label implies, the memory arithmetic from config.json, every published build with its exact size, and which machines can hold it. I did not run it on my own hardware. Every fit verdict here is computed, and every speed figure names who measured it. Sources were read on October 2, 2026.

770B total · 49B active 256 experts · top-8 + 1 shared Gated DSA · top-2,048 1M context license Apache-2.0

A preview checkpoint

Tencent ships two repositories: the BF16 weights at tencent/Hy4-preview and an MXFP8 copy at tencent/Hy4-preview-FP8. Both appeared at about 08:55 UTC on August 27. The model card is direct about the status. It calls this "an early version of Hy4" with "real headroom left in both pre-training and post-training." It names two known issues: the model spends "longer than necessary reasoning through complex tasks" and has "a tendency to over-verify its own work."

There is a precedent for what happens next. Tencent published Hy3-preview on April 13, 2026, and Hy3 on July 2, 2026, about eleven weeks later. Both repositories report the same parameter count, 298.8B. If Hy4 follows the same path, a final checkpoint will replace this one. For a local setup, that has four practical consequences:

The license is not provisional. The LICENSE file states that "Tencent Hy4 preview is licensed under the Apache-2.0," followed by the standard Apache 2.0 text. I found no acceptable-use policy and no user-count clause in the repository. That differs from many earlier Hunyuan releases, so check the file yourself if your use case depends on it.

The architecture

The model type is hy_v4 and the class is HYV4ForCausalLM. The values below come from the raw config.json. The parameter split comes from the safetensors headers of all 131 BF16 shards, which I summed per tensor class.

ComponentValue
Parameters, backbone769.9B total, 49.1B active per token
MTP draft layer10.05B, about 0.7B active (card)
Layers78: layer 0 dense FFN, layers 1 to 77 MoE
Hidden size6,144
Experts per MoE layer256 routed + 1 shared, top-8, sigmoid scores
Expert FFN width2,048 (dense layer 0: 18,432)
AttentionMLA, 64 heads, q rank 2,048, kv rank 512
Head dims192 no-RoPE + 64 RoPE for keys, 256 for values
Output gateelement-wise sigmoid gate, 16,384 x 6,144 per layer
Sparse attentionDSA indexer, 32 heads x 128 dims, top-k 2,048
Indexer layers21 "full", 57 "shared" (IndexCache)
ResidualiHC, 4 residual streams
Context1,048,576, RoPE base 10,000,000
Vocabulary120,832

Four parts deserve a plain description. MLA (multi-head latent attention) caches one compressed 512-value latent plus a 64-value RoPE key per token per layer. It does not cache full keys and values per head. The gate multiplies each attention output by a learned sigmoid before the output projection. That gate matrix is as large as the output projection. It adds about 7.9B parameters across the 78 layers and nothing to the cache. DSA adds a small indexer that scores every earlier token and keeps the top 2,048 for the real attention. IndexCache lets most layers reuse the picks of an earlier layer instead of scoring again. The sections below put numbers on each.

The parameter count also shows where the weight is. Routed experts hold 744.1B of the 769.9B backbone, which is 96.6%. Everything else is 25.8B. Attention is 20.7B, the shared experts are 2.9B, and the embeddings and output head are 1.5B. Norms, routers and indexers make up the rest. Per token, the model reads all 25.8B of that, plus 8 of 256 routed experts in each of 77 layers, which is 23.3B. The sum is the 49B active figure on the card.

256 experts, 8 at a time

Each routed expert is a small SwiGLU block of three 6,144 by 2,048 matrices, 37.7M parameters. A router scores all 256 for the current token and keeps the 8 best. The shared expert runs for every token. For one token alone, the model reads 8 experts per layer, 3.1% of the routed weights.

That fraction holds for one token at a time only. A forward pass that carries several tokens reads the union of their picks. Two tokens rarely pick the same 8 experts. With uniform routing, a pass of B tokens touches 256 x (1 − (248/256)B) distinct experts per layer. That is 30.5 experts for 4 tokens, 102 for 16, and nearly all 256 for a 512-token prefill chunk. This matters twice for a local run. Prefill reads almost the full model per chunk. Speculative decoding verifies several drafted tokens in one pass, and that pass reads more experts than a single decode step. Drive the meter to see the union grow.

Fig. 1 · the expert union meter

One MoE layer, blk.40, as 256 routed expert cells. Each token in the pass picks 8 cells. A cell darkens as more tokens pick it. The shared expert runs once per pass for all tokens. The picks are simulated with a fixed seed, and the skewed option is an assumption, not a measured Hy4 router trace. The byte counts use measured file sizes.

forward pass
tokens in this pass = 1
build
expert 0expert 255
shared expertruns once
distinct experts, this pass8
uniform formula8.0
not picked1 token2 to 3 tokens4 or more
routed bytes, this pass0
all weights, this pass0
weights per token0
pass vs 1-token step1.00x

Routed bytes = distinct experts per layer, averaged over all 77 simulated layers, x 77 layers x bytes per expert. Bytes per expert are the routed tensor bytes of each GGUF divided by 19,712 experts: IQ4_XS and Q2_K from the qtum shard headers, STQ1_0 from the anemll sidecar (807,665,664 bytes per expert slot across 77 layers). MXFP8 is computed from the format: 8 bits per weight plus one 8-bit scale per 32 weights. The other weights (attention, shared experts, embeddings) are read once per pass. Skewed routing draws picks from a Zipf-like popularity with exponent 0.8, an assumption for illustration.

The meter explains a line in the SGLang cookbook for this model. It offers two recipes: low latency with the MTP drafter on, and high throughput with it off. The cookbook says that "at saturation the draft+verify overhead outweighs the speedup." With one user, a 4-token verify pass reads about 2.0 times the weight bytes of a plain decode step at IQ4_XS. On a machine limited by memory bandwidth, the drafter must land about 2.0 tokens per pass to break even. With many users, each pass already touches a large share of the experts. The drafted tokens then add compute and save little bandwidth.

The cache arithmetic

MLA is why the per-token cache is small for a model of this size. Each layer stores 576 values per token: the 512-value latent and the 64-value RoPE key. In BF16 that is 1,152 bytes per layer.

llama.cpp stores the index keys with the same type as the main cache, F16 by default, as I read its source. That gives 21 x 256 = 5,376 bytes, and 95,232 bytes per token in total. Without IndexCache, all 78 layers would keep index keys: 78 x 128 = 9,984 bytes in FP8. The shared indexers save 7,296 bytes per token.

ContextSGLang layoutllama.cpp F16llama.cpp q8_0
8,1920.76 GB0.78 GB0.41 GB
32,7683.03 GB3.12 GB1.66 GB
131,07212.13 GB12.48 GB6.63 GB
262,14424.26 GB24.96 GB13.26 GB
1,048,57697.04 GB99.86 GB53.06 GB

The q8_0 column applies 34 bytes per 32 values to both caches. Next to the weights these numbers are small. At 131,072 tokens the cache is about 3% of the IQ4_XS file. The weights decide what you can run, and the context is a minor term until you pass 256K.

What sparse attention reads

Sparse attention does not shrink the cache. Every token stays cached, as the table above shows. What it changes is how much of the cache each new token reads. In dense MLA, a decode step at 131,072 tokens reads all 131,072 latent rows in all 78 layers, 11.78 GB. With DSA, the indexer first scores every cached token using its small 128-dim keys. Then the real attention reads only the 2,048 best rows per layer, 184 MB in total.

The scoring pass is not free. It reads one index key per cached token in every layer that owns an indexer. This is where IndexCache earns its place. Hy4 owns indexers in layers 0, 1, 5, 9 and every fourth layer after that, up to 77. The 57 layers in between reuse the picks of the nearest owner above them. Tap any layer to see whose picks it uses.

Fig. 2 · the index relay board

All 78 layers of Hy4 preview, numbered as in config.json. Blue layers own a DSA indexer and pick 2,048 positions. Grey layers reuse the picks of the owner shown under the number. Set a context length and an attention design, then read what one decode step costs on the attention side.

attention design
context = 131,072 tokens

latent rows read per layer2,048
latent bytes per token0
index scan per token0
cache held at this context0

Latent row = 576 BF16 values, 1,152 bytes. Index key = 128 FP8 bytes, the SGLang layout. Reads per decode step: latent rows x 1,152 x 78 layers, plus context x 128 bytes x indexer layers. Owner layers from indexer_types: 0, 1, 5, 9, 13 and every fourth layer to 77. Weight reads and compute are left out.

Two readings follow from the board. Past 2,048 tokens, the latent read per step stops growing, and the index scan becomes the part that scales with context. At 1M tokens, the shipped design scans 2.82 GB of index keys per step, against 10.47 GB if every layer kept its own indexer. Dense MLA would read 94.2 GB per step at that length, which no single machine streams at a usable rate. The second reading is about memory. You still need room for the full cache, so a 1M session needs about 97 GB for the cache alone.

The gated part of "Gated DSA" is the output gate from the architecture table. It changes the math of each attention output, not what is read or stored. For sizing, you can ignore it.

Every build, with its size

Sizes are sums of the files in each Hugging Face repository, in decimal gigabytes, read on October 2. The KL divergence and top-1 agreement figures are measurements by qtum against the BF16 weights on 8x H100. Compare them only within this table.

GGUF

BuildFilesQuality vs BF16Runs on
AngelSlim STQ1_0229.41 GBnot publishedpatched llama.cpp only
qtum IQ2_XS235.11 GBKLD 0.599, top-1 73.4%llama.cpp b10813+
AngelSlim UD-IQ1_M235.35 GBnot publishedbuilt with the patch, see below
avar6 IQ2_XS 2.73 bpw262.43 GBnot publishedno model card
qtum Q2_K289.54 GBKLD 0.345, top-1 80.4%llama.cpp b10813+
qtum IQ4_XS416.97 GBKLD 0.079, top-1 90.5%llama.cpp b10813+
AngelSlim Q4_K_M467.29 GBnot publishedbuilt with the patch, see below

Three notes on that table. The 6block/Hy4-preview-GGUF repository has the same three files as qtum, at the same sizes. AngelSlim is Tencent's own compression group. Its card says that "neither file runs on stock llama.cpp" because the architecture was not upstream at the time. The upstream support that came later was "developed from" that patch, per its pull request. I have not confirmed that the AngelSlim Q4_K_M and UD-IQ1_M files load on a stock build. The STQ1_0 file cannot, because the STQ1_0 type is in llama.cpp PR #22836, which is still open.

MLX, NVFP4, FP8 and BF16

BuildFilesEngine
mlx-community 4bit433.36 GBmlx-lm from a branch
inferencerlabs MLX-Q4i456.03 GBInferencer app 2.3.7
0xTank NVFP4 W4A16 + MTP453.01 GBcustom vLLM image, 4x DGX Spark
Oxmiq NVFP4 W4A16490.36 GBvLLM hy4-preview image, 8x 96 GB Blackwell
tencent Hy4-preview-FP8 (MXFP8)813.77 GBvLLM 0.29+, SGLang 0.5.20+
tencent Hy4-preview (BF16)1,559.98 GBvLLM, SGLang, multi-node

The NVFP4 builds quantize only the routed experts and keep everything else at its original precision. The FP8 repository is MXFP8 made with NVIDIA ModelOpt, with one shared 8-bit scale per 32 weights. The vLLM recipe notes that MI325X GPUs dequantize MXFP8 weights to BF16 at load, so plan memory for BF16 on that hardware.

Which machines fit

The rule I use: weights plus cache plus 4 GB of buffers must fit in usable memory. Usable means 92% of GPU memory, 80% of system RAM for offloaded experts, and 75% of a Mac's unified memory at the default wired limit. Cache is the llama.cpp F16 figure at 32,768 tokens, 3.12 GB. None of these are my measurements. Where someone published a measured speed, the last column names them.

MachineUsableBest build that fits, needVerdict
Laptop, 64 GB unified48 GBnone, smallest needs 242.2 GBno
RTX 4090 24 GB + 128 GB RAM22.1 + 102.4 GBnone, 124.5 GB short of 242.2 GBno
RTX 4090 24 GB + 256 GB RAM22.1 + 204.8 GBnone, 226.9 GB is below 242.2 GBno
48 GB GPU + 256 GB RAM44.2 + 204.8 GBqtum IQ2_XS, 242.2 GBfits with expert offload, 6.8 GB spare
RTX 4090 24 GB + 384 GB RAM22.1 + 307.2 GBqtum Q2_K, 296.7 GBfits with expert offload, tight on the GPU
Mac Studio 256 GB192 GBnone at the default limitno
Mac Studio 512 GB384 GBqtum Q2_K, 296.7 GBfits; IQ4_XS needs the limit raised to 425 GB
8x H100 80 GB588.8 GBqtum IQ4_XS, 424.1 GBfits, about 165 GB for more context
8x H20 96 GB706.6 GBAngelSlim STQ1_0 measured by AngelSlim20.47 tok/s decode, 204.56 tok/s prefill
8x RTX PRO 6000 96 GB706.6 GBOxmiq NVFP4, 490.36 GBfits, per the Oxmiq card
4x DGX Spark4 x 128 GB0xTank NVFP4 + MTP14.71 tok/s decode on prose, measured by 0xTank
8x B300 or 16x B200vendor recipeMXFP8, 813.77 GBthe vLLM recipe target

For a ceiling on decode speed, divide memory bandwidth by bytes read per token. At IQ4_XS a single decode step reads about 34.0 GB of weights, by the arithmetic behind Fig. 1. An M3 Ultra at 819 GB/s therefore tops out near 24 tokens per second, before any compute or overhead. The MLX-Q4i card from inferencerlabs reports about 10.3 tokens per second on that chip at 1,000 tokens of context. Those two numbers agree, because real engines usually reach a third to a half of the bandwidth ceiling.

llama.cpp and Ollama

Upstream support arrived in PR #28127, "add Tencent Hy 4 (hy_v4) preview architecture support," merged on September 4. The first release that contains it is b10813. The implementation covers iHC, the gated MLA, the learnable sink, the MoE, and the DSA indexer with shared layers. It drops the MTP layer, so speculative decoding with the built-in drafter does not run in llama.cpp.

# 8 GPUs that hold the whole model (qtum IQ4_XS, 417 GB)
llama-server -m Hy4-preview-IQ4_XS-00001-of-00039.gguf \
  --jinja -fa on -c 65536 \
  --temp 0.9 --top-p 1.0 \
  --host 127.0.0.1 --port 8080

# One GPU plus system RAM: attention on the card, routed experts in RAM
llama-server -m Hy4-preview-Q2_K-00001-of-00039.gguf \
  --jinja -fa on -c 32768 -ngl 999 \
  -ot "blk\..*\.ffn_.*_exps\..*=CPU" \
  -ctk q8_0 \
  --temp 0.9 --top-p 1.0

Download all 39 shards into one folder and name only the first on the command line. The qtum card advises against passing -ngl for the multi-GPU case, because llama.cpp fits layers to free VRAM by itself. The -ot regex in the second command sends every routed expert tensor to the CPU. What stays on the GPU is the 25.8B non-routed part, 20.18 GB at Q2_K by the shard headers. On a 24 GB card that leaves under 2 GB for cache and compute buffers, so keep the context short. A 48 GB card has room to spare.

The --jinja flag is required. AngelSlim's card notes that the chat template "matches no llama.cpp built-in family." Keep the GGUF on a local disk. The same card reports that loading over NFS ran at about 12 MB/s and turned "a 1-minute load into hours."

Ollama has no official library entry, and ollama.com/library/hy4-preview returns 404. A community upload linked in the Ollama issue thread, frob/hy4-preview, offers q2_k (280 GB), q4_K_M (465 GB) and q8_0 (818 GB) tags. Ollama synced past b10813 on September 10, and the first release with that sync is v0.34.1.

ollama run frob/hy4-preview:770b-a49b-q2_k

vLLM and SGLang

vLLM support merged in PR #54160 on August 29 and first shipped in v0.29.0 on September 9. The vLLM recipe targets 16x B200, 8x B300 and 8x MI355X. Its command with the MTP drafter on:

export VLLM_ENABLE_HPC_OPS=1
vllm serve tencent/Hy4-preview-FP8 \
  --tensor-parallel-size 8 \
  --speculative-config '{"num_speculative_tokens":3,"method":"mtp"}' \
  --attention-backend FLASHMLA_SPARSE \
  --tool-call-parser hy_v4 \
  --reasoning-parser hy_v4 \
  --enable-auto-tool-choice \
  --served-model-name hy4-preview --port 8000

The recipe calls 262,144 tokens "the default verified window" on MI355X. It reports that MXFP8 also starts at the full 1,048,576 at a memory utilization of 0.92. For the prebuilt container, use vllm/vllm-openai:hy4-preview.

SGLang support merged in PR #36805 on September 5 and first shipped in v0.5.20 on September 18. The SGLang cookbook covers BF16 on H200, B200, B300 and GB300, and MXFP8 on Blackwell. It suggests a context of 262,144, or 131,072 on H200. The model card's command:

sglang serve --model-path tencent/Hy4-preview-FP8 \
  --tp-size 8 \
  --reasoning-parser auto --tool-call-parser auto \
  --speculative-algorithm NEXTN \
  --speculative-num-steps 3 \
  --speculative-eagle-topk 1 \
  --speculative-num-draft-tokens 4 \
  --served-model-name hy4-preview --port 8000

Drop the four speculative flags for batch work, as Fig. 1 explains. The auto parsers resolve Hy4's suffixed special tokens, such as <think:opensource>, from the tokenizer at run time. Once several people or an agent fleet share the server, the vLLM internals guide explains how paged KV and continuous batching change the arithmetic.

On a Mac

Only the 512 GB Mac Studio holds a standard build in memory. Two paths exist. The llama.cpp path is the Metal build of b10813 or later with the qtum Q2_K at the default wired limit. For IQ4_XS, raise the limit first, for example sudo sysctl iogpu.wired_limit_mb=440000, and leave the rest of the machine idle.

The MLX path is ahead of its library. The mlx-community/Hy4-preview-4bit build was converted with mlx-lm 0.32.0. Its card says the model code is in a branch "not yet included in an mlx-lm release":

pip install git+https://github.com/kernelpool/mlx-lm.git@add-hy4-preview

mlx_lm.generate --model mlx-community/Hy4-preview-4bit \
  --max-tokens 2048 --temp 0.9 --top-p 1.0 \
  --prompt "Summarize the failing test output below."

On a 128 GB Mac, one experimental path exists. anemll/Hy4-preview-FlashMoE-STQ1_0 splits the AngelSlim STQ1_0 build into a 22.65 GB dense file and per-layer expert files. A patched llama.cpp fork streams experts from the SSD into a bank of slots in memory. The uploader measured 4.8 generated tokens per second on an M5 Max with 128 GB and a 96-slot bank, using about 93.3 GiB. That is a measurement on one machine with a 2,048-token context. It needs a fast internal SSD, and it is the only path below 192 GB.

Reasoning, sampling, tools

Sampling. Tencent recommends temperature=0.9 and top_p=1.0, and generation_config.json sets the same values. The SGLang cookbook advises against hardcoding sampling in client code, because the server applies these defaults.

Reasoning. The chat template accepts one variable, reasoning_effort, with two legal values: high (the default) and no_think. Any other value raises a template error, so low or medium from another model will fail. Send it in the request body:

{"model": "hy4-preview",
 "messages": [{"role": "user", "content": "Rename this column."}],
 "chat_template_kwargs": {"reasoning_effort": "no_think"}}

In llama.cpp, set it for the whole server with --chat-template-kwargs '{"reasoning_effort":"no_think"}'. Given the known over-reasoning issue, no_think is the better default for short edits and lookups.

Tools. Tool calls use suffixed XML-style tokens: <tool_calls:opensource>, <tool_call:opensource>, and an <arg_key> and <arg_value> pair per argument. When tools are present, the template keeps earlier reasoning in the history by default (preserved_thinking). Use --tool-call-parser hy_v4 in vLLM and auto in SGLang. For an agent, serve an OpenAI-compatible endpoint and point the harness at http://127.0.0.1:8000/v1.

Benchmarks, vendor-reported

Tencent publishes its scores as a chart image on the model card. I read the values below from that chart. They use Tencent's harness choices, and I found no independent evaluation of the open weights.

BenchmarkHy4 previewHy3GLM 5.3Kimi K3Claude Opus 5
Terminal Bench 2.185.470.888.385.785.4
DeepSWE64.328.068.174.074.7
SWE Atlas Refactoring53.332.951.937.460.0
Toolathlon-Verified74.156.273.874.776.5
HLE, text only, no tools43.434.442.346.653.2

The jump over Hy3 is large on every row, and largest on DeepSWE, where the score more than doubles. Against the other open models in the chart, Hy4 preview is level or close, and it trails Claude Opus 5 on most rows. The card also reports a blind side-by-side test: 163 Tencent experts rated outputs on 203 engineering tasks. Hy4 preview averaged 2.99, against 2.92 for GLM 5.3 and 2.94 for Kimi K3.

Failure modes and fixes

FAQ

Can I run Hy4 preview on a laptop or a single 24 GB GPU?

No for any standard build. The smallest GGUF is 229 GB, so a 24 GB card needs about 230 GB of system RAM beside it. On a 128 GB Mac, only the experimental SSD-streaming fork runs it, at 4.8 tokens per second as measured by its uploader.

What does the preview label mean for Hy4?

Tencent calls it an early version with known issues, such as long reasoning and repeated self-verification. Hy3 preview became Hy3 in about eleven weeks with the same parameter count. Expect a replacement checkpoint and plan to download again.

Does sparse attention make the 1M context cheap in memory?

No. DSA cuts the rows each token reads to 2,048 per layer, but every token stays cached. At about 92.5 KB per token, 1M tokens need about 97 GB of cache.

Which build should I download?

IQ4_XS from qtum at 417 GB if it fits, because it has the lowest measured KL divergence of the published GGUFs. Q2_K at 290 GB is the next step down. The STQ1_0 build needs a patched llama.cpp.

Does Hy4 preview have a license restriction?

The weights are under the Apache License 2.0, with a Tencent copyright notice in the LICENSE file. There is no separate use policy or user-count limit in the repository.

rg
Rohit Ghumare

CNCF Ambassador and Google Developer Expert. I build agent infrastructure and write about the fundamentals underneath the AI stack. Config values, tensor sizes, and file sizes here come from the Hugging Face repositories of each uploader named above. Engine status comes from the llama.cpp, vLLM, SGLang and Ollama repositories, the vLLM recipe, and the SGLang cookbook, read on October 2, 2026. I did not run this model on my own hardware.

DeepSeek V4 · Hardware guide · More guides · X