How to Run MiniMax-M3 Locally
MiniMax published MiniMax-M3 on Hugging Face on June 2, 2026. It is a 428B mixture-of-experts model with about 23B parameters active per token. It reads text, images and video, and its config allows 1,048,576 tokens of context. Two things make it unusual to run at home. Its attention is sparse by training: each layer reads 16 blocks of 128 tokens, not the whole cache. And at 4 bits the weights are about 250 GB, so most machines must spread the experts across GPU, RAM and SSD. This guide covers the license, the config, the cache arithmetic, every quant size, computed fit per machine, and the commands per engine. I checked everything on October 2, 2026. I did not run this model myself, so every speed here is a computed ceiling or a cited number from someone else.
What MiniMax released
The release is one checkpoint in two precisions. The main repository holds BF16 weights in 59 safetensors shards, 854.18 GB in total. The safetensors metadata counts 427,040,140,160 parameters, which the card rounds to "~428B." An official MXFP8 variant went up the same day at 443.75 GB. The architecture class is MiniMaxM3SparseForConditionalGeneration, and the vision tower ships inside the same files.
The card names three selling points. The model trains on text, images and video "from the very first step." It introduces MiniMax Sparse Attention (MSA), which the card says gives "9x prefill and 15x decode speedups compared to M2 at 1M context." And it targets long agentic coding work. The method has its own paper, arXiv 2606.13392, posted on June 11. MiniMax also open-sourced the MSA kernels under MIT, but they target NVIDIA SM100 (Blackwell datacenter) GPUs only.
One caveat on context. The config sets max_position_embeddings to 1,048,576 with no RoPE scaling. The vLLM recipe notes for AMD say the native context is 512K and show a YaRN override for 1M. Ollama's cloud page promises "a guaranteed minimum of 512K tokens." I could not reconcile these from primary files, so plan local work around 512K or less.
The license, clause by clause
The repository's LICENSE file is the MiniMax Community License. It is MIT-shaped text with a different grant, and the differences matter if you ship anything.
- The base grant is non-commercial. You may use, copy, modify and distribute the model "for non-commercial purposes." Running it at home for your own work fits this grant.
- Commercial use has two conditions. First, show "Built with MiniMax M3" on a related site, interface, blog post or product documentation. Second, tell MiniMax.
- The notice depends on revenue. Under 20 million US dollars a year, send a one-time notice to api@minimax.io. Above that, you need prior written authorization before you use it.
- "Commercial Use" is broad. It includes paid products that rely on the model, commercial API use, and commercial deployment of any fine-tuned or modified version.
- An appendix bans some uses outright. The list includes any military purpose, harm to minors, and harmful disinformation.
The threshold is yearly revenue, not monthly users. This is my reading of the text, not legal advice. Ollama states that its cloud tag is "officially licensed with MiniMax for commercial usage." That is a separate agreement, and it does not cover your own deployment.
The architecture
The values below come from the raw config.json. I checked the tensor shapes against the safetensors headers of shards 1 and 3.
| Component | Value |
|---|---|
| Layers | 60: layers 0 to 2 dense with full attention, layers 3 to 59 MoE with MSA |
| Hidden size | 6,144 |
| Attention | 64 query heads, 4 KV heads, head dim 128, per-head QK norm |
| RoPE | 64 of 128 dims rotated, theta 5,000,000 |
| Routed experts | 128 per layer, 4 per token, FFN width 3,072, sigmoid scoring with a routing bias |
| Shared expert | 1 per MoE layer, FFN width 3,072 |
| Dense FFN (layers 0 to 2) | 12,288 |
| MSA indexer | 4 index query heads and 1 index key of 128 dims per layer |
| MSA selection | blocks of 128 tokens, top 16 blocks per GQA group, local block forced |
| Vocab / context | 200,064 / 1,048,576 |
| Vision | 32-layer ViT, width 1,280, patch 14, 2x2 patch merge, up to 2,016 px |
The parameter split explains the whole offload story. One routed expert is three 3,072 by 6,144 matrices, 56.6M parameters. Across 57 layers and 128 experts that is 413.1B, or 96.9% of the text model. The rest is small: 6.4B of attention, 3.2B of shared experts, 2.5B of embeddings, 0.7B of dense FFN. Per token, the four routed experts in each layer add 12.9B. Add everything that always runs and my count is 24.7B, including the output head. The card says about 23B.
The config lists num_mtp_modules: 7, but the weight index has no MTP tensors. The llama.cpp PR author confirms "MTP heads are not present in the released weights." A third-party EAGLE-3 drafter, Inferact/MiniMax-M3-EAGLE3 at 6.53 GB, fills that gap for vLLM.
Sparse attention and the cache
MSA is grouped-query attention with a scout in front. In each of the 57 MSA layers, a small indexer compares the current query with one 128-dim index key per cached token. It max-pools those scores into 128-token blocks. Each of the four GQA groups then keeps its own top 16 blocks, and the block that holds the current token is always included. The 16 query heads in that group attend to those 2,048 tokens only. The three dense layers still attend to everything.
This is a trained behavior. The llama.cpp PR that added MSA says it plainly: "MSA is not an optional speed optimization, as the model is trained with sparse attention." Dense attention on this model is an approximation that degrades output.
What the cache stores
MSA does not make the cache smaller. Every token still writes keys and values in all 60 layers, and an index key in the 57 MSA layers. The arithmetic for llama.cpp, which keeps index keys in F32:
- Keys and values: 60 layers x 4 KV heads x 128 dims x 2 tensors x 2 bytes = 122,880 bytes per token.
- Index keys: 57 layers x 128 dims x 4 bytes = 29,184 bytes per token.
- Total: 152,064 bytes per token at F16. With a q8_0 cache the K and V part drops to 65,280 bytes, so 94,464 in total.
| Context | F16 K/V + F32 index | q8_0 K/V + F32 index | K/V only, FP8 (vLLM) |
|---|---|---|---|
| 32,768 | 4.98 GB | 3.10 GB | 2.01 GB |
| 65,536 | 9.97 GB | 6.19 GB | 4.03 GB |
| 131,072 | 19.93 GB | 12.38 GB | 8.05 GB |
| 262,144 | 39.86 GB | 24.76 GB | 16.11 GB |
| 524,288 | 79.73 GB | 49.53 GB | 32.21 GB |
| 1,048,576 | 159.45 GB | 99.05 GB | 64.42 GB |
The vLLM column leaves out index keys because I could not confirm their storage type there. The vLLM engine requires --block-size 128 so its cache pages line up with MSA blocks.
What one decode step reads
The saving is in the read. A dense layer reads 2,048 bytes of K and V per cached token. An MSA layer reads 2,048 tokens of K and V, plus the index key of every cached token. At 131,072 tokens that is 4.87 GB per step against 16.11 GB for a dense model of the same shape. The picker below lets you drive it.
One MSA layer, one decode step, inside a coding-agent session. Pick the question the agent is answering. Each GQA group keeps its own 16 blocks, and the black tick is the forced local block. The block scores are an illustration with a fixed seed. The byte counts are exact for the config.
The current questionRead per step = 3 dense layers x context x 2,048 bytes, plus 57 MSA layers x (2,048 tokens x 2,048 bytes + context x index-key bytes). Cache stored = context x (122,880 + 57 x 128 x index-key bytes). Index keys are F32 (4 bytes) in llama.cpp. The 2-byte option shows a half-width layout for comparison.
Two things stand out. First, MSA makes attention reads grow slowly, not stay flat. The index keys still scale with context, and at 1M tokens in llama.cpp they are 30.6 GB of a 37.3 GB read. Second, the dense fallback is not a quality-neutral switch. Old GGUFs without the indexer tensors, or llama.cpp with flash attention off, take that path and read every block.
The SGLang cookbook uses the same structure for cache offload. Its HiSparse mode keeps the three dense layers' cache on the GPU and moves the 57 sparse layers' K and V to pinned host memory. It then copies in only the selected blocks.
Every quant, with sizes
File sizes are sums of shard sizes from the Hugging Face tree API on October 2, in decimal gigabytes.
Serving formats (vLLM, SGLang)
| Repository | Format | Size | Target |
|---|---|---|---|
MiniMaxAI/MiniMax-M3 | BF16 | 854.18 GB | 8x H200 or 8x H20, TP8 |
MiniMaxAI/MiniMax-M3-MXFP8 | MXFP8 | 443.75 GB | Blackwell, MI350X/MI355X, KTransformers |
EmbeddedLLM/MiniMax-M3-FP8-dynamic | FP8 | 430.69 GB | community |
cyankiwi/MiniMax-M3-AWQ-INT4 | AWQ INT4 | 258.02 GB | community |
nvidia/MiniMax-M3-NVFP4 | NVFP4 | 250.10 GB | B200, vLLM PR #46380 |
amd/MiniMax-M3-MXFP4 | MXFP4 | 242.67 GB | MI355X |
GGUF that loads on mainline llama.cpp
AesSedai's builds come with measured perplexity and KL divergence against BF16. They keep shared tensors at Q8_0 or Q6_K and squeeze only the expert FFNs.
| AesSedai build | Files | PPL vs BF16 | KLD |
|---|---|---|---|
IQ2_S | 143.90 GB | +27.99% | 0.398 |
IQ3_S | 158.96 GB | +23.76% | 0.321 |
IQ4_XS | 203.07 GB | +9.10% | 0.151 |
Q4_K_M | 264.26 GB | +1.88% | 0.069 |
Q5_K_M | 316.98 GB | +0.65% | 0.042 |
bartowski's builds were made with llama.cpp b10141, the first release with MSA. They have no perplexity table, but they cover more sizes.
| bartowski build | Files | bartowski build | Files |
|---|---|---|---|
IQ1_S | 90.53 GB | IQ4_XS | 229.69 GB |
IQ1_M | 100.74 GB | Q4_K_S | 251.36 GB |
IQ2_XXS | 116.61 GB | Q4_K_M | 261.28 GB |
IQ2_M | 145.58 GB | Q5_K_M | 305.35 GB |
Q2_K_L | 153.09 GB | Q6_K | 369.40 GB |
IQ3_XXS | 180.01 GB | Q8_0 | 453.61 GB |
Q3_K_M | 196.61 GB | mmproj f16 | 1.73 GB |
The vision projector is a separate file. AesSedai also publishes a Q8_0 projector at 0.93 GB. The unsloth repository lists 21 builds, from UD-IQ1_M at 128.42 GB to BF16 at 852.00 GB. Its files were uploaded on June 12 from preliminary PR #24523, which had no sparse attention. The merged PR states that GGUFs need the indexer tensors and that other conversions "will fail to load." I have not loaded the unsloth files, so treat them as unverified on current builds.
MLX
| Repository | Size | Notes |
|---|---|---|
pipenetwork/MiniMax-M3-MLX-3bit | 186.39 GB | 256 GB Mac with a raised wired limit |
pipenetwork/MiniMax-M3-MLX-4bit | 239.62 GB | 512 GB Mac |
mlx-community/MiniMax-M3-4bit | 241.48 GB | converted with mlx-vlm 0.6.3 |
pipenetwork/MiniMax-M3-MLX-8bit | 452.58 GB | 512 GB Mac, short context |
Where the experts go
With 7,296 routed expert slots (57 layers x 128) and only 228 used per token, placement decides speed. The rule for llama.cpp on a discrete GPU follows from the arithmetic above. Put everything that runs on every token on the GPU: attention, shared experts, dense layers, the cache and the compute buffer. Then fill leftover VRAM with whole expert layers. System RAM holds the rest. Anything that does not fit in RAM pages in from the SSD through mmap, which llama.cpp uses by default.
Each token then costs three reads. The GPU reads the always-on weights and the attention bytes. The CPU reads its share of the 228 routed experts from RAM. Missing experts come from the SSD. Divide each by its bandwidth and add them, and you get a ceiling. Real engines land below it. The tape below runs that model on eight simulated tokens.
Each token routes to 4 of 128 experts in all 57 MoE layers, blk.3 to blk.59. Each cell is one routed expert, colored by where it sits. Tap a token row to see its 228 picks. The ceiling is exact for the placement. The 8-token tape samples routes with a fixed seed.
Usable memory: GPUs 92% of VRAM, system RAM 85%, Macs 75% of unified memory (the default wired limit). Always-on = non-expert weights at the build's base type (Q8_0, or Q6_K for IQ4_XS and below) + cache + 4.94 GB compute buffer (4.6 GiB, from PR #24908). RAM is dual-channel DDR5-5600 on the two consumer boxes and 8-channel DDR5-4800 on the workstation. Bandwidths are peak spec values. SSD rates are typical for the PCIe generation, and 6 GB/s for the Mac SSD is an assumption. Skewed routing is an assumed Zipf curve (s = 0.8). MiniMax has not published expert-use statistics. Hottest-experts placement is the idea behind KTransformers' --kt-num-gpu-experts, whose documented runs use the MXFP8 weights, not these GGUF files.
The computed results for the presets, at 32K context, f16 cache, uniform routing and whole-layer placement:
| Machine | Build | Placement | Ceiling |
|---|---|---|---|
| RTX 4090 24 GB + 128 GB DDR5 | IQ2_S, q8_0 cache | 1 expert layer on GPU, 16.9% of expert reads from SSD | 6.7 tok/s |
| RTX 4090 24 GB + 128 GB DDR5 | Q4_K_M | always-on part needs 23.8 GB, 22.1 GB usable | does not start |
| RTX 5090 32 GB + 256 GB DDR5 | IQ4_XS | 2 layers on GPU, no SSD reads | 14.1 tok/s |
| RTX 5090 32 GB + 256 GB DDR5 | Q4_K_M | 1 layer on GPU, 11.5% from SSD | 6.8 tok/s |
| RTX PRO 6000 96 GB + 512 GB 8-channel | Q4_K_M | 14 layers on GPU, no SSD reads | 35.6 tok/s |
| Mac Studio M3 Ultra 256 GB | IQ3_S | all resident | 52.1 tok/s |
| Mac Studio M3 Ultra 256 GB | IQ4_XS | 10.9% from SSD | 7.7 tok/s |
| Mac Studio M3 Ultra 512 GB | Q5_K_M | all resident | 35.0 tok/s |
Read the table with three facts in mind. First, the SSD tier is expensive as soon as it appears. On the 5090 box, 11.5% of expert reads take 64 ms per token, nearly as long as the 76 ms for the 86.7% in RAM. Second, the 24 GB card fails on the always-on part alone, because the Q8_0 non-expert tensors are 13.9 GB before any cache. Third, IQ3_S on a 256 GB Mac is fast but carries a measured 23.76% perplexity cost. For a measured point, the MSA PR author reports 7.7 tok/s decode at 5K and 7.67 at 62K on an expert-offload setup. The PR does not name the hardware.
Smaller machines can still load something. A 128 GB DGX Spark has about 115 GB usable at 90%, enough for bartowski's IQ1_M at 100.74 GB with a q8_0 cache at 32K. Nobody has published the quality cost of a 1-bit M3, so test it on your own tasks first. At the other end, the BF16 weights need a node of eight 141 GB H200s, and that is the configuration both vLLM and SGLang validate.
llama.cpp
Three merges matter. PR #24908 added the model with MSA on July 26 and first shipped in b10141. Vision support (#25113) shipped in b10142. A dedicated tool-call parser (#26210) shipped in b10164. Quantized K and V caches (#26180) first worked in b10213. Use b10213 or newer. I checked the flags below against master at b11342.
# 1. Download one build (here AesSedai Q4_K_M, 7 shards)
hf download AesSedai/MiniMax-M3-GGUF \
--include "Q4_K_M/*" --local-dir models/MiniMax-M3
# 2. Serve it on a 96 GB GPU with 512 GB of RAM
llama-server \
-m models/MiniMax-M3/Q4_K_M/MiniMax-M3-Q4_K_M-00001-of-00007.gguf \
--jinja -fa on -ngl 999 -c 32768 \
--n-cpu-moe 46 \
-np 1 \
--temp 1.0 --top-p 0.95 --top-k 40 \
--host 127.0.0.1 --port 8080
--n-cpu-moe N keeps the expert weights of the first N layers in system RAM. The figure prints the value for your placement: 60 minus the number of expert layers that fit on the GPU. Start there, then lower N until VRAM is nearly full. For a 24 GB card, add -ctk q8_0 -ctv q8_0 and use --cpu-moe, which keeps all experts on the CPU.
Three flags are not optional here. -fa on is required, because MSA calls flash attention directly. With flash attention off, the model falls back to dense attention and warns once. --jinja enables the chat template and the M3 tool-call parser. Keep -np 1, or set several slots without --kv-unified. The PR explains that a unified cache with several sequences "would silently break block anchoring." For images, add --mmproj with one of the projector files.
Ollama has no local build. Its library page lists only minimax-m3:cloud, with a 512K context, which runs on Ollama's servers.
vLLM, SGLang, KTransformers
For a GPU server, the vLLM recipe uses a dedicated image, vllm/vllm-openai:minimax-m3, because support "has not yet shipped in a stable vLLM release." The block size flag is mandatory on every platform:
vllm serve MiniMaxAI/MiniMax-M3 \
--tensor-parallel-size 8 \
--block-size 128 \
--max-model-len 131072 \
--kv-cache-dtype fp8 \
--tool-call-parser minimax_m3 \
--reasoning-parser minimax_m3 \
--enable-auto-tool-choice
The recipe calls the FP8 cache "lossless in our testing" and says it gives about 1.5x the KV pool. Add --language-model-only for text work to skip the vision encoder. SGLang publishes per-GPU images such as lmsysorg/sglang:dev-cu12-minimax-m3 for H200. On Hopper it serves the BF16 weights, because its MXFP8 kernels are Blackwell-only:
python3 -m sglang.launch_server \
--model-path MiniMaxAI/MiniMax-M3 \
--trust-remote-code --tp 8 \
--mem-fraction-static 0.75 \
--reasoning-parser auto --tool-call-parser auto \
--host 0.0.0.0 --port 30000
KTransformers is the documented route for one big GPU plus a lot of RAM. It runs the MXFP8 weights through an SGLang fork. A set number of experts per layer stay on the GPU, and the CPU computes the rest. Its tutorial says "a single 96 GB Hopper card is sufficient if most routed experts stay on CPU." It suggests 2 to 8 GPU experts per layer at TP1. The MXFP8 files are 443.75 GB, so plan for about 512 GB of RAM:
python -m sglang.launch_server \
--model-path /models/MiniMax-M3-MXFP8 \
--kt-weight-path /models/MiniMax-M3-MXFP8 \
--kt-method MXFP8 --kt-cpuinfer 64 \
--kt-threadpool-count 2 --kt-num-gpu-experts 4 \
--kt-gpu-prefill-token-threshold 500 \
--tp-size 1 --quantization mxfp8 \
--moe-runner-backend triton --trust-remote-code \
--mem-fraction-static 0.85 --chunked-prefill-size 4096 \
--cuda-graph-max-bs 1 \
--tool-call-parser minimax-m3 --reasoning-parser minimax-m3 \
--served-model-name MiniMax-M3 --port 8000
The tutorial lists Hopper (SM90) as the supported GPU family for this hybrid path. Its own advice for out-of-memory errors is to lower --kt-num-gpu-experts, --mem-fraction-static, or --chunked-prefill-size. Note the different parser spelling: minimax-m3 here, minimax_m3 in vLLM.
On a Mac
The MLX path runs through mlx-vlm, not mlx-lm, because the model type is minimax_m3_vl. The mlx-vlm model README says its implementation covers text, image and video, MSA, thinking tags, tool calls and the EAGLE-3 drafter. The mlx-community 4-bit build was converted with mlx-vlm 0.6.3. Update to a current mlx-vlm before you load it.
pip install -U mlx-vlm
mlx_vlm.generate --model mlx-community/MiniMax-M3-4bit \
--thinking-mode enabled --max-tokens 1024 \
--prompt "Explain the failing test in this log."
The 4-bit build is 241.48 GB, so it needs a 512 GB machine. On a 256 GB Mac, the 3-bit build at 186.39 GB fits once you raise the GPU wired limit with sudo sysctl iogpu.wired_limit_mb. The llama.cpp GGUF path also runs on Metal in principle. The MSA PR lists Metal as "untested," and I found no Metal-specific report either way.
Thinking, sampling, tools
Sampling. The card recommends temperature=1.0 and top_p=0.95. The vLLM recipe and the unsloth guide add top_k=40. The vLLM recipe also gives a default system prompt: "You are a helpful assistant. Your name is MiniMax-M3 and is built by MiniMax."
Thinking. The chat template takes a thinking_mode of enabled, disabled or adaptive. Adaptive is the default, and the model then decides per turn. Send it per request as "chat_template_kwargs": {"thinking_mode": "disabled"}. Reasoning arrives between <mm:think> tags, and the parsers move it into reasoning_content. Current vLLM names that field reasoning and keeps the old name as an alias.
Tools. M3 writes tool calls as namespaced XML and writes nested arguments as nested XML. That is why llama.cpp needed its own parser in b10164. The minimax_m3 parsers in vLLM and SGLang convert these calls to the OpenAI tool_calls array. Every engine above exposes an OpenAI-compatible endpoint, so point your harness at http://127.0.0.1:8080/v1 or the engine's port.
Benchmarks, vendor-reported
These are MiniMax's numbers from the card's benchmark chart, as transcribed in the repository's .eval_results file. MiniMax ran the coding rows with Claude Code as the agent scaffold. I found no independent evaluation of the open weights.
| Benchmark | MiniMax-M3 | Notes from the source |
|---|---|---|
| SWE-bench Verified | 80.5 | average of 4 runs |
| SWE-bench Pro | 59.0 | Claude Code scaffold |
| Terminal-Bench 2.1 | 66.0 | quoted in the vLLM recipe |
| MMMU-Pro | 78.1 | variant not stated |
| VideoMME (with subtitles) | 85.4 | |
| Claw-Eval | 74.5 | overall score |
| Apex-Agents | 27.7 |
The quantized builds do not inherit these scores. AesSedai's table above is the only published measure of quant loss so far, and it is perplexity on text, not task success.
Failure modes and fixes
- A GGUF fails to load with missing tensors. It predates the MSA converter. Use a build made with b10141 or later, such as bartowski or AesSedai.
- The log warns about dense attention. Flash attention is off, or you combined
--kv-unifiedwith several slots. Pass-fa onand drop the unified cache. - q8_0 or q4 cache types crash at startup. That was issue #26155. Update to b10213 or newer.
- Long chats go wrong after a context trim. Block selection is tied to absolute cache slots. Context shift is not supported, and erasing a prefix breaks selection silently. Start a new session instead of shifting.
- vLLM crashes with "No common block size for 16". Add
--block-size 128. - Tool calls come back as plain text. Use the
minimax_m3parsers, not theminimax_m2ones from the previous release. In llama.cpp, use b10164 or newer with--jinja. - Decode crawls below 2 tokens per second. Experts are paging from the SSD. Pick a smaller build, add RAM, or move fewer expert layers to the GPU so the cache fits.
- Images are ignored. Pass
--mmprojin llama.cpp. On AMD, SGLang treats vision as unvalidated.
FAQ
Can I run MiniMax-M3 on one consumer GPU?
Only with a lot of system RAM. The always-on weights, cache and buffer need about 19 to 24 GB of VRAM at 32K context. The routed experts then live in RAM. A 24 GB card with 128 GB of RAM runs IQ2_S with a q8_0 cache at a computed ceiling of about 6.7 tokens per second.
Does sparse attention make the KV cache smaller?
No. Every token still stores keys and values in all 60 layers, plus an index key in 57 layers. That is 152,064 bytes per token in llama.cpp. MSA cuts what each decode step reads, because each GQA group attends to only 16 blocks of 128 tokens.
Which GGUF files work with mainline llama.cpp?
Files converted with the MSA support from PR #24908, first shipped in b10141. The bartowski and AesSedai repositories were converted after that merge. The unsloth files came from an earlier preliminary branch, and the merged PR says such files fail to load.
Can I use MiniMax-M3 commercially?
Under conditions. The base grant is non-commercial. Commercial use needs a visible "Built with MiniMax M3" notice and a one-time notice email to MiniMax. Above 20 million US dollars of yearly revenue, you need prior written authorization.
Is there an Ollama model for MiniMax-M3?
Only a cloud tag, minimax-m3:cloud, with a 512K context and no local download. For your own hardware, use llama.cpp, vLLM, SGLang, KTransformers or mlx-vlm.
Keep reading