How to Run MiniMax-M3 Locally

MiniMax published MiniMax-M3 on Hugging Face on June 2, 2026. It is a 428B mixture-of-experts model with about 23B parameters active per token. It reads text, images and video, and its config allows 1,048,576 tokens of context. Two things make it unusual to run at home. Its attention is sparse by training: each layer reads 16 blocks of 128 tokens, not the whole cache. And at 4 bits the weights are about 250 GB, so most machines must spread the experts across GPU, RAM and SSD. This guide covers the license, the config, the cache arithmetic, every quant size, computed fit per machine, and the commands per engine. I checked everything on October 2, 2026. I did not run this model myself, so every speed here is a computed ceiling or a cited number from someone else.

428B total · ~23B active 128 experts · 4 routed + 1 shared MSA 16 x 128-token blocks license MiniMax Community llama.cpp b10213+

What MiniMax released

The release is one checkpoint in two precisions. The main repository holds BF16 weights in 59 safetensors shards, 854.18 GB in total. The safetensors metadata counts 427,040,140,160 parameters, which the card rounds to "~428B." An official MXFP8 variant went up the same day at 443.75 GB. The architecture class is MiniMaxM3SparseForConditionalGeneration, and the vision tower ships inside the same files.

The card names three selling points. The model trains on text, images and video "from the very first step." It introduces MiniMax Sparse Attention (MSA), which the card says gives "9x prefill and 15x decode speedups compared to M2 at 1M context." And it targets long agentic coding work. The method has its own paper, arXiv 2606.13392, posted on June 11. MiniMax also open-sourced the MSA kernels under MIT, but they target NVIDIA SM100 (Blackwell datacenter) GPUs only.

One caveat on context. The config sets max_position_embeddings to 1,048,576 with no RoPE scaling. The vLLM recipe notes for AMD say the native context is 512K and show a YaRN override for 1M. Ollama's cloud page promises "a guaranteed minimum of 512K tokens." I could not reconcile these from primary files, so plan local work around 512K or less.

The license, clause by clause

The repository's LICENSE file is the MiniMax Community License. It is MIT-shaped text with a different grant, and the differences matter if you ship anything.

The threshold is yearly revenue, not monthly users. This is my reading of the text, not legal advice. Ollama states that its cloud tag is "officially licensed with MiniMax for commercial usage." That is a separate agreement, and it does not cover your own deployment.

The architecture

The values below come from the raw config.json. I checked the tensor shapes against the safetensors headers of shards 1 and 3.

ComponentValue
Layers60: layers 0 to 2 dense with full attention, layers 3 to 59 MoE with MSA
Hidden size6,144
Attention64 query heads, 4 KV heads, head dim 128, per-head QK norm
RoPE64 of 128 dims rotated, theta 5,000,000
Routed experts128 per layer, 4 per token, FFN width 3,072, sigmoid scoring with a routing bias
Shared expert1 per MoE layer, FFN width 3,072
Dense FFN (layers 0 to 2)12,288
MSA indexer4 index query heads and 1 index key of 128 dims per layer
MSA selectionblocks of 128 tokens, top 16 blocks per GQA group, local block forced
Vocab / context200,064 / 1,048,576
Vision32-layer ViT, width 1,280, patch 14, 2x2 patch merge, up to 2,016 px

The parameter split explains the whole offload story. One routed expert is three 3,072 by 6,144 matrices, 56.6M parameters. Across 57 layers and 128 experts that is 413.1B, or 96.9% of the text model. The rest is small: 6.4B of attention, 3.2B of shared experts, 2.5B of embeddings, 0.7B of dense FFN. Per token, the four routed experts in each layer add 12.9B. Add everything that always runs and my count is 24.7B, including the output head. The card says about 23B.

The config lists num_mtp_modules: 7, but the weight index has no MTP tensors. The llama.cpp PR author confirms "MTP heads are not present in the released weights." A third-party EAGLE-3 drafter, Inferact/MiniMax-M3-EAGLE3 at 6.53 GB, fills that gap for vLLM.

Sparse attention and the cache

MSA is grouped-query attention with a scout in front. In each of the 57 MSA layers, a small indexer compares the current query with one 128-dim index key per cached token. It max-pools those scores into 128-token blocks. Each of the four GQA groups then keeps its own top 16 blocks, and the block that holds the current token is always included. The 16 query heads in that group attend to those 2,048 tokens only. The three dense layers still attend to everything.

This is a trained behavior. The llama.cpp PR that added MSA says it plainly: "MSA is not an optional speed optimization, as the model is trained with sparse attention." Dense attention on this model is an approximation that degrades output.

What the cache stores

MSA does not make the cache smaller. Every token still writes keys and values in all 60 layers, and an index key in the 57 MSA layers. The arithmetic for llama.cpp, which keeps index keys in F32:

ContextF16 K/V + F32 indexq8_0 K/V + F32 indexK/V only, FP8 (vLLM)
32,7684.98 GB3.10 GB2.01 GB
65,5369.97 GB6.19 GB4.03 GB
131,07219.93 GB12.38 GB8.05 GB
262,14439.86 GB24.76 GB16.11 GB
524,28879.73 GB49.53 GB32.21 GB
1,048,576159.45 GB99.05 GB64.42 GB

The vLLM column leaves out index keys because I could not confirm their storage type there. The vLLM engine requires --block-size 128 so its cache pages line up with MSA blocks.

What one decode step reads

The saving is in the read. A dense layer reads 2,048 bytes of K and V per cached token. An MSA layer reads 2,048 tokens of K and V, plus the index key of every cached token. At 131,072 tokens that is 4.87 GB per step against 16.11 GB for a dense model of the same shape. The picker below lets you drive it.

Fig. 1 · per-group block picker

One MSA layer, one decode step, inside a coding-agent session. Pick the question the agent is answering. Each GQA group keeps its own 16 blocks, and the black tick is the forced local block. The block scores are an illustration with a fixed seed. The byte counts are exact for the config.

The current question
context = 131,072 tokens
context
blocks per layer0
tokens attended0
read per step0
cache stored0

Read per step = 3 dense layers x context x 2,048 bytes, plus 57 MSA layers x (2,048 tokens x 2,048 bytes + context x index-key bytes). Cache stored = context x (122,880 + 57 x 128 x index-key bytes). Index keys are F32 (4 bytes) in llama.cpp. The 2-byte option shows a half-width layout for comparison.

Two things stand out. First, MSA makes attention reads grow slowly, not stay flat. The index keys still scale with context, and at 1M tokens in llama.cpp they are 30.6 GB of a 37.3 GB read. Second, the dense fallback is not a quality-neutral switch. Old GGUFs without the indexer tensors, or llama.cpp with flash attention off, take that path and read every block.

The SGLang cookbook uses the same structure for cache offload. Its HiSparse mode keeps the three dense layers' cache on the GPU and moves the 57 sparse layers' K and V to pinned host memory. It then copies in only the selected blocks.

Every quant, with sizes

File sizes are sums of shard sizes from the Hugging Face tree API on October 2, in decimal gigabytes.

Serving formats (vLLM, SGLang)

RepositoryFormatSizeTarget
MiniMaxAI/MiniMax-M3BF16854.18 GB8x H200 or 8x H20, TP8
MiniMaxAI/MiniMax-M3-MXFP8MXFP8443.75 GBBlackwell, MI350X/MI355X, KTransformers
EmbeddedLLM/MiniMax-M3-FP8-dynamicFP8430.69 GBcommunity
cyankiwi/MiniMax-M3-AWQ-INT4AWQ INT4258.02 GBcommunity
nvidia/MiniMax-M3-NVFP4NVFP4250.10 GBB200, vLLM PR #46380
amd/MiniMax-M3-MXFP4MXFP4242.67 GBMI355X

GGUF that loads on mainline llama.cpp

AesSedai's builds come with measured perplexity and KL divergence against BF16. They keep shared tensors at Q8_0 or Q6_K and squeeze only the expert FFNs.

AesSedai buildFilesPPL vs BF16KLD
IQ2_S143.90 GB+27.99%0.398
IQ3_S158.96 GB+23.76%0.321
IQ4_XS203.07 GB+9.10%0.151
Q4_K_M264.26 GB+1.88%0.069
Q5_K_M316.98 GB+0.65%0.042

bartowski's builds were made with llama.cpp b10141, the first release with MSA. They have no perplexity table, but they cover more sizes.

bartowski buildFilesbartowski buildFiles
IQ1_S90.53 GBIQ4_XS229.69 GB
IQ1_M100.74 GBQ4_K_S251.36 GB
IQ2_XXS116.61 GBQ4_K_M261.28 GB
IQ2_M145.58 GBQ5_K_M305.35 GB
Q2_K_L153.09 GBQ6_K369.40 GB
IQ3_XXS180.01 GBQ8_0453.61 GB
Q3_K_M196.61 GBmmproj f161.73 GB

The vision projector is a separate file. AesSedai also publishes a Q8_0 projector at 0.93 GB. The unsloth repository lists 21 builds, from UD-IQ1_M at 128.42 GB to BF16 at 852.00 GB. Its files were uploaded on June 12 from preliminary PR #24523, which had no sparse attention. The merged PR states that GGUFs need the indexer tensors and that other conversions "will fail to load." I have not loaded the unsloth files, so treat them as unverified on current builds.

MLX

RepositorySizeNotes
pipenetwork/MiniMax-M3-MLX-3bit186.39 GB256 GB Mac with a raised wired limit
pipenetwork/MiniMax-M3-MLX-4bit239.62 GB512 GB Mac
mlx-community/MiniMax-M3-4bit241.48 GBconverted with mlx-vlm 0.6.3
pipenetwork/MiniMax-M3-MLX-8bit452.58 GB512 GB Mac, short context

Where the experts go

With 7,296 routed expert slots (57 layers x 128) and only 228 used per token, placement decides speed. The rule for llama.cpp on a discrete GPU follows from the arithmetic above. Put everything that runs on every token on the GPU: attention, shared experts, dense layers, the cache and the compute buffer. Then fill leftover VRAM with whole expert layers. System RAM holds the rest. Anything that does not fit in RAM pages in from the SSD through mmap, which llama.cpp uses by default.

Each token then costs three reads. The GPU reads the always-on weights and the attention bytes. The CPU reads its share of the 228 routed experts from RAM. Missing experts come from the SSD. Divide each by its bandwidth and add them, and you get a ceiling. Real engines land below it. The tape below runs that model on eight simulated tokens.

Fig. 2 · expert hit tape across VRAM, RAM, SSD

Each token routes to 4 of 128 experts in all 57 MoE layers, blk.3 to blk.59. Each cell is one routed expert, colored by where it sits. Tap a token row to see its 228 picks. The ceiling is exact for the placement. The 8-token tape samples routes with a fixed seed.

Machine Build
VRAMsystem RAMpaged from SSD
always-on, GPU0
expert reads / token0
ceiling0
tape average0

Usable memory: GPUs 92% of VRAM, system RAM 85%, Macs 75% of unified memory (the default wired limit). Always-on = non-expert weights at the build's base type (Q8_0, or Q6_K for IQ4_XS and below) + cache + 4.94 GB compute buffer (4.6 GiB, from PR #24908). RAM is dual-channel DDR5-5600 on the two consumer boxes and 8-channel DDR5-4800 on the workstation. Bandwidths are peak spec values. SSD rates are typical for the PCIe generation, and 6 GB/s for the Mac SSD is an assumption. Skewed routing is an assumed Zipf curve (s = 0.8). MiniMax has not published expert-use statistics. Hottest-experts placement is the idea behind KTransformers' --kt-num-gpu-experts, whose documented runs use the MXFP8 weights, not these GGUF files.

The computed results for the presets, at 32K context, f16 cache, uniform routing and whole-layer placement:

MachineBuildPlacementCeiling
RTX 4090 24 GB + 128 GB DDR5IQ2_S, q8_0 cache1 expert layer on GPU, 16.9% of expert reads from SSD6.7 tok/s
RTX 4090 24 GB + 128 GB DDR5Q4_K_Malways-on part needs 23.8 GB, 22.1 GB usabledoes not start
RTX 5090 32 GB + 256 GB DDR5IQ4_XS2 layers on GPU, no SSD reads14.1 tok/s
RTX 5090 32 GB + 256 GB DDR5Q4_K_M1 layer on GPU, 11.5% from SSD6.8 tok/s
RTX PRO 6000 96 GB + 512 GB 8-channelQ4_K_M14 layers on GPU, no SSD reads35.6 tok/s
Mac Studio M3 Ultra 256 GBIQ3_Sall resident52.1 tok/s
Mac Studio M3 Ultra 256 GBIQ4_XS10.9% from SSD7.7 tok/s
Mac Studio M3 Ultra 512 GBQ5_K_Mall resident35.0 tok/s

Read the table with three facts in mind. First, the SSD tier is expensive as soon as it appears. On the 5090 box, 11.5% of expert reads take 64 ms per token, nearly as long as the 76 ms for the 86.7% in RAM. Second, the 24 GB card fails on the always-on part alone, because the Q8_0 non-expert tensors are 13.9 GB before any cache. Third, IQ3_S on a 256 GB Mac is fast but carries a measured 23.76% perplexity cost. For a measured point, the MSA PR author reports 7.7 tok/s decode at 5K and 7.67 at 62K on an expert-offload setup. The PR does not name the hardware.

Smaller machines can still load something. A 128 GB DGX Spark has about 115 GB usable at 90%, enough for bartowski's IQ1_M at 100.74 GB with a q8_0 cache at 32K. Nobody has published the quality cost of a 1-bit M3, so test it on your own tasks first. At the other end, the BF16 weights need a node of eight 141 GB H200s, and that is the configuration both vLLM and SGLang validate.

llama.cpp

Three merges matter. PR #24908 added the model with MSA on July 26 and first shipped in b10141. Vision support (#25113) shipped in b10142. A dedicated tool-call parser (#26210) shipped in b10164. Quantized K and V caches (#26180) first worked in b10213. Use b10213 or newer. I checked the flags below against master at b11342.

# 1. Download one build (here AesSedai Q4_K_M, 7 shards)
hf download AesSedai/MiniMax-M3-GGUF \
  --include "Q4_K_M/*" --local-dir models/MiniMax-M3

# 2. Serve it on a 96 GB GPU with 512 GB of RAM
llama-server \
  -m models/MiniMax-M3/Q4_K_M/MiniMax-M3-Q4_K_M-00001-of-00007.gguf \
  --jinja -fa on -ngl 999 -c 32768 \
  --n-cpu-moe 46 \
  -np 1 \
  --temp 1.0 --top-p 0.95 --top-k 40 \
  --host 127.0.0.1 --port 8080

--n-cpu-moe N keeps the expert weights of the first N layers in system RAM. The figure prints the value for your placement: 60 minus the number of expert layers that fit on the GPU. Start there, then lower N until VRAM is nearly full. For a 24 GB card, add -ctk q8_0 -ctv q8_0 and use --cpu-moe, which keeps all experts on the CPU.

Three flags are not optional here. -fa on is required, because MSA calls flash attention directly. With flash attention off, the model falls back to dense attention and warns once. --jinja enables the chat template and the M3 tool-call parser. Keep -np 1, or set several slots without --kv-unified. The PR explains that a unified cache with several sequences "would silently break block anchoring." For images, add --mmproj with one of the projector files.

Ollama has no local build. Its library page lists only minimax-m3:cloud, with a 512K context, which runs on Ollama's servers.

vLLM, SGLang, KTransformers

For a GPU server, the vLLM recipe uses a dedicated image, vllm/vllm-openai:minimax-m3, because support "has not yet shipped in a stable vLLM release." The block size flag is mandatory on every platform:

vllm serve MiniMaxAI/MiniMax-M3 \
  --tensor-parallel-size 8 \
  --block-size 128 \
  --max-model-len 131072 \
  --kv-cache-dtype fp8 \
  --tool-call-parser minimax_m3 \
  --reasoning-parser minimax_m3 \
  --enable-auto-tool-choice

The recipe calls the FP8 cache "lossless in our testing" and says it gives about 1.5x the KV pool. Add --language-model-only for text work to skip the vision encoder. SGLang publishes per-GPU images such as lmsysorg/sglang:dev-cu12-minimax-m3 for H200. On Hopper it serves the BF16 weights, because its MXFP8 kernels are Blackwell-only:

python3 -m sglang.launch_server \
  --model-path MiniMaxAI/MiniMax-M3 \
  --trust-remote-code --tp 8 \
  --mem-fraction-static 0.75 \
  --reasoning-parser auto --tool-call-parser auto \
  --host 0.0.0.0 --port 30000

KTransformers is the documented route for one big GPU plus a lot of RAM. It runs the MXFP8 weights through an SGLang fork. A set number of experts per layer stay on the GPU, and the CPU computes the rest. Its tutorial says "a single 96 GB Hopper card is sufficient if most routed experts stay on CPU." It suggests 2 to 8 GPU experts per layer at TP1. The MXFP8 files are 443.75 GB, so plan for about 512 GB of RAM:

python -m sglang.launch_server \
  --model-path /models/MiniMax-M3-MXFP8 \
  --kt-weight-path /models/MiniMax-M3-MXFP8 \
  --kt-method MXFP8 --kt-cpuinfer 64 \
  --kt-threadpool-count 2 --kt-num-gpu-experts 4 \
  --kt-gpu-prefill-token-threshold 500 \
  --tp-size 1 --quantization mxfp8 \
  --moe-runner-backend triton --trust-remote-code \
  --mem-fraction-static 0.85 --chunked-prefill-size 4096 \
  --cuda-graph-max-bs 1 \
  --tool-call-parser minimax-m3 --reasoning-parser minimax-m3 \
  --served-model-name MiniMax-M3 --port 8000

The tutorial lists Hopper (SM90) as the supported GPU family for this hybrid path. Its own advice for out-of-memory errors is to lower --kt-num-gpu-experts, --mem-fraction-static, or --chunked-prefill-size. Note the different parser spelling: minimax-m3 here, minimax_m3 in vLLM.

On a Mac

The MLX path runs through mlx-vlm, not mlx-lm, because the model type is minimax_m3_vl. The mlx-vlm model README says its implementation covers text, image and video, MSA, thinking tags, tool calls and the EAGLE-3 drafter. The mlx-community 4-bit build was converted with mlx-vlm 0.6.3. Update to a current mlx-vlm before you load it.

pip install -U mlx-vlm

mlx_vlm.generate --model mlx-community/MiniMax-M3-4bit \
  --thinking-mode enabled --max-tokens 1024 \
  --prompt "Explain the failing test in this log."

The 4-bit build is 241.48 GB, so it needs a 512 GB machine. On a 256 GB Mac, the 3-bit build at 186.39 GB fits once you raise the GPU wired limit with sudo sysctl iogpu.wired_limit_mb. The llama.cpp GGUF path also runs on Metal in principle. The MSA PR lists Metal as "untested," and I found no Metal-specific report either way.

Thinking, sampling, tools

Sampling. The card recommends temperature=1.0 and top_p=0.95. The vLLM recipe and the unsloth guide add top_k=40. The vLLM recipe also gives a default system prompt: "You are a helpful assistant. Your name is MiniMax-M3 and is built by MiniMax."

Thinking. The chat template takes a thinking_mode of enabled, disabled or adaptive. Adaptive is the default, and the model then decides per turn. Send it per request as "chat_template_kwargs": {"thinking_mode": "disabled"}. Reasoning arrives between <mm:think> tags, and the parsers move it into reasoning_content. Current vLLM names that field reasoning and keeps the old name as an alias.

Tools. M3 writes tool calls as namespaced XML and writes nested arguments as nested XML. That is why llama.cpp needed its own parser in b10164. The minimax_m3 parsers in vLLM and SGLang convert these calls to the OpenAI tool_calls array. Every engine above exposes an OpenAI-compatible endpoint, so point your harness at http://127.0.0.1:8080/v1 or the engine's port.

Benchmarks, vendor-reported

These are MiniMax's numbers from the card's benchmark chart, as transcribed in the repository's .eval_results file. MiniMax ran the coding rows with Claude Code as the agent scaffold. I found no independent evaluation of the open weights.

BenchmarkMiniMax-M3Notes from the source
SWE-bench Verified80.5average of 4 runs
SWE-bench Pro59.0Claude Code scaffold
Terminal-Bench 2.166.0quoted in the vLLM recipe
MMMU-Pro78.1variant not stated
VideoMME (with subtitles)85.4
Claw-Eval74.5overall score
Apex-Agents27.7

The quantized builds do not inherit these scores. AesSedai's table above is the only published measure of quant loss so far, and it is perplexity on text, not task success.

Failure modes and fixes

FAQ

Can I run MiniMax-M3 on one consumer GPU?

Only with a lot of system RAM. The always-on weights, cache and buffer need about 19 to 24 GB of VRAM at 32K context. The routed experts then live in RAM. A 24 GB card with 128 GB of RAM runs IQ2_S with a q8_0 cache at a computed ceiling of about 6.7 tokens per second.

Does sparse attention make the KV cache smaller?

No. Every token still stores keys and values in all 60 layers, plus an index key in 57 layers. That is 152,064 bytes per token in llama.cpp. MSA cuts what each decode step reads, because each GQA group attends to only 16 blocks of 128 tokens.

Which GGUF files work with mainline llama.cpp?

Files converted with the MSA support from PR #24908, first shipped in b10141. The bartowski and AesSedai repositories were converted after that merge. The unsloth files came from an earlier preliminary branch, and the merged PR says such files fail to load.

Can I use MiniMax-M3 commercially?

Under conditions. The base grant is non-commercial. Commercial use needs a visible "Built with MiniMax M3" notice and a one-time notice email to MiniMax. Above 20 million US dollars of yearly revenue, you need prior written authorization.

Is there an Ollama model for MiniMax-M3?

Only a cloud tag, minimax-m3:cloud, with a 512K context and no local download. For your own hardware, use llama.cpp, vLLM, SGLang, KTransformers or mlx-vlm.

rg
Rohit Ghumare

CNCF Ambassador and Google Developer Expert. I build agent infrastructure and write about the fundamentals underneath the AI stack. Config values, tensor shapes and file sizes come from the MiniMaxAI, AesSedai, bartowski, unsloth, mlx-community and pipenetwork repositories on Hugging Face. I read them on October 2, 2026. Engine status comes from the llama.cpp pull requests, the vLLM recipe, the SGLang cookbook and the KTransformers tutorial. I did not run this model locally. Every speed figure is a computed ceiling or a cited number.

Qwen3.8-Flash-Next · Hardware guide · More guides · X