How to Run LFM2.5 Locally

On July 28, 2026, Liquid AI put LFM2.5-2.6B on Hugging Face. It is the flagship of a family made for machines that run on a battery: phones, laptops without a GPU, a Raspberry Pi, and Macs. The family also has an 8B mixture-of-experts model with 1.5B active parameters, a 3B vision model, and two small models at 350M and 230M. Most layers in these models are short convolutions, so the memory that grows with a conversation stays small. This guide covers the license as the LICENSE file states it and the cache arithmetic from config.json. It also covers every file size and the commands for each runtime that Liquid AI documents. It ends with numbers I measured on an M1 Max. I checked every source on October 2, 2026.

2.6B 30 layers · 8 attention 8B-A1B 8.3B · 1.5B active VL-3B 2.6B + SigLIP2 350M · 230M 32K ctx license LFM Open v1.0 llama.cpp b10180+ for DSpark

Five models, one layout

LFM2.5 is a set of separate releases spread over five months, and they share one design. The dates below are the creation times of the Hugging Face repositories. Liquid's blog post for the 2.6B carries the date August 4, a week after the repository went up.

ModelRepo createdParametersContextLiquid's stated use
LFM2.5-350MMar 31, 2026350M32,768extraction, structured output, tool use
LFM2.5-8B-A1BMay 28, 20268.3B total, 1.5B active128,000on-device assistant, tool chains
LFM2.5-230MJun 24, 2026230M32,768extraction, light agent pipelines
LFM2.5-2.6BJul 28, 20262.69B131,072agents, tool use, RAG, long context
LFM2.5-VL-3BAug 11, 2026about 3.1B32,768OCR, grounding, single-turn vision

Each model card also says what the model is bad at. The 2.6B card advises against "agentic coding and knowledge-heavy tasks." The 8B-A1B card says it is "not the best fit for heavy programming or knowledge-intensive question answering without retrieval." Take that at face value. These models follow instructions and call tools well for their size, and they do not know many facts. Give them retrieval and narrow jobs.

The 2.6B and the 8B-A1B are reasoning models. Both write a chain of thought before the answer, and the 2.6B cannot skip it. The 350M and 230M answer directly. The VL-3B also answers directly, which Liquid says keeps its time to first token low.

The license, read from the file

All five repositories carry the same LICENSE file. I compared them, and they differ only in trailing whitespace. The file is the LFM Open License v1.0. Its text is Apache 2.0 with a revenue limit added to it. These are the terms that matter, in the file's own words where the words decide the outcome.

Two cautions. Liquid's plain-language license page says "under $10M" and "exceeds," while the file says "$10,000,000 or more." The two texts read differently at exactly $10M, so ask Liquid if you are near the line. Also, the 230M blog post says "deploy without restrictions." The LICENSE file is what binds you, and it has a restriction.

Convolutions and a few attention layers

Every model here is an LFM2 design. Most layers are what the cards call "double-gated short convolution blocks." The block projects the input into three streams. One stream gates the input, a depthwise convolution mixes each channel over the last three positions, and the third stream gates the result. The kernel size is conv_L_cache: 3 in every config. A few grouped-query attention layers sit between the convolutions and give the model a view of the whole context.

From config.json230M350M2.6B and VL-3B8B-A1B
ClassLfm2ForCausalLMLfm2ForCausalLMLfm2ForCausalLMLfm2MoeForCausalLM
Layers, conv + attention14 = 8 + 616 = 10 + 630 = 22 + 824 = 18 + 6
Hidden size1,0241,0242,0482,048
Query / KV heads16 / 816 / 832 / 832 / 8
Head dimension64646464
Feed-forward2,5606,65610,75232 experts of 1,792, 4 active; first 2 layers dense
Vocabulary65,53665,536128,000128,000
RoPE theta1,000,0001,000,00010,000,000 (VL: 1,000,000)5,000,000
Training tokens (card)19T28T34T38T

The attention layers sit at fixed positions in layer_types. In the 2.6B they are layers 2, 5, 9, 13, 17, 21, 24 and 27. The VL-3B uses the same 2.6B backbone with a SigLIP2 NaFlex vision encoder of about 400M parameters. The 8B-A1B routes each token to 4 of 32 experts. Its first two layers are dense. One config detail can confuse a reader: the 350M and 230M configs list max_position_embeddings: 128000, but their cards state 32,768. Plan around the card.

The cache arithmetic

A long conversation costs memory in two places. Attention layers keep keys and values for every past token. Convolution layers keep only the last two positions they need for the next step. In llama.cpp that state is n_embd × (L_cache − 1) values in float32 per layer. For the 2.6B with a 16-bit cache:

The other models use the same per-layer figure of 2,048 bytes, because every one has 8 KV heads of 64 dimensions. The 8B-A1B, 350M and 230M each have 6 attention layers, so each costs 12,288 bytes per token. The 350M and the 8B-A1B therefore grow their caches at the same rate.

KV cache, 16-bit4,096 tokens32,768 tokensCard maximumQ4_0 weights
230M50 MB403 MB403 MB at 32,768149 MB
350M50 MB403 MB403 MB at 32,768219 MB
2.6B and VL-3B67 MB537 MB2.15 GB at 131,072 (VL: 537 MB)1.59 GB
8B-A1B50 MB403 MB1.57 GB at 128,0004.84 GB

Look at the small models. At a full 32K context, the 230M's cache is 2.7 times the size of its own Q4_0 weights. On a phone, the context length you allow decides the memory of the small models, and the quant hardly matters. For the 2.6B, the cache passes the weights only past about 100K tokens. A q8_0 cache in llama.cpp stores 34 bytes per 32 values, which is 0.53 of the 16-bit size. Drive the stepper below to watch where the state lives.

Fig. 1 · the shift-register stepper

A real prompt, tokenized by the 128K LFM2.5 tokenizer. Each row is a run of layers from config.json. Convolution rows keep a window of three positions that shifts. Attention rows append every token. A dot marks a leading space.

jump to position 0
conv window, fixed sizeattention cache, grows per token
conv state, fixed0 B
attention KV0 B
all-attention would hold0 B

Bytes use a 16-bit KV cache (2,048 bytes per attention layer per token) and llama.cpp's float32 convolution state (2 positions × hidden size × 4 bytes per conv layer). The 350M and 230M use a 65,536-token vocabulary, so the same text splits differently on those models. Positions past the prompt repeat its tokens.

Every file and its size

Liquid publishes its own GGUF, MLX and ONNX builds, so you do not need a third-party quantizer. Sizes come from the Hugging Face tree API in decimal gigabytes. The model cards round some of them differently.

GGUF (LiquidAI)Q4_0Q4_K_MQ5_K_MQ6_KQ8_0BF16
230M0.1490.1530.1720.1910.2470.462
350M0.2190.2290.2600.2930.3790.711
2.6B1.5941.6741.9402.2222.8755.403
VL-3B (text part)1.5941.6741.9402.2222.8755.403
8B-A1B4.8455.1566.0306.9609.01016.947

The edge budget

On a phone, the operating system does not give one app all of its RAM. It also kills apps that hold too much. Liquid's mobile guide makes one point about this. llama.cpp memory-maps the weights, so iOS jetsam and Android's low-memory killer see them as file-backed pages. The guide says the OS treats those pages "far more leniently" than other app memory. The KV cache and the compute buffers do not get that treatment. A phone budget therefore has two parts: whether the bytes fit, and how much battery each reply costs.

Battery cost is simple arithmetic. Power draw in watts divided by decode speed in tokens per second gives joules per token. Reasoning models write long replies, so a 1,000-token reply is a fair unit. Nobody publishes the power draw of these models on a phone, so the sheet below makes that an input. Measure your own device and set it.

Fig. 2 · the pocket budget sheet

Pick a device, a model, a build and a context. The memory map fills with the real file sizes and the computed cache. The battery line prices one reply from the decode speed and a power draw you set.

context = 8,192 tokens
OS and other apps keep 4.0 GBassumed
0 GB

Battery

decode speed 30 tok/svendor
power while decoding 6.0 Wassumed
reply length 1,000 tokens
battery capacity 19.3 Wh

Weights are exact GGUF sizes. KV and conv state use the arithmetic above. Compute buffers for the 2.6B are 0.21 GB, derived from the measured server memory below. The VL-3B and 8B-A1B use 0.3 GB and the small models 0.1 GB, an estimate. Vendor speeds come from Liquid's model cards and blogs, which do not name the quant. The measured speed is mine, from the M1 Max run below, for Q4_K_M. Battery capacities are spec-sheet values where named and assumptions elsewhere. Set the battery to 0 Wh for wall power.

Some quick picks fall out of the sheet. Assume the system leaves your app about 3 GB of an 8 GB phone. Then the 2.6B at Q4_0 fits with an 8K context, and the 8B-A1B does not. A 12 GB phone holds the 2.6B at Q8_0 with 32K of context. A Raspberry Pi 5 with 8 GB runs the 8B-A1B at Q4_K_M. A 16 GB laptop runs every model here, the 8B-A1B at Q8_0 up to about 32K of context. The 230M and 350M fit anywhere, as long as you cap their context.

Measured on an M1 Max

I ran the 2.6B on an Apple M1 Max with 64 GB of unified memory on October 2, 2026. The runtime was the Homebrew llama.cpp build 10330 on Metal, with every layer on the GPU and a 16-bit cache. Speeds are the median of three llama-bench runs.

BuildPrefill 512Prefill 4,096Decode 128First token, 1,000-token promptServer RSS at 8K
LFM2.5-2.6B-Q4_K_M.gguf, 1.67 GB1,391.6 tok/s1,024.4 tok/s90.5 tok/s871 ms2.02 GB

I did not measure the 8B-A1B. Its download failed repeatedly during this session, so the 8B-A1B numbers in this guide are Liquid's. For scale, Liquid reports 220 tokens per second for the 2.6B on an M5 Max. Its DSpark card gives a baseline of about 61 on an M4 Max.

llama.cpp on a laptop or a Pi

llama.cpp is the runtime Liquid points to for CPUs, edge boxes and phones. The LFM2 and LFM2-MoE architectures have been in mainline for a long time, so any recent build loads every model here. DSpark speculative decoding arrived later, in PR #25173, merged on July 28, 2026. Build b10180 of July 29 includes it. Install with brew install llama.cpp or take a release binary.

# The 2.6B as an OpenAI-compatible server on :8080
llama-server -hf LiquidAI/LFM2.5-2.6B-GGUF:Q4_K_M \
  --jinja -c 32768 -ngl 99 \
  --temp 0.1 --top-k 50 --repeat-penalty 1.1 \
  --host 127.0.0.1 --port 8080

# The 8B-A1B with its own sampling preset
llama-server -hf LiquidAI/LFM2.5-8B-A1B-GGUF:Q4_K_M \
  --jinja -c 32768 -ngl 99 \
  --temp 0.2 --top-k 80 --repeat-penalty 1.05

# The VL-3B; -hf also fetches the mmproj file
llama-server -hf LiquidAI/LFM2.5-VL-3B-GGUF:Q4_K_M \
  -c 8192 -ngl 99 \
  --temp 0.2 --top-k 50 --repeat-penalty 1.0

On a CPU-only machine, -ngl 99 does nothing and you can drop it. To get the QAD build, download the file by name, because :Q4_0 matches two files in the repo:

hf download LiquidAI/LFM2.5-2.6B-GGUF LFM2.5-2.6B-QAD-Q4_0.gguf --local-dir .
llama-server -m LFM2.5-2.6B-QAD-Q4_0.gguf --jinja -c 8192

DSpark pairs a 328M drafter with the target model. The target checks every drafted token, so the output is the same as the target alone. Liquid's card for the 2.6B drafter reports a mean of 4.81 accepted tokens per step and a 2.27x speedup on an M4 Max through Metal. The command from the sidecar card:

llama-server -m LFM2.5-2.6B-F16.gguf \
  -md LFM2.5-2.6B-DSpark-F16.gguf \
  --spec-type draft-dspark --spec-draft-n-max 10 --spec-draft-n-min 0 \
  -fa on -ngl 99

Any 2.6B GGUF works as the target. The sidecar card says the target quant is "the main speed/quality lever," and it calls Q4_K_M the smallest drafter worth using.

Raspberry Pi 5 and other Arm boards

Liquid reports 42 tokens per second for the 230M on a Raspberry Pi 5. It ran flash attention on for the Pi and off for the Snapdragon phone, to get the best prefill on each. Measure your own board before you build an app on it:

llama-bench -m LFM2.5-230M-Q4_0.gguf -p 512 -n 128 -fa 1
llama-bench -m LFM2.5-230M-Q4_0.gguf -p 512 -n 128 -fa 0

On a phone

Liquid used to ship a mobile wrapper called the LEAP SDK. Its docs now mark it as deprecated: "no longer receives new releases." The current advice is to link llama.cpp into the app and load GGUF files from Hugging Face. Existing LEAP artifacts stay on Maven Central, and Liquid has a migration page that maps each LEAP call to its llama.cpp form.

  1. On iOS, add llama.xcframework from a llama.cpp release. It has Metal and the multimodal library built in.
  2. On Android, build llama.cpp through the NDK for arm64-v8a. Start from examples/llama.android in the llama.cpp repo.
  3. Download the GGUF at first launch into app storage. Do not ship it as a compressed asset.
  4. Use Q4_0 or the QAD Q4_0 file. llama.cpp repacks Q4_0 into Arm kernels at load time.
  5. Set n_gpu_layers = 99 on iOS for Metal. Keep the CPU on Android unless you test Vulkan or OpenCL on your device.
  6. Set the thread count to the core count minus two, as the guide suggests.
  7. Set n_ctx as small as your task allows. The cache is the memory the OS does not forgive.

For a fast check before you write app code, the release page has llama-*-bin-android-arm64 archives with llama-bench inside. Run them over adb or in Termux. Liquid's phone numbers are 213 tokens per second for the 230M on a Galaxy S25 Ultra and 188 for the 350M on the same chip. It reports about 30 for the 2.6B and the 8B-A1B "on a phone," without naming the phone. The VL-3B gets 20 on a Galaxy S26 Ultra.

On Apple Silicon

Two paths work. The llama.cpp commands above run on Metal with -ngl 99, and that is the path I measured. The MLX path uses Liquid's own MLX builds and mlx-lm, which has lfm2, lfm2_moe and lfm2-vl model files. The current release on PyPI is 0.32.0.

pip install -U mlx-lm

mlx_lm.generate --model LiquidAI/LFM2.5-2.6B-MLX-4bit \
  --temp 0.1 --top-k 50 --max-tokens 4096 \
  --prompt "List the steps to rotate an API key."

mlx_lm.server --model LiquidAI/LFM2.5-8B-A1B-MLX-4bit --port 8080

The MLX cards apply the repetition penalty through make_logits_processors(repetition_penalty=1.05) in Python. Use that form if you need the full preset. The one-repo-many-folders LiquidAI/LFM2.5-2.6B-MLX does not load directly, because mlx_lm.load does not resolve subfolders. Use the standalone repos such as LFM2.5-2.6B-MLX-4bit. For the VL-3B, use mlx-vlm with LiquidAI/LFM2.5-VL-3B-MLX-8bit. Liquid's vendor numbers for Apple Silicon come from an M5 Max. They are 220 tokens per second for the 2.6B, 253 for the 8B-A1B, and 228 for the VL-3B.

Ollama, LM Studio, vLLM, SGLang

Ollama runs the GGUF repos straight from Hugging Face. Liquid's Ollama page warns that v0.17.0 fails on the MoE architecture with missing tensor 'output_norm.weight', and that you need v0.17.1-rc0 or later. The current release is v0.35.0.

ollama run hf.co/LiquidAI/LFM2.5-2.6B-GGUF:Q4_K_M
ollama run hf.co/LiquidAI/LFM2.5-8B-A1B-GGUF:Q4_K_M

LM Studio loads the same GGUF files through its search panel and serves them on port 1234.

vLLM 0.23.0 and later has all three classes built in, so you need no --trust-remote-code. The vLLM recipe for LFM2.5 uses the qwen3 reasoning parser and the lfm2 tool parser. It notes that the 8B-A1B needs a 24 GB GPU in BF16, because every expert stays resident:

vllm serve LiquidAI/LFM2.5-8B-A1B \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser lfm2

The recipe lists the 8B-A1B command. The 2.6B uses the same parsers, because it emits the same think tags and tool tokens. SGLang serves the family too, and the DSpark drafter runs there with a build that includes PR #31041, merged August 31. The 2.6B drafter card reports a 2.67x mean speedup on an H100:

python -m sglang.launch_server \
  --model-path LiquidAI/LFM2.5-2.6B \
  --speculative-algorithm DSPARK \
  --speculative-draft-model-path LiquidAI/LFM2.5-2.6B-DSpark \
  --speculative-draft-attention-backend flashinfer \
  --disable-radix-cache --mem-fraction-static 0.75 --port 30000

Sampling, thinking, tools

Sampling. Each model has its own preset in its card and generation_config.json. Send it on every request. The vLLM recipe calls the presets "per-request client defaults, not server flags."

Modeltemperaturetop_krepetition_penalty
2.6B0.1501.1
8B-A1B0.2801.05
VL-3B (text)0.2501.0
350M and 230M0.1501.05

Over the OpenAI API, top_k and repetition_penalty are extras. Put them in extra_body with the Python client.

Chat template. The format is ChatML-like: <|startoftext|>, then <|im_start|>role blocks closed by <|im_end|>. In llama.cpp, pass --jinja so the server uses the template from the GGUF.

Thinking. The 2.6B template always appends <think> to the assistant turn, so there is no switch to turn reasoning off. The template has one option, preserve_thinking, which keeps past reasoning in the history. It is off by default, and older thinking is stripped from earlier turns. The 8B-A1B template does not add the tag itself, and the model writes it. Do not set a low max_tokens. The vLLM recipe warns that a low cap cuts the reasoning off before the answer.

Tools. The model writes Pythonic calls between <|tool_call_start|> and <|tool_call_end|>. An example is [get_candidate_status(candidate_id="12345")]. With --jinja, llama-server parses LFM2 and LFM2.5 calls into OpenAI tool_calls, per Liquid's docs. vLLM and SGLang use the lfm2 parser. If your harness only reads JSON calls, the cards say you can ask for JSON in the system prompt. The 2.6B card gives setup lines for three agent harnesses: Hermes, OpenClaw and Pi.

Benchmarks, vendor-reported

These numbers come from Liquid's model cards, under Liquid's own harness. I picked rows that test the stated use cases. No independent evaluation was in the sources I read.

BenchmarkLFM2.5-2.6BLFM2.5-8B-A1BQwen3.5-4BGemma-4-E4B
IFBench59.1756.4748.40 / 50.3839.24 / 39.48
Multi-IF80.0779.9355.67 / 67.4377.35 / 77.58
BFCLv456.8849.7350.56 / 54.0146.39 / 33.92
AIME2551.8742.5349.33 / 54.2834.27 / 34.33
AA-Omniscience accuracy8.138.6717.63 / 17.208.33 / 8.10

Where a cell has two values, the first comes from the 2.6B card and the second from the 8B-A1B card. The two cards measured the same rival models and got different numbers. Read that as the noise in vendor tables. The pattern is stable: strong instruction following and tool use, weak recall of facts. The VL-3B card reports 80.7 on ScreenSpot-v2 and 87.9 on RefCOCO, up from 57.1 on RefCOCO for LFM2-VL-3B.

Failure modes and fixes

FAQ

Can I use LFM2.5 commercially?

Yes, if your Legal Entity has annual revenue below $10 million. That entity includes the companies that control you, that you control, or that share control with you. Above that line, the license does not cover commercial use, and you need a separate deal with Liquid AI. Research and non-commercial use have no revenue limit.

Which model should I run on a phone?

Start with the 2.6B at Q4_0 or QAD Q4_0, which is 1.59 GB, with a context near 8K. Use the 350M or 230M for extraction and routing, or on phones with 8 GB of RAM or less.

Why is the cache so small?

Only 8 of the 2.6B's 30 layers use attention. The other 22 keep a fixed state of two past positions. The cache grows by 16,384 bytes per token, so 32,768 tokens cost 537 MB, against 2.01 GB if all 30 layers were attention.

Can I turn off thinking on the 2.6B?

No. The card calls it a pure reasoning model, and its template always opens the reply with a think tag. Budget output tokens for the reasoning.

Is the LEAP SDK still the way to ship on mobile?

No. Liquid marks it as deprecated with no new releases. Link llama.cpp into the app with the XCFramework on iOS or the NDK build on Android, and load GGUF files from Hugging Face.

rg
Rohit Ghumare

CNCF Ambassador and Google Developer Expert. I build agent infrastructure and write about the fundamentals underneath the AI stack. Config values, file sizes, license text, and runtime status come from Liquid AI's Hugging Face repositories, docs, and blog. Other sources are the vLLM recipe and the llama.cpp and SGLang trackers, read on October 2, 2026. The M1 Max numbers are my own runs.

MiMo-V2.6 · Hardware guide · More guides · X