How to Run LFM2.5 Locally
On July 28, 2026, Liquid AI put LFM2.5-2.6B on Hugging Face. It is the flagship of a family made for machines that run on a battery: phones, laptops without a GPU, a Raspberry Pi, and Macs. The family also has an 8B mixture-of-experts model with 1.5B active parameters, a 3B vision model, and two small models at 350M and 230M. Most layers in these models are short convolutions, so the memory that grows with a conversation stays small. This guide covers the license as the LICENSE file states it and the cache arithmetic from config.json. It also covers every file size and the commands for each runtime that Liquid AI documents. It ends with numbers I measured on an M1 Max. I checked every source on October 2, 2026.
Five models, one layout
LFM2.5 is a set of separate releases spread over five months, and they share one design. The dates below are the creation times of the Hugging Face repositories. Liquid's blog post for the 2.6B carries the date August 4, a week after the repository went up.
| Model | Repo created | Parameters | Context | Liquid's stated use |
|---|---|---|---|---|
LFM2.5-350M | Mar 31, 2026 | 350M | 32,768 | extraction, structured output, tool use |
LFM2.5-8B-A1B | May 28, 2026 | 8.3B total, 1.5B active | 128,000 | on-device assistant, tool chains |
LFM2.5-230M | Jun 24, 2026 | 230M | 32,768 | extraction, light agent pipelines |
LFM2.5-2.6B | Jul 28, 2026 | 2.69B | 131,072 | agents, tool use, RAG, long context |
LFM2.5-VL-3B | Aug 11, 2026 | about 3.1B | 32,768 | OCR, grounding, single-turn vision |
Each model card also says what the model is bad at. The 2.6B card advises against "agentic coding and knowledge-heavy tasks." The 8B-A1B card says it is "not the best fit for heavy programming or knowledge-intensive question answering without retrieval." Take that at face value. These models follow instructions and call tools well for their size, and they do not know many facts. Give them retrieval and narrow jobs.
The 2.6B and the 8B-A1B are reasoning models. Both write a chain of thought before the answer, and the 2.6B cannot skip it. The 350M and 230M answer directly. The VL-3B also answers directly, which Liquid says keeps its time to first token low.
The license, read from the file
All five repositories carry the same LICENSE file. I compared them, and they differ only in trailing whitespace. The file is the LFM Open License v1.0. Its text is Apache 2.0 with a revenue limit added to it. These are the terms that matter, in the file's own words where the words decide the outcome.
- The revenue line. Section 1 defines the Threshold as "annual revenue of 10 million United States dollars ($10,000,000) or more." Section 5(a) makes commercial rights conditional on "not exceeding the Threshold."
- Above the line, no commercial license. Section 5(b) says commercial use by an entity above the Threshold "is not licensed under this Agreement." Liquid's docs direct those companies to sales@liquid.ai.
- Your group counts, not only your company. "Legal Entity" includes every entity that controls you, that you control, or that shares control with you. Control includes 50% ownership. A small subsidiary of a large company counts the parent's revenue.
- Research has no revenue limit. Section 5 limits commercial use only. Section 5(c) also states that the Threshold does not apply to a qualified non-profit for non-commercial or research purposes.
- Commercial use is broad. The file defines it as "any use of the Work for direct or indirect commercial advantage or monetary compensation." An internal tool at a company counts.
- Distribution. Give recipients a copy of the license, mark the files you changed, and keep the notices. This is the Apache 2.0 text.
- Fine-tunes. You can license your own changes on your own terms. The base work stays under this license, so the revenue limit follows your fine-tune.
- Termination. Section 11 ends the license "automatically and immediately" if you break any term. Then you must delete all copies. Apache 2.0 has no such clause.
Two cautions. Liquid's plain-language license page says "under $10M" and "exceeds," while the file says "$10,000,000 or more." The two texts read differently at exactly $10M, so ask Liquid if you are near the line. Also, the 230M blog post says "deploy without restrictions." The LICENSE file is what binds you, and it has a restriction.
Convolutions and a few attention layers
Every model here is an LFM2 design. Most layers are what the cards call "double-gated short convolution blocks." The block projects the input into three streams. One stream gates the input, a depthwise convolution mixes each channel over the last three positions, and the third stream gates the result. The kernel size is conv_L_cache: 3 in every config. A few grouped-query attention layers sit between the convolutions and give the model a view of the whole context.
| From config.json | 230M | 350M | 2.6B and VL-3B | 8B-A1B |
|---|---|---|---|---|
| Class | Lfm2ForCausalLM | Lfm2ForCausalLM | Lfm2ForCausalLM | Lfm2MoeForCausalLM |
| Layers, conv + attention | 14 = 8 + 6 | 16 = 10 + 6 | 30 = 22 + 8 | 24 = 18 + 6 |
| Hidden size | 1,024 | 1,024 | 2,048 | 2,048 |
| Query / KV heads | 16 / 8 | 16 / 8 | 32 / 8 | 32 / 8 |
| Head dimension | 64 | 64 | 64 | 64 |
| Feed-forward | 2,560 | 6,656 | 10,752 | 32 experts of 1,792, 4 active; first 2 layers dense |
| Vocabulary | 65,536 | 65,536 | 128,000 | 128,000 |
| RoPE theta | 1,000,000 | 1,000,000 | 10,000,000 (VL: 1,000,000) | 5,000,000 |
| Training tokens (card) | 19T | 28T | 34T | 38T |
The attention layers sit at fixed positions in layer_types. In the 2.6B they are layers 2, 5, 9, 13, 17, 21, 24 and 27. The VL-3B uses the same 2.6B backbone with a SigLIP2 NaFlex vision encoder of about 400M parameters. The 8B-A1B routes each token to 4 of 32 experts. Its first two layers are dense. One config detail can confuse a reader: the 350M and 230M configs list max_position_embeddings: 128000, but their cards state 32,768. Plan around the card.
The cache arithmetic
A long conversation costs memory in two places. Attention layers keep keys and values for every past token. Convolution layers keep only the last two positions they need for the next step. In llama.cpp that state is n_embd × (L_cache − 1) values in float32 per layer. For the 2.6B with a 16-bit cache:
- Per attention layer, per token: 8 KV heads × 64 dimensions × 2 (key and value) × 2 bytes = 2,048 bytes.
- Whole model, per token: 8 attention layers × 2,048 = 16,384 bytes, or 16 KB per token.
- Convolution state, fixed: 22 layers × 2 positions × 2,048 channels × 4 bytes = 360,448 bytes. That is 0.36 MB, the same as 22 tokens of attention cache, and it never grows.
- The counterfactual: if all 30 layers were attention, the cache would be 61,440 bytes per token, 3.75 times larger.
The other models use the same per-layer figure of 2,048 bytes, because every one has 8 KV heads of 64 dimensions. The 8B-A1B, 350M and 230M each have 6 attention layers, so each costs 12,288 bytes per token. The 350M and the 8B-A1B therefore grow their caches at the same rate.
| KV cache, 16-bit | 4,096 tokens | 32,768 tokens | Card maximum | Q4_0 weights |
|---|---|---|---|---|
| 230M | 50 MB | 403 MB | 403 MB at 32,768 | 149 MB |
| 350M | 50 MB | 403 MB | 403 MB at 32,768 | 219 MB |
| 2.6B and VL-3B | 67 MB | 537 MB | 2.15 GB at 131,072 (VL: 537 MB) | 1.59 GB |
| 8B-A1B | 50 MB | 403 MB | 1.57 GB at 128,000 | 4.84 GB |
Look at the small models. At a full 32K context, the 230M's cache is 2.7 times the size of its own Q4_0 weights. On a phone, the context length you allow decides the memory of the small models, and the quant hardly matters. For the 2.6B, the cache passes the weights only past about 100K tokens. A q8_0 cache in llama.cpp stores 34 bytes per 32 values, which is 0.53 of the 16-bit size. Drive the stepper below to watch where the state lives.
A real prompt, tokenized by the 128K LFM2.5 tokenizer. Each row is a run of layers from config.json. Convolution rows keep a window of three positions that shifts. Attention rows append every token. A dot marks a leading space.
Bytes use a 16-bit KV cache (2,048 bytes per attention layer per token) and llama.cpp's float32 convolution state (2 positions × hidden size × 4 bytes per conv layer). The 350M and 230M use a 65,536-token vocabulary, so the same text splits differently on those models. Positions past the prompt repeat its tokens.
Every file and its size
Liquid publishes its own GGUF, MLX and ONNX builds, so you do not need a third-party quantizer. Sizes come from the Hugging Face tree API in decimal gigabytes. The model cards round some of them differently.
| GGUF (LiquidAI) | Q4_0 | Q4_K_M | Q5_K_M | Q6_K | Q8_0 | BF16 |
|---|---|---|---|---|---|---|
| 230M | 0.149 | 0.153 | 0.172 | 0.191 | 0.247 | 0.462 |
| 350M | 0.219 | 0.229 | 0.260 | 0.293 | 0.379 | 0.711 |
| 2.6B | 1.594 | 1.674 | 1.940 | 2.222 | 2.875 | 5.403 |
| VL-3B (text part) | 1.594 | 1.674 | 1.940 | 2.222 | 2.875 | 5.403 |
| 8B-A1B | 4.845 | 5.156 | 6.030 | 6.960 | 9.010 | 16.947 |
- QAD Q4_0. The 230M, 350M and 2.6B repos also ship a
QAD-Q4_0file of the same size as plain Q4_0. QAD means quantization-aware distillation. Liquid trained these weights for Q4_0 deployment, and the card says its published QAD results "apply to the Q4_0 GGUF" only. - VL projector. The vision encoder is a separate
mmprojfile: 0.583 GB at Q8_0 and 0.856 GB at BF16. - DSpark drafters. Speculative-decoding drafters exist for the 2.6B, 8B-A1B and VL-3B. The GGUF sidecars are 0.200 GB at Q4_K_M, 0.356 GB at Q8_0 and 0.664 GB at F16. Each sidecar borrows the target's embeddings, so it only works with its own target.
- MLX. The 2.6B MLX repo holds eight precisions. The 4-bit build is 1.583 GB, 5-bit 1.888 GB, 6-bit 2.192 GB, and 8-bit 2.866 GB. MXFP4 is 1.564 GB, NVFP4 1.640 GB, MXFP8 2.782 GB, and BF16 5.394 GB. The 8B-A1B MLX builds are 4.834 GB at 4-bit and 9.002 GB at 8-bit. The VL-3B MLX 8-bit build is 3.719 GB. The 350M and 230M 4-bit builds are 0.222 GB and 0.146 GB.
- Original weights. BF16 safetensors: 5.39 GB for the 2.6B, 16.94 GB for the 8B-A1B, 6.25 GB for the VL-3B.
- ONNX. Every model has an ONNX repo for ONNX Runtime and WebGPU. The 2.6B
q4f16build is 1.53 GB.
The edge budget
On a phone, the operating system does not give one app all of its RAM. It also kills apps that hold too much. Liquid's mobile guide makes one point about this. llama.cpp memory-maps the weights, so iOS jetsam and Android's low-memory killer see them as file-backed pages. The guide says the OS treats those pages "far more leniently" than other app memory. The KV cache and the compute buffers do not get that treatment. A phone budget therefore has two parts: whether the bytes fit, and how much battery each reply costs.
Battery cost is simple arithmetic. Power draw in watts divided by decode speed in tokens per second gives joules per token. Reasoning models write long replies, so a 1,000-token reply is a fair unit. Nobody publishes the power draw of these models on a phone, so the sheet below makes that an input. Measure your own device and set it.
Pick a device, a model, a build and a context. The memory map fills with the real file sizes and the computed cache. The battery line prices one reply from the decode speed and a power draw you set.
Battery
Weights are exact GGUF sizes. KV and conv state use the arithmetic above. Compute buffers for the 2.6B are 0.21 GB, derived from the measured server memory below. The VL-3B and 8B-A1B use 0.3 GB and the small models 0.1 GB, an estimate. Vendor speeds come from Liquid's model cards and blogs, which do not name the quant. The measured speed is mine, from the M1 Max run below, for Q4_K_M. Battery capacities are spec-sheet values where named and assumptions elsewhere. Set the battery to 0 Wh for wall power.
Some quick picks fall out of the sheet. Assume the system leaves your app about 3 GB of an 8 GB phone. Then the 2.6B at Q4_0 fits with an 8K context, and the 8B-A1B does not. A 12 GB phone holds the 2.6B at Q8_0 with 32K of context. A Raspberry Pi 5 with 8 GB runs the 8B-A1B at Q4_K_M. A 16 GB laptop runs every model here, the 8B-A1B at Q8_0 up to about 32K of context. The 230M and 350M fit anywhere, as long as you cap their context.
Measured on an M1 Max
I ran the 2.6B on an Apple M1 Max with 64 GB of unified memory on October 2, 2026. The runtime was the Homebrew llama.cpp build 10330 on Metal, with every layer on the GPU and a 16-bit cache. Speeds are the median of three llama-bench runs.
| Build | Prefill 512 | Prefill 4,096 | Decode 128 | First token, 1,000-token prompt | Server RSS at 8K |
|---|---|---|---|---|---|
LFM2.5-2.6B-Q4_K_M.gguf, 1.67 GB | 1,391.6 tok/s | 1,024.4 tok/s | 90.5 tok/s | 871 ms | 2.02 GB |
- Decode. The three samples were 85.7, 90.5 and 94.6 tokens per second. Background downloads ran on the network during the run, so treat the spread as real. At 90.5 tokens per second, the 1.67 GB of weights are read about 151 GB per second.
- Prefill. Speed fell by 26% from a 512-token prompt to a 4,096-token prompt. The attention layers do more work per token as the prompt grows. The convolution layers do not.
- Memory. The server held 2.02 GB of resident memory at an 8K context. The weights and the computed cache add up to 1.81 GB, which leaves about 0.21 GB of buffers. The macOS "footprint" figure was only 0.3 GB, because the mapped weights count as file pages. That is the same effect that helps on a phone.
- Thinking. The reply always started with reasoning. Passing
thinking: falseinchat_template_kwargschanged nothing. llama-server returned the reasoning inreasoning_content. - A sanity prompt. A bat-and-ball question plus a one-line Python task came back correct after 323 tokens and 3.6 seconds. About 1,000 characters of that were reasoning.
I did not measure the 8B-A1B. Its download failed repeatedly during this session, so the 8B-A1B numbers in this guide are Liquid's. For scale, Liquid reports 220 tokens per second for the 2.6B on an M5 Max. Its DSpark card gives a baseline of about 61 on an M4 Max.
llama.cpp on a laptop or a Pi
llama.cpp is the runtime Liquid points to for CPUs, edge boxes and phones. The LFM2 and LFM2-MoE architectures have been in mainline for a long time, so any recent build loads every model here. DSpark speculative decoding arrived later, in PR #25173, merged on July 28, 2026. Build b10180 of July 29 includes it. Install with brew install llama.cpp or take a release binary.
# The 2.6B as an OpenAI-compatible server on :8080
llama-server -hf LiquidAI/LFM2.5-2.6B-GGUF:Q4_K_M \
--jinja -c 32768 -ngl 99 \
--temp 0.1 --top-k 50 --repeat-penalty 1.1 \
--host 127.0.0.1 --port 8080
# The 8B-A1B with its own sampling preset
llama-server -hf LiquidAI/LFM2.5-8B-A1B-GGUF:Q4_K_M \
--jinja -c 32768 -ngl 99 \
--temp 0.2 --top-k 80 --repeat-penalty 1.05
# The VL-3B; -hf also fetches the mmproj file
llama-server -hf LiquidAI/LFM2.5-VL-3B-GGUF:Q4_K_M \
-c 8192 -ngl 99 \
--temp 0.2 --top-k 50 --repeat-penalty 1.0
On a CPU-only machine, -ngl 99 does nothing and you can drop it. To get the QAD build, download the file by name, because :Q4_0 matches two files in the repo:
hf download LiquidAI/LFM2.5-2.6B-GGUF LFM2.5-2.6B-QAD-Q4_0.gguf --local-dir .
llama-server -m LFM2.5-2.6B-QAD-Q4_0.gguf --jinja -c 8192
DSpark pairs a 328M drafter with the target model. The target checks every drafted token, so the output is the same as the target alone. Liquid's card for the 2.6B drafter reports a mean of 4.81 accepted tokens per step and a 2.27x speedup on an M4 Max through Metal. The command from the sidecar card:
llama-server -m LFM2.5-2.6B-F16.gguf \
-md LFM2.5-2.6B-DSpark-F16.gguf \
--spec-type draft-dspark --spec-draft-n-max 10 --spec-draft-n-min 0 \
-fa on -ngl 99
Any 2.6B GGUF works as the target. The sidecar card says the target quant is "the main speed/quality lever," and it calls Q4_K_M the smallest drafter worth using.
Raspberry Pi 5 and other Arm boards
Liquid reports 42 tokens per second for the 230M on a Raspberry Pi 5. It ran flash attention on for the Pi and off for the Snapdragon phone, to get the best prefill on each. Measure your own board before you build an app on it:
llama-bench -m LFM2.5-230M-Q4_0.gguf -p 512 -n 128 -fa 1
llama-bench -m LFM2.5-230M-Q4_0.gguf -p 512 -n 128 -fa 0
On a phone
Liquid used to ship a mobile wrapper called the LEAP SDK. Its docs now mark it as deprecated: "no longer receives new releases." The current advice is to link llama.cpp into the app and load GGUF files from Hugging Face. Existing LEAP artifacts stay on Maven Central, and Liquid has a migration page that maps each LEAP call to its llama.cpp form.
- On iOS, add
llama.xcframeworkfrom a llama.cpp release. It has Metal and the multimodal library built in. - On Android, build llama.cpp through the NDK for
arm64-v8a. Start fromexamples/llama.androidin the llama.cpp repo. - Download the GGUF at first launch into app storage. Do not ship it as a compressed asset.
- Use Q4_0 or the QAD Q4_0 file. llama.cpp repacks Q4_0 into Arm kernels at load time.
- Set
n_gpu_layers = 99on iOS for Metal. Keep the CPU on Android unless you test Vulkan or OpenCL on your device. - Set the thread count to the core count minus two, as the guide suggests.
- Set
n_ctxas small as your task allows. The cache is the memory the OS does not forgive.
For a fast check before you write app code, the release page has llama-*-bin-android-arm64 archives with llama-bench inside. Run them over adb or in Termux. Liquid's phone numbers are 213 tokens per second for the 230M on a Galaxy S25 Ultra and 188 for the 350M on the same chip. It reports about 30 for the 2.6B and the 8B-A1B "on a phone," without naming the phone. The VL-3B gets 20 on a Galaxy S26 Ultra.
On Apple Silicon
Two paths work. The llama.cpp commands above run on Metal with -ngl 99, and that is the path I measured. The MLX path uses Liquid's own MLX builds and mlx-lm, which has lfm2, lfm2_moe and lfm2-vl model files. The current release on PyPI is 0.32.0.
pip install -U mlx-lm
mlx_lm.generate --model LiquidAI/LFM2.5-2.6B-MLX-4bit \
--temp 0.1 --top-k 50 --max-tokens 4096 \
--prompt "List the steps to rotate an API key."
mlx_lm.server --model LiquidAI/LFM2.5-8B-A1B-MLX-4bit --port 8080
The MLX cards apply the repetition penalty through make_logits_processors(repetition_penalty=1.05) in Python. Use that form if you need the full preset. The one-repo-many-folders LiquidAI/LFM2.5-2.6B-MLX does not load directly, because mlx_lm.load does not resolve subfolders. Use the standalone repos such as LFM2.5-2.6B-MLX-4bit. For the VL-3B, use mlx-vlm with LiquidAI/LFM2.5-VL-3B-MLX-8bit. Liquid's vendor numbers for Apple Silicon come from an M5 Max. They are 220 tokens per second for the 2.6B, 253 for the 8B-A1B, and 228 for the VL-3B.
Ollama, LM Studio, vLLM, SGLang
Ollama runs the GGUF repos straight from Hugging Face. Liquid's Ollama page warns that v0.17.0 fails on the MoE architecture with missing tensor 'output_norm.weight', and that you need v0.17.1-rc0 or later. The current release is v0.35.0.
ollama run hf.co/LiquidAI/LFM2.5-2.6B-GGUF:Q4_K_M
ollama run hf.co/LiquidAI/LFM2.5-8B-A1B-GGUF:Q4_K_M
LM Studio loads the same GGUF files through its search panel and serves them on port 1234.
vLLM 0.23.0 and later has all three classes built in, so you need no --trust-remote-code. The vLLM recipe for LFM2.5 uses the qwen3 reasoning parser and the lfm2 tool parser. It notes that the 8B-A1B needs a 24 GB GPU in BF16, because every expert stays resident:
vllm serve LiquidAI/LFM2.5-8B-A1B \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser lfm2
The recipe lists the 8B-A1B command. The 2.6B uses the same parsers, because it emits the same think tags and tool tokens. SGLang serves the family too, and the DSpark drafter runs there with a build that includes PR #31041, merged August 31. The 2.6B drafter card reports a 2.67x mean speedup on an H100:
python -m sglang.launch_server \
--model-path LiquidAI/LFM2.5-2.6B \
--speculative-algorithm DSPARK \
--speculative-draft-model-path LiquidAI/LFM2.5-2.6B-DSpark \
--speculative-draft-attention-backend flashinfer \
--disable-radix-cache --mem-fraction-static 0.75 --port 30000
Sampling, thinking, tools
Sampling. Each model has its own preset in its card and generation_config.json. Send it on every request. The vLLM recipe calls the presets "per-request client defaults, not server flags."
| Model | temperature | top_k | repetition_penalty |
|---|---|---|---|
| 2.6B | 0.1 | 50 | 1.1 |
| 8B-A1B | 0.2 | 80 | 1.05 |
| VL-3B (text) | 0.2 | 50 | 1.0 |
| 350M and 230M | 0.1 | 50 | 1.05 |
Over the OpenAI API, top_k and repetition_penalty are extras. Put them in extra_body with the Python client.
Chat template. The format is ChatML-like: <|startoftext|>, then <|im_start|>role blocks closed by <|im_end|>. In llama.cpp, pass --jinja so the server uses the template from the GGUF.
Thinking. The 2.6B template always appends <think> to the assistant turn, so there is no switch to turn reasoning off. The template has one option, preserve_thinking, which keeps past reasoning in the history. It is off by default, and older thinking is stripped from earlier turns. The 8B-A1B template does not add the tag itself, and the model writes it. Do not set a low max_tokens. The vLLM recipe warns that a low cap cuts the reasoning off before the answer.
Tools. The model writes Pythonic calls between <|tool_call_start|> and <|tool_call_end|>. An example is [get_candidate_status(candidate_id="12345")]. With --jinja, llama-server parses LFM2 and LFM2.5 calls into OpenAI tool_calls, per Liquid's docs. vLLM and SGLang use the lfm2 parser. If your harness only reads JSON calls, the cards say you can ask for JSON in the system prompt. The 2.6B card gives setup lines for three agent harnesses: Hermes, OpenClaw and Pi.
Benchmarks, vendor-reported
These numbers come from Liquid's model cards, under Liquid's own harness. I picked rows that test the stated use cases. No independent evaluation was in the sources I read.
| Benchmark | LFM2.5-2.6B | LFM2.5-8B-A1B | Qwen3.5-4B | Gemma-4-E4B |
|---|---|---|---|---|
| IFBench | 59.17 | 56.47 | 48.40 / 50.38 | 39.24 / 39.48 |
| Multi-IF | 80.07 | 79.93 | 55.67 / 67.43 | 77.35 / 77.58 |
| BFCLv4 | 56.88 | 49.73 | 50.56 / 54.01 | 46.39 / 33.92 |
| AIME25 | 51.87 | 42.53 | 49.33 / 54.28 | 34.27 / 34.33 |
| AA-Omniscience accuracy | 8.13 | 8.67 | 17.63 / 17.20 | 8.33 / 8.10 |
Where a cell has two values, the first comes from the 2.6B card and the second from the 8B-A1B card. The two cards measured the same rival models and got different numbers. Read that as the noise in vendor tables. The pattern is stable: strong instruction following and tool use, weak recall of facts. The VL-3B card reports 80.7 on ScreenSpot-v2 and 87.9 on RefCOCO, up from 57.1 on RefCOCO for LFM2-VL-3B.
Failure modes and fixes
- The 8B-A1B fails in Ollama with
missing tensor 'output_norm.weight'. Update Ollama past v0.17.0. - Replies stop mid-thought. A low
max_tokenscut the reasoning. Raise the cap or leave it unset. - The phone app gets killed after long chats. The KV cache grew past what the OS allows. Lower
n_ctx, switch to a q8_0 cache, or use a smaller model. -hf ...:Q4_0picks the wrong file. The 2.6B, 350M and 230M repos have bothQ4_0andQAD-Q4_0. Download the file by name and load it with-m.- The DSpark sidecar fails to load. It needs its own target and a build at b10180 or later. A 2.6B drafter does not work with the 8B-A1B.
- MLX cannot find weights in
LFM2.5-2.6B-MLX. That repo stores precisions in subfolders. Use the standalone-MLX-4bitstyle repos. - The VL model ignores the image. The mmproj file did not load. Pass
--mmprojwhen you load local files. - Answers invent facts. Liquid says these models are weak at knowledge tasks. Add retrieval and put the facts in the prompt.
- Repetition at long outputs. Check that the repetition penalty from the preset reached the server. Many clients drop extra fields unless you use
extra_body.
FAQ
Can I use LFM2.5 commercially?
Yes, if your Legal Entity has annual revenue below $10 million. That entity includes the companies that control you, that you control, or that share control with you. Above that line, the license does not cover commercial use, and you need a separate deal with Liquid AI. Research and non-commercial use have no revenue limit.
Which model should I run on a phone?
Start with the 2.6B at Q4_0 or QAD Q4_0, which is 1.59 GB, with a context near 8K. Use the 350M or 230M for extraction and routing, or on phones with 8 GB of RAM or less.
Why is the cache so small?
Only 8 of the 2.6B's 30 layers use attention. The other 22 keep a fixed state of two past positions. The cache grows by 16,384 bytes per token, so 32,768 tokens cost 537 MB, against 2.01 GB if all 30 layers were attention.
Can I turn off thinking on the 2.6B?
No. The card calls it a pure reasoning model, and its template always opens the reply with a think tag. Budget output tokens for the reasoning.
Is the LEAP SDK still the way to ship on mobile?
No. Liquid marks it as deprecated with no new releases. Link llama.cpp into the app with the XCFramework on iOS or the NDK build on Android, and load GGUF files from Hugging Face.
Keep reading