nicolasembleton/LFM2.5-2.6B-GGUF overview
LFM2.5 2.6B GGUF Quantized GGUF versions of LiquidAI/LFM2.5 2.6B https://huggingface.co/LiquidAI/LFM2.5 2.6B for efficient local inference via llama.cpp, LM St…
Runs locally from ~1.35 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| LFM2.5-2.6B-BF16.gguf | GGUF | BF16 | 5.03 GB | Download |
| LFM2.5-2.6B-F16.gguf | GGUF | F16 | 5.03 GB | Download |
| LFM2.5-2.6B-Q3_K_L.gguf | GGUF | Q3_K_L | 1.35 GB | Download |
| LFM2.5-2.6B-Q4_K_M.gguf | GGUF | Q4_K_M | 1.56 GB | Download |
| LFM2.5-2.6B-Q5_K_M.gguf | GGUF | Q5_K_M | 1.81 GB | Download |
| LFM2.5-2.6B-Q6_K.gguf | GGUF | Q6_K | 2.07 GB | Download |
| LFM2.5-2.6B-Q8_0.gguf | GGUF | Q8_0 | 2.68 GB | Download |
Model Details
| Model ID | nicolasembleton/LFM2.5-2.6B-GGUF |
|---|---|
| Author | nicolasembleton |
| Pipeline | text-generation |
| License | other |
| Base model | LiquidAI/LFM2.5-2.6B |
| Last modified | 2026-08-05T12:41:22.000Z |
Model README
---
license: other
license_name: lfm1.0
library_name: gguf
base_model: LiquidAI/LFM2.5-2.6B
pipeline_tag: text-generation
tags:
- gguf
- quantized
- llama.cpp
- lfm2
- lfm2.5
- liquid
- on-device
- webgpu
---
LFM2.5-2.6B-GGUF
Quantized GGUF versions of LiquidAI/LFM2.5-2.6B for efficient local inference via llama.cpp, LM Studio, and Ollama.
LFM2.5 is a hybrid (conv + attention) model designed for on-device agentic deployment. 2.6B parameters, 128K context, optimized for sub-2.5 GB running memory.
Quantization overview
This repo ships best quality per compression band — no Q2, no I-quants, no XL variants. Just the cleanest K-quant in each size band plus the lossless baselines.
| File | Size | Bits/weight | Use case |
|------|------|-------------|----------|
| LFM2.5-2.6B-F16.gguf | ~5.2 GB | 16 | Full precision, lossless |
| LFM2.5-2.6B-BF16.gguf | ~2.6 GB | 16 (bfloat16) | Faster loading, equivalent quality |
| LFM2.5-2.6B-Q8_0.gguf | ~2.9 GB | 8 | Near-lossless |
| LFM2.5-2.6B-Q6_K.gguf | ~2.4 GB | 6 | Excellent quality |
| LFM2.5-2.6B-Q5_K_M.gguf | ~2.1 GB | ~5.5 | High quality |
| LFM2.5-2.6B-Q4_K_M.gguf | ~1.8 GB | ~4.5 | Recommended default |
| LFM2.5-2.6B-Q3_K_L.gguf | ~1.5 GB | ~3.5 | Tight memory, lowest viable quality |
All K-quants use an importance matrix (imatrix) calibrated against Project Gutenberg text for better quality at low bit-widths.
Running
llama.cpp (CLI)
llama-cli -m LFM2.5-2.6B-Q4_K_M.gguf -c 4096 --color -i --temp 0.1 --top-k 50 --repeat-penalty 1.1
llama.cpp (one-liner via HF)
llama-cli -hf nicolasembleton/LFM2.5-2.6B-GGUF:Q4_K_M -c 4096 --color -i
Python (llama-cpp-python)
from llama_cpp import Llama
llm = Llama(
model_path="LFM2.5-2.6B-Q4_K_M.gguf",
n_ctx=4096,
n_threads=8,
n_gpu_layers=99, # offload all layers to GPU if available
)
print(llm("Hello, how are you?", max_tokens=256)["choices"][0]["text"])
Ollama
Create a Modelfile:
FROM ./LFM2.5-2.6B-Q4_K_M.gguf
Then:
ollama create lfm2.5-2.6b -f Modelfile
ollama run lfm2.5-2.6b
In-browser (Transformers.js + ONNX Runtime Web)
For browser-based inference, use the official ONNX export from Liquid AI:
LiquidAI/LFM2.5-2.6B-ONNX— multiple precision variants for all browsers
WebGPU path (Chrome, Firefox, Edge — fastest)
import { pipeline } from "@huggingface/transformers";
const generator = await pipeline("text-generation", "LiquidAI/LFM2.5-2.6B-ONNX", {
device: "webgpu",
dtype: "q4f16", // or "q4", "fp16"
});
const output = await generator("Hello, how are you?", { max_new_tokens: 256 });
WASM / Apple Safari path (no WebGPU needed)
import { pipeline } from "@huggingface/transformers";
const generator = await pipeline("text-generation", "LiquidAI/LFM2.5-2.6B-ONNX", {
device: "wasm",
dtype: "q8", // or "q4" for smaller download
});
const output = await generator("Hello, how are you?", { max_new_tokens: 256 });
The official ONNX repo ships FP32, FP16, Q4, Q4F16, and Q8 variants covering every browser configuration including older Apple Safari.
> Note: This GGUF repo is for native/server-side inference (llama.cpp, Ollama, LM Studio). For browser inference, use the ONNX repo above. The architectures are different export targets — both load the same underlying model.
Architecture
Lfm2ForCausalLM — hybrid model with 30 layers alternating conv/attention (config includes per-layer layer_types). 32 heads, 8 KV heads, 2048 hidden, 128K vocab, 128K context.
Built with llama.cpp b10276 (Aug 2026) — the first release to include LFM2 architecture support.
Files
*.gguf— quantized model filesREADME.md— this file
License
Inherited: LFM 1.0 license (see LiquidAI/LFM2.5-2.6B).
Citation
@misc{lfm25-2.6b-gguf,
title = {{LFM2.5-2.6B-GGUF}},
author = {{Liquid AI, quantizations by nicolasembleton}},
year = {{2026}},
howpublished = {{Hugging Face}},
note = {{GGUF quantizations of LFM2.5-2.6B; for browser inference see LiquidAI/LFM2.5-2.6B-ONNX}},
}}Run nicolasembleton/LFM2.5-2.6B-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models