GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

nicolasembleton/LFM2.5-2.6B-GGUF overview

LFM2.5 2.6B GGUF Quantized GGUF versions of LiquidAI/LFM2.5 2.6B https://huggingface.co/LiquidAI/LFM2.5 2.6B for efficient local inference via llama.cpp, LM St…

ggufquantizedllama.cpplfm2lfm2.5liquidon-devicewebgputext-generationbase_model:LiquidAI/LFM2.5-2.6Bbase_model:quantized:LiquidAI/LFM2.5-2.6Blicense:otherendpoints_compatibleregion:usconversational

Runs locally from ~1.35 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation

Repository Files & Downloads

7 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
LFM2.5-2.6B-BF16.ggufGGUFBF165.03 GBDownload
LFM2.5-2.6B-F16.ggufGGUFF165.03 GBDownload
LFM2.5-2.6B-Q3_K_L.ggufGGUFQ3_K_L1.35 GBDownload
LFM2.5-2.6B-Q4_K_M.ggufGGUFQ4_K_M1.56 GBDownload
LFM2.5-2.6B-Q5_K_M.ggufGGUFQ5_K_M1.81 GBDownload
LFM2.5-2.6B-Q6_K.ggufGGUFQ6_K2.07 GBDownload
LFM2.5-2.6B-Q8_0.ggufGGUFQ8_02.68 GBDownload

Model Details

Model IDnicolasembleton/LFM2.5-2.6B-GGUF
Authornicolasembleton
Pipelinetext-generation
Licenseother
Base modelLiquidAI/LFM2.5-2.6B
Last modified2026-08-05T12:41:22.000Z

Model README

---

license: other

license_name: lfm1.0

library_name: gguf

base_model: LiquidAI/LFM2.5-2.6B

pipeline_tag: text-generation

tags:

  • gguf
  • quantized
  • llama.cpp
  • lfm2
  • lfm2.5
  • liquid
  • on-device
  • webgpu

---

LFM2.5-2.6B-GGUF

Quantized GGUF versions of LiquidAI/LFM2.5-2.6B for efficient local inference via llama.cpp, LM Studio, and Ollama.

LFM2.5 is a hybrid (conv + attention) model designed for on-device agentic deployment. 2.6B parameters, 128K context, optimized for sub-2.5 GB running memory.

Quantization overview

This repo ships best quality per compression band — no Q2, no I-quants, no XL variants. Just the cleanest K-quant in each size band plus the lossless baselines.

| File | Size | Bits/weight | Use case |

|------|------|-------------|----------|

| LFM2.5-2.6B-F16.gguf | ~5.2 GB | 16 | Full precision, lossless |

| LFM2.5-2.6B-BF16.gguf | ~2.6 GB | 16 (bfloat16) | Faster loading, equivalent quality |

| LFM2.5-2.6B-Q8_0.gguf | ~2.9 GB | 8 | Near-lossless |

| LFM2.5-2.6B-Q6_K.gguf | ~2.4 GB | 6 | Excellent quality |

| LFM2.5-2.6B-Q5_K_M.gguf | ~2.1 GB | ~5.5 | High quality |

| LFM2.5-2.6B-Q4_K_M.gguf | ~1.8 GB | ~4.5 | Recommended default |

| LFM2.5-2.6B-Q3_K_L.gguf | ~1.5 GB | ~3.5 | Tight memory, lowest viable quality |

All K-quants use an importance matrix (imatrix) calibrated against Project Gutenberg text for better quality at low bit-widths.

Running

llama.cpp (CLI)

llama-cli -m LFM2.5-2.6B-Q4_K_M.gguf -c 4096 --color -i   --temp 0.1 --top-k 50 --repeat-penalty 1.1

llama.cpp (one-liner via HF)

llama-cli -hf nicolasembleton/LFM2.5-2.6B-GGUF:Q4_K_M -c 4096 --color -i

Python (llama-cpp-python)

from llama_cpp import Llama

llm = Llama(
    model_path="LFM2.5-2.6B-Q4_K_M.gguf",
    n_ctx=4096,
    n_threads=8,
    n_gpu_layers=99,  # offload all layers to GPU if available
)
print(llm("Hello, how are you?", max_tokens=256)["choices"][0]["text"])

Ollama

Create a Modelfile:

FROM ./LFM2.5-2.6B-Q4_K_M.gguf

Then:

ollama create lfm2.5-2.6b -f Modelfile
ollama run lfm2.5-2.6b

In-browser (Transformers.js + ONNX Runtime Web)

For browser-based inference, use the official ONNX export from Liquid AI:

WebGPU path (Chrome, Firefox, Edge — fastest)

import { pipeline } from "@huggingface/transformers";

const generator = await pipeline("text-generation", "LiquidAI/LFM2.5-2.6B-ONNX", {
  device: "webgpu",
  dtype: "q4f16",  // or "q4", "fp16"
});
const output = await generator("Hello, how are you?", { max_new_tokens: 256 });

WASM / Apple Safari path (no WebGPU needed)

import { pipeline } from "@huggingface/transformers";

const generator = await pipeline("text-generation", "LiquidAI/LFM2.5-2.6B-ONNX", {
  device: "wasm",
  dtype: "q8",  // or "q4" for smaller download
});
const output = await generator("Hello, how are you?", { max_new_tokens: 256 });

The official ONNX repo ships FP32, FP16, Q4, Q4F16, and Q8 variants covering every browser configuration including older Apple Safari.

> Note: This GGUF repo is for native/server-side inference (llama.cpp, Ollama, LM Studio). For browser inference, use the ONNX repo above. The architectures are different export targets — both load the same underlying model.

Architecture

Lfm2ForCausalLM — hybrid model with 30 layers alternating conv/attention (config includes per-layer layer_types). 32 heads, 8 KV heads, 2048 hidden, 128K vocab, 128K context.

Built with llama.cpp b10276 (Aug 2026) — the first release to include LFM2 architecture support.

Files

  • *.gguf — quantized model files
  • README.md — this file

License

Inherited: LFM 1.0 license (see LiquidAI/LFM2.5-2.6B).

Citation

@misc{lfm25-2.6b-gguf,
  title = {{LFM2.5-2.6B-GGUF}},
  author = {{Liquid AI, quantizations by nicolasembleton}},
  year = {{2026}},
  howpublished = {{Hugging Face}},
  note = {{GGUF quantizations of LFM2.5-2.6B; for browser inference see LiquidAI/LFM2.5-2.6B-ONNX}},
}}

Run nicolasembleton/LFM2.5-2.6B-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models