GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

LiquidAI/LFM2.5-2.6B-DSpark-GGUF overview

<div align="center" <img src="https://cdn uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/2b08LKpev0DNEk6DlnWkY.png" alt="Liquid AI" style="…

llama.cppggufspeculative-decodingdsparklfm2draft-modeltext-generationbase_model:LiquidAI/LFM2.5-2.6B-DSparkbase_model:quantized:LiquidAI/LFM2.5-2.6B-DSparklicense:otherendpoints_compatibleregion:usconversational

Runs locally from ~190.4 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
940
Likes
29
Pipeline
text-generation
Author

Repository Files & Downloads

3 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
LFM2.5-2.6B-DSpark-F16.ggufGGUFF16632.9 MBDownload
LFM2.5-2.6B-DSpark-Q4_K_M.ggufGGUFQ4_K_M190.4 MBDownload
LFM2.5-2.6B-DSpark-Q8_0.ggufGGUFQ8_0340.0 MBDownload

Model Details

Model IDLiquidAI/LFM2.5-2.6B-DSpark-GGUF
AuthorLiquidAI
Pipelinetext-generation
Licenseother
Base modelLiquidAI/LFM2.5-2.6B-DSpark
Last modified2026-08-21T10:28:48.000Z

Model README

---

library_name: llama.cpp

base_model: LiquidAI/LFM2.5-2.6B-DSpark

license: other

license_name: lfm1.0

license_link: LICENSE

pipeline_tag: text-generation

tags:

  • speculative-decoding
  • dspark
  • lfm2
  • draft-model
  • gguf

---

<div align="center">

<img

src="https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/2b08LKpev0DNEk6DlnWkY.png"

alt="Liquid AI"

style="width: 100%; max-width: 100%; height: auto; display: inline-block; margin-bottom: 0.5em; margin-top: 0.5em;"

/>

<div style="display: flex; justify-content: center; gap: 0.5em; margin-bottom: 1em;">

<a href="https://playground.liquid.ai/"><strong>Try LFM</strong></a> •

<a href="https://docs.liquid.ai/lfm/getting-started/welcome"><strong>Docs</strong></a> •

<a href="https://leap.liquid.ai/"><strong>LEAP</strong></a> •

<a href="https://discord.com/invite/liquid-ai"><strong>Discord</strong></a>

</div>

</div>

LFM2.5-2.6B-DSpark-GGUF

GGUF build of LiquidAI/LFM2.5-2.6B-DSpark for llama.cpp (DSpark speculative decoding is in mainline, ggml-org/llama.cpp #25173).

This is a standalone draft sidecar: it carries only the drafter (5 attention layers, rank-256 Markov head, confidence head, block size 9). Token embeddings and the LM head are shared from the target model at load time, so it must be paired with a LFM2.5-2.6B-GGUF target file.

Find more information about LFM2.5-DSpark in our blog post.

📦 Files

| file | quant | size | notes |

|---|---|---:|---|

| LFM2.5-2.6B-DSpark-F16.gguf | F16 | 664 MB | best accept length, recommended when memory allows |

| LFM2.5-2.6B-DSpark-Q8_0.gguf | Q8_0 | 349 MB | accept length −2% vs F16 |

| LFM2.5-2.6B-DSpark-Q4_K_M.gguf | Q4_K_M | 191 MB | accept length −3% vs F16, smallest recommended — sub-4-bit draft quants measurably hurt both accept length and throughput |

Draft quantization changes speed only marginally (the drafter is a small share of each cycle); choose by memory budget. The target model quant is the main speed/quality lever and is independent of this file.

🏃 How to run (llama.cpp)

llama-server -m LFM2.5-2.6B-F16.gguf \
  -md LFM2.5-2.6B-DSpark-F16.gguf \
  --spec-type draft-dspark --spec-draft-n-max 10 --spec-draft-n-min 0 \
  -fa on -ngl 99

The block size is read from the sidecar metadata (n-max is clamped to it). Speculative decoding is exact: the target verifies every proposed token, so greedy output equals the target alone; per-response timings report draft_n / draft_n_accepted.

Other models in the LFM2.5-DSpark GGUF family:

| Draft (GGUF) | Target (GGUF) |

|---|---|

| LFM2.5-1.2B-Instruct-DSpark-GGUF | LFM2.5-1.2B-Instruct-GGUF |

| LFM2.5-2.6B-DSpark-GGUF | LFM2.5-2.6B-GGUF |

| LFM2.5-8B-A1B-DSpark-GGUF | LFM2.5-8B-A1B-GGUF |

📊 Acceptance and benchmarks

See LiquidAI/LFM2.5-2.6B-DSpark for acceptance-length tables (H100 and Apple silicon) and target benchmarks.

📬 Contact

Citation

@article{liquidAI202626B,
  author  = {Liquid AI},
  title   = {LFM2.5-2.6B: Agents Everywhere},
  journal = {Liquid AI Blog},
  year    = {2026},
  note    = {www.liquid.ai/blog/lfm2-5-2-6b},
}
@article{liquidAI2026dspark,
  author = {Liquid AI},
  title = {LFM2.5-DSpark: Up to 3.2x Faster Inference from H100 to MacBook},
  journal = {Liquid AI Blog},
  year = {2026},
  note = {www.liquid.ai/blog/lfm2.5-dspark},
}

Run LiquidAI/LFM2.5-2.6B-DSpark-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models