GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

WhiskyAKM/LFM2.5-2.6B-GGUF overview

LFM2.5 2.6B GGUF GGUF quantized versions of LiquidAI/LFM2.5 2.6B https://huggingface.co/LiquidAI/LFM2.5 2.6B , a high performance hybrid model designed for on …

llama-cppggufliquidlfm2.5quantizedtext-generationconversationalbase_model:LiquidAI/LFM2.5-2.6B-Basebase_model:quantized:LiquidAI/LFM2.5-2.6B-Baselicense:otherendpoints_compatibleregion:us

Runs locally from ~1.48 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

8 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
lfm2.5-2.6b-Q4_0.ggufGGUFQ4_01.48 GBDownload
lfm2.5-2.6b-Q4_K_M.ggufGGUFQ4_K_M1.56 GBDownload
lfm2.5-2.6b-Q4_K_S.ggufGGUFQ4_K_S1.49 GBDownload
lfm2.5-2.6b-Q5_K_M.ggufGGUFQ5_K_M1.81 GBDownload
lfm2.5-2.6b-Q5_K_S.ggufGGUFQ5_K_S1.77 GBDownload
lfm2.5-2.6b-Q6_K.ggufGGUFQ6_K2.07 GBDownload
lfm2.5-2.6b-Q8_0.ggufGGUFQ8_02.68 GBDownload
lfm2.5-2.6b.ggufGGUFGGUF5.03 GBDownload

Model Details

Model IDWhiskyAKM/LFM2.5-2.6B-GGUF
AuthorWhiskyAKM
Pipelinetext-generation
Licenseother
Base modelLiquidAI/LFM2.5-2.6B-Base
Last modified2026-08-09T16:33:41.000Z

Model README

---

pipeline_tag: text-generation

base_model: LiquidAI/LFM2.5-2.6B-Base

license: other

license_name: lfm1.0

library_name: llama-cpp

tags:

  • liquid
  • lfm2.5
  • gguf
  • quantized

languages:

  • en
  • ar
  • zh
  • fr
  • de
  • hi
  • id
  • it
  • ja
  • ko
  • pl
  • pt
  • ru
  • es
  • th
  • vi

---

LFM2.5-2.6B GGUF

GGUF quantized versions of LiquidAI/LFM2.5-2.6B, a high-performance hybrid model designed for on-device deployment, featuring a 128K context window and advanced agentic capabilities.

Model Overview

LFM2.5-2.6B is part of the LFM2.5 family, building on the LFM2 architecture to provide best-in-class performance for its size. It is specifically optimized for agentic workloads, tool use, and long-context workflows, offering competitive performance against models 4x its size.

Key features include:

  • Agentic Post-Training: Trained using agentic reinforcement learning for improved tool use and instruction following.
  • Efficient Inference: Designed for high-speed execution on both CPU and GPU.
  • Reasoning Capabilities: A pure reasoning model that utilizes a <think> tag to reason before answering.
  • Massive Context: Supports up to 131,072 tokens.

Model Architecture

| Property | Value |

| :----------------------- | :--------- |

| Architecture | LFM2 |

| Parameters | 2.69B |

| Layers | 30 (22 conv + 8 GQA) |

| Context Length | 131,072 |

| Vocabulary Size | 128,000 |

| Training Budget | 34 Trillion Tokens |

| Supported Languages | English, Arabic, Chinese, French, German, Italian, Japanese, Korean, Portuguese, Spanish, Vietnamese, Thai, Indonesian, Hindi, Russian, Polish |

Available GGUF Files

| File | Quantization | Use Case |

| :------------------------- | :----------- | :----------------------------------------- |

| lfm2.5-2.6b.gguf | FP16/BF16 | Max precision, reference model |

| lfm2.5-2.6b-Q8_0.gguf | Q8_0 | Near-lossless, high fidelity |

| lfm2.5-2.6b-Q6_K.gguf | Q6_K | Very high quality, recommended for quality |

| lfm2.5-2.6b-Q5_K_M.gguf | Q5_K_M | High quality, balanced |

| lfm2.5-2.6b-Q5_K_S.gguf | Q5_K_S | High quality, slightly smaller |

| lfm2.5-2.6b-Q4_K_M.gguf | Q4_K_M | Good quality, recommended default |

| lfm2.5-2.6b-Q4_K_S.gguf | Q4_K_S | Smaller, acceptable quality |

| lfm2.5-2.6b-Q4_0.gguf | Q4_0 | Legacy quant, fastest inference |

> Recommended: Q4_K_M or Q5_K_M offer the best quality-to-size trade-off for most use cases.

Usage

llama.cpp CLI

./llama-cli \
  -m lfm2.5-2.6b-Q4_K_M.gguf \
  -p "What is the capital of France?" \
  --temp 0.1 --top-k 50 --repeat-penalty 1.1

llama-server (OpenAI-compatible API)

./llama-server \
  -m lfm2.5-2.6b-Q4_K_M.gguf \
  --host 0.0.0.0 --port 8080

Chat Template & Reasoning

LFM2.5 uses a ChatML-like format. It is a reasoning model that automatically adds a <think> tag when starting an assistant answer to process its logic before providing the final response.

Example format:

<|startoftext|><|im_start|>system
You are a helpful assistant trained by Liquid AI.<|im_end|>
<|im_start|>user
What is C. elegans?<|im_end|>
<|im_start|>assistant
<think>
... reasoning process ...
</think>
C. elegans is a species of small roundworm...<|im_end|>

Tool Calling

LFM2.5 supports Pythonic function calling. It outputs function calls between <|tool_call_start|> and <|tool_call_end|> tokens.

Generation Parameters

Recommended parameters for optimal performance:

| Parameter | Value |

| :------------------ | :---- |

| Temperature | 0.1 |

| Top-K | 50 |

| Repetition Penalty | 1.1 |

Quantization

These GGUF files were created using llama.cpp tools to enable efficient local deployment on CPUs and GPUs with reduced memory footprints.

Acknowledgements

License

LFM 1.0 License

Run WhiskyAKM/LFM2.5-2.6B-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models