GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

WhiskyAKM/LFM2.5-2.6B-NVFP4-GGUF overview

LFM2.5 2.6B GGUF GGUF quantized versions of LiquidAI/LFM2.5 2.6B https://huggingface.co/LiquidAI/LFM2.5 2.6B , a high performance hybrid model designed for on …

llama-cppggufliquidlfm2.5quantizedtext-generationconversationalbase_model:LiquidAI/LFM2.5-2.6B-Basebase_model:quantized:LiquidAI/LFM2.5-2.6B-Baselicense:otherendpoints_compatibleregion:us

Runs locally from ~1.48 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
lfm2.5-2.6b-nvfp4.ggufGGUFGGUF1.48 GBDownload

Model Details

Model IDWhiskyAKM/LFM2.5-2.6B-NVFP4-GGUF
AuthorWhiskyAKM
Pipelinetext-generation
Licenseother
Base modelLiquidAI/LFM2.5-2.6B-Base
Last modified2026-08-09T16:47:30.000Z

Model README

---

pipeline_tag: text-generation

base_model: LiquidAI/LFM2.5-2.6B-Base

license: other

license_name: lfm1.0

library_name: llama-cpp

tags:

  • liquid
  • lfm2.5
  • gguf
  • quantized

languages:

  • en
  • ar
  • zh
  • fr
  • de
  • hi
  • id
  • it
  • ja
  • ko
  • pl
  • pt
  • ru
  • es
  • th
  • vi

---

LFM2.5-2.6B GGUF

GGUF quantized versions of LiquidAI/LFM2.5-2.6B, a high-performance hybrid model designed for on-device deployment, featuring a 128K context window and advanced agentic capabilities.

Model Overview

LFM2.5-2.6B is part of the LFM2.5 family, building on the LFM2 architecture to provide best-in-class performance for its size. It is specifically optimized for agentic workloads, tool use, and long-context workflows, offering competitive performance against models 4x its size.

Key features include:

  • Agentic Post-Training: Trained using agentic reinforcement learning for improved tool use and instruction following.
  • Efficient Inference: Designed for high-speed execution on both CPU and GPU.
  • Reasoning Capabilities: A pure reasoning model that utilizes a <think> tag to reason before answering.
  • Massive Context: Supports up to 131,072 tokens.

Model Architecture

| Property | Value |

| :----------------------- | :--------- |

| Architecture | LFM2 |

| Parameters | 2.69B |

| Layers | 30 (22 conv + 8 GQA) |

| Context Length | 131,072 |

| Vocabulary Size | 128,000 |

| Training Budget | 34 Trillion Tokens |

| Supported Languages | English, Arabic, Chinese, French, German, Italian, Japanese, Korean, Portuguese, Spanish, Vietnamese, Thai, Indonesian, Hindi, Russian, Polish |

Available GGUF Files

| File | Quantization | Use Case |

| :------------------------- | :----------- | :----------------------------------------- |

| lfm2.5-2.6b-nvfp4.gguf | NVFP4 | Optimized 4-bit precision |

Usage

llama.cpp CLI

./llama-cli \
  -m lfm2.5-2.6b-nvfp4.gguf \
  -p "What is the capital of France?" \
  --temp 0.1 --top-k 50 --repeat-penalty 1.1

llama-server (OpenAI-compatible API)

./llama-server \
  -m lfm2.5-2.6b-nvfp4.gguf \
  --host 0.0.0.0 --port 8080

Chat Template & Reasoning

LFM2.5 uses a ChatML-like format. It is a reasoning model that automatically adds a <think> tag when starting an assistant answer to process its logic before providing the final response.

Example format:

<|startoftext|><|im_start|>system
You are a helpful assistant trained by Liquid AI.<|im_end|>
<|im_start|>user
What is C. elegans?<|im_end|>
<|im_start|>assistant
<think>
... reasoning process ...
</think>
C. elegans is a species of small roundworm...<|im_end|>

Tool Calling

LFM2.5 supports Pythonic function calling. It outputs function calls between <|tool_call_start|> and <|tool_call_end|> tokens.

Generation Parameters

Recommended parameters for optimal performance:

| Parameter | Value |

| :------------------ | :---- |

| Temperature | 0.1 |

| Top-K | 50 |

| Repetition Penalty | 1.1 |

Quantization

These GGUF files were created using llama.cpp tools to enable efficient local deployment on CPUs and GPUs with reduced memory footprints.

Acknowledgements

License

LFM 1.0 License

Run WhiskyAKM/LFM2.5-2.6B-NVFP4-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models