GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

petr567/LFM2.5-2.6B-Windows-RTX-CUDA-GGUF overview

LFM2.5 2.6B Q4 K M Fast — Windows CUDA This repository contains a directly runnable Q4 K M GGUF of LiquidAI/LFM2.5 2.6B GGUF https://huggingface.co/LiquidAI/LF…

llama.cppgguflfm2lfm2.5windowscudanvidiaspeculative-decodingfast-inferencetext-generationbase_model:LiquidAI/LFM2.5-2.6B-GGUFbase_model:quantized:LiquidAI/LFM2.5-2.6B-GGUFlicense:otherendpoints_compatibleregion:usconversational

Runs locally from ~1.56 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
LFM2.5-2.6B-Q4_K_M.ggufGGUFQ4_K_M1.56 GBDownload

Model Details

Model IDpetr567/LFM2.5-2.6B-Windows-RTX-CUDA-GGUF
Authorpetr567
Pipelinetext-generation
Licenseother
Base modelLiquidAI/LFM2.5-2.6B-GGUF
Last modified2026-08-06T04:55:41.000Z

Model README

---

base_model: LiquidAI/LFM2.5-2.6B-GGUF

license: other

license_name: lfm-open-license-v1.0

license_link: https://www.liquid.ai/lfm-license

library_name: llama.cpp

pipeline_tag: text-generation

tags:

- gguf

- llama.cpp

- lfm2

- lfm2.5

- windows

- cuda

- nvidia

- speculative-decoding

- fast-inference

- text-generation

---

LFM2.5-2.6B Q4_K_M Fast — Windows CUDA

This repository contains a directly runnable Q4_K_M GGUF of

LiquidAI/LFM2.5-2.6B-GGUF

and a validated Fast single-request Windows/CUDA profile for llama.cpp.

The model weights are not modified. The same verified GGUF is published in both paired repositories; only the tested runtime profile differs.

> Fast profile: inference is accelerated relative to the same Q4_K_M baseline without the Fast runtime settings. The frozen validation gate detected no quality regression and no new failures. This is a measured result for the documented hardware, workloads, and single-request setup—not a universal guarantee for every prompt or runtime.

Measured Fast result

Primary metric: wall-clock decoded tokens per second for one request, without batching. The profile validation used three workloads with five repetitions each (15 runs total, 256 generated tokens per run).

| Workload | Baseline, tok/s | Fast, tok/s | Fast vs baseline |

|---|---:|---:|---:|

| Code copy | 114.74 | 226.27 | 1.972× (+97.2%) |

| Editorial rewrite | 112.09 | 141.26 | 1.260× (+26.0%) |

| Technical summary | 113.30 | 133.67 | 1.180× (+18.0%) |

| All 15 runs, mean ± SD | 113.38 ± 1.56 | 167.07 ± 43.51 | 1.474× (+47.4%) |

| Independent quality gate | Baseline | Fast | Regression |

|---|---:|---:|---:|

| Passed tasks | 10/12 | 10/12 | None measured |

The larger Fast standard deviation reflects the deliberately mixed workload set: repetitive code benefits more than free-form editing and summarization. These are profile-validation measurements, not the pending frozen cross-machine benchmark.

Choose the matching profile

| Platform | Hardware/backend | Repository |

|---|---|---|

| Windows 11 | NVIDIA RTX / CUDA | This repository |

| Ubuntu | AMD Strix Halo / Vulkan | LFM2.5-2.6B-Ubuntu-Strix-Halo-Vulkan-GGUF |

Included weight

| File | Quantization | Size | SHA-256 |

|---|---:|---:|---|

| LFM2.5-2.6B-Q4_K_M.gguf | Q4_K_M | 1,674,454,848 bytes (1.56 GiB) | 79fdf00351b46cf26f020aead28d01889886be87c55fa0eb907e6f9b00bfee14 |

Source revision: b22e29ebf6249a8c9fcdda36914743e9980595c4.

Tested setup

  • Windows 11 laptop
  • NVIDIA GeForce RTX 4060 Laptop GPU, 8 GiB VRAM
  • Intel Core i7-13650HX, 64 GiB RAM
  • NVIDIA driver 591.74
  • official llama.cpp CUDA server container, build 10066 (86a9c79f8)
  • context 8,192, one parallel slot, continuous batching disabled

Download and verify

Install the Hugging Face CLI once, or download the GGUF with the file link above.

py -m pip install -U huggingface_hub

$ModelDir = "C:\Models\LFM2.5-2.6B"
New-Item -ItemType Directory -Force -Path $ModelDir | Out-Null

hf download petr567/LFM2.5-2.6B-Windows-RTX-CUDA-GGUF `
  LFM2.5-2.6B-Q4_K_M.gguf `
  --local-dir $ModelDir

$Expected = "79fdf00351b46cf26f020aead28d01889886be87c55fa0eb907e6f9b00bfee14"
$Actual = (Get-FileHash "$ModelDir\LFM2.5-2.6B-Q4_K_M.gguf" -Algorithm SHA256).Hash.ToLower()
if ($Actual -ne $Expected) { throw "GGUF SHA-256 mismatch" }

Run the validated Fast Windows/CUDA profile

Docker Desktop must be configured for NVIDIA GPU access.

$ModelDir = "C:\Models\LFM2.5-2.6B"

docker run --rm --gpus all `
  -p 127.0.0.1:8080:8080 `
  -v "${ModelDir}:/models:ro" `
  --entrypoint /app/llama-server `
  ghcr.io/ggml-org/llama.cpp@sha256:1b3d1458ccda7287feab41b8001311acc03e24cde99ec0a2908fe83830562f38 `
  -m /models/LFM2.5-2.6B-Q4_K_M.gguf `
  --alias lfm2.5-2.6b-q4_k_m `
  --host 0.0.0.0 --port 8080 `
  -c 8192 -np 1 -ngl 99 `
  -t 14 -tb 14 -b 2048 -ub 512 `
  -fa auto -ctk q8_0 -ctv q8_0 `
  --no-cont-batching --no-cache-prompt --cache-ram 0 `
  --slot-prompt-similarity 0 --jinja --no-webui `
  --spec-type ngram-simple `
  --spec-ngram-simple-size-n 8 `
  --spec-ngram-simple-size-m 64 `
  --spec-ngram-simple-min-hits 1 `
  --spec-draft-n-max 64

The OpenAI-compatible endpoint is available at http://127.0.0.1:8080/v1.

$Body = @{
  model = "lfm2.5-2.6b-q4_k_m"
  messages = @(@{ role = "user"; content = "Write a short hello-world function in Python." })
  max_tokens = 128
  temperature = 0.2
} | ConvertTo-Json -Depth 5

Invoke-RestMethod -Method Post `
  -Uri "http://127.0.0.1:8080/v1/chat/completions" `
  -ContentType "application/json" `
  -Body $Body

LM Studio

The GGUF itself can also be opened in LM Studio. Use an 8,192-token context, maximum GPU offload, and one parallel request. The exact ngram-simple profile above requires a compatible llama.cpp server build; do not assume an arbitrary GUI runtime exposes the same acceleration controls.

Release scope

This release contains the runnable weight and the final launch recipe. The frozen cross-machine benchmark package and its results will be attached in a later revision after verification.

Attribution and license

The license includes a commercial-use revenue threshold. Review the included license before use or redistribution.

Run petr567/LFM2.5-2.6B-Windows-RTX-CUDA-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models