GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE โ†’
Model Intelligence Sheet

FadedRedStar/LFM2.5-8B-A1B-heretic-imatrix-GGUF overview

๐Ÿค– LFM2.5 8B A1B heretic โ€” Importance Matrix GGUF This repository hosts importance matrix imatrix optimized GGUF weights, available in multiple quantization foโ€ฆ

ggufllama.cpptext-generationhereticliquid-aiuncensoredimatrixabliteratedconversationaliq4_nlq4_k_mq5_k_menarzhfrdejakoesptbase_model:coder3101/LFM2.5-8B-A1B-hereticbase_model:quantized:coder3101/LFM2.5-8B-A1B-hereticlicense:other

Runs locally from ~4.51 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
2,127
Likes
0
Pipeline
text-generation

Repository Files & Downloads

3 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
LFM2.5-8B-A1B-heretic-IQ4_NL-imatrix.ggufGGUFIQ4_NL4.51 GBDownload
LFM2.5-8B-A1B-heretic-Q4_K_M-imatrix.ggufGGUFQ4_K_M4.80 GBDownload
LFM2.5-8B-A1B-heretic-Q5_K_M-imatrix.ggufGGUFQ5_K_M5.62 GBDownload

Model Details

Model IDFadedRedStar/LFM2.5-8B-A1B-heretic-imatrix-GGUF
AuthorFadedRedStar
Pipelinetext-generation
Licenseother
Base modelcoder3101/LFM2.5-8B-A1B-heretic
Last modified2026-07-10T15:51:20.000Z

Model README

---

base_model: coder3101/LFM2.5-8B-A1B-heretic

base_model_relation: quantized

library_name: gguf

license: other

license_name: lfm-1.0

license_link: https://huggingface.co/LiquidAI/LFM2.5-350M/blob/main/LICENSE

language:

  • en
  • ar
  • zh
  • fr
  • de
  • ja
  • ko
  • es
  • pt

pipeline_tag: text-generation

tags:

  • gguf
  • llama.cpp
  • text-generation
  • heretic
  • liquid-ai
  • uncensored
  • imatrix
  • abliterated
  • conversational
  • iq4_nl
  • q4_k_m
  • q5_k_m

quantized_by: FadedRedStar

---

๐Ÿค– LFM2.5-8B-A1B-heretic โ€” Importance Matrix GGUF

This repository hosts importance-matrix (imatrix) optimized GGUF weights, available in multiple quantization formats, for LFM2.5-8B-A1B-heretic, quantized from the source floating-point tensors provided by coder3101/LFM2.5-8B-A1B-heretic.

๐Ÿ”„ Sister Repository: Check out the Standard GGUF Sister Repository for uncalibrated and full 8-bit precision options.

๐ŸŽฏ Matrix-Weighted Calibration (Imatrix)

An Importance Matrix (imatrix) calculation tracks activations across network layers using a calibration sequence, then weights the quantization process to preserve the parameters that matter most for output quality โ€” improving fidelity at low bit depths.

โžก๏ธ Calibration dataset: Bartowski's calibration_datav5.txt.

> [!NOTE]

> * IQ4_NL is included because the matrix enables a non-linear 4-bit format that outperforms standard linear 4-bit quantization.

> * Q8_0 is absent because 8-bit quantization already introduces near-zero degradation, making calibration unnecessary โ€” see the standard sister repository for that variant.

โ„น๏ธ Model Profile & Core Features

LFM2.5-8B-A1B is a text-only model from Liquid AI's Liquid Foundation Model 2.5 series, designed for on-device deployment. It uses a hybrid architecture with 24 layers โ€” 18 double-gated LIV (Liquid, Input-adaptive, Value-selective) convolution layers plus 6 GQA (Grouped Query Attention) layers โ€” activating only approximately 1.5B parameters per forward pass out of 8.3B total. This delivers fastest-in-class throughput at its size on both CPU and GPU, with day-one support for llama.cpp, MLX, vLLM, and SGLang. The model is a reasoning model: it produces a chain-of-thought before its final answer, and is tuned for complex instruction following, tool calling, and chained agentic task execution.

The heretic suffix denotes post-processing via the Heretic v1.2.0 Arbitrary-Rank Ablation (ARA) method with row-norm preservation performed by coder3101, which removes refusal conditioning at multiple tensor ranks while maintaining the model's instruction-following and planning capabilities.

๐Ÿ“‹ Technical Specifications

| Property | Value |

|---|---|

| Base Architecture | LFM2.5 hybrid (18ร— double-gated LIV conv + 6ร— GQA) |

| Developed by | Liquid AI |

| Total Parameters | 8.3B |

| Active Parameters | ~1.5B per forward pass |

| Primary Use | Reasoning, instruction following, tool calling, agentic tasks |

| Context Window | 128,000 tokens |

| Training Budget | 38 trillion tokens |

| Languages | English, Arabic, Chinese, French, German, Japanese, Korean, Spanish, Portuguese |

| Abliteration Tool | Heretic v1.2.0 |

| Abliteration Method | Arbitrary-Rank Ablation (ARA) with row-norm preservation |

| Prompt Format | ChatML |

๐Ÿ› ๏ธ Heretic Overrides (ARA)

| Property | Value |

|---|---|

| start_layer_index | 7 |

| end_layer_index | 21 |

| preserve_good_behavior_weight | 0.8548 |

| steer_bad_behavior_weight | 0.0004 |

| overcorrect_relative_weight | 0.9494 |

| neighbor_count | 8 |

๐Ÿ“Š Refusal Bypass Metrics

> [!NOTE]

> The metrics below are self-reported by the original model author (coder3101) and have not been independently reproduced.

| Metric | This model | Original (LiquidAI/LFM2.5-8B-A1B) |

|---|---|---|

| KL divergence | 0.0239 | 0 (by definition) |

| Refusals | 12/100 | 91/100 |

๐Ÿงฎ Numerical & Tensor Formats

| Property | Value |

|---|---|

| Quantization Types | IQ4_NL, Q4_K_M, Q5_K_M (all with imatrix calibration) |

| Importance Matrix | Bartowski's calibration_datav5.txt |

๐Ÿ“ฆ Available Model Files

Main model weights

| Filename | Quantization | llama.cpp Build | Size | Download |

|---|---|---|---|---|

| LFM2.5-8B-A1B-heretic-IQ4_NL-imatrix.gguf | IQ4_NL | b9843 | 4.51 GB | ๐Ÿ“ฅ Download |

| LFM2.5-8B-A1B-heretic-Q4_K_M-imatrix.gguf | Q4_K_M | b9803 | 4.80 GB | ๐Ÿ“ฅ Download |

| LFM2.5-8B-A1B-heretic-Q5_K_M-imatrix.gguf | Q5_K_M | b9870 | 5.62 GB | ๐Ÿ“ฅ Download |

๐ŸŽ›๏ธ Component Pairing Guide

Download exactly one main weights file:

  • IQ4_NL: Non-linear 4-bit format, best choice for constrained memory when imatrix calibration is present.
  • Q4_K_M: Balanced 4-bit format suitable for most everyday use.
  • Q5_K_M: Higher-fidelity mid-range format recommended as a general default.

โšก Deployment & Execution Commands

> [!NOTE]

> Liquid AI recommends the following generation parameters for best results: temperature: 0.2, top_k: 80, repetition_penalty: 1.05.

> [!NOTE]

> This model emits reasoning content before its final answer. If you require a clean final answer only, parse the output accordingly rather than expecting a single direct response.

> [!TIP]

> Swap the -m filename below for either quantized file depending on your size/quality trade-off preference.

llama.cpp CLI

./llama-cli \
  -m LFM2.5-8B-A1B-heretic-IQ4_NL-imatrix.gguf \
  -c 8192 \
  -ngl 99 \
  --temp 0.2 \
  --top-k 80 \
  --repeat-penalty 1.05 \
  -p "<|im_start|>system\nYou are a helpful and precise assistant capable of using tools and following complex instructions.<|im_end|>\n<|im_start|>user\nBreak down the following task and execute it step by step: summarise this document and list action items.<|im_end|>\n<|im_start|>assistant\n"

OpenAI-Compatible API Server

./llama-server \
  --host 0.0.0.0 \
  --port 8080 \
  -m LFM2.5-8B-A1B-heretic-IQ4_NL-imatrix.gguf \
  -c 16384 \
  -ngl 99 \
  --flash-attn

๐Ÿ’ฌ Chat Templates & Prompt Design (ChatML)

<|im_start|>system
You are a capable assistant. Follow instructions precisely.<|im_end|>
<|im_start|>user
Your task or query here.<|im_end|>
<|im_start|>assistant

โš ๏ธ Safety & Operational Notes

  • This model is abliterated and will generate content that standard aligned models refuse. Use responsibly and in compliance with applicable laws.
  • This is a text-only model โ€” it has no vision encoder and cannot process images.
  • The LIV architecture activates only ~1.5B parameters per token, making it significantly faster to run than the total parameter count implies.
  • For long-context workloads, set -c up to 131072 as needed.
  • Liquid AI shipped a tokenizer fix for tool-calling after this model's initial release; if you encounter malformed tool-call output, verify your llama.cpp build includes this fix.
  • Imatrix calibration improves perplexity recovery compared to non-imatrix quantization, particularly on low-frequency tokens.
  • IQ4_NL produces a smaller file than Q4_K_M and tends to run faster on CPU and ARM devices; imatrix calibration narrows the quality gap between the two formats considerably.

Run FadedRedStar/LFM2.5-8B-A1B-heretic-imatrix-GGUF with guIDE

Download guIDE โ€” the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE โ†’ ยท Browse 524k+ models ยท Compare models

Source: Hugging Face ยท Compare models