GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE โ†’
Model Intelligence Sheet

FadedRedStar/LFM2.5-8B-A1B-heretic-GGUF overview

๐Ÿค– LFM2.5 8B A1B heretic โ€” GGUF This repository hosts GGUF weights for LFM2.5 8B A1B heretic , quantized from the source floating point tensors provided by codโ€ฆ

ggufllama.cpptext-generationhereticliquid-aiuncensoredabliteratedconversationalq4_k_mq5_k_mq8_0enarzhfrdejakoesptbase_model:coder3101/LFM2.5-8B-A1B-hereticbase_model:quantized:coder3101/LFM2.5-8B-A1B-hereticlicense:otherendpoints_compatible

Runs locally from ~4.80 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
558
Likes
0
Pipeline
text-generation

Repository Files & Downloads

3 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
LFM2.5-8B-A1B-heretic-Q4_K_M.ggufGGUFQ4_K_M4.80 GBDownload
LFM2.5-8B-A1B-heretic-Q5_K_M.ggufGGUFQ5_K_M5.62 GBDownload
LFM2.5-8B-A1B-heretic-Q8_0.ggufGGUFQ8_08.39 GBDownload

Model Details

Model IDFadedRedStar/LFM2.5-8B-A1B-heretic-GGUF
AuthorFadedRedStar
Pipelinetext-generation
Licenseother
Base modelcoder3101/LFM2.5-8B-A1B-heretic
Last modified2026-07-10T15:51:28.000Z

Model README

---

base_model: coder3101/LFM2.5-8B-A1B-heretic

base_model_relation: quantized

library_name: gguf

license: other

license_name: lfm-1.0

license_link: https://huggingface.co/LiquidAI/LFM2.5-350M/blob/main/LICENSE

language:

  • en
  • ar
  • zh
  • fr
  • de
  • ja
  • ko
  • es
  • pt

pipeline_tag: text-generation

tags:

  • gguf
  • llama.cpp
  • text-generation
  • heretic
  • liquid-ai
  • uncensored
  • abliterated
  • conversational
  • q4_k_m
  • q5_k_m
  • q8_0

quantized_by: FadedRedStar

---

๐Ÿค– LFM2.5-8B-A1B-heretic โ€” GGUF

This repository hosts GGUF weights for LFM2.5-8B-A1B-heretic, quantized from the source floating-point tensors provided by coder3101/LFM2.5-8B-A1B-heretic.

๐Ÿ”„ Sister Repository: Check out the Imatrix Sister Repository for enhanced precision at lower bit fractions.

> [!NOTE]

> If you plan on using 4-bit or 5-bit variants, consider the imatrix sister repository instead โ€” importance matrix calibration improves logic retention at those bit depths. This repository is best suited if you want the near-lossless Q8_0 build.

โ„น๏ธ Model Profile & Core Features

LFM2.5-8B-A1B is a text-only model from Liquid AI's Liquid Foundation Model 2.5 series, designed for on-device deployment. It uses a hybrid architecture with 24 layers โ€” 18 double-gated LIV (Liquid, Input-adaptive, Value-selective) convolution layers plus 6 GQA (Grouped Query Attention) layers โ€” activating only approximately 1.5B parameters per forward pass out of 8.3B total. This delivers fastest-in-class throughput at its size on both CPU and GPU, with day-one support for llama.cpp, MLX, vLLM, and SGLang. The model is a reasoning model: it produces a chain-of-thought before its final answer, and is tuned for complex instruction following, tool calling, and chained agentic task execution.

The heretic suffix denotes post-processing via the Heretic v1.2.0 Arbitrary-Rank Ablation (ARA) method with row-norm preservation performed by coder3101, which removes refusal conditioning at multiple tensor ranks while maintaining the model's instruction-following and planning capabilities.

๐Ÿ“‹ Technical Specifications

| Property | Value |

|---|---|

| Base Architecture | LFM2.5 hybrid (18ร— double-gated LIV conv + 6ร— GQA) |

| Developed by | Liquid AI |

| Total Parameters | 8.3B |

| Active Parameters | ~1.5B per forward pass |

| Primary Use | Reasoning, instruction following, tool calling, agentic tasks |

| Context Window | 128,000 tokens |

| Training Budget | 38 trillion tokens |

| Languages | English, Arabic, Chinese, French, German, Japanese, Korean, Spanish, Portuguese |

| Abliteration Tool | Heretic v1.2.0 |

| Abliteration Method | Arbitrary-Rank Ablation (ARA) with row-norm preservation |

| Prompt Format | ChatML |

๐Ÿ› ๏ธ Heretic Overrides (ARA)

| Property | Value |

|---|---|

| start_layer_index | 7 |

| end_layer_index | 21 |

| preserve_good_behavior_weight | 0.8548 |

| steer_bad_behavior_weight | 0.0004 |

| overcorrect_relative_weight | 0.9494 |

| neighbor_count | 8 |

๐Ÿ“Š Refusal Bypass Metrics

> [!NOTE]

> The metrics below are self-reported by the original model author (coder3101) and have not been independently reproduced.

| Metric | This model | Original (LiquidAI/LFM2.5-8B-A1B) |

|---|---|---|

| KL divergence | 0.0239 | 0 (by definition) |

| Refusals | 12/100 | 91/100 |

๐Ÿงฎ Numerical & Tensor Formats

| Property | Value |

|---|---|

| Quantization Type | Q4_K_M, Q5_K_M, Q8_0 |

๐Ÿ“ฆ Available Model Files

Main model weights

| Filename | Quantization | llama.cpp Build | Size | Download |

|---|---|---|---|---|

| LFM2.5-8B-A1B-heretic-Q4_K_M.gguf | Q4_K_M | b9803 | 4.80 GB | ๐Ÿ“ฅ Download |

| LFM2.5-8B-A1B-heretic-Q5_K_M.gguf | Q5_K_M | b9870 | 5.62 GB | ๐Ÿ“ฅ Download |

| LFM2.5-8B-A1B-heretic-Q8_0.gguf | Q8_0 | b9870 | 8.39 GB | ๐Ÿ“ฅ Download |

๐ŸŽ›๏ธ Component Pairing Guide

Download exactly one main weights file:

  • Q4_K_M: Balanced 4-bit format suitable for most everyday use.
  • Q5_K_M: Higher-fidelity mid-range format recommended as a general default.
  • Q8_0: Near-lossless 8-bit format for when memory is not a constraint.

โšก Deployment & Execution Commands

> [!NOTE]

> Liquid AI recommends the following generation parameters for best results: temperature: 0.2, top_k: 80, repetition_penalty: 1.05.

> [!NOTE]

> This model emits reasoning content before its final answer. If you require a clean final answer only, parse the output accordingly rather than expecting a single direct response.

> [!TIP]

> Swap the -m filename below for either quantized file depending on your size/quality trade-off preference.

llama.cpp CLI

./llama-cli \
  -m LFM2.5-8B-A1B-heretic-Q4_K_M.gguf \
  -c 8192 \
  -ngl 99 \
  --temp 0.2 \
  --top-k 80 \
  --repeat-penalty 1.05 \
  -p "<|im_start|>system\nYou are a helpful and precise assistant capable of using tools and following complex instructions.<|im_end|>\n<|im_start|>user\nBreak down the following task and execute it step by step: summarise this document and list action items.<|im_end|>\n<|im_start|>assistant\n"

OpenAI-Compatible API Server

./llama-server \
  --host 0.0.0.0 \
  --port 8080 \
  -m LFM2.5-8B-A1B-heretic-Q4_K_M.gguf \
  -c 16384 \
  -ngl 99 \
  --flash-attn

๐Ÿ’ฌ Chat Templates & Prompt Design (ChatML)

<|im_start|>system
You are a capable assistant. Follow instructions precisely.<|im_end|>
<|im_start|>user
Your task or query here.<|im_end|>
<|im_start|>assistant

โš ๏ธ Safety & Operational Notes

  • This model is abliterated and will generate content that standard aligned models refuse. Use responsibly and in compliance with applicable laws.
  • This is a text-only model โ€” it has no vision encoder and cannot process images.
  • The LIV architecture activates only ~1.5B parameters per token, making it significantly faster to run than the total parameter count implies.
  • For long-context workloads, set -c up to 131072 as needed.
  • Liquid AI shipped a tokenizer fix for tool-calling after this model's initial release; if you encounter malformed tool-call output, verify your llama.cpp build includes this fix.
  • For better output quality at this quantization level, consider the imatrix variant in the companion repository.

Run FadedRedStar/LFM2.5-8B-A1B-heretic-GGUF with guIDE

Download guIDE โ€” the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE โ†’ ยท Browse 524k+ models ยท Compare models

Source: Hugging Face ยท Compare models