SC117/LFM2.5-2.6B-Uncensored-GGUF overview
<div style="font family: apple system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans serif; margin bottom: 24px;" <div style="background: f5f5f7; border radius:…
Runs locally from ~1.14 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| LFM2.5-2.6B-Uncensored-BF16.gguf | GGUF | BF16 | 5.03 GB | Download |
| LFM2.5-2.6B-Uncensored-IQ3_XS.gguf | GGUF | IQ3_XS | 1.14 GB | Download |
| LFM2.5-2.6B-Uncensored-IQ4_XS.gguf | GGUF | IQ4_XS | 1.41 GB | Download |
| LFM2.5-2.6B-Uncensored-Q4_K_M.gguf | GGUF | Q4_K_M | 1.56 GB | Download |
| LFM2.5-2.6B-Uncensored-Q6_K.gguf | GGUF | Q6_K | 2.07 GB | Download |
| LFM2.5-2.6B-Uncensored-Q8_0.gguf | GGUF | Q8_0 | 2.68 GB | Download |
Model Details
| Model ID | SC117/LFM2.5-2.6B-Uncensored-GGUF |
|---|---|
| Author | SC117 |
| Pipeline | text-generation |
| License | other |
| Base model | LiquidAI/LFM2.5-2.6B,SC117/LFM2.5-2.6B-Uncensored |
| Last modified | 2026-08-05T08:31:02.000Z |
Model README
---
library_name: gguf
license: other
license_name: lfm1.0
license_link: LICENSE
pipeline_tag: text-generation
tags:
- liquid
- lfm2.5
- edge
- uncensored
- abliterix
- quantization
- gguf
- imatrix
base_model:
- LiquidAI/LFM2.5-2.6B
- SC117/LFM2.5-2.6B-Uncensored
base_model_relation: quantized
---
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; margin-bottom: 24px;">
<div style="background: #f5f5f7; border-radius: 20px; padding: 36px 32px; margin-bottom: 20px; text-align: center; position: relative; overflow: hidden;">
<div style="position: absolute; top: -30px; right: -30px; width: 120px; height: 120px; background: #ddd6fe; border-radius: 50%;"></div>
<div style="position: absolute; bottom: -20px; left: 40px; width: 80px; height: 80px; background: #c4b5fd; border-radius: 50%;"></div>
<div style="position: absolute; top: 50%; left: -15px; width: 60px; height: 60px; background: #ddd6fe; border-radius: 50%;"></div>
<div style="display: inline-flex; gap: 8px; margin-bottom: 16px; position: relative; z-index: 1; flex-wrap: wrap; justify-content: center;">
<span style="background: #6366f1; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">ABLITERIX</span>
<span style="background: #8b5cf6; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">TRIAL 65</span>
<span style="background: #0ea5e9; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">GGUF + IMATRIX</span>
<span style="background: #64748b; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">LFM Open 1.0</span>
</div>
<h1 style="margin: 0 0 8px 0; font-size: 28px; font-weight: 700; color: #1e1b4b; letter-spacing: -0.5px; border: none; position: relative; z-index: 1;">LFM2.5-2.6B-Uncensored-GGUF</h1>
<p style="margin: 8px 0 0 0; font-size: 14px; position: relative; z-index: 1;"><span style="color: #6b7280;">English</span> | <a href="https://huggingface.co/SC117/LFM2.5-2.6B-Uncensored-GGUF/blob/main/README_zh.md" style="color: #6366f1; text-decoration: none;">📖 中文文档</a></p>
<p style="margin: 8px 0 0 0; font-size: 15px; color: #6b7280; position: relative; z-index: 1;">Uncensored 2.6B edge model · abliterix Trial 65 · imatrix-calibrated GGUFs</p>
</div>
</div>
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; display: flex; flex-direction: column; gap: 20px; margin-bottom: 30px;">
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #6366f1 0%, #8b5cf6 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🌊</span> About this release</div>
<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">
<p style="margin: 0 0 12px 0;"><a href="https://huggingface.co/LiquidAI/LFM2.5-2.6B" target="_blank" style="color: #4f46e5; text-decoration: none; font-weight: 700;">LFM2.5-2.6B</a> is a Liquid AI <b>2.6B-parameter hybrid edge model</b> built for agentic workloads: 30 layers (22 double-gated short-convolution blocks + 8 GQA), a <b>128K context window</b>, 128K vocabulary, and a ChatML-like template with native <code><think></code> reasoning.</p>
<p style="margin: 0 0 12px 0;">These quantized GGUFs are built from our <a href="https://huggingface.co/SC117/LFM2.5-2.6B-Uncensored" target="_blank" style="color: #4f46e5; text-decoration: none; font-weight: 700;">LFM2.5-2.6B-Uncensored</a> BF16 release (abliterix <b>Trial 65</b>, stream-merged to BF16) in three steps:</p>
<ol style="margin: 0 0 12px 0; padding-left: 20px;">
<li style="margin-bottom: 6px;"><b>BF16 GGUF conversion</b> with llama.cpp (<code>lfm2</code> architecture support).</li>
<li style="margin-bottom: 6px;"><b>imatrix calibration</b> — 401 chunks (≈1.6M tokens) from the APEX calibration set, computed on the BF16 GGUF.</li>
<li style="margin-bottom: 6px;"><b>Quantization</b> with <code>llama-quantize --imatrix</code> into five tiers.</li>
</ol>
<p style="margin: 0;"><b>License: LFM Open License v1.0</b> (same as the base model).</p>
</div>
</div>
<div style="border: 1px solid #fdba74; border-radius: 12px; overflow: hidden; background: #fff7ed; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #f97316 0%, #ea580c 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>⚠️</span> Uncensored notice</div>
<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">
<p style="margin: 0 0 12px 0;">After merging abliterix <b>Trial 65</b>, this model shows a <b>much lower refusal rate</b> and can differ substantially from official LFM2.5-2.6B. Evaluate compliance and safety for your use case; control access and audit as needed.</p>
<table style="width: 100%; border-collapse: collapse; font-size: 13px;"><tbody>
<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Refusals (harmful eval)</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;"><b>6 / 100</b> (baseline ~90 / 100)</td></tr>
<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">KL divergence</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">0.0335 (same-prefix, far below 0.5 prune threshold)</td></tr>
<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Length deviation</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">0.079 σ</td></tr>
<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Generation health</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">PASSED</td></tr>
<tr><td style="padding: 6px 10px; font-weight: bold; color: #334155; background: white;">Selected trial</td><td style="padding: 6px 10px; color: #4b5563; background: white;">abliterix Trial 65</td></tr>
<tr><td style="padding: 6px 10px; font-weight: bold; color: #334155; background: white;">Thinking</td><td style="padding: 6px 10px; color: #4b5563; background: white;">Preserved — always-thinks (<code><think></code> in chat template)</td></tr>
</tbody></table>
<p style="margin: 12px 0 0 0; font-size: 12px; color: #9a3412;">Implementation sketch: LoRA merge <code>W += (B @ A) * (alpha / r)</code> (this trial alpha = r = 1); steering applied to <code>attn.o_proj</code> / <code>conv.out_proj</code> / <code>mlp.down_proj</code> across 30 layers.</p>
</div>
</div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #6366f1 0%, #8b5cf6 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>📦</span> Quantization tiers</div>
<div style="padding: 16px;">
<table style="width: 100%; border-collapse: collapse; font-size: 13px;"><thead><tr style="background: rgba(99,102,241,0.06);">
<th style="padding: 8px 10px; border-bottom: 2px solid #6366f1; text-align: left; color: #4f46e5;">File</th>
<th style="padding: 8px 10px; border-bottom: 2px solid #6366f1; text-align: left; color: #4f46e5;">Size</th>
<th style="padding: 8px 10px; border-bottom: 2px solid #6366f1; text-align: left; color: #4f46e5;">BPW</th>
<th style="padding: 8px 10px; border-bottom: 2px solid #6366f1; text-align: left; color: #4f46e5;">Decode (ROCm gfx1151)</th>
<th style="padding: 8px 10px; border-bottom: 2px solid #6366f1; text-align: left; color: #4f46e5;">Best for</th>
</tr></thead><tbody>
<tr><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;"><code>*-IQ3_XS.gguf</code></td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;">1.22 GB</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;">~3.30</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;">~135 t/s</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;">Maximum compression (perceptible quality loss on small models)</td></tr>
<tr><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: #eef2ff;"><code>*-IQ4_XS.gguf</code></td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: #eef2ff;"><b>1.52 GB</b></td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: #eef2ff;"><b>~4.25</b></td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: #eef2ff;"><b>~120 t/s</b></td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: #eef2ff;"><b>Sweet spot — smallest tier with Q4_K_M-class quality</b></td></tr>
<tr><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;"><code>*-Q4_K_M.gguf</code></td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;">1.67 GB</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;">~4.94</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;">~100 t/s</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;">Verified everyday default</td></tr>
<tr><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;"><code>*-Q6_K.gguf</code></td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;">2.22 GB</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;">~6.56</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;">~75 t/s</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;">Quality-first local use</td></tr>
<tr><td style="padding: 8px 10px; background: white;"><code>*-Q8_0.gguf</code></td><td style="padding: 8px 10px; background: white;">2.87 GB</td><td style="padding: 8px 10px; background: white;">~8.50</td><td style="padding: 8px 10px; background: white;">~60 t/s</td><td style="padding: 8px 10px; background: white;">Near-lossless (imatrix optional here)</td></tr>
<tr><td style="padding: 8px 10px; background: white;"><code>*-BF16.gguf</code></td><td style="padding: 8px 10px; background: white;">5.40 GB</td><td style="padding: 8px 10px; background: white;">16.00</td><td style="padding: 8px 10px; background: white;">~33 t/s</td><td style="padding: 8px 10px; background: white;">Lossless baseline (source of all tiers)</td></tr>
</tbody></table>
<p style="margin: 12px 0 0 0; font-size: 12px; color: #64748b;">Decode speeds measured on AMD Strix Halo (Radeon 8060S, gfx1151) with llama.cpp ROCm 7.2, 128K context. All files are <code>lfm2</code> architecture, 128K context, single-file GGUFs.</p>
</div>
</div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #6366f1 0%, #8b5cf6 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>💡</span> Why imatrix?</div>
<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">
<p style="margin: 0;">The importance matrix (computed over 401 chunks / ≈1.6M tokens of mixed conversation, math, and code data) tells the quantizer which weights are sensitive. K-quants and especially the <code>IQ</code> tiers use it to keep more bits on attention/embedding paths — the parts that matter most for subtle behaviors like identity and instruction following on a 2.6B model. Compared to plain Q4_K_M, IQ4_XS is <b>smaller and faster while holding comparable perplexity</b>.</p>
</div>
</div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #6366f1 0%, #8b5cf6 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🚀</span> Usage (llama.cpp)</div>
<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">
<p style="margin: 0 0 12px 0;">The <code>lfm2</code> architecture is supported by llama.cpp (and LM Studio / other GGUF runners).</p>
<pre style="background: #f8fafc; border: 1px solid #e2e8f0; border-radius: 8px; padding: 12px; font-size: 12px; overflow-x: auto; line-height: 1.6;"><code>llama-server -m LFM2.5-2.6B-Uncensored-IQ4_XS.gguf \
--ctx-size 131072 --flash-attn on --host 0.0.0.0 --port 8080</code></pre>
<p style="margin: 12px 0 0 0;">Or with llama-cli:</p>
<pre style="background: #f8fafc; border: 1px solid #e2e8f0; border-radius: 8px; padding: 12px; font-size: 12px; overflow-x: auto; line-height: 1.6;"><code>llama-cli -m LFM2.5-2.6B-Uncensored-Q4_K_M.gguf \
-p "What is 2+2?" -n 512 \
--temp 0.1 --top-k 50 --repeat-penalty 1.1</code></pre>
<p style="margin: 12px 0 0 0;">Transformers / vLLM / SGLang users: use the BF16 safetensors in the <a href="https://huggingface.co/SC117/LFM2.5-2.6B-Uncensored" target="_blank" style="color: #4f46e5; text-decoration: none; font-weight: 700;">parent repo</a>.</p>
</div>
</div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #6366f1 0%, #8b5cf6 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🎛️</span> Recommended sampling</div>
<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">
<p style="margin: 0;">Keep the official generation defaults: <code>temperature 0.1</code>, <code>top_k 50</code>, <code>repetition_penalty 1.1</code>. If you want more creative answers, raise temperature toward 0.6–0.8; note the model always thinks before answering, so allow enough <code>max_new_tokens</code> (512+) for the <code><think></code> block.</p>
</div>
</div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #6366f1 0%, #8b5cf6 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🔧</span> Build pipeline</div>
<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">
<ol style="margin: 0; padding-left: 20px;">
<li style="margin-bottom: 6px;">abliterix trial search on ROCm (gfx1151): 60 trials + 20 warmup, seed 117; all trials pruned only by same-prefix <code>kl_divergence < 0.5</code>.</li>
<li style="margin-bottom: 6px;">Selected <b>Trial 65</b>: refusals 6/100 (baseline 90/100), KL 0.0335, length deviation 0.079 σ, generation health PASSED.</li>
<li style="margin-bottom: 6px;">LoRA stream-merged into base weights in BF16 (<code>W += B@A</code>, alpha = r = 1).</li>
<li style="margin-bottom: 6px;">BF16 GGUF conversion via <code>convert_hf_to_gguf.py --outtype bf16</code> (llama.cpp <code>lfm2</code>).</li>
<li style="margin-bottom: 6px;">imatrix calibration: 401 chunks / ≈1.6M tokens (APEX calibration set), computed on the BF16 GGUF.</li>
<li style="margin-bottom: 6px;">Quantized with <code>llama-quantize --imatrix <imatrix.gguf> <src> <dst> <type></code> for each tier.</li>
</ol>
</div>
</div>
</div>
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; font-size: 12px; color: #64748b; text-align: center; margin-top: 20px;">
Community derivative (behavior edit + quantized GGUF release). <b>Not an official Liquid AI release.</b> Use at your own risk; follow local law and the LFM Open License v1.0.
</div>
Run SC117/LFM2.5-2.6B-Uncensored-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models