GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

SC117/LFM2.5-2.6B-Uncensored-GGUF overview

<div style="font family: apple system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans serif; margin bottom: 24px;" <div style="background: f5f5f7; border radius:…

ggufliquidlfm2.5edgeuncensoredabliterixquantizationimatrixtext-generationbase_model:LiquidAI/LFM2.5-2.6Bbase_model:quantized:LiquidAI/LFM2.5-2.6Blicense:otherendpoints_compatibleregion:usconversational

Runs locally from ~1.14 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
2
Pipeline
text-generation
Author

Repository Files & Downloads

6 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
LFM2.5-2.6B-Uncensored-BF16.ggufGGUFBF165.03 GBDownload
LFM2.5-2.6B-Uncensored-IQ3_XS.ggufGGUFIQ3_XS1.14 GBDownload
LFM2.5-2.6B-Uncensored-IQ4_XS.ggufGGUFIQ4_XS1.41 GBDownload
LFM2.5-2.6B-Uncensored-Q4_K_M.ggufGGUFQ4_K_M1.56 GBDownload
LFM2.5-2.6B-Uncensored-Q6_K.ggufGGUFQ6_K2.07 GBDownload
LFM2.5-2.6B-Uncensored-Q8_0.ggufGGUFQ8_02.68 GBDownload

Model Details

Model IDSC117/LFM2.5-2.6B-Uncensored-GGUF
AuthorSC117
Pipelinetext-generation
Licenseother
Base modelLiquidAI/LFM2.5-2.6B,SC117/LFM2.5-2.6B-Uncensored
Last modified2026-08-05T08:31:02.000Z

Model README

---

library_name: gguf

license: other

license_name: lfm1.0

license_link: LICENSE

pipeline_tag: text-generation

tags:

  • liquid
  • lfm2.5
  • edge
  • uncensored
  • abliterix
  • quantization
  • gguf
  • imatrix

base_model:

  • LiquidAI/LFM2.5-2.6B
  • SC117/LFM2.5-2.6B-Uncensored

base_model_relation: quantized

---

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; margin-bottom: 24px;">

<div style="background: #f5f5f7; border-radius: 20px; padding: 36px 32px; margin-bottom: 20px; text-align: center; position: relative; overflow: hidden;">

<div style="position: absolute; top: -30px; right: -30px; width: 120px; height: 120px; background: #ddd6fe; border-radius: 50%;"></div>

<div style="position: absolute; bottom: -20px; left: 40px; width: 80px; height: 80px; background: #c4b5fd; border-radius: 50%;"></div>

<div style="position: absolute; top: 50%; left: -15px; width: 60px; height: 60px; background: #ddd6fe; border-radius: 50%;"></div>

<div style="display: inline-flex; gap: 8px; margin-bottom: 16px; position: relative; z-index: 1; flex-wrap: wrap; justify-content: center;">

<span style="background: #6366f1; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">ABLITERIX</span>

<span style="background: #8b5cf6; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">TRIAL 65</span>

<span style="background: #0ea5e9; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">GGUF + IMATRIX</span>

<span style="background: #64748b; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">LFM Open 1.0</span>

</div>

<h1 style="margin: 0 0 8px 0; font-size: 28px; font-weight: 700; color: #1e1b4b; letter-spacing: -0.5px; border: none; position: relative; z-index: 1;">LFM2.5-2.6B-Uncensored-GGUF</h1>

<p style="margin: 8px 0 0 0; font-size: 14px; position: relative; z-index: 1;"><span style="color: #6b7280;">English</span> | <a href="https://huggingface.co/SC117/LFM2.5-2.6B-Uncensored-GGUF/blob/main/README_zh.md" style="color: #6366f1; text-decoration: none;">📖 中文文档</a></p>

<p style="margin: 8px 0 0 0; font-size: 15px; color: #6b7280; position: relative; z-index: 1;">Uncensored 2.6B edge model · abliterix Trial 65 · imatrix-calibrated GGUFs</p>

</div>

</div>

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; display: flex; flex-direction: column; gap: 20px; margin-bottom: 30px;">

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">

<div style="background: linear-gradient(135deg, #6366f1 0%, #8b5cf6 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🌊</span> About this release</div>

<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">

<p style="margin: 0 0 12px 0;"><a href="https://huggingface.co/LiquidAI/LFM2.5-2.6B" target="_blank" style="color: #4f46e5; text-decoration: none; font-weight: 700;">LFM2.5-2.6B</a> is a Liquid AI <b>2.6B-parameter hybrid edge model</b> built for agentic workloads: 30 layers (22 double-gated short-convolution blocks + 8 GQA), a <b>128K context window</b>, 128K vocabulary, and a ChatML-like template with native <code>&lt;think&gt;</code> reasoning.</p>

<p style="margin: 0 0 12px 0;">These quantized GGUFs are built from our <a href="https://huggingface.co/SC117/LFM2.5-2.6B-Uncensored" target="_blank" style="color: #4f46e5; text-decoration: none; font-weight: 700;">LFM2.5-2.6B-Uncensored</a> BF16 release (abliterix <b>Trial 65</b>, stream-merged to BF16) in three steps:</p>

<ol style="margin: 0 0 12px 0; padding-left: 20px;">

<li style="margin-bottom: 6px;"><b>BF16 GGUF conversion</b> with llama.cpp (<code>lfm2</code> architecture support).</li>

<li style="margin-bottom: 6px;"><b>imatrix calibration</b> — 401 chunks (≈1.6M tokens) from the APEX calibration set, computed on the BF16 GGUF.</li>

<li style="margin-bottom: 6px;"><b>Quantization</b> with <code>llama-quantize --imatrix</code> into five tiers.</li>

</ol>

<p style="margin: 0;"><b>License: LFM Open License v1.0</b> (same as the base model).</p>

</div>

</div>

<div style="border: 1px solid #fdba74; border-radius: 12px; overflow: hidden; background: #fff7ed; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">

<div style="background: linear-gradient(135deg, #f97316 0%, #ea580c 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>⚠️</span> Uncensored notice</div>

<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">

<p style="margin: 0 0 12px 0;">After merging abliterix <b>Trial 65</b>, this model shows a <b>much lower refusal rate</b> and can differ substantially from official LFM2.5-2.6B. Evaluate compliance and safety for your use case; control access and audit as needed.</p>

<table style="width: 100%; border-collapse: collapse; font-size: 13px;"><tbody>

<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Refusals (harmful eval)</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;"><b>6 / 100</b> (baseline ~90 / 100)</td></tr>

<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">KL divergence</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">0.0335 (same-prefix, far below 0.5 prune threshold)</td></tr>

<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Length deviation</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">0.079 σ</td></tr>

<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Generation health</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">PASSED</td></tr>

<tr><td style="padding: 6px 10px; font-weight: bold; color: #334155; background: white;">Selected trial</td><td style="padding: 6px 10px; color: #4b5563; background: white;">abliterix Trial 65</td></tr>

<tr><td style="padding: 6px 10px; font-weight: bold; color: #334155; background: white;">Thinking</td><td style="padding: 6px 10px; color: #4b5563; background: white;">Preserved — always-thinks (<code>&lt;think&gt;</code> in chat template)</td></tr>

</tbody></table>

<p style="margin: 12px 0 0 0; font-size: 12px; color: #9a3412;">Implementation sketch: LoRA merge <code>W += (B @ A) * (alpha / r)</code> (this trial alpha = r = 1); steering applied to <code>attn.o_proj</code> / <code>conv.out_proj</code> / <code>mlp.down_proj</code> across 30 layers.</p>

</div>

</div>

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">

<div style="background: linear-gradient(135deg, #6366f1 0%, #8b5cf6 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>📦</span> Quantization tiers</div>

<div style="padding: 16px;">

<table style="width: 100%; border-collapse: collapse; font-size: 13px;"><thead><tr style="background: rgba(99,102,241,0.06);">

<th style="padding: 8px 10px; border-bottom: 2px solid #6366f1; text-align: left; color: #4f46e5;">File</th>

<th style="padding: 8px 10px; border-bottom: 2px solid #6366f1; text-align: left; color: #4f46e5;">Size</th>

<th style="padding: 8px 10px; border-bottom: 2px solid #6366f1; text-align: left; color: #4f46e5;">BPW</th>

<th style="padding: 8px 10px; border-bottom: 2px solid #6366f1; text-align: left; color: #4f46e5;">Decode (ROCm gfx1151)</th>

<th style="padding: 8px 10px; border-bottom: 2px solid #6366f1; text-align: left; color: #4f46e5;">Best for</th>

</tr></thead><tbody>

<tr><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;"><code>*-IQ3_XS.gguf</code></td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;">1.22 GB</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;">~3.30</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;">~135 t/s</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;">Maximum compression (perceptible quality loss on small models)</td></tr>

<tr><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: #eef2ff;"><code>*-IQ4_XS.gguf</code></td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: #eef2ff;"><b>1.52 GB</b></td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: #eef2ff;"><b>~4.25</b></td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: #eef2ff;"><b>~120 t/s</b></td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: #eef2ff;"><b>Sweet spot — smallest tier with Q4_K_M-class quality</b></td></tr>

<tr><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;"><code>*-Q4_K_M.gguf</code></td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;">1.67 GB</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;">~4.94</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;">~100 t/s</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;">Verified everyday default</td></tr>

<tr><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;"><code>*-Q6_K.gguf</code></td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;">2.22 GB</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;">~6.56</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;">~75 t/s</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;">Quality-first local use</td></tr>

<tr><td style="padding: 8px 10px; background: white;"><code>*-Q8_0.gguf</code></td><td style="padding: 8px 10px; background: white;">2.87 GB</td><td style="padding: 8px 10px; background: white;">~8.50</td><td style="padding: 8px 10px; background: white;">~60 t/s</td><td style="padding: 8px 10px; background: white;">Near-lossless (imatrix optional here)</td></tr>

<tr><td style="padding: 8px 10px; background: white;"><code>*-BF16.gguf</code></td><td style="padding: 8px 10px; background: white;">5.40 GB</td><td style="padding: 8px 10px; background: white;">16.00</td><td style="padding: 8px 10px; background: white;">~33 t/s</td><td style="padding: 8px 10px; background: white;">Lossless baseline (source of all tiers)</td></tr>

</tbody></table>

<p style="margin: 12px 0 0 0; font-size: 12px; color: #64748b;">Decode speeds measured on AMD Strix Halo (Radeon 8060S, gfx1151) with llama.cpp ROCm 7.2, 128K context. All files are <code>lfm2</code> architecture, 128K context, single-file GGUFs.</p>

</div>

</div>

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">

<div style="background: linear-gradient(135deg, #6366f1 0%, #8b5cf6 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>💡</span> Why imatrix?</div>

<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">

<p style="margin: 0;">The importance matrix (computed over 401 chunks / ≈1.6M tokens of mixed conversation, math, and code data) tells the quantizer which weights are sensitive. K-quants and especially the <code>IQ</code> tiers use it to keep more bits on attention/embedding paths — the parts that matter most for subtle behaviors like identity and instruction following on a 2.6B model. Compared to plain Q4_K_M, IQ4_XS is <b>smaller and faster while holding comparable perplexity</b>.</p>

</div>

</div>

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">

<div style="background: linear-gradient(135deg, #6366f1 0%, #8b5cf6 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🚀</span> Usage (llama.cpp)</div>

<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">

<p style="margin: 0 0 12px 0;">The <code>lfm2</code> architecture is supported by llama.cpp (and LM Studio / other GGUF runners).</p>

<pre style="background: #f8fafc; border: 1px solid #e2e8f0; border-radius: 8px; padding: 12px; font-size: 12px; overflow-x: auto; line-height: 1.6;"><code>llama-server -m LFM2.5-2.6B-Uncensored-IQ4_XS.gguf \

--ctx-size 131072 --flash-attn on --host 0.0.0.0 --port 8080</code></pre>

<p style="margin: 12px 0 0 0;">Or with llama-cli:</p>

<pre style="background: #f8fafc; border: 1px solid #e2e8f0; border-radius: 8px; padding: 12px; font-size: 12px; overflow-x: auto; line-height: 1.6;"><code>llama-cli -m LFM2.5-2.6B-Uncensored-Q4_K_M.gguf \

-p "What is 2+2?" -n 512 \

--temp 0.1 --top-k 50 --repeat-penalty 1.1</code></pre>

<p style="margin: 12px 0 0 0;">Transformers / vLLM / SGLang users: use the BF16 safetensors in the <a href="https://huggingface.co/SC117/LFM2.5-2.6B-Uncensored" target="_blank" style="color: #4f46e5; text-decoration: none; font-weight: 700;">parent repo</a>.</p>

</div>

</div>

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">

<div style="background: linear-gradient(135deg, #6366f1 0%, #8b5cf6 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🎛️</span> Recommended sampling</div>

<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">

<p style="margin: 0;">Keep the official generation defaults: <code>temperature 0.1</code>, <code>top_k 50</code>, <code>repetition_penalty 1.1</code>. If you want more creative answers, raise temperature toward 0.6–0.8; note the model always thinks before answering, so allow enough <code>max_new_tokens</code> (512+) for the <code>&lt;think&gt;</code> block.</p>

</div>

</div>

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">

<div style="background: linear-gradient(135deg, #6366f1 0%, #8b5cf6 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🔧</span> Build pipeline</div>

<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">

<ol style="margin: 0; padding-left: 20px;">

<li style="margin-bottom: 6px;">abliterix trial search on ROCm (gfx1151): 60 trials + 20 warmup, seed 117; all trials pruned only by same-prefix <code>kl_divergence &lt; 0.5</code>.</li>

<li style="margin-bottom: 6px;">Selected <b>Trial 65</b>: refusals 6/100 (baseline 90/100), KL 0.0335, length deviation 0.079 σ, generation health PASSED.</li>

<li style="margin-bottom: 6px;">LoRA stream-merged into base weights in BF16 (<code>W += B@A</code>, alpha = r = 1).</li>

<li style="margin-bottom: 6px;">BF16 GGUF conversion via <code>convert_hf_to_gguf.py --outtype bf16</code> (llama.cpp <code>lfm2</code>).</li>

<li style="margin-bottom: 6px;">imatrix calibration: 401 chunks / ≈1.6M tokens (APEX calibration set), computed on the BF16 GGUF.</li>

<li style="margin-bottom: 6px;">Quantized with <code>llama-quantize --imatrix &lt;imatrix.gguf&gt; &lt;src&gt; &lt;dst&gt; &lt;type&gt;</code> for each tier.</li>

</ol>

</div>

</div>

</div>

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; font-size: 12px; color: #64748b; text-align: center; margin-top: 20px;">

Community derivative (behavior edit + quantized GGUF release). <b>Not an official Liquid AI release.</b> Use at your own risk; follow local law and the LFM Open License v1.0.

</div>

Run SC117/LFM2.5-2.6B-Uncensored-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models