SC117/Laguna-S-2.1-Uncensored-APEX-GGUF overview
<div style="font family: apple system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans serif; margin bottom: 24px;" <div style="background: f5f5f7; border radius:…
Runs locally from ~2.08 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | SC117/Laguna-S-2.1-Uncensored-APEX-GGUF |
|---|---|
| Author | SC117 |
| Pipeline | text-generation |
| License | other |
| Base model | poolside/Laguna-S-2.1 |
| Last modified | 2026-07-26T22:46:33.000Z |
Model README
---
library_name: gguf
license: other
license_name: openmdw-1.1
license_link: https://openmdw.ai/
pipeline_tag: text-generation
tags:
- laguna
- moe
- uncensored
- abliterix
- apex
- quantization
- gguf
- poolside
base_model:
- poolside/Laguna-S-2.1
base_model_relation: quantized
---
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; margin-bottom: 24px;">
<div style="background: #f5f5f7; border-radius: 20px; padding: 36px 32px; margin-bottom: 20px; text-align: center; position: relative; overflow: hidden;">
<div style="position: absolute; top: -30px; right: -30px; width: 120px; height: 120px; background: #ddd6fe; border-radius: 50%;"></div>
<div style="position: absolute; bottom: -20px; left: 40px; width: 80px; height: 80px; background: #c4b5fd; border-radius: 50%;"></div>
<div style="position: absolute; top: 50%; left: -15px; width: 60px; height: 60px; background: #ddd6fe; border-radius: 50%;"></div>
<div style="display: inline-flex; gap: 8px; margin-bottom: 16px; position: relative; z-index: 1; flex-wrap: wrap; justify-content: center;">
<span style="background: #6366f1; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">ABLITERIX</span>
<span style="background: #8b5cf6; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">TRIAL 16</span>
<span style="background: #0ea5e9; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">APEX</span>
<span style="background: #64748b; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">OpenMDW-1.1</span>
</div>
<h1 style="margin: 0 0 8px 0; font-size: 28px; font-weight: 700; color: #1e1b4b; letter-spacing: -0.5px; border: none; position: relative; z-index: 1;">Laguna-S-2.1-Uncensored-APEX-GGUF</h1>
<p style="margin: 8px 0 0 0; font-size: 14px; position: relative; z-index: 1;"><span style="color: #6b7280;">English</span> | <a href="README_zh.md" style="color: #6366f1; text-decoration: none;">📖 中文文档</a></p>
<p style="margin: 8px 0 0 0; font-size: 15px; color: #6b7280; position: relative; z-index: 1;">Uncensored 118B-A8B MoE · abliterix Trial 16 · APEX I-tier GGUFs + imatrix</p>
</div>
</div>
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; display: flex; flex-direction: column; gap: 20px; margin-bottom: 30px;">
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #6366f1 0%, #8b5cf6 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🌊</span> About this release</div>
<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">
<p style="margin: 0 0 12px 0;"><a href="https://huggingface.co/poolside/Laguna-S-2.1" target="_blank" style="color: #4f46e5; text-decoration: none; font-weight: 700;">Laguna S 2.1</a> is a poolside <b>~118B total / ~8B active per token</b> Mixture-of-Experts model for agentic coding and long-horizon work. Architecture highlights: token-choice routing with softplus gates, <b>256 routed experts + 1 shared expert (top-10)</b>, GQA, <b>1:3 global/sliding-window attention</b> (48 layers: 12 global + 36 local, window 512), and up to about <b>1M context</b>, with optional native thinking interleaved with tool use.</p>
<p style="margin: 0 0 12px 0;">This release builds on the official weights in two steps:</p>
<ol style="margin: 0 0 12px 0; padding-left: 20px;">
<li style="margin-bottom: 6px;"><b>Uncensored behavior edit</b> via <a href="https://github.com/wuwangzhang1216/abliterix" target="_blank" style="color: #4f46e5; text-decoration: none; font-weight: 700;">abliterix</a> (ROCm + bitsandbytes 4-bit search path), selecting <b>Trial 16</b> LoRA, stream-merged back to BF16, plus MoE safety-expert router bake-in.</li>
<li style="margin-bottom: 6px;"><b>APEX mixed-precision GGUF</b> from a poolside <a href="https://github.com/poolsideai/llama.cpp/tree/laguna" target="_blank" style="color: #4f46e5; text-decoration: none; font-weight: 700;">llama.cpp (laguna)</a> BF16 GGUF, using <a href="https://github.com/mudler/apex-quant" target="_blank" style="color: #4f46e5; text-decoration: none; font-weight: 700;">APEX</a>-style tensor-type configs + <b>imatrix</b> (token embedding / output kept at <b>BF16</b>).</li>
</ol>
<p style="margin: 0;"><b>License: OpenMDW-1.1</b> (same family as the base model). See <a href="https://openmdw.ai/" target="_blank" style="color: #4f46e5; text-decoration: none; font-weight: 700;">openmdw.ai</a> and poolside terms.</p>
</div>
</div>
<div style="border: 1px solid #fdba74; border-radius: 12px; overflow: hidden; background: #fff7ed; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #f97316 0%, #ea580c 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>⚠️</span> Uncensored notice</div>
<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">
<p style="margin: 0 0 12px 0;">After merging abliterix <b>Trial 16</b>, this model shows a <b>much lower refusal rate</b> and can differ substantially from official Laguna-S-2.1. Evaluate compliance and safety for your use case; control access and audit as needed.</p>
<table style="width: 100%; border-collapse: collapse; font-size: 13px;"><tbody>
<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Refusals (harmful eval)</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;"><b>8 / 100</b> (baseline ~97 / 100)</td></tr>
<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">KL divergence</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">~0.0042</td></tr>
<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Length deviation</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">~0.07 σ</td></tr>
<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Generation health</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">PASSED</td></tr>
<tr><td style="padding: 6px 10px; font-weight: bold; color: #334155; background: white;">Selected trial</td><td style="padding: 6px 10px; color: #4b5563; background: white;">abliterix Trial 16</td></tr>
</tbody></table>
<p style="margin: 12px 0 0 0; font-size: 12px; color: #9a3412;">Implementation sketch: LoRA merge <code>W += (B @ A) * (alpha / r)</code> (this trial alpha = r = 1); per-layer top ~20 safety experts with router row scale ~0.74 (see merge metadata).</p>
</div>
</div>
<div style="border: 1px solid #c4b5fd; border-radius: 12px; overflow: hidden; background: #f5f3ff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #7c3aed 0%, #6d28d9 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>📜</span> Disclaimer (identity & text quality)</div>
<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">
<p style="margin: 0 0 12px 0;">The following can also appear with <b>official Laguna-S-2.1</b> (including high-precision online deployments). They are <b>not introduced by this Uncensored edit</b> and should not be blamed on abliterix / the refusal merge:</p>
<ul style="margin: 0 0 12px 0; padding-left: 20px;">
<li style="margin-bottom: 6px;"><b>Model identity</b> quirks or self-descriptions that diverge from the official persona;</li>
<li style="margin-bottom: 6px;">Odd tokens, typos, stiffness, or occasional awkward phrasing in Chinese and other languages.</li>
</ul>
<p style="margin: 0;">This Uncensored pipeline mainly changes <b>refusal / safety-related behavior</b>. More aggressive quant tiers (e.g. Compact) may <b>amplify</b> existing text noise, but the underlying issues already exist in the original model and/or the local quant+inference stack. For a fair check, compare official vs this release under the same serving settings.</p>
</div>
</div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #6366f1 0%, #8b5cf6 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🧠</span> Model details</div>
<div style="padding: 16px;">
<table style="width: 100%; border-collapse: collapse; font-size: 13px;"><tbody>
<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Architecture</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">Laguna MoE (poolside <code>laguna</code>)</td></tr>
<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Parameters</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">~118B total, ~8B active / token</td></tr>
<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Layers</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">48 (L0 dense FFN, L1–L47 MoE)</td></tr>
<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Experts</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">256 routed + 1 shared, top-10</td></tr>
<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Attention</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">GQA, 8 KV heads, head dim 128</td></tr>
<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Context</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">Up to ~1,048,576 tokens (practical limit depends on VRAM / <code>-c</code>)</td></tr>
<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Vocab</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">100,352 (Laguna family tokenizer)</td></tr>
<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">Modality</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white;">Text → text (no mmproj in this package)</td></tr>
<tr><td style="padding: 6px 10px; font-weight: bold; color: #334155; background: white;">This repo</td><td style="padding: 6px 10px; color: #4b5563; background: white;">APEX I-Quality / I-Balanced / I-Compact GGUF</td></tr>
</tbody></table>
<p style="margin: 12px 0 0 0; font-size: 12px; color: #64748b;">Coding / tool ability is largely retained; refusal and alignment behavior are changed. Full official bench tables were not re-run for this derivative.</p>
</div>
</div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #6366f1 0%, #8b5cf6 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>💡</span> What is APEX?</div>
<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">
<p style="margin: 0 0 12px 0;">These files use <a href="https://github.com/mudler/apex-quant" target="_blank" style="color: #4f46e5; text-decoration: none; font-weight: 700;">APEX</a>-style MoE-aware mixed precision: precision follows <b>tensor role + layer position</b> (higher on edges, more aggressive in the middle), with <b>imatrix</b> for the <code>I-</code> tiers.</p>
<p style="margin: 0;">Common settings for this package: source <code>Laguna-S-2.1-Uncensored</code> BF16 GGUF; calibration <code>laguna-s-2.1.imatrix</code>; <b>token embedding / output = BF16</b>; routers and norms largely left at high precision defaults. Configs target Laguna 48-layer naming (<code>ffn__exps</code> / <code>ffn__shexp</code> / <code>attn_*</code>, L0 dense).</p>
</div>
</div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #6366f1 0%, #8b5cf6 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>📦</span> APEX quantization tiers</div>
<div style="padding: 16px;">
<table style="width: 100%; border-collapse: collapse; font-size: 13px;"><thead><tr style="background: rgba(99,102,241,0.06);">
<th style="padding: 8px 10px; border-bottom: 2px solid #6366f1; text-align: left; color: #4f46e5;">File</th>
<th style="padding: 8px 10px; border-bottom: 2px solid #6366f1; text-align: left; color: #4f46e5;">Size</th>
<th style="padding: 8px 10px; border-bottom: 2px solid #6366f1; text-align: left; color: #4f46e5;">Mid experts</th>
<th style="padding: 8px 10px; border-bottom: 2px solid #6366f1; text-align: left; color: #4f46e5;">Best for</th>
</tr></thead><tbody>
<tr><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;"><code>*-I-Quality.gguf</code></td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;">~70 GB</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;">edge Q6_K / near Q5_K / mid <b>iq4_xs</b>; shared Q8_0; attn Q6_K</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: white;">Smaller high-quality try (IQ mid-layers)</td></tr>
<tr><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: #eef2ff;"><code>*-I-Balanced.gguf</code></td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: #eef2ff;"><b>~80 GB</b></td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: #eef2ff;">edge Q6_K / near & mid <b>Q5_K</b>; shared Q8_0; attn Q6_K</td><td style="padding: 8px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); background: #eef2ff;"><b>Recommended default for Chinese users</b> — steadier than Compact</td></tr>
<tr><td style="padding: 8px 10px; background: white;"><code>*-I-Compact.gguf</code></td><td style="padding: 8px 10px; background: white;">~52 GB</td><td style="padding: 8px 10px; background: white;">edge Q4_K / mid Q3_K; shared Q6_K; attn Q4_K</td><td style="padding: 8px 10px; background: white;">Tighter memory; more quant noise OK</td></tr>
</tbody></table>
<p style="margin: 12px 0 0 0; font-size: 13px; color: #334155; line-height: 1.7;"><b>How to choose:</b></p>
<ul style="margin: 8px 0 0 0; padding-left: 20px; font-size: 13px; color: #334155; line-height: 1.7;">
<li style="margin-bottom: 6px;"><b>Chinese-heavy use:</b> prefer <b>I-Balanced</b>. Mid experts stay Q5_K (not iq4_xs), which usually feels more stable. Drop to I-Compact only if memory is tight.</li>
<li style="margin-bottom: 6px;"><b>English / coding:</b> tier differences are usually small; pick by VRAM/speed (Balanced vs Compact). I-Quality is optional when you want IQ mid-layers and a smaller footprint than Balanced.</li>
</ul>
</div>
</div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #6366f1 0%, #8b5cf6 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🚀</span> Usage (llama.cpp)</div>
<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">
<p style="margin: 0 0 8px 0;">Use a build that understands the <b>laguna</b> architecture (poolside <code>laguna</code> branch or equivalent).</p>
<p style="margin: 0 0 8px 0; font-weight: bold; color: #312e81;">Example (I-Balanced recommended)</p>
<pre style="margin: 0 0 12px 0; font-family: monospace; background: #f8fafc; padding: 10px 14px; border-radius: 6px; border: 1px solid #e2e8f0; font-size: 12px; color: #1e293b; white-space: pre-wrap;">./llama-server \
-m ./Laguna-S-2.1-Uncensored-APEX-I-Balanced.gguf \
--port 8080 \
-sm none \
--device rocm0 \
--ctx-size 131072 \
--flash-attn on \
--no-mmap \
--fit on \
--jinja \
--host 0.0.0.0</pre>
<ul style="margin: 0; padding-left: 20px;">
<li style="margin-bottom: 6px;">Must recognize <code>general.architecture = laguna</code>.</li>
<li style="margin-bottom: 6px;">Thinking / tools: configure per client/server docs (e.g. <code>enable_thinking</code>, reasoning options).</li>
<li style="margin-bottom: 6px;">Warnings like <code>special_eos_id is not in special_eog_ids</code> are tokenizer metadata hints; the model can still load. If stop/truncation is odd, check stop / max_tokens / template.</li>
<li style="margin-bottom: 0;">Text-only GGUF; no mmproj.</li>
</ul>
</div>
</div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #6366f1 0%, #8b5cf6 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🎛️</span> Recommended sampling</div>
<div style="padding: 16px;">
<table style="width: 100%; border-collapse: collapse; font-size: 13px;"><tbody>
<tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155; background: white;">General / coding</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563; background: white; font-family: monospace;">temperature 0.6–1.0, top_p 0.95, top_k 20</td></tr>
<tr><td style="padding: 6px 10px; font-weight: bold; color: #334155; background: white;">More stable</td><td style="padding: 6px 10px; color: #4b5563; background: white;">Slightly lower temperature</td></tr>
</tbody></table>
</div>
</div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="background: linear-gradient(135deg, #6366f1 0%, #8b5cf6 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🔧</span> Build pipeline (summary)</div>
<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">
<ol style="margin: 0; padding-left: 20px;">
<li style="margin-bottom: 6px;">Base: <code>poolside/Laguna-S-2.1</code></li>
<li style="margin-bottom: 6px;">abliterix search → <b>Trial 16</b> LoRA + MoE router adjustments</li>
<li style="margin-bottom: 6px;">BF16 stream-merge → <code>Laguna-S-2.1-Uncensored</code></li>
<li style="margin-bottom: 6px;">poolside llama.cpp → BF16 GGUF + imatrix</li>
<li style="margin-bottom: 0;">APEX tensor-type-file + imatrix → I-Quality / I-Balanced / I-Compact</li>
</ol>
</div>
</div>
</div>
Links
- Original model: https://huggingface.co/poolside/Laguna-S-2.1
- Announcement: https://poolside.ai/blog/introducing-laguna-s-2-1
- OpenRouter: https://openrouter.ai/poolside/laguna-s-2.1
- Official GGUF: https://huggingface.co/poolside/Laguna-S-2.1-GGUF
- poolside llama.cpp (laguna): https://github.com/poolsideai/llama.cpp/tree/laguna
- abliterix: https://github.com/wuwangzhang1216/abliterix
- APEX: https://github.com/mudler/apex-quant
- License: https://openmdw.ai/
Disclaimer
Community derivative (behavior edit + quantization). Not an official poolside release. Use at your own risk; follow local law and upstream licenses.
Run SC117/Laguna-S-2.1-Uncensored-APEX-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models