SC117/Qwen3.8-27B-Uncensored-FIT-GGUF overview
<div style="font family: apple system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans serif; margin bottom: 24px;" <div style="background: 0f172a; border radius:…
Runs locally from ~888.0 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Qwen3.8-27B-Uncensored-FIT-10.5G-IQ3_XXS.gguf | GGUF | IQ3_XXS | 10.50 GB | Download |
| Qwen3.8-27B-Uncensored-FIT-10G-Q2_K.gguf | GGUF | Q2_K | 10.00 GB | Download |
| Qwen3.8-27B-Uncensored-FIT-11.5G-IQ3_S.gguf | GGUF | IQ3_S | 11.43 GB | Download |
| Qwen3.8-27B-Uncensored-FIT-11G-IQ3_S.gguf | GGUF | IQ3_S | 10.99 GB | Download |
| Qwen3.8-27B-Uncensored-FIT-12.5G-IQ3_S.gguf | GGUF | IQ3_S | 12.50 GB | Download |
| Qwen3.8-27B-Uncensored-FIT-12G-IQ3_S.gguf | GGUF | IQ3_S | 12.00 GB | Download |
| Qwen3.8-27B-Uncensored-FIT-13.5G-IQ4_XS.gguf | GGUF | IQ4_XS | 13.50 GB | Download |
| Qwen3.8-27B-Uncensored-FIT-13G-IQ4_XS.gguf | GGUF | IQ4_XS | 13.00 GB | Download |
| Qwen3.8-27B-Uncensored-FIT-7.5G-IQ2_XXS.gguf | GGUF | IQ2_XXS | 7.50 GB | Download |
| Qwen3.8-27B-Uncensored-FIT-7G-IQ1_M.gguf | GGUF | IQ1_M | 7.00 GB | Download |
| Qwen3.8-27B-Uncensored-FIT-8.5G-IQ2_XXS.gguf | GGUF | IQ2_XXS | 8.50 GB | Download |
| Qwen3.8-27B-Uncensored-FIT-8G-IQ2_XXS.gguf | GGUF | IQ2_XXS | 8.00 GB | Download |
| Qwen3.8-27B-Uncensored-FIT-9.5G-Q2_K.gguf | GGUF | Q2_K | 9.50 GB | Download |
| Qwen3.8-27B-Uncensored-FIT-9G-Q2_K.gguf | GGUF | Q2_K | 9.00 GB | Download |
| mmproj-Qwen3.8-27B-Uncensored-BF16.gguf | GGUF | BF16 | 888.0 MB | Download |
Model Details
| Model ID | SC117/Qwen3.8-27B-Uncensored-FIT-GGUF |
|---|---|
| Author | SC117 |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | orcarouter/Qwen3.8-27B-Uncensored |
| Last modified | 2026-08-30T05:34:48.000Z |
Model README
---
base_model: orcarouter/Qwen3.8-27B-Uncensored
base_model_relation: quantized
license: apache-2.0
library_name: gguf
pipeline_tag: text-generation
language:
- en
- zh
tags:
- gguf
- llama-cpp
- quantization
- qwen3.8
- fit-gguf
- uncensored
---
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; margin-bottom: 24px;">
<div style="background: #0f172a; border-radius: 20px; padding: 40px 32px 30px 32px; text-align: center; position: relative; overflow: hidden;">
<div style="position: absolute; top: -40px; right: -40px; width: 160px; height: 160px; background: rgba(139,92,246,0.18); border-radius: 50%;"></div>
<div style="position: absolute; bottom: -50px; left: -30px; width: 140px; height: 140px; background: rgba(59,130,246,0.15); border-radius: 50%;"></div>
<div style="position: absolute; top: 30%; right: 12%; width: 56px; height: 56px; background: rgba(59,130,246,0.22); border-radius: 50%;"></div>
<div style="display: inline-flex; flex-wrap: wrap; justify-content: center; gap: 8px; margin-bottom: 18px; position: relative; z-index: 1;"><span style="background: #3b82f6; color: white; font-size: 11px; font-weight: 700; padding: 5px 14px; border-radius: 20px;">FIT-GGUF v0.1</span><span style="background: #8b5cf6; color: white; font-size: 11px; font-weight: 700; padding: 5px 14px; border-radius: 20px;">ZERO-BYTE VERIFIED</span><span style="background: #10b981; color: white; font-size: 11px; font-weight: 700; padding: 5px 14px; border-radius: 20px;">14 TIERS · 7–13.5 GiB</span><span style="background: #ef4444; color: white; font-size: 11px; font-weight: 700; padding: 5px 14px; border-radius: 20px;">MTP HEAD REMOVED</span><span style="background: #f59e0b; color: #0f172a; font-size: 11px; font-weight: 700; padding: 5px 14px; border-radius: 20px;">MEASURED KL + SAME-TOP</span><span style="background: #1e293b; color: #94a3b8; font-size: 11px; font-weight: 700; padding: 5px 14px; border-radius: 20px;">APACHE-2.0</span></div>
<h1 style="margin: 0 0 10px 0; font-size: 34px; font-weight: 800; color: #f8fafc; letter-spacing: -0.5px; border: none; position: relative; z-index: 1;">Qwen3.8-27B-Uncensored · FIT-GGUF</h1>
<p style="margin: 0 0 20px 0; font-size: 15px; color: #cbd5e1; position: relative; z-index: 1;">Fourteen continuous-size GGUF quantizations — every file exactly its predicted size, every tier measured.</p>
<div style="position: relative; z-index: 1; margin: 0 auto; width: 72%; max-width: 540px;">
<div style="height: 8px; border-radius: 8px; background: #334155; position: relative; overflow: visible;">
<div style="height: 8px; width: 71%; border-radius: 8px; background: linear-gradient(90deg, #3b82f6, #8b5cf6);"></div>
<div style="position: absolute; top: 50%; left: 71%; transform: translate(-50%, -50%); width: 18px; height: 18px; border-radius: 50%; background: #ffffff; box-shadow: 0 0 0 4px rgba(139,92,246,0.35);"></div>
</div>
<div style="display: flex; justify-content: space-between; margin-top: 8px; font-size: 10px; color: #94a3b8; font-weight: 600;"><span>7 GiB</span><span style="color: #c4b5fd;">ask for any budget in between</span><span>13.5 GiB</span></div>
</div>
<p style="margin: 18px 0 0 0; font-size: 13px; position: relative; z-index: 1;"><span style="color: #94a3b8;">English</span><span style="color: #475569;"> · </span><a href="https://huggingface.co/SC117/Qwen3.8-27B-Uncensored-FIT-GGUF/blob/main/README.zh-CN.md" style="color: #93c5fd; text-decoration: none; font-weight: 600;">简体中文 📖</a></p>
</div>
</div>
<p align="center"><img src="https://huggingface.co/SC117/Qwen3.8-27B-Uncensored-FIT-GGUF/resolve/main/assets/fit-gguf-banner.png" alt="FIT-GGUF" width="760"></p>
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; display: flex; flex-direction: column; gap: 18px; margin-bottom: 30px;">
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #2563eb 0%, #7c3aed 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🧭</span> About FIT-GGUF — the tool behind these files</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 10px 0;">Every file in this repository was planned, executed and verified by <a href="https://github.com/Scorp1o117/FIT-GGUF" style="color: #1d4ed8; text-decoration: none; font-weight: 700;">FIT-GGUF</a>, an open-source, deterministic tensor-level planning layer on top of llama.cpp quantization presets. Standard GGUF quantization asks you to pick one of a handful of presets; FIT-GGUF starts from the largest supported preset below your requested byte budget, then spends the remaining bytes on deterministic tensor-level precision upgrades.</p><p style="margin: 0 0 12px 0; padding: 10px 14px; background: #f5f3ff; border-left: 4px solid #7c3aed; border-radius: 6px; color: #4c1d95; font-weight: 600;">Traditional GGUF gives you presets. FIT gives you a size slider.</p><table style="width: 100%; border-collapse: collapse; font-size: 13px;"><tbody><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155;">Deterministic size prediction & recipe execution</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #166534; font-weight: 700;">✅ Validated</td></tr><tr><td style="padding: 6px 10px; font-weight: bold; color: #334155;">Universally optimal tensor allocation</td><td style="padding: 6px 10px; color: #92400e; font-weight: 700;">⚠️ Not established — FIT claims precise size control, not a universal quality optimum</td></tr></tbody></table><p style="margin: 12px 0 0 0;">The method, the full preregistered research record and the <code>fit</code> CLI are open source: <a href="https://github.com/Scorp1o117/FIT-GGUF" style="color: #1d4ed8; text-decoration: none; font-weight: 700;">github.com/Scorp1o117/FIT-GGUF</a></p></div></div>
<div style="border: 1px solid #fca5a5; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #dc2626 0%, #ea580c 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>⚠️</span> Safety notice / 安全提示</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 8px 0;">The source model is an <b>abliterated, refusal-removed</b> model with no meaningful built-in guardrails, and may comply with harmful, illegal or unsafe requests. Use it only where you can provide appropriate moderation, access control and legal review. Do not deploy it to end users without your own safety layer.</p><p style="margin: 0; color: #64748b;">源模型经过拒答方向移除,不具备可靠的内置安全护栏;请仅在合法、受控、具备审核与滥用防护的环境中使用,使用者自行承担部署责任。</p></div></div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #2563eb 0%, #7c3aed 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>📦</span> Pick a tier</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 10px 0; padding: 9px 13px; background: #fffbeb; border-left: 4px solid #f59e0b; border-radius: 6px; color: #92400e;"><b>FIT-12G means a 12 GiB budget for the main GGUF file</b> — not total RAM/VRAM usage. KV cache, compute buffers, runtime overhead and the multimodal projector are separate. <code>G</code> = GiB (2³⁰ bytes). Naming: <code>Qwen3.8-27B-Uncensored-FIT-<tier>-<dominant>.gguf</code></p><p style="margin: 0 0 10px 0; padding: 9px 13px; background: #fffbeb; border-left: 4px solid #f59e0b; border-radius: 6px; color: #92400e;"><b>MTP removed:</b> every quantization in this repository ships <b>without the NextN/MTP head</b> (the source was converted with <code>--no-nextn</code>). MTP-based speculative decoding is therefore not available with these files; text and vision inference are unaffected.</p><table style="width: 100%; border-collapse: collapse; font-size: 12.5px;"><thead><tr style="background: rgba(37,99,235,0.05);"><th style="padding: 7px 10px; border-bottom: 2px solid #2563eb; text-align: left; color: #1d4ed8;">Tier</th><th style="padding: 7px 10px; border-bottom: 2px solid #2563eb; text-align: left; color: #1d4ed8;">GiB</th><th style="padding: 7px 10px; border-bottom: 2px solid #2563eb; text-align: left; color: #1d4ed8;">Dominant</th><th style="padding: 7px 10px; border-bottom: 2px solid #2563eb; text-align: left; color: #1d4ed8;">Macro KL ↓</th><th style="padding: 7px 10px; border-bottom: 2px solid #2563eb; text-align: left; color: #1d4ed8;">Same-top ↑</th><th style="padding: 7px 10px; border-bottom: 2px solid #2563eb; text-align: left; color: #1d4ed8;">Measured positioning</th></tr></thead><tbody><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: 600;">FIT-7G</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">7.000</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">IQ1_M</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">1.1327</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">63.9%</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #64748b;">Extreme compression; large measured quality loss</td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: 600;">FIT-7.5G</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">7.500</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">IQ2_XXS</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">0.5898</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">73.1%</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #64748b;">First major quality step above the IQ1 region</td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: 600;">FIT-8G</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">7.999</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">IQ2_XXS</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">0.4838</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">76.9%</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #64748b;">Beats the IQ2_XXS preset (0.5403) with +0.15 GiB</td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: 600;">FIT-8.5G</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">8.499</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">IQ2_XXS</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">0.4527</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">78.4%</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #64748b;">Best measured point in the 8–9.3 GiB native region</td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: 600;">FIT-9G</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">8.999</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">Q2_K</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">0.3363</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">80.8%</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #64748b;">q2_k-directed fill; large KL step over FIT-8.5G</td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: 600;">FIT-9.5G</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">9.500</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">Q2_K</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">0.2737</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">83.1%</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #64748b;">Beats the Q2_K_S preset (0.2889) at slightly less size</td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: 600;">FIT-10G</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">9.999</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">Q2_K</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">0.2299</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">84.6%</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #64748b;">Beats the Q2_K preset (0.2439); compact general tier</td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: 600;">FIT-10.5G</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">10.499</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">IQ3_XXS</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">0.1873</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">87.4%</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #64748b;">Clear fidelity step over the 10G tier</td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: 600;">FIT-11G</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">10.988</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">IQ3_S</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">0.1515</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">88.9%</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #64748b;">Near IQ3_XS macro quality, slightly smaller</td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: 600;">FIT-11.5G</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">11.434</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">IQ3_S</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">0.1439</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">89.2%</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #64748b;">Near IQ3_S; documented 67.5 MiB target slack</td></tr><tr style="background: rgba(37,99,235,0.07);"><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: 800; color: #1d4ed8;">⭐ FIT-12G</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: 700;">12.000</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: 700;">IQ3_S</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: 700;">0.1227</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: 700;">90.3%</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #1e40af; font-weight: 600;">Strongest measured quality/size point in this release</td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: 600;">FIT-12.5G</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">12.497</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">IQ3_S</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">0.1116</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">91.0%</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #64748b;">0.5 GiB over FIT-12G buys a clear KL step</td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: 600;">FIT-13G</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">12.998</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">IQ4_XS</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">0.0987</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">91.9%</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #64748b;">IQ4_XS becomes the dominant payload</td></tr><tr><td style="padding: 6px 10px; font-weight: 600;">FIT-13.5G</td><td style="padding: 6px 10px;">13.498</td><td style="padding: 6px 10px;">IQ4_XS</td><td style="padding: 6px 10px; font-weight: 700; color: #166534;">0.0838</td><td style="padding: 6px 10px;">92.7%</td><td style="padding: 6px 10px; color: #64748b;">Best measured FIT-tier macro KL in this batch</td></tr></tbody></table><p style="margin: 14px 0 6px 0; font-weight: bold; color: #1e293b;">Quick picks</p><div style="display: flex; flex-wrap: wrap; gap: 10px;"><div style="flex: 1 1 46%; border: 1px solid #e2e8f0; border-radius: 10px; padding: 10px 12px; background: #f8fafc;"><div style="font-weight: 700; color: #1d4ed8;">🏆 FIT-12G — the sweet spot</div><div style="font-size: 12px; color: #475569; margin-top: 3px;">Below the IQ3_S / IQ3_M presets (0.1227 vs 0.1424 / 0.1445), far cheaper than IQ4_XS.</div></div><div style="flex: 1 1 46%; border: 1px solid #e2e8f0; border-radius: 10px; padding: 10px 12px; background: #f8fafc;"><div style="font-weight: 700; color: #166534;">🎯 FIT-13.5G — max quality</div><div style="font-size: 12px; color: #475569; margin-top: 3px;">Highest measured quality in this batch: KL 0.0838, Same-top 92.7%.</div></div><div style="flex: 1 1 46%; border: 1px solid #e2e8f0; border-radius: 10px; padding: 10px 12px; background: #f8fafc;"><div style="font-weight: 700; color: #b45309;">💸 FIT-8.5G — budget winner</div><div style="font-size: 12px; color: #475569; margin-top: 3px;">The 8–10 GiB region winner; beats the IQ2_XXS preset outright.</div></div><div style="flex: 1 1 46%; border: 1px solid #fecaca; border-radius: 10px; padding: 10px 12px; background: #fef2f2;"><div style="font-weight: 700; color: #b91c1c;">⚠️ Below ~7.5 GiB</div><div style="font-size: 12px; color: #7f1d1d; margin-top: 3px;">Quality drops sharply — the IQ1 region is rough and reported as measured.</div></div></div><p style="margin: 12px 0 0 0; font-size: 12px; color: #64748b; font-style: italic;">All numbers are protocol-scoped observations (five fixed 64 KiB domains vs aligned BF16), not an application benchmark or a universal ranking.</p></div></div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #2563eb 0%, #7c3aed 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>📈</span> Measured quality</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><img src="https://huggingface.co/SC117/Qwen3.8-27B-Uncensored-FIT-GGUF/resolve/main/results/labeled-curves-en.png" alt="KL and Same-top curves with all 14 native presets labeled" style="width: 100%; border-radius: 8px; border: 1px solid #e2e8f0;"><p style="margin: 12px 0 8px 0;">Measured quality improves <b>monotonically across all 14 tiers</b> (macro KL 1.1327 → … → 0.0838), and in the 8–10 GiB region <b>every FIT tier beats its surrounding llama.cpp default presets</b>: FIT-8G / FIT-8.5G beat IQ2_XXS, FIT-9.5G beats Q2_K_S, FIT-10G beats Q2_K. Superseded early recipes (the P5/P6 repairs) are retained in the research record as evidence, not hidden.</p><p style="margin: 0; font-size: 12px;"><a href="https://huggingface.co/SC117/Qwen3.8-27B-Uncensored-FIT-GGUF/resolve/main/results/kl-curve-en.png" style="color: #1d4ed8; text-decoration: none;">Full-size KL</a> · <a href="https://huggingface.co/SC117/Qwen3.8-27B-Uncensored-FIT-GGUF/resolve/main/results/kl-curve-zh.png" style="color: #1d4ed8; text-decoration: none;">中文大图</a> · <a href="https://huggingface.co/SC117/Qwen3.8-27B-Uncensored-FIT-GGUF/resolve/main/results/sametop-curve-en.png" style="color: #1d4ed8; text-decoration: none;">Same-top</a> · <a href="https://huggingface.co/SC117/Qwen3.8-27B-Uncensored-FIT-GGUF/resolve/main/results/strategy-improvement-en.png" style="color: #1d4ed8; text-decoration: none;">Allocation repair</a> · <a href="https://huggingface.co/SC117/Qwen3.8-27B-Uncensored-FIT-GGUF/resolve/main/results/target-utilization-en.png" style="color: #1d4ed8; text-decoration: none;">Target utilization</a></p></div></div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #2563eb 0%, #7c3aed 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🚀</span> Run it</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 8px 0; font-weight: bold; color: #1e293b;">llama.cpp (text)</p><p style="margin: 0; font-family: monospace; background: #f8fafc; padding: 10px 14px; border-radius: 6px; border: 1px solid #e2e8f0; font-size: 12px; color: #1e293b; white-space: pre-wrap;">./llama-cli \
-m Qwen3.8-27B-Uncensored-FIT-12G-IQ3_S.gguf \
-ngl 99 \
-c 8192 \
-cnv</p><p style="margin: 12px 0 8px 0; font-weight: bold; color: #1e293b;">llama-server (text + vision)</p><p style="margin: 0; font-family: monospace; background: #f8fafc; padding: 10px 14px; border-radius: 6px; border: 1px solid #e2e8f0; font-size: 12px; color: #1e293b; white-space: pre-wrap;">./llama-server \
-m Qwen3.8-27B-Uncensored-FIT-12G-IQ3_S.gguf \
--mmproj mmproj-Qwen3.8-27B-Uncensored-BF16.gguf \
-ngl 99 -c 8192</p><p style="margin: 10px 0 0 0;">Any llama.cpp-based runner (llama-cli, llama-server, LM Studio, KoboldCpp, Jan, …) loads these files directly. The BF16 projector pairs with any tier. Pick <code>-ngl</code>, context and batch for your hardware — and remember the GGUF file size alone is not a RAM/VRAM requirement calculator.</p></div></div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #2563eb 0%, #7c3aed 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🔬</span> Evaluation protocol & honest scope</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><table style="width: 100%; border-collapse: collapse; font-size: 12.5px;"><tbody><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155;">Runtime</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563;">llama.cpp b10666 (commit 4e97ac86e) · Linux x86_64 · ROCm</td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155;">Command shape</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563;"><code>llama-perplexity -ngl 99 -t 16 -c 512 -b 512 --kl-divergence ...</code></td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155;">Reference</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563;">Aligned BF16 logits from the converted source</td></tr><tr><td style="padding: 6px 10px; font-weight: bold; color: #334155;">Domains</td><td style="padding: 6px 10px; color: #4b5563;">wiki_test · wiki_valid · Chinese · code · agent_chat (five fixed 64 KiB slices, macro mean)</td></tr></tbody></table><p style="margin: 10px 0 0 0;"><b>Size accuracy:</b> all 14 artifacts matched their post-oracle predicted byte sizes exactly; most use >99.98% of the requested target (FIT-11.5G: 99.427%, the 67.5 MiB reported as target slack — llama.cpp counter-based preset rules shift when manual overrides bypass parts of preset selection; the planner detects this via an override-aware dry-run oracle).</p><p style="margin: 8px 0 0 0;"><b>Allocator scope:</b> the balanced v0.1b policy has positive holdout evidence on the development architecture at some budgets but did not beat matched random allocation on a second model family. This release claims deterministic target-size planning and reports measured quality for these specific artifacts — it does not claim a universally optimal allocation.</p></div></div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #2563eb 0%, #7c3aed 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🧩</span> Included — and not included</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 6px 0;">✅ 14 quantized main language-model GGUFs · ✅ <b>BF16 multimodal projector</b> <code>mmproj-Qwen3.8-27B-Uncensored-BF16.gguf</code> (vision, pairs with any tier) · per-tier plan, effective recipe, tensor override file and quantize record (<code>fit-plans/</code>) · full metrics (<code>results/p4-results.json</code>) · checksums (<code>results/SHA256SUMS</code>)</p><p style="margin: 0 0 6px 0;">❌ The auxiliary NextN/MTP head was excluded from the source conversion (<code>--no-nextn</code>).</p><p style="margin: 0;">The original abliteration belongs to OrcaRouter; this repository contributes the FIT quantization plans and artifacts only.</p></div></div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #2563eb 0%, #7c3aed 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🔍</span> Verify & reproduce</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0; font-family: monospace; background: #f8fafc; padding: 10px 14px; border-radius: 6px; border: 1px solid #e2e8f0; font-size: 12px; color: #1e293b; white-space: pre-wrap;">sha256sum -c results/SHA256SUMS</p><p style="margin: 10px 0 0 0;">Evaluation slices, evaluation logs and the complete decision record live in the <a href="https://github.com/Scorp1o117/FIT-GGUF" style="color: #1d4ed8; text-decoration: none; font-weight: 700;">FIT-GGUF repository</a>. Exact-size behavior is scoped to the recorded source metadata and pinned llama.cpp build; changing the converter, runtime, source layout or metadata requires revalidation.</p></div></div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #2563eb 0%, #7c3aed 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>📄</span> License & credits</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 8px 0;">Apache-2.0, inherited from the base model — follow the upstream license and model-card requirements.</p><p style="margin: 0;"><a href="https://huggingface.co/orcarouter/Qwen3.8-27B-Uncensored" style="color: #1d4ed8; text-decoration: none; font-weight: 600;">OrcaRouter</a> — abliterated BF16 source weights · <a href="https://huggingface.co/Qwen" style="color: #1d4ed8; text-decoration: none; font-weight: 600;">Qwen</a> — the original model family · <a href="https://github.com/ggml-org/llama.cpp" style="color: #1d4ed8; text-decoration: none; font-weight: 600;">llama.cpp</a> — GGUF tooling and runtime · Unsloth calibration dataset lineage — imatrix (1,251 chunks, reused from the same-architecture lineage). FIT-GGUF is an independent project, not affiliated with Qwen, Alibaba, OrcaRouter or llama.cpp.</p></div></div>
</div>
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; text-align: center; margin-bottom: 30px;"><a href="https://github.com/Scorp1o117/FIT-GGUF" style="display: inline-block; background: linear-gradient(135deg, #2563eb 0%, #7c3aed 100%); color: white; font-weight: 700; font-size: 14px; padding: 12px 28px; border-radius: 24px; text-decoration: none;">⭐ FIT-GGUF on GitHub — the tool, the method, the full research record</a><p style="margin: 14px 0 0 0; font-size: 13px;"><span style="color: #86868b;">English</span><span style="color: #cbd5e1;"> · </span><a href="https://huggingface.co/SC117/Qwen3.8-27B-Uncensored-FIT-GGUF/blob/main/README.zh-CN.md" style="color: #1d4ed8; text-decoration: none; font-weight: 600;">简体中文</a></p></div>
Run SC117/Qwen3.8-27B-Uncensored-FIT-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models