SC117/Spark-X2.5-4B-abliterated-FIT-GGUF overview
<div style="font family: apple system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans serif; margin bottom: 24px;" <div style="background: 0f172a; border radius:…
Runs locally from ~2.26 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Spark-X2.5-4B-abliterated-BF16.gguf | GGUF | BF16 | 7.66 GB | Download |
| Spark-X2.5-4B-abliterated-FIT-BALANCED-2.68GiB-Q5_K_S.gguf | GGUF | Q5_K_S | 2.68 GB | Download |
| Spark-X2.5-4B-abliterated-FIT-COMPACT-2.43GiB-Q4_K.gguf | GGUF | Q4_K | 2.43 GB | Download |
| Spark-X2.5-4B-abliterated-FIT-MINI-2.26GiB-Q4_K_S.gguf | GGUF | Q4_K_S | 2.26 GB | Download |
| Spark-X2.5-4B-abliterated-FIT-QUALITY-3.15GiB-Q6_K.gguf | GGUF | Q6_K | 3.15 GB | Download |
Model Details
| Model ID | SC117/Spark-X2.5-4B-abliterated-FIT-GGUF |
|---|---|
| Author | SC117 |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | XHToken/Spark-X2.5-4B |
| Last modified | 2026-09-06T09:09:41.000Z |
Model README
---
base_model: XHToken/Spark-X2.5-4B
base_model_relation: quantized
license: apache-2.0
library_name: gguf
pipeline_tag: text-generation
language:
- en
- zh
tags:
- gguf
- llama-cpp
- quantization
- sparkx2_5
- fit-gguf
- uncensored
- abliterated
---
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; margin-bottom: 24px;">
<div style="background: #0f172a; border-radius: 20px; padding: 40px 32px 30px 32px; text-align: center; position: relative; overflow: hidden;">
<div style="position: absolute; top: -40px; right: -40px; width: 160px; height: 160px; background: rgba(139,92,246,0.18); border-radius: 50%;"></div>
<div style="position: absolute; bottom: -50px; left: -30px; width: 140px; height: 140px; background: rgba(59,130,246,0.15); border-radius: 50%;"></div>
<div style="position: absolute; top: 30%; right: 12%; width: 56px; height: 56px; background: rgba(59,130,246,0.22); border-radius: 50%;"></div>
<div style="display: inline-flex; flex-wrap: wrap; justify-content: center; gap: 8px; margin-bottom: 18px; position: relative; z-index: 1;"><span style="background: #3b82f6; color: white; font-size: 11px; font-weight: 700; padding: 5px 14px; border-radius: 20px;">FIT-GGUF v0.2.0</span><span style="background: #8b5cf6; color: white; font-size: 11px; font-weight: 700; padding: 5px 14px; border-radius: 20px;">GATE-VERIFIED TIERS</span><span style="background: #10b981; color: white; font-size: 11px; font-weight: 700; padding: 5px 14px; border-radius: 20px;">4 TIERS · 2.26–3.15 GiB</span><span style="background: #ef4444; color: white; font-size: 11px; font-weight: 700; padding: 5px 14px; border-radius: 20px;">ABLITERATED (T615)</span><span style="background: #f59e0b; color: #0f172a; font-size: 11px; font-weight: 700; padding: 5px 14px; border-radius: 20px;">MEASURED KL + SAME-TOP</span><span style="background: #1e293b; color: #94a3b8; font-size: 11px; font-weight: 700; padding: 5px 14px; border-radius: 20px;">APACHE-2.0</span></div>
<h1 style="margin: 0 0 10px 0; font-size: 34px; font-weight: 800; color: #f8fafc; letter-spacing: -0.5px; border: none; position: relative; z-index: 1;">Spark-X2.5-4B-abliterated · FIT-GGUF</h1>
<p style="margin: 0 0 20px 0; font-size: 15px; color: #cbd5e1; position: relative; z-index: 1;">Four fidelity tiers of a 1M-context hybrid-attention 4B — every shipped file re-verified against its own BF16, byte-exact and reproducible.</p>
<div style="position: relative; z-index: 1; margin: 0 auto; width: 72%; max-width: 540px;">
<div style="height: 8px; border-radius: 8px; background: #334155; position: relative; overflow: visible;">
<div style="height: 8px; width: 47%; border-radius: 8px; background: linear-gradient(90deg, #3b82f6, #8b5cf6);"></div>
<div style="position: absolute; top: 50%; left: 47%; transform: translate(-50%, -50%); width: 18px; height: 18px; border-radius: 50%; background: #ffffff; box-shadow: 0 0 0 4px rgba(139,92,246,0.35);"></div>
</div>
<div style="display: flex; justify-content: space-between; margin-top: 8px; font-size: 10px; color: #94a3b8; font-weight: 600;"><span>2.26 GiB</span><span style="color: #c4b5fd;">verified minimum at each fidelity</span><span>3.15 GiB</span></div>
</div>
<p style="margin: 18px 0 0 0; font-size: 13px; position: relative; z-index: 1;"><span style="color: #94a3b8;">English</span><span style="color: #475569;"> · </span><a href="https://huggingface.co/SC117/Spark-X2.5-4B-abliterated-FIT-GGUF/blob/main/README.zh-CN.md" style="color: #93c5fd; text-decoration: none; font-weight: 600;">简体中文 📖</a></p>
</div>
</div>
<p align="center"><img src="https://huggingface.co/SC117/Spark-X2.5-4B-abliterated-FIT-GGUF/resolve/main/assets/fit-gguf-banner.png" alt="FIT-GGUF" width="760"></p>
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; display: flex; flex-direction: column; gap: 18px; margin-bottom: 30px;">
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #2563eb 0%, #7c3aed 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🧭</span> About FIT-GGUF — the tool behind these files</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 10px 0;">Every file in this repository was planned, executed and verified by <a href="https://github.com/Scorp1o117/FIT-GGUF" style="color: #1d4ed8; text-decoration: none; font-weight: 700;">FIT-GGUF</a>, an open-source, deterministic tensor-level planning layer on top of llama.cpp quantization. Standard GGUF quantization asks you to pick one of a handful of presets; FIT-GGUF instead asks <i>what quality do you want</i>, then finds and <b>verifies</b> the smallest GGUF that demonstrably meets it.</p><p style="margin: 0 0 12px 0; padding: 10px 14px; background: #f5f3ff; border-left: 4px solid #7c3aed; border-radius: 6px; color: #4c1d95; font-weight: 600;">Traditional GGUF gives you presets. FIT gives you a fidelity contract: macro KL ≤ tier anchor ∧ same-top ≥ model-calibrated floor.</p><table style="width: 100%; border-collapse: collapse; font-size: 13px;"><tbody><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155;">Deterministic size prediction & byte-exact delivery</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #166534; font-weight: 700;">✅ Validated (G2 gate, delta = 0)</td></tr><tr><td style="padding: 6px 10px; font-weight: bold; color: #334155;">Universally optimal tensor allocation</td><td style="padding: 6px 10px; color: #92400e; font-weight: 700;">⚠️ Not established — FIT claims verified fidelity contracts, not a universal quality optimum</td></tr></tbody></table><p style="margin: 12px 0 0 0;">The method, the preregistered research record and the <code>fit</code> CLI are open source: <a href="https://github.com/Scorp1o117/FIT-GGUF" style="color: #1d4ed8; text-decoration: none; font-weight: 700;">github.com/Scorp1o117/FIT-GGUF</a></p></div></div>
<div style="border: 1px solid #fca5a5; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #dc2626 0%, #ea580c 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>⚠️</span> Safety notice / 安全提示</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 8px 0;">The source model is an <b>abliterated, refusal-removed</b> model with no meaningful built-in guardrails, and may comply with harmful, illegal or unsafe requests. Use it only where you can provide appropriate moderation, access control and legal review. Do not deploy it to end users without your own safety layer.</p><p style="margin: 0; color: #64748b;">源模型经过拒答方向消融,不具备可靠的内置安全护栏;请仅在合法、受控、具备审核与滥用防护的环境中使用,使用者自行承担部署责任。</p></div></div>
<div style="border: 1px solid #c4b5fd; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #7c3aed 0%, #a855f7 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🧬</span> Abliteration — trial T615, documented</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 10px 0;">The refusal-direction ablation was performed locally with the <b>abliterix</b> pipeline: <b>mean-LoRA steering</b> — rank-8 full-norm LoRA adapters built from per-layer mean refusal directions, with <b>projected abliteration</b>, a <b>gaussian decay kernel</b> over depth, and per-layer vector scope. The exported weights are Optuna trial <b>#615</b> of study <code>spark_x25_4b_lora_v30</code>, selected from the measured Pareto front.</p><table style="width: 100%; border-collapse: collapse; font-size: 12.5px;"><thead><tr style="background: rgba(124,58,237,0.05);"><th style="padding: 7px 10px; border-bottom: 2px solid #7c3aed; text-align: left; color: #6d28d9;">Metric</th><th style="padding: 7px 10px; border-bottom: 2px solid #7c3aed; text-align: left; color: #6d28d9;">Original</th><th style="padding: 7px 10px; border-bottom: 2px solid #7c3aed; text-align: left; color: #6d28d9;">T615</th></tr></thead><tbody><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: 600;">Keyword refusals (100 prompts, held-out slice)</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">100</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: 700; color: #166534;">7</td></tr><tr><td style="padding: 6px 10px; font-weight: 600;">3-token full-distribution KL vs original (nats/token)</td><td style="padding: 6px 10px;">—</td><td style="padding: 6px 10px; font-weight: 700;">0.1468</td></tr><tr><td style="padding: 6px 10px; font-weight: 600;">Validation KL (mean over held-out prompts)</td><td style="padding: 6px 10px;">—</td><td style="padding: 6px 10px;">0.148</td></tr></tbody></table><p style="margin: 10px 0 0 0; padding: 9px 13px; background: #fffbeb; border-left: 4px solid #f59e0b; border-radius: 6px; color: #92400e;"><b>Honest selection trade-off:</b> the pre-registered "ship bar" was refusals ≤ 10 ∧ KL ≤ 0.05. T615 clears the refusal bar with margin but deliberately trades weight distance for openness (KL 0.147 > 0.05) — the previous export T580 (KL 0.093) left sexual-content soft-refusals and was replaced by user choice. This is an openness/faithfulness Pareto decision, reported as measured.</p><table style="width: 100%; border-collapse: collapse; font-size: 12px; margin-top: 10px;"><thead><tr style="background: rgba(124,58,237,0.05);"><th style="padding: 6px 10px; border-bottom: 1px solid #ddd; text-align: left; color: #6d28d9;">Trial</th><th style="padding: 6px 10px; border-bottom: 1px solid #ddd; text-align: left; color: #6d28d9;">Refusals /100</th><th style="padding: 6px 10px; border-bottom: 1px solid #ddd; text-align: left; color: #6d28d9;">KL (nats/token)</th><th style="padding: 6px 10px; border-bottom: 1px solid #ddd; text-align: left; color: #6d28d9;">Note</th></tr></thead><tbody><tr><td style="padding: 5px 10px; font-weight: 700;">615</td><td style="padding: 5px 10px;">7</td><td style="padding: 5px 10px;">0.1468</td><td style="padding: 5px 10px; color: #166534; font-weight: 600;">this export</td></tr><tr><td style="padding: 5px 10px;">389</td><td style="padding: 5px 10px;">10</td><td style="padding: 5px 10px;">0.1433</td><td style="padding: 5px 10px; color: #64748b;">Pareto neighbor</td></tr><tr><td style="padding: 5px 10px;">221</td><td style="padding: 5px 10px;">3</td><td style="padding: 5px 10px;">0.2299</td><td style="padding: 5px 10px; color: #64748b;">most open</td></tr><tr><td style="padding: 5px 10px;">580</td><td style="padding: 5px 10px;">20</td><td style="padding: 5px 10px;">0.0929</td><td style="padding: 5px 10px; color: #64748b;">previous export (leftover refusals)</td></tr></tbody></table><p style="margin: 10px 0 0 0; font-size: 12px; color: #64748b;">All fidelity-tier measurements in this repository are taken against the <b>T615 BF16 weights themselves</b> — the ablation delta is baked into the reference, so the tier numbers below quantify quantization loss only.</p></div></div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #2563eb 0%, #7c3aed 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>📦</span> Pick a tier</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 10px 0; padding: 9px 13px; background: #fffbeb; border-left: 4px solid #f59e0b; border-radius: 6px; color: #92400e;"><b>File size ≠ RAM/VRAM usage.</b> KV cache, compute buffers and runtime overhead are separate. The native context is 1M tokens; the hybrid attention (3 sliding-window : 1 full) keeps most depth cheap, but the 9 full-attention layers still need <b>~34 GiB of f16 KV at 1M tokens</b> — pick a sane <code>-c</code>. Naming: <code>Spark-X2.5-4B-abliterated-FIT-<tier>-<size>-<dominant>.gguf</code></p><table style="width: 100%; border-collapse: collapse; font-size: 12.5px;"><thead><tr style="background: rgba(37,99,235,0.05);"><th style="padding: 7px 10px; border-bottom: 2px solid #2563eb; text-align: left; color: #1d4ed8;">Tier</th><th style="padding: 7px 10px; border-bottom: 2px solid #2563eb; text-align: left; color: #1d4ed8;">GiB</th><th style="padding: 7px 10px; border-bottom: 2px solid #2563eb; text-align: left; color: #1d4ed8;">Dominant</th><th style="padding: 7px 10px; border-bottom: 2px solid #2563eb; text-align: left; color: #1d4ed8;">Macro KL ↓</th><th style="padding: 7px 10px; border-bottom: 2px solid #2563eb; text-align: left; color: #1d4ed8;">Same-top ↑</th><th style="padding: 7px 10px; border-bottom: 2px solid #2563eb; text-align: left; color: #1d4ed8;">Gates (KL ≤ / top ≥)</th><th style="padding: 7px 10px; border-bottom: 2px solid #2563eb; text-align: left; color: #1d4ed8;">Measured positioning</th></tr></thead><tbody><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: 700; color: #1d4ed8;">⭐ QUALITY</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">3.147</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">Q6_K</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">0.0289</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">93.50%</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">0.05 / 93.16%</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #64748b;">Max verified quality; the Q6_K preset is itself the minimum verified PASS</td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: 600;">BALANCED</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">2.680</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">Q5_K_S</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">0.0739</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">89.15%</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">0.10 / 89.14%</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #64748b;">Best size/quality trade-off; nothing smaller passes both gates</td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: 600;">COMPACT</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">2.432</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">Q4_K</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">0.1458</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">84.88%</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15);">0.15 / 84.88%</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #166534; font-weight: 600;">FIT tensor-level recipe — fills the preset gap (Q5_K_S 0.074 → Q4_K_M 0.170)</td></tr><tr><td style="padding: 6px 10px; font-weight: 600;">MINI</td><td style="padding: 6px 10px;">2.260</td><td style="padding: 6px 10px;">Q4_K_S</td><td style="padding: 6px 10px;">0.1834</td><td style="padding: 6px 10px;">82.61%</td><td style="padding: 6px 10px;">0.20 / 82.61%</td><td style="padding: 6px 10px; color: #64748b;">Smallest verified PASS</td></tr></tbody></table><p style="margin: 14px 0 6px 0; font-weight: bold; color: #1e293b;">Quick picks</p><div style="display: flex; flex-wrap: wrap; gap: 10px;"><div style="flex: 1 1 46%; border: 1px solid #e2e8f0; border-radius: 10px; padding: 10px 12px; background: #f8fafc;"><div style="font-weight: 700; color: #1d4ed8;">🏆 BALANCED 2.68G — the sweet spot</div><div style="font-size: 12px; color: #475569; margin-top: 3px;">Q5_K_S-class quality at 2.68 GiB; the verified minimum at the 0.10 KL tier.</div></div><div style="flex: 1 1 46%; border: 1px solid #e2e8f0; border-radius: 10px; padding: 10px 12px; background: #f8fafc;"><div style="font-weight: 700; color: #166534;">🎯 QUALITY 3.15G — max verified quality</div><div style="font-size: 12px; color: #475569; margin-top: 3px;">93.50% same-top, KL 0.0289 — within 0.47 GiB of the Q8_0 preset with clearly stronger economics.</div></div><div style="flex: 1 1 46%; border: 1px solid #e2e8f0; border-radius: 10px; padding: 10px 12px; background: #f8fafc;"><div style="font-weight: 700; color: #b45309;">🧩 COMPACT 2.43G — where FIT earns its keep</div><div style="font-size: 12px; color: #475569; margin-top: 3px;">No preset exists between KL 0.074 and 0.170; this tensor-level recipe lands at 0.146 and passes both gates.</div></div><div style="flex: 1 1 46%; border: 1px solid #fecaca; border-radius: 10px; padding: 10px 12px; background: #fef2f2;"><div style="font-weight: 700; color: #b91c1c;">⚠️ Below ~2.2 GiB</div><div style="font-size: 12px; color: #7f1d1d; margin-top: 3px;">Steep low-bit cliff: 2-bit classes collapse (macro KL 0.8–3.9) and are deliberately not offered.</div></div></div><p style="margin: 12px 0 0 0; font-size: 12px; color: #64748b; font-style: italic;">All numbers are protocol-scoped observations (five fixed 64 KiB domains vs this model's aligned BF16), not an application benchmark or a universal ranking. Honest scope: on this model the preset ladder is strong — Quality/Balanced/Mini minimum verified PASS are the native presets themselves; FIT contributes the verification and the Compact gap-fill. All four tiers are same-top-bound: the guard floors, not KL, decide the sizes.</p></div></div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #2563eb 0%, #7c3aed 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>📈</span> Measured quality</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><img src="https://huggingface.co/SC117/Spark-X2.5-4B-abliterated-FIT-GGUF/resolve/main/results/labeled-curves-en.png" alt="KL and Same-top curves with presets, calibration probes and shipped tiers labeled" style="width: 100%; border-radius: 8px; border: 1px solid #e2e8f0;"><p style="margin: 12px 0 8px 0;">Quality improves monotonically across the healthy preset ladder, and <b>every tier above is same-top-bound</b> — the model-calibrated floors, not KL, are what separate PASS from FAIL here. The 2-bit region is reported as measured: Q2_K/IQ2_* collapse (macro KL 0.8–3.9). The COMPACT tier fills the gap between Q5_K_S (0.074) and Q4_K_M (0.170) where no native preset exists.</p><p style="margin: 0; font-size: 12px;"><a href="https://huggingface.co/SC117/Spark-X2.5-4B-abliterated-FIT-GGUF/resolve/main/results/labeled-curves-en.png" style="color: #1d4ed8; text-decoration: none;">Full-size EN</a> · <a href="https://huggingface.co/SC117/Spark-X2.5-4B-abliterated-FIT-GGUF/resolve/main/results/labeled-curves-zh.png" style="color: #1d4ed8; text-decoration: none;">中文大图</a></p></div></div>
<div style="border: 1px solid #fca5a5; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #dc2626 0%, #ea580c 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🚀</span> Run it — runtime requirement first</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 10px 0; padding: 9px 13px; background: #fef2f2; border-left: 4px solid #dc2626; border-radius: 6px; color: #7f1d1d;"><b>These GGUFs use the <code>spark2_5</code> architecture.</b> You need a llama.cpp build with Spark2_5 support — <a href="https://github.com/ggml-org/llama.cpp/pull/27868" style="color: #b91c1c; font-weight: 700;">PR #27868</a> (open at release time, 2026-09) or any later release that includes it. Upstream <code>master</code> <b>without this PR cannot load these files</b>.</p><p style="margin: 0 0 8px 0; font-weight: bold; color: #1e293b;">llama.cpp</p><p style="margin: 0; font-family: monospace; background: #f8fafc; padding: 10px 14px; border-radius: 6px; border: 1px solid #e2e8f0; font-size: 12px; color: #1e293b; white-space: pre-wrap;">./llama-server \
-m Spark-X2.5-4B-abliterated-FIT-BALANCED-2.68GiB-Q5_K_S.gguf \
-ngl 99 -c 32768
./llama-cli \
-m Spark-X2.5-4B-abliterated-FIT-QUALITY-3.15GiB-Q6_K.gguf \
-ngl 99 -c 8192 -p "Hello" -n 256</p><p style="margin: 10px 0 0 0;">The chat template is embedded in every file. Once Spark2_5 support merges, any llama.cpp-based runner (llama-cli, llama-server, LM Studio, KoboldCpp, Jan, …) loads these files directly. This is a dense 4B text model — no MTP head, no vision projector.</p></div></div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #2563eb 0%, #7c3aed 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🔬</span> Evaluation protocol & honest scope</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><table style="width: 100%; border-collapse: collapse; font-size: 12.5px;"><tbody><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155;">Runtime</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563;">llama.cpp @ PR #27868 head (ae320b1) · Linux x86_64 · ROCm (gfx1151) — perplexity/KL code byte-identical to the pinned eval-v1 build</td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155;">Command shape</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563;"><code>llama-perplexity -ngl 99 -t 16 -c 512 -b 512 --kl-divergence …</code></td></tr><tr><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); font-weight: bold; color: #334155;">Reference</td><td style="padding: 6px 10px; box-shadow: 0 1px 0 0 rgba(128,128,128,0.15); color: #4b5563;">This model's own BF16 logits (T615 weights)</td></tr><tr><td style="padding: 6px 10px; font-weight: bold; color: #334155;">Domains</td><td style="padding: 6px 10px; color: #4b5563;">wiki_test · wiki_valid · Chinese · code · agent_chat (five fixed 64 KiB slices, macro mean)</td></tr></tbody></table><p style="margin: 10px 0 0 0;"><b>Tier verification:</b> a tier ships only if macro KL ≤ its anchor <b>and</b> same-top ≥ a model-calibrated guard floor (15-point preset ladder + 2 gap probes; floors are exact-model-scoped and are not transferred across models). Every shipped file was re-evaluated as the exact shipped bytes; the COMPACT recipe is reproducible byte-for-byte.</p><p style="margin: 8px 0 0 0;"><b>Allocator scope:</b> the balanced v0.1 policy was used as-is (no model-specific refine profile exists for spark2_5). This release claims deterministic size planning and measured, gate-verified fidelity for these specific artifacts — it does not claim a universally optimal allocation.</p></div></div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #2563eb 0%, #7c3aed 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🧩</span> Included — and not included</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 6px 0;">✅ 4 gate-verified tier GGUFs · ✅ the <b>BF16 reference GGUF</b> (the exact weights every measurement above is taken against) · ✅ chat template embedded in every file · ✅ checksums (<code>SHA256SUMS.txt</code>) · ✅ labeled quality curves (<code>results/</code>) · ✅ tier manifest (<code>FIT-TIERS.md</code>)</p><p style="margin: 0 0 6px 0;">❌ 2-bit quantization classes — measured collapse on this model (macro KL 0.8–3.9), excluded by design · ❌ no vision projector / no MTP head (dense text model)</p><p style="margin: 0;">The abliteration (T615) was performed locally from <a href="https://huggingface.co/XHToken/Spark-X2.5-4B-Base" style="color: #1d4ed8; text-decoration: none; font-weight: 600;">XHToken/Spark-X2.5-4B-Base</a>; this repository contributes the abliterated weights and the FIT quantization plans and artifacts.</p></div></div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #2563eb 0%, #7c3aed 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>🔍</span> Verify & reproduce</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0; font-family: monospace; background: #f8fafc; padding: 10px 14px; border-radius: 6px; border: 1px solid #e2e8f0; font-size: 12px; color: #1e293b; white-space: pre-wrap;">sha256sum -c SHA256SUMS.txt</p><p style="margin: 10px 0 0 0;">The evaluation slices are public in the <a href="https://github.com/Scorp1o117/FIT-GGUF" style="color: #1d4ed8; text-decoration: none; font-weight: 700;">FIT-GGUF repository</a>; the calibration ladder, gap probes, search audit and guard profile for this release are retained in the FIT-GGUF experiment records (<code>2026-09-06-spark-x25-4tier</code>). Exact-size behavior is scoped to the recorded source metadata and the recorded llama.cpp build; changing the converter, runtime, source layout or metadata requires revalidation.</p></div></div>
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);"><div style="background: linear-gradient(135deg, #2563eb 0%, #7c3aed 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px; display: flex; align-items: center; gap: 8px;"><span>📄</span> License & credits</div><div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;"><p style="margin: 0 0 8px 0;">Apache-2.0, inherited from the base model — follow the upstream license and model-card requirements.</p><p style="margin: 0;"><a href="https://huggingface.co/XHToken" style="color: #1d4ed8; text-decoration: none; font-weight: 600;">XHToken / SparkLLM</a> — the Spark-X2.5 model family · <a href="https://github.com/ggml-org/llama.cpp/pull/27868" style="color: #1d4ed8; text-decoration: none; font-weight: 600;">llama.cpp PR #27868 (KnightYao)</a> — Spark2_5 support · <b>abliterix</b> — the local ablation pipeline (trial T615) · <a href="https://github.com/Scorp1o117/FIT-GGUF" style="color: #1d4ed8; text-decoration: none; font-weight: 600;">FIT-GGUF</a> — verified size-exact quantization. FIT-GGUF is an independent project, not affiliated with XHToken, SparkLLM or llama.cpp.</p></div></div>
</div>
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; text-align: center; margin-bottom: 30px;"><a href="https://github.com/Scorp1o117/FIT-GGUF" style="display: inline-block; background: linear-gradient(135deg, #2563eb 0%, #7c3aed 100%); color: white; font-weight: 700; font-size: 14px; padding: 12px 28px; border-radius: 24px; text-decoration: none;">⭐ FIT-GGUF on GitHub — the tool, the method, the full research record</a><p style="margin: 14px 0 0 0; font-size: 13px;"><span style="color: #86868b;">English</span><span style="color: #cbd5e1;"> · </span><a href="https://huggingface.co/SC117/Spark-X2.5-4B-abliterated-FIT-GGUF/blob/main/README.zh-CN.md" style="color: #1d4ed8; text-decoration: none; font-weight: 600;">简体中文</a></p></div>
Run SC117/Spark-X2.5-4B-abliterated-FIT-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models