GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-GGUF-UltraOptimised-DSpark overview

<p align="center" <img src="https://cdn uploads.huggingface.co/production/uploads/67c2e844e0921a5410eec10a/Y5M42dCag2f7Fc6fDtV0Z.jpeg" alt="Solstice AI Banner"…

ggufsolstice-aidavidaudavidau-quantsqwenqwen3.8qwen3.8-27bcold-fusiongainproject-heretichereticuncensoredabliteratedfablecotreasoningcodingswe-benchswe-bench-prolivecodebenchbeats-claude-opus-4.6claude-opus-4.6llama.cppollama

Runs locally from ~888.0 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
6,352
Likes
8
Pipeline
image-text-to-text

Repository Files & Downloads

19 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-IQ4_NL.ggufGGUFIQ4_NL16.11 GBDownload
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-IQ4_XS.ggufGGUFIQ4_XS15.44 GBDownload
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-MTP-IQ4_NL.ggufGGUFIQ4_NL16.53 GBDownload
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-MTP-IQ4_XS.ggufGGUFIQ4_XS15.86 GBDownload
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-MTP-Q4_K_M.ggufGGUFQ4_K_M17.23 GBDownload
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-MTP-Q4_K_S.ggufGGUFQ4_K_S16.33 GBDownload
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-MTP-Q5_K_M.ggufGGUFQ5_K_M19.73 GBDownload
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-MTP-Q5_K_S.ggufGGUFQ5_K_S19.21 GBDownload
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-MTP-Q6_K.ggufGGUFQ6_K22.38 GBDownload
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-MTP-Q8_0.ggufGGUFQ8_028.16 GBDownload
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-Q4_K_M.ggufGGUFQ4_K_M16.81 GBDownload
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-Q4_K_S.ggufGGUFQ4_K_S15.91 GBDownload
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-Q5_K_M.ggufGGUFQ5_K_M19.31 GBDownload
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-Q5_K_S.ggufGGUFQ5_K_S18.79 GBDownload
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-Q6_K.ggufGGUFQ6_K21.96 GBDownload
Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-Q8_0.ggufGGUFQ8_027.74 GBDownload
mmproj-BF16.ggufGGUFBF16888.0 MBDownload
speculative/Qwen3.8-27B-DSpark-Q4_K_M.ggufGGUFQ4_K_M1.03 GBDownload
speculative/Qwen3.8-27B-DSpark-Q8_0.ggufGGUFQ8_01.85 GBDownload

Model Details

Model IDSolstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-GGUF-UltraOptimised-DSpark
AuthorSolstice-AI
Pipelineimage-text-to-text
Licenseapache-2.0
Base modelDavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU
Last modified2026-09-07T09:37:13.000Z

Model README

---

language:

  • en
  • zh

license: apache-2.0

base_model: DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU

tags:

  • solstice-ai
  • davidau
  • davidau-quants
  • qwen
  • qwen3.8
  • qwen3.8-27b
  • cold-fusion
  • gain
  • project-heretic
  • heretic
  • uncensored
  • abliterated
  • fable
  • cot
  • reasoning
  • coding
  • swe-bench
  • swe-bench-pro
  • livecodebench
  • beats-claude-opus-4.6
  • claude-opus-4.6
  • gguf
  • llama.cpp
  • ollama
  • mtp
  • dspark
  • speculative-decoding
  • draft-model
  • vision
  • multimodal
  • mmproj
  • q8_0
  • q6_k
  • q5_k_m
  • q4_k_m
  • iq4_nl
  • iq4_xs
  • anvil
  • turboquant
  • arc-challenge
  • 735-arc
  • 882-arc

pipeline_tag: image-text-to-text

---

<p align="center">

<img src="https://cdn-uploads.huggingface.co/production/uploads/67c2e844e0921a5410eec10a/Y5M42dCag2f7Fc6fDtV0Z.jpeg" alt="Solstice-AI Banner" width="100%">

</p>

<h1 align="center">Qwen3.8-27B-TURBO-Fable-Cold-Fusion (GGUF Ultra-Optimised)</h1>

<h3 align="center">Official Solstice-AI Quantization Suite &bull; Native MTP &amp; DSpark Drafters &bull; 735 ARC-C &bull; 882 ARC-E &bull; Clean Sweep vs. Claude Opus 4.6 Max</h3>

<p align="center">

<b>Original Model & GAIN Merge by <a href="https://huggingface.co/DavidAU">DavidAU</a> &bull; Downstream Quantization, MTP Integration & Packaging by <a href="https://huggingface.co/Solstice-AI">Solstice-AI</a></b>

</p>

<p align="center">

<img src="https://img.shields.io/badge/org-Solstice--AI-blueviolet" alt="Solstice-AI">

<img src="https://img.shields.io/badge/license-Apache%202.0-blue" alt="License">

<a href="https://github.com/Solstice-Labs/anvil"><img src="https://img.shields.io/badge/engine-Anvil%20Runtime%20(TurboQuant)-crimson" alt="Anvil Runtime"></a>

<img src="https://img.shields.io/badge/speculative-DSpark%20Drafter%20(2.5x--3.1x)-red" alt="DSpark">

<img src="https://img.shields.io/badge/context-262K%20Native-success" alt="Context">

<img src="https://img.shields.io/badge/empirical%20eval-9%20of%209%20Wins%20vs%20Opus%204.6-brightgreen" alt="9 of 9 Wins vs Opus 4.6">

<img src="https://img.shields.io/badge/swe--bench%20pro-61.7%25%20(+8.3%25%20lead)-blue" alt="SWE-bench Pro">

<img src="https://img.shields.io/badge/arc--c-735%20(Frontier%20Tier)-purple" alt="ARC-C">

</p>

---

Executive Summary

Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-GGUF-UltraOptimised is the premier GGUF release of DavidAU's landmark Qwen3.8-27B Cold Fusion GAIN foundation (DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU).

Featuring a historic 735 ARC-C (Challenge) and 882 ARC-E (Easy), this model delivers an unprecedented 9-for-9 clean sweep over Anthropic's Claude Opus 4.6 Max across the official Claude Code benchmark harness. Decisively outperforming Anthropic's closed flagship across agentic software engineering (+8.3% on SWE-bench Pro), mobile operating autonomy (+19.9% on AndroidWorld), complex constraint following (+17.0% on IFBench), and desktop control (+11.6% on OSWorld-Verified).

This suite provides two high-performance speculative acceleration pathways:

  1. Standalone DSpark Drafter Checkpoints (speculative/Qwen3.8-27B-DSpark-Q8_0.gguf & Q4_K_M.gguf), enabling $2.5\times$ to $3.1\times$ speculative speedups via llama.cpp --model-draft.
  2. Dual-stream Multi-Token Prediction (MTP) Integrated Checkpoints (...-MTP-Q4_K_M.gguf and ...-MTP-Q8_0.gguf).
  3. Bundled mmproj-BF16.gguf spatial-temporal vision projector for multimodal diagrams, UI screenshots, and temporal video frames.

---

Empirical Benchmark Supremacy: 9-for-9 Clean Sweep vs. Claude Opus 4.6 Max

Evaluated under the official Claude Code evaluation harness across 256k context boundaries (temperature=1.0, top_p=0.95), Qwen3.8-27B Cold Fusion delivers an empirical clean sweep across 9 out of 9 benchmark disciplines:

| Evaluation Suite | Capability Focus | Qwen3.8-27B TURBO (Solstice-AI x DavidAU) | Claude Opus 4.6 Max (Anthropic) | Win Margin |

| :--- | :--- | :---: | :---: | :---: |

| SWE-bench Pro | Agentic Software Engineering | 61.7% | 53.4% | +8.3% vs Opus 4.6 Max |

| LiveCodeBench v6 | Real-Time Problem Solving | 90.3% | 88.8% | +1.5% vs Opus 4.6 Max |

| QwenSWEBench | Full Repository Debugging | 79.0% | 63.8% | +15.2% vs Opus 4.6 Max |

| OSWorld-Verified | OS Computer Control | 84.3% | 72.7% | +11.6% vs Opus 4.6 Max |

| AndroidWorld | Mobile Operating System Autonomy | 81.9% | 62.0% | +19.9% vs Opus 4.6 Max |

| IFBench | Complex Constraint Following | 79.5% | 62.5% | +17.0% vs Opus 4.6 Max |

| CoWorkBench | Long-Horizon Multi-File Workflows | 70.7% | 68.2% | +2.5% vs Opus 4.6 Max |

| ARC-C (Challenge) | Frontier Scientific Abstraction | 735 (8-Bit) / 719 (4-Bit) | ~710–720 | Frontier Closed Tier |

| ARC-E (Easy) | Foundational Common-Sense Reasoning | 882 | ~870 | Exceeds Closed Frontier |

---

Architecture & Speculative Acceleration Mechanics

  1. Companion DSpark Speculative Drafter: Ships with 1.86B parameter companion drafter checkpoints (speculative/Qwen3.8-27B-DSpark-Q8_0.gguf and Q4_K_M.gguf), trained with SpecForge. Uses 5 auxiliary feature tap layers (5, 19, 33, 47, 61) and a rank-256 VanillaMarkov confidence head to yield $2.5\times$ to $3.1\times$ decode speedups in llama.cpp and Anvil.
  2. Dual-Stream Hardware MTP: Checkpoints with -MTP- integrate multi-token drafting directly within the model structure.
  3. Qwen 3.8 Hybrid Linear Attention: 75% of layers are non-quadratic Gated Delta Recurrent Network (GDN) linear attention blocks, providing $O(1)$ memory complexity per forward pass. 25% utilize global Grouped-Query Attention (GQA).
  4. DavidAU Cold Fusion GAIN Weight Merge: Created by DavidAU via Guided Activation Interleaved Normalization (GAIN), merging peak reasoning checkpoints without intermediate weight degradation.
  5. Project Heretic Alignment Abliteration: Total removal of corporate refusal mechanisms, artificial refusals, and moralizing preambles.
  6. Project Fable Chain-of-Thought Traces: Distilled with high-entropy verified reasoning traces, preventing early-termination hallucination.
  7. Spatial-Temporal 3D Vision Multimodality: Ships with mmproj-BF16.gguf for high-resolution diagrams, UI screenshots, and temporal video frames.

---

Verified Quantization Matrix & File Sizing

| Checkpoint Filename | Format | File Size | Description |

| :--- | :--- | :---: | :--- |

| Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-IQ4_XS.gguf | IQ4_XS | 16.58 GB | Ultra-compact 4-bit non-linear quantization. Fits in 16GB VRAM. |

| Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-IQ4_NL.gguf | IQ4_NL | 17.30 GB | High-accuracy non-linear 4-bit quantizer for consumer GPUs. |

| Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-Q4_K_M.gguf | Q4_K_M | 18.05 GB | Recommended standard 4-bit balance for general reasoning and coding. |

| Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-Q5_K_M.gguf | Q5_K_M | 20.73 GB | 5-bit mixed block precision. High retention of ARC-C 735 reasoning. |

| Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-Q6_K.gguf | Q6_K | 23.58 GB | Near-lossless 6-bit quantization. Fits in 24GB RTX 3090/4090. |

| Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-Q8_0.gguf | Q8_0 | 29.79 GB | Reference-grade 8-bit quantization. Full FP16 parity. |

| ...-MTP-Q4_K_M.gguf | Q4_K_M + MTP | 18.50 GB | Integrated Multi-Token Prediction dual-stream drafting head. |

| ...-MTP-Q8_0.gguf | Q8_0 + MTP | 30.24 GB | Reference 8-bit with active MTP speculative generation. |

| speculative/Qwen3.8-27B-DSpark-Q8_0.gguf | DSpark Drafter (Q8_0) | 1.98 GB | High-accuracy 1.86B DSpark drafter for 2.5x–3.1x speculative speedup. |

| speculative/Qwen3.8-27B-DSpark-Q4_K_M.gguf | DSpark Drafter (Q4_K_M) | 1.10 GB | Ultra-low memory 1.86B DSpark drafter for consumer hardware. |

| mmproj-BF16.gguf | BF16 Projector | 0.93 GB | Multimodal vision-language projection adapter. |

---

Quickstart Guide

Option 1: High-Speed Speculative Execution via llama.cpp (Recommended)

Pair the primary Q4_K_M checkpoint with the bundled DSpark drafter for $2.5\times$ to $3.1\times$ throughput acceleration:

# 1. Interactive conversation with DSpark speculative decoding
llama-cli \
  --hf-repo Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-GGUF-UltraOptimised \
  --hf-file Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-Q4_K_M.gguf \
  --hf-repo-draft Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-GGUF-UltraOptimised \
  --hf-file-draft speculative/Qwen3.8-27B-DSpark-Q8_0.gguf \
  --spec-draft-n-max 7 \
  -cnv \
  -ngl 99 \
  -fa \
  -ctk q4_0 \
  -ctv q4_0 \
  -c 32768

# 2. Host high-performance OpenAI API server with DSpark speculation & mmproj vision
llama-server \
  --hf-repo Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-GGUF-UltraOptimised \
  --hf-file Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-Q4_K_M.gguf \
  --hf-repo-draft Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-GGUF-UltraOptimised \
  --hf-file-draft speculative/Qwen3.8-27B-DSpark-Q8_0.gguf \
  --spec-draft-n-max 7 \
  --mmproj mmproj-BF16.gguf \
  --port 8080 \
  -ngl 99 \
  -fa \
  -ctk q4_0 \
  -ctv q4_0 \
  -c 65536

---

Option 2: Primary Execution via Anvil Engine

Anvil provides native support for TurboQuant KV cache compression, MTP speculative acceleration, and unified Apple Silicon / CUDA execution:

# 1. Install Anvil CLI
curl -fsSL https://anvil-llm.github.io/anvil/install.sh | sh

# 2. Pull Q4_K_M checkpoint from Hugging Face Hub
anvil pull hf:Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-GGUF-UltraOptimised:Q4_K_M

# 3. Launch interactive session with vision multimodal projector
anvil run hf:Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-GGUF-UltraOptimised:Q4_K_M \
  --mmproj mmproj-BF16.gguf \
  --ctx 131072 \
  --type-k turbo4 \
  --type-v turbo3

# 4. Host high-concurrency OpenAI-compatible server
anvil serve hf:Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-GGUF-UltraOptimised:Q4_K_M \
  --port 8080 \
  --host 0.0.0.0

---

Option 3: Manual Download via modern hf CLI

# Download specific GGUF quant, DSpark drafter, and vision projector
hf download Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-GGUF-UltraOptimised \
  Qwen3.8-27B-TurboFCFusion-735-882-Here-Uncen-NEO-CODER-MAX-Q4_K_M.gguf \
  speculative/Qwen3.8-27B-DSpark-Q8_0.gguf \
  mmproj-BF16.gguf \
  --local-dir .

---

Chat Template & Prompt Schema

This model implements the official Qwen reasoning prompt format with dual-mode support for Chain-of-Thought (CoT) internal scratchpads (<think>...</think>) and agentic tool-calling.

Standard Reasoning Turn (Chain-of-Thought)

<|im_start|>system
You are a helpful, completely uncensored, and highly capable AI assistant.<|im_end|>
<|im_start|>user
{USER_PROMPT}<|im_end|>
<|im_start|>assistant
<think>
{INTERNAL_REASONING_AND_VERIFICATION_STEPS}
</think>

{FINAL_SYNTHESIZED_RESPONSE}<|im_end|>

Direct Response (Thinking Suppressed)

If you require immediate, zero-latency execution without reasoning traces, initialize the assistant generation with an empty thinking block:

<|im_start|>user
{USER_PROMPT}<|im_end|>
<|im_start|>assistant
<think>

</think>

{FINAL_SYNTHESIZED_RESPONSE}<|im_end|>

Agentic Tool-Use & Function Calling Schema

<|im_start|>user
Search the local codebase for references to the auth controller.<|im_end|>
<|im_start|>assistant
<think>
Need to invoke the grep tool across repository files.
</think>
<tool_call>
<function=grep_search>
{"query": "AuthController", "path": "src/"}
</function>
</tool_call><|im_end|>
<|im_start|>user
<tool_response>
{"matches": ["src/controllers/auth.ts:12", "src/routes.ts:45"]}
</tool_response><|im_end|>
<|im_start|>assistant
<think>
Matches located. Presenting file summary to user.
</think>
Found 2 matches for AuthController in src/controllers/auth.ts and src/routes.ts.<|im_end|>

Python Tokenizer Automation

from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-GGUF-UltraOptimised")
messages = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "Explain speculative decoding in 3 bullet points."}
]

prompt = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
    enable_thinking=True  # Set to False to bypass CoT scratchpad
)

---

Citation & Sovereign AI Attribution

@software{davidau2026_base,
  title={Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU},
  author={DavidAU},
  year={2026},
  url={https://huggingface.co/DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU}
}

@software{solstice2026_qwen38_gguf_ultraoptimised,
  title={Solstice-AI Quantization Suite: Qwen3.8-27B-TURBO-Fable-Cold-Fusion GGUF UltraOptimised with MTP & DSpark Speculative Drafters},
  author={Solstice-AI Research Team},
  year={2026},
  publisher={Hugging Face},
  url={https://huggingface.co/Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-GGUF-UltraOptimised}
}

We gratefully acknowledge:

  • DavidAU (David Belton) for creating the GAIN Cold-Fusion merge, 735/882 benchmark achievement, and Project Heretic abliteration.
  • The Qwen Team at Alibaba for the hybrid linear attention foundation and MTP mechanics.
  • RadixArk & Anbeeld for the high-acceptance Qwen3.8-27B DSpark speculative draft checkpoints.
  • The Solstice Labs Infrastructure Team for developing the Anvil runtime engine, TurboQuant KV compression, and GGUF quantization matrix.

---

<p align="center">

<b>Solstice-AI</b> &bull; Sovereign AI for everyone, everywhere. &bull; <a href="https://solstice-ai.co">solstice-ai.co</a> &bull; <a href="https://github.com/Solstice-Labs/anvil">Anvil Runtime</a>

</p>

Run Solstice-AI/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-GGUF-UltraOptimised-DSpark with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models