GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

satgeze/QwenPaw-Flash-9B-heretic-1M-GGUF overview

<div style="font family: apple system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans serif; margin bottom: 30px;" <div style="background: f5f5f7; border radius:…

gguflong-contextyarnqwen3.5hereticuncensoredmtpspeculative-decodingvisionagenttool-callllama.cppollamatext-generationenzhbase_model:Qwen/Qwen3.5-9Bbase_model:quantized:Qwen/Qwen3.5-9Blicense:apache-2.0region:us

Runs locally from ~879.0 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
130
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

3 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
mmproj-qwenpaw-9b.ggufGGUFGGUF879.0 MBDownload
qwenpaw-9b-1M-MTP-Q4_K_M.ggufGGUFQ4_K_M5.38 GBDownload
qwenpaw-9b-1M-MTP-Q8_0.ggufGGUFQ8_09.11 GBDownload

Model Details

Model IDsatgeze/QwenPaw-Flash-9B-heretic-1M-GGUF
Authorsatgeze
Pipelinetext-generation
Licenseapache-2.0
Base modelQwen/Qwen3.5-9B,agentscope-ai/QwenPaw-Flash-9B,SC117/QwenPaw-Flash-9B-heretic-MTP-GGUF
Last modified2026-07-10T22:59:15.000Z

Model README

---

language:

  • en
  • zh

license: apache-2.0

pipeline_tag: text-generation

base_model:

  • Qwen/Qwen3.5-9B
  • agentscope-ai/QwenPaw-Flash-9B
  • SC117/QwenPaw-Flash-9B-heretic-MTP-GGUF

tags:

  • gguf
  • long-context
  • yarn
  • qwen3.5
  • heretic
  • uncensored
  • mtp
  • speculative-decoding
  • vision
  • agent
  • tool-call
  • llama.cpp
  • ollama

---

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; margin-bottom: 30px;">

<div style="background: #f5f5f7; border-radius: 20px; padding: 36px 32px; margin-bottom: 20px; text-align: center; position: relative; overflow: hidden;">

<div style="position: absolute; top: -30px; right: -30px; width: 120px; height: 120px; background: #ffedd5; border-radius: 50%;"></div>

<div style="position: absolute; bottom: -20px; left: 40px; width: 80px; height: 80px; background: #fed7aa; border-radius: 50%;"></div>

<div style="position: absolute; top: 50%; left: -15px; width: 60px; height: 60px; background: #ffedd5; border-radius: 50%;"></div>

<div style="display: inline-flex; gap: 8px; margin-bottom: 16px; position: relative; z-index: 1;">

<span style="background: #34c759; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">1M CONTEXT</span><span style="background: #007aff; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">MTP</span><span style="background: #af52de; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">VISION</span><span style="background: #ff3b30; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">UNCENSORED</span>

</div>

<h1 style="margin: 0 0 8px 0; font-size: 32px; font-weight: 700; color: #1d1d1f; letter-spacing: -0.5px; border: none; position: relative; z-index: 1;">QwenPaw-Flash-9B-heretic-1M</h1>

<p style="margin: 8px 0 0 0; font-size: 14px; color: #86868b; position: relative; z-index: 1;">Agent-optimized 9B, uncensored, with a needle-verified million-token context</p>

</div>

</div>

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; display: flex; flex-direction: column; gap: 20px; margin-bottom: 30px;">

<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">

<div style="padding: 16px;">

<p style="margin: 0 0 12px 0; font-size: 13px; color: #334155; line-height: 1.7;"><a href="https://huggingface.co/SC117/QwenPaw-Flash-9B-heretic-MTP-GGUF" target="_blank" style="color: #c2410c; text-decoration: none; font-weight: 700;">SC117's heretic build</a> of <a href="https://huggingface.co/agentscope-ai/QwenPaw-Flash-9B" target="_blank" style="color: #c2410c; text-decoration: none; font-weight: 700;">QwenPaw-Flash-9B</a> (agent-trajectory finetune of Qwen3.5-9B) with YaRN rope scaling <b>baked into the GGUF metadata for a 1,048,576-token context window</b>, 4x the native 262,144. Weights are bit-identical to SC117's release, which already carries the official Qwen3.5-9B MTP layer injected back after the QwenPaw fine-tuning stripped it.</p>

<p style="margin: 0; font-size: 13px; color: #334155; line-height: 1.7;"><i>Uncensored (Heretic v1.3.0) · MTP baked in · Vision tower included · 1M needle-verified</i></p>

</div>

</div>

</div>

Verified on these exact files

| Capability | Result |

|---|---|

| 1M context | 10 needles per rung, depths 5 to 95 percent, temp 0, Q8_0 + f16 KV: 10/10 at every rung from 64K through 524K on an RTX 5090; 786K and 1M rungs running on a 128 GB M3 Max, card updates when they land |

| MTP speculative decoding | 217.9 to 273.2 tok/s (+25 percent), draft acceptance 0.702, output identical by construction |

| Vision | mmproj tower reads image text and identifies objects correctly |

| Coherence | Q8_0 and Q4_K_M both pass the repetition-collapse gate |

<img src="niah_heatmap.png" width="640"/>

<img src="mtp_speedup.png" width="480"/>

Raw per-needle records including every run: results.jsonl.

Files

| File | Size | Pick it when |

|---|---|---|

| qwenpaw-9b-1M-MTP-Q8_0.gguf | 9.8 GB | Max quality. On a 128 GB Mac this runs the full 1M with f16 KV (~43 GB total) |

| qwenpaw-9b-1M-MTP-Q4_K_M.gguf | 5.8 GB | 32 GB GPUs. Full 1M fits with q8_0 KV (budget config); ~524K at f16 KV |

| mmproj-qwenpaw-9b.gguf | 0.9 GB | Vision, attach with --mmproj |

No other quants on purpose: the 9B is small enough that Q8_0 is the sensible default and Q4_K_M covers the budget case.

Every file, every mirror

Nothing was discontinued: every quant is one click away. Hugging Face carries the curated picks, ModelScope always carries everything, and Ollama serves ready-to-run tags.

On Ollama every tag ships with the vision tower bundled and the 1M rope metadata baked in.

| File | Size | Hugging Face | ModelScope | Ollama |

|---|---|---|---|---|

| mmproj-qwenpaw-9b.gguf | 922 MB | download | download | bundled in every tag |

| qwenpaw-9b-1M-MTP-Q4_K_M.gguf | 5.8 GB | download | download | ollama run satgeze/qwenpaw-9b-heretic-1m:q4_k_m |

| qwenpaw-9b-1M-MTP-Q8_0.gguf | 9.8 GB | download | download | ollama run satgeze/qwenpaw-9b-heretic-1m |

Run it

llama-server -m qwenpaw-9b-1M-MTP-Q8_0.gguf \
  -c 1048576 -np 1 --jinja \
  --spec-type draft-mtp --spec-draft-n-max 3 \
  --mmproj mmproj-qwenpaw-9b.gguf

Ollama (1M and vision work; no speculative decoding in Ollama yet):

FROM ./qwenpaw-9b-1M-MTP-Q8_0.gguf
RENDERER qwen3.5
PARSER qwen3.5
PARAMETER num_ctx 262144

How this was built

YaRN rope-scaling metadata (factor 4.0 over native 262,144) written into the GGUF header with gguf-py. No weight changes, no fine-tuning by us. Certification: multi-needle harness against llama-server, f16 KV only for cert runs. Method and tooling: github.com/satindergrewal/aviary-1m.

Note on upstream speed claims: SC117 reports up to 4.1x on time-scored agent benchmarks. Our controlled A/B on identical prompts measures +25 percent decode; speculative decoding cannot change model outputs at temperature 0, so treat benchmark-score deltas from MTP as timing artifacts.

How to actually use a 1M-context model

Habits that measurably help, from our RULER, hop and adherence testing across this fleet:

  1. Re-state standing instructions near the end of long prompts; recency beats depth.
  2. One big reference dump beats a long accumulated conversation. Fresh session per task.
  3. After any compaction or summarization, repeat your active rules yourself.
  4. Prefill at 500K+ takes real time on any hardware; stage your questions accordingly.
  5. Know your quant: the results tables on this card show what each quant actually holds at depth; pick the strongest one your memory allows.

Credits

Base: Qwen (Apache-2.0), including the MTP layer and vision tower. Agent fine-tune: agentscope-ai. Heretic abliteration and MTP re-injection: SC117. 1M YaRN extension and certification: SatGeze.

Mirrors: Hugging Face | ModelScope. Sister repos: Uncensored 1M collection

Run satgeze/QwenPaw-Flash-9B-heretic-1M-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models