satgeze/QwenPaw-Flash-9B-heretic-1M-GGUF overview
<div style="font family: apple system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans serif; margin bottom: 30px;" <div style="background: f5f5f7; border radius:…
Runs locally from ~879.0 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | satgeze/QwenPaw-Flash-9B-heretic-1M-GGUF |
|---|---|
| Author | satgeze |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | Qwen/Qwen3.5-9B,agentscope-ai/QwenPaw-Flash-9B,SC117/QwenPaw-Flash-9B-heretic-MTP-GGUF |
| Last modified | 2026-07-10T22:59:15.000Z |
Model README
---
language:
- en
- zh
license: apache-2.0
pipeline_tag: text-generation
base_model:
- Qwen/Qwen3.5-9B
- agentscope-ai/QwenPaw-Flash-9B
- SC117/QwenPaw-Flash-9B-heretic-MTP-GGUF
tags:
- gguf
- long-context
- yarn
- qwen3.5
- heretic
- uncensored
- mtp
- speculative-decoding
- vision
- agent
- tool-call
- llama.cpp
- ollama
---
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; margin-bottom: 30px;">
<div style="background: #f5f5f7; border-radius: 20px; padding: 36px 32px; margin-bottom: 20px; text-align: center; position: relative; overflow: hidden;">
<div style="position: absolute; top: -30px; right: -30px; width: 120px; height: 120px; background: #ffedd5; border-radius: 50%;"></div>
<div style="position: absolute; bottom: -20px; left: 40px; width: 80px; height: 80px; background: #fed7aa; border-radius: 50%;"></div>
<div style="position: absolute; top: 50%; left: -15px; width: 60px; height: 60px; background: #ffedd5; border-radius: 50%;"></div>
<div style="display: inline-flex; gap: 8px; margin-bottom: 16px; position: relative; z-index: 1;">
<span style="background: #34c759; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">1M CONTEXT</span><span style="background: #007aff; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">MTP</span><span style="background: #af52de; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">VISION</span><span style="background: #ff3b30; color: white; font-size: 11px; font-weight: 600; padding: 5px 14px; border-radius: 20px;">UNCENSORED</span>
</div>
<h1 style="margin: 0 0 8px 0; font-size: 32px; font-weight: 700; color: #1d1d1f; letter-spacing: -0.5px; border: none; position: relative; z-index: 1;">QwenPaw-Flash-9B-heretic-1M</h1>
<p style="margin: 8px 0 0 0; font-size: 14px; color: #86868b; position: relative; z-index: 1;">Agent-optimized 9B, uncensored, with a needle-verified million-token context</p>
</div>
</div>
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; display: flex; flex-direction: column; gap: 20px; margin-bottom: 30px;">
<div style="border: 1px solid #cbd5e1; border-radius: 12px; overflow: hidden; background: #ffffff; box-shadow: 0 2px 4px rgba(0,0,0,0.02);">
<div style="padding: 16px;">
<p style="margin: 0 0 12px 0; font-size: 13px; color: #334155; line-height: 1.7;"><a href="https://huggingface.co/SC117/QwenPaw-Flash-9B-heretic-MTP-GGUF" target="_blank" style="color: #c2410c; text-decoration: none; font-weight: 700;">SC117's heretic build</a> of <a href="https://huggingface.co/agentscope-ai/QwenPaw-Flash-9B" target="_blank" style="color: #c2410c; text-decoration: none; font-weight: 700;">QwenPaw-Flash-9B</a> (agent-trajectory finetune of Qwen3.5-9B) with YaRN rope scaling <b>baked into the GGUF metadata for a 1,048,576-token context window</b>, 4x the native 262,144. Weights are bit-identical to SC117's release, which already carries the official Qwen3.5-9B MTP layer injected back after the QwenPaw fine-tuning stripped it.</p>
<p style="margin: 0; font-size: 13px; color: #334155; line-height: 1.7;"><i>Uncensored (Heretic v1.3.0) · MTP baked in · Vision tower included · 1M needle-verified</i></p>
</div>
</div>
</div>
Verified on these exact files
| Capability | Result |
|---|---|
| 1M context | 10 needles per rung, depths 5 to 95 percent, temp 0, Q8_0 + f16 KV: 10/10 at every rung from 64K through 524K on an RTX 5090; 786K and 1M rungs running on a 128 GB M3 Max, card updates when they land |
| MTP speculative decoding | 217.9 to 273.2 tok/s (+25 percent), draft acceptance 0.702, output identical by construction |
| Vision | mmproj tower reads image text and identifies objects correctly |
| Coherence | Q8_0 and Q4_K_M both pass the repetition-collapse gate |
<img src="niah_heatmap.png" width="640"/>
<img src="mtp_speedup.png" width="480"/>
Raw per-needle records including every run: results.jsonl.
Files
| File | Size | Pick it when |
|---|---|---|
| qwenpaw-9b-1M-MTP-Q8_0.gguf | 9.8 GB | Max quality. On a 128 GB Mac this runs the full 1M with f16 KV (~43 GB total) |
| qwenpaw-9b-1M-MTP-Q4_K_M.gguf | 5.8 GB | 32 GB GPUs. Full 1M fits with q8_0 KV (budget config); ~524K at f16 KV |
| mmproj-qwenpaw-9b.gguf | 0.9 GB | Vision, attach with --mmproj |
No other quants on purpose: the 9B is small enough that Q8_0 is the sensible default and Q4_K_M covers the budget case.
Every file, every mirror
Nothing was discontinued: every quant is one click away. Hugging Face carries the curated picks, ModelScope always carries everything, and Ollama serves ready-to-run tags.
On Ollama every tag ships with the vision tower bundled and the 1M rope metadata baked in.
| File | Size | Hugging Face | ModelScope | Ollama |
|---|---|---|---|---|
| mmproj-qwenpaw-9b.gguf | 922 MB | download | download | bundled in every tag |
| qwenpaw-9b-1M-MTP-Q4_K_M.gguf | 5.8 GB | download | download | ollama run satgeze/qwenpaw-9b-heretic-1m:q4_k_m |
| qwenpaw-9b-1M-MTP-Q8_0.gguf | 9.8 GB | download | download | ollama run satgeze/qwenpaw-9b-heretic-1m |
Run it
llama-server -m qwenpaw-9b-1M-MTP-Q8_0.gguf \
-c 1048576 -np 1 --jinja \
--spec-type draft-mtp --spec-draft-n-max 3 \
--mmproj mmproj-qwenpaw-9b.gguf
Ollama (1M and vision work; no speculative decoding in Ollama yet):
FROM ./qwenpaw-9b-1M-MTP-Q8_0.gguf
RENDERER qwen3.5
PARSER qwen3.5
PARAMETER num_ctx 262144
How this was built
YaRN rope-scaling metadata (factor 4.0 over native 262,144) written into the GGUF header with gguf-py. No weight changes, no fine-tuning by us. Certification: multi-needle harness against llama-server, f16 KV only for cert runs. Method and tooling: github.com/satindergrewal/aviary-1m.
Note on upstream speed claims: SC117 reports up to 4.1x on time-scored agent benchmarks. Our controlled A/B on identical prompts measures +25 percent decode; speculative decoding cannot change model outputs at temperature 0, so treat benchmark-score deltas from MTP as timing artifacts.
How to actually use a 1M-context model
Habits that measurably help, from our RULER, hop and adherence testing across this fleet:
- Re-state standing instructions near the end of long prompts; recency beats depth.
- One big reference dump beats a long accumulated conversation. Fresh session per task.
- After any compaction or summarization, repeat your active rules yourself.
- Prefill at 500K+ takes real time on any hardware; stage your questions accordingly.
- Know your quant: the results tables on this card show what each quant actually holds at depth; pick the strongest one your memory allows.
Credits
Base: Qwen (Apache-2.0), including the MTP layer and vision tower. Agent fine-tune: agentscope-ai. Heretic abliteration and MTP re-injection: SC117. 1M YaRN extension and certification: SatGeze.
Mirrors: Hugging Face | ModelScope. Sister repos: Uncensored 1M collection
Run satgeze/QwenPaw-Flash-9B-heretic-1M-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models