GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

pearsonkyle/tmax-27b-imatrix-MTP-GGUF overview

<div style="font family: apple system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans serif; border: 1px solid bfdbfe; border radius: 16px; box shadow: 0 10px 15…

ggufllama.cpptext-generationquantizationquantizedimatriximportance-matrixlow-bit2-bit3-bit4-bitiq2_xsiq2_mq2_k_siq3_miq4_xsq5_k_m5-bitmtpmulti-token-predictionspeculative-decodingqwen327btool-use

Runs locally from ~8.89 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
2,982
Likes
1
Pipeline
text-generation

Repository Files & Downloads

6 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
tmax-27b-IQ2_M.ggufGGUFIQ2_M9.74 GBDownload
tmax-27b-IQ2_XS.ggufGGUFIQ2_XS8.89 GBDownload
tmax-27b-IQ3_M.ggufGGUFIQ3_M12.14 GBDownload
tmax-27b-IQ4_XS.ggufGGUFIQ4_XS14.47 GBDownload
tmax-27b-Q2_K_S.ggufGGUFQ2_K_S9.96 GBDownload
tmax-27b-Q5_K_M.ggufGGUFQ5_K_M18.33 GBDownload

Model Details

Model IDpearsonkyle/tmax-27b-imatrix-MTP-GGUF
Authorpearsonkyle
Pipelinetext-generation
Licenseapache-2.0
Base modelallenai/tmax-27b
Last modified2026-07-02T00:49:36.000Z

Model README

---

library_name: gguf

base_model:

  • allenai/tmax-27b

tags:

  • gguf
  • llama.cpp
  • text-generation
  • quantization
  • quantized
  • imatrix
  • importance-matrix
  • low-bit
  • 2-bit
  • 3-bit
  • 4-bit
  • iq2_xs
  • iq2_m
  • q2_k_s
  • iq3_m
  • iq4_xs
  • q5_k_m
  • 5-bit
  • mtp
  • multi-token-prediction
  • speculative-decoding
  • qwen3
  • 27b
  • tool-use
  • function-calling

license: apache-2.0

language:

  • en

pipeline_tag: text-generation

---

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #bfdbfe; border-radius: 16px; box-shadow: 0 10px 15px -3px rgba(0,0,0,0.05), 0 4px 6px -2px rgba(0,0,0,0.05); overflow: hidden; background: #ffffff; margin-bottom: 30px;">

<div style="background: linear-gradient(135deg, #1d4ed8 0%, #1e3a8a 100%); padding: 24px; color: white;">

<div style="display: flex; align-items: center; justify-content: space-between; flex-wrap: wrap; gap: 10px;">

<h1 style="margin: 0; font-size: 26px; font-weight: 800; display: flex; align-items: center; gap: 12px; color: white; border: none;">🧊 allenai/tmax-27b — imatrix GGUF (2 → 5-bit) + MTP</h1>

<span style="background: #f59e0b; color: #1c1917; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; text-transform: uppercase; letter-spacing: 0.5px;">imatrix + MTP</span>

</div>

</div>

<div style="display: flex; gap: 8px; flex-wrap: wrap; padding: 12px 24px; background: #f8fafc; border-bottom: 1px solid #e2e8f0;">

<span style="background: #dbeafe; color: #1e40af; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #bfdbfe;">📦 8.47 – 17.91 GiB</span>

<span style="background: #d1fae5; color: #065f46; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #a7f3d0;"> IQ2_XS → Q5_K_M · 7 quants</span>

<span style="background: #ede9fe; color: #5b21b6; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #ddd6fe;">⚡ MTP grafted (Q8) · 95.6% accept · @n=1</span>

<span style="background: #fce7f3; color: #9d174d; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #fbcfe8;">🏗️ Qwen3.6-27B derivative</span>

<span style="background: #fef3c7; color: #92400e; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #fde68a;">🏅 KLD 0.0014 · top_p 95.1% (Q5_K_M)</span>

</div>

<div style="padding: 24px; display: flex; flex-direction: column; gap: 20px;">

<div style="background: #eff6ff; border-left: 5px solid #2563eb; padding: 16px; border-radius: 0 8px 8px 0;">

<h3 style="margin: 0 0 8px 0; font-size: 15px; color: #1e3a8a; font-weight: 700; display: flex; align-items: center; gap: 6px;"><span>🧊</span> What this is</h3>

<p style="margin: 0; font-size: 13px; color: #334155; line-height: 1.7;">A bunch of importance-matrix quantizations of <b>allenai/tmax-27b</b>, from aggressive 2-bit up to near-lossless 5-bit, each calibrated with a <b>hybrid importance matrix</b> from real agentic-coding usage logs + wiki text — and each shipping a <b>Multi-Token-Prediction (MTP) draft head bundled in at Q8_0</b> for built-in speculative decoding (<code>--spec-type draft-mtp</code>). The 2-bit row is benchmarked <b>head-to-head against Qwopus3.6</b> below; pick a higher-bit row when you have the VRAM.</p>

</div>

<div style="display: grid; grid-template-columns: repeat(auto-fit, minmax(200px, 1fr)); gap: 15px; margin-top: 10px;">

<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa; box-shadow: inset 0 2px 4px rgba(0,0,0,0.02);"><span style="font-weight: 700; color: #1e3a8a; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase; letter-spacing: 0.5px;">📉 ~2–6× smaller on disk</span><span style="font-size: 13px; color: #4b5563; line-height: 1.5;">8.47–17.91 GiB on disk (incl. the bundled MTP head) vs ~54 GiB for FP16. Tuned for English + Python agentic-coding workloads (see calibration scope below).</span></div>

<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa; box-shadow: inset 0 2px 4px rgba(0,0,0,0.02);"><span style="font-weight: 700; color: #1e3a8a; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase; letter-spacing: 0.5px;">⚡ MTP speculative decoding</span><span style="font-size: 13px; color: #4b5563; line-height: 1.5;">Grafted Qwopus3.6 nextn head at Q8_0, 95.6% draft acceptance at n=1 on IQ4_XS. Built-in — no separate draft model required.</span></div>

<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa; box-shadow: inset 0 2px 4px rgba(0,0,0,0.02);"><span style="font-weight: 700; color: #1e3a8a; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase; letter-spacing: 0.5px;">🤖 SWE-rebench 70% pass</span><span style="font-size: 13px; color: #4b5563; line-height: 1.5;">IQ2_XS / Q2_K_S / IQ3_M / IQ4_XS all resolve 7/10 SWE-rebench instances — strong even at 2-bit. Full results in §1.</span></div>

</div>

</div>

</div>

🧰 1. Files & comparison

| | Q2_K (plain) | IQ2_XS | IQ2_M | Q2_K_S | IQ3_M | IQ4_XS | Q5_K_M |

|---|---|---|---|---|---|---|---|

| Technique | plain | hybrid imatrix | hybrid imatrix | hybrid imatrix | hybrid imatrix | hybrid imatrix | hybrid imatrix |

| File | Q2_K | IQ2_XS | IQ2_M | Q2_K_S | IQ3_M | IQ4_XS | Q5_K_M |

| Quality | ❌ | ⭐⭐ | ⭐ | ⭐⭐ | ⭐⭐⭐ | ⭐⭐⭐ | ⭐⭐⭐ |

| Size (GiB) | 9.98 | 8.47 | 9.32 | 9.54 | 11.72 | 14.05 | 17.91 |

| BPW | 3.186 | 2.704 | 2.976 | 3.048 | 3.742 | 4.486 | 5.720 |

| PPL (general) | 7.6005 | 25.5923 | 20.1178 | 18.1105 | 20.0037 | 13.5009 | 14.1304 |

| KLD med (general) | 0.1727 | 0.1345 | 0.0767 | 0.0825 | 0.0265 | 0.0059 | 0.0014 |

| top_p (general) | 73.03% | 72.89% | 78.21% | 78.34% | 83.72% | 91.50% | 95.04% |

Head-to-head vs Qwopus3.6-27B-Coder

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #bfdbfe; border-radius: 12px; overflow: hidden; background: #ffffff; margin: 16px 0;">

<div style="background: linear-gradient(135deg, #2563eb 0%, #1e3a8a 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px;">⚔️ IQ2_M head-to-head</div>

<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">

<p style="margin: 0 0 10px 0;">Both models are Qwen3.6-27B derivatives benched on the <b>identical</b> general eval corpus and the <b>same</b> llama.cpp build (lower KLD / higher top_p is better). Qwopus3.6 numbers are re-measured here, not quoted from its README, so the comparison is apples-to-apples.</p>

<table style="width:100%; border-collapse:collapse; font-size:12px; margin:0;">

<thead><tr style="background:#eff6ff;"><th style="padding:8px 10px; border:1px solid #bfdbfe; text-align:left; color:#1e3a8a;">Metric (IQ2_M)</th><th style="padding:8px 10px; border:1px solid #bfdbfe; text-align:left; color:#1e3a8a;">tmax-27b</th><th style="padding:8px 10px; border:1px solid #bfdbfe; text-align:left; color:#1e3a8a;">Qwopus3.6-27B-Coder</th></tr></thead>

<tbody>

<tr><td style="padding:8px 10px; border:1px solid #bfdbfe;"><b>KLD med (general)</b></td><td style="padding:8px 10px; border:1px solid #bfdbfe;">0.0767</td><td style="padding:8px 10px; border:1px solid #bfdbfe;">0.0535</td></tr>

<tr><td style="padding:8px 10px; border:1px solid #bfdbfe;"><b>top_p (general)</b></td><td style="padding:8px 10px; border:1px solid #bfdbfe;">78.21%</td><td style="padding:8px 10px; border:1px solid #bfdbfe;">83.23%</td></tr>

<tr><td style="padding:8px 10px; border:1px solid #bfdbfe;"><b>PPL (general)</b></td><td style="padding:8px 10px; border:1px solid #bfdbfe;">20.1178</td><td style="padding:8px 10px; border:1px solid #bfdbfe;">8.5961</td></tr>

</tbody>

</table>

</div>

</div>

Agentic coding across quants (SWE-rebench)

Every quant run as a coding agent (mini-swe-agent) over the **same 10 held-out

SWE-rebench instances**, one clean Docker container each. pass_rate = fraction whose

patch makes the gold FAIL_TO_PASS tests pass; patch_rate = fraction that produced a

non-empty diff. Token/step counts are per instance.

| Metric | Q2_K | IQ2_XS | IQ2_M | Q2_K_S | IQ3_M | IQ4_XS |

|---|---|---|---|---|---|---|

| pass_rate | 50% | 70% | 60% | 70% | 70% | 70% |

| patch_rate | 100% | 100% | 100% | 100% | 100% | 100% |

| resolved | 5/10 | 7/10 | 6/10 | 7/10 | 7/10 | 7/10 |

| tokens | 621,931 | 784,972 | 596,658 | 529,560 | 770,113 | 791,474 |

| steps | 38.7 | 49.8 | 40.9 | 37.1 | 47.5 | 48.3 |

| tool-err | 11% | 9% | 10% | 12% | 10% | 9% |

> Same 10-instance holdout as the head-to-head, no speculative decoding. At 2-bit

> the quants stay surprisingly capable agents; higher-bit rows trade size for headroom.

> Only one repetition is reported above and uncertainties on the pass rate can vary ~5-10%

> due to sampling variance.

>

> ⚠️ _These agentic numbers were measured on the pre-recalibration quants (see the

> "Recalibrated" note in §1). General-eval fidelity is unchanged by the recalibration, so

> they remain broadly indicative, but a fresh agentic run on the recalibrated GGUFs has not

> yet been done._

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #fde68a; border-radius: 12px; background: #fffbeb; padding: 16px; margin: 16px 0; color: #92400e; font-size: 13px; line-height: 1.7;">

<b>⚠️ Caveat.</b> The 2-bit rows (Q2_K / IQ2_*) are aggressive — strong for their size

but rougher on tmax than on Qwopus3.6 (see the head-to-head). Prefer <b>IQ3_M</b> or

<b>IQ4_XS</b> when you have the VRAM; reach for the 2-bit rows only when memory is the

binding constraint.

</div>

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #bfdbfe; border-radius: 12px; background: #eff6ff; padding: 16px; margin: 16px 0; color: #1e3a8a; font-size: 13px; line-height: 1.7;">

<b>📋 Calibration scope — English &amp; Python, agentic coding.</b> The importance

matrix (and the windowed packing that shaped it) was calibrated on <b>real

agentic-coding sessions that are overwhelmingly English-language and

Python-centric</b> (Claude Code, opencode, qwen code). Expect <b>weaker fidelity on

other natural languages, non-Python ecosystems, and general-chat workloads</b>.

</div>

---

🔬 2. How they were made

<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; display: flex; flex-direction: column; gap: 16px; margin-bottom: 24px;">

<div style="border: 1px solid #bfdbfe; border-radius: 12px; overflow: hidden; background: #ffffff;">

<div style="background: linear-gradient(135deg, #2563eb 0%, #1e3a8a 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px;">🧮 2.1 Hybrid importance matrix</div>

<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">

<p style="margin: 0 0 10px 0;">At 2 bits the quantizer must decide <i>where</i> to spend its limited precision. An <b>importance matrix</b> measures, per input channel, how much each channel drives a layer's output on a calibration corpus, and tells <code>llama-quantize</code> to preserve the high-impact channels. This release blends activation energy <code>E[a²]</code> with weight-column energy <code>‖W[:, c]‖² · E[a²]</code>, collected at ctx=4096. Linear-attention / SSM tensors (Qwen3.6 is a hybrid architecture) pass through with raw <code>E[a²]</code>. The output is a standard GGUF with no runtime overhead.</p>

</div>

</div>

<div style="border: 1px solid #ddd6fe; border-radius: 12px; overflow: hidden; background: #ffffff;">

<div style="background: linear-gradient(135deg, #7c3aed 0%, #5b21b6 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px;">⚡ 2.2 Bundled MTP (grafted)</div>

<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">

<p style="margin: 0 0 10px 0;">tmax-27b's released checkpoint dropped the Qwen3.6 MTP draft head (its <code>mtp_num_hidden_layers</code> flag is vestigial). But tmax and <b>Jackrong/Qwopus3.6-27B-Coder</b> are <b>byte-for-byte architecture-identical</b> finetunes of the same Qwen3.6-27B base (hidden 5120, 64 layers, 24/4 GQA, head_dim 256, vocab 248320, identical hybrid attention layout), so Qwopus3.6's trained nextn head is dimensionally compatible. We <b>graft it onto tmax's trunk</b> as <code>blk.64</code> and keep it near-lossless at <b>Q8_0</b> in every quant; llama.cpp serves it as built-in speculative decoding (<code>--spec-type draft-mtp</code>).</p>

<p style="margin: 0 0 10px 0;">Despite being a <i>cross-finetune</i> graft, it drafts at high acceptance — measured on the IQ4_XS trunk over agentic-coding prompts (draft acceptance = fraction of drafted tokens the trunk confirms):</p>

<table style="width:100%; border-collapse:collapse; font-size:12px; margin:0;">

<thead><tr style="background:#f5f3ff;"><th style="padding:8px 10px; border:1px solid #ddd6fe; text-align:left; color:#5b21b6;"><code>--spec-draft-n-max</code></th><th style="padding:8px 10px; border:1px solid #ddd6fe; text-align:left; color:#5b21b6;">Draft acceptance</th></tr></thead>

<tbody>

<tr><td style="padding:8px 10px; border:1px solid #ddd6fe;"><b>1</b> (recommended)</td><td style="padding:8px 10px; border:1px solid #ddd6fe;"><b>95.6%</b></td></tr>

<tr><td style="padding:8px 10px; border:1px solid #ddd6fe;">2</td><td style="padding:8px 10px; border:1px solid #ddd6fe;">84.7%</td></tr>

<tr><td style="padding:8px 10px; border:1px solid #ddd6fe;">3</td><td style="padding:8px 10px; border:1px solid #ddd6fe;">74.6%</td></tr>

<tr><td style="padding:8px 10px; border:1px solid #ddd6fe;">4</td><td style="padding:8px 10px; border:1px solid #ddd6fe;">74.0%</td></tr>

</tbody>

</table>

<p style="margin: 10px 0 0 0;"><b><code>--spec-draft-n-max 1</code> is recommended</b> (one nextn layer; higher <code>n</code> drafts longer chains at falling acceptance). Speedup is GPU-bandwidth-dependent — spec decoding helps most on memory-bound CUDA GPUs; the win is smaller on unified-memory/Metal. Acceptance is also prompt-dependent (higher on predictable code, lower on open-ended text).</p>

</div>

</div>

<div style="border: 1px solid #bfdbfe; border-radius: 12px; overflow: hidden; background: #ffffff;">

<div style="background: linear-gradient(135deg, #2563eb 0%, #1e3a8a 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px;">📚 2.3 Calibration & evaluation data</div>

<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">

<p style="margin: 0 0 10px 0;">Calibration and every eval corpus are disjoint by construction — the tool-call eval is the held-out 10% of sessions, windowed exactly like calibration but never seen by it — so §1 measures generalization, not fit. All shipped under <code>calibration_data/</code>.</p>

<table style="width:100%; border-collapse:collapse; font-size:12px; margin:0;">

<thead><tr style="background:#eff6ff;"><th style="padding:8px 10px; border:1px solid #bfdbfe; text-align:left; color:#1e3a8a;">Corpus</th><th style="padding:8px 10px; border:1px solid #bfdbfe; text-align:left; color:#1e3a8a;">Source</th><th style="padding:8px 10px; border:1px solid #bfdbfe; text-align:left; color:#1e3a8a;">Used for</th></tr></thead>

<tbody>

<tr><td style="padding:8px 10px; border:1px solid #bfdbfe;"><b>Calibration</b></td><td style="padding:8px 10px; border:1px solid #bfdbfe;">~500k tokens of usage-log text (windowed) + all of <code>wiki.test.raw</code></td><td style="padding:8px 10px; border:1px solid #bfdbfe;">hybrid imatrix collection</td></tr>

<tr><td style="padding:8px 10px; border:1px solid #bfdbfe;"><b>Eval — tools</b> (in-distribution)</td><td style="padding:8px 10px; border:1px solid #bfdbfe;">held-out logtrain session slice (10%), windowed like calibration but disjoint from it</td><td style="padding:8px 10px; border:1px solid #bfdbfe;"><b>§1 tools columns (PPL · KLD · top_p)</b></td></tr>

<tr><td style="padding:8px 10px; border:1px solid #bfdbfe;"><b>Eval — general</b></td><td style="padding:8px 10px; border:1px solid #bfdbfe;"><code>combined_en_tiny</code> (broad English) from <code>eaddario/imatrix-calibration</code></td><td style="padding:8px 10px; border:1px solid #bfdbfe;"><b>§1 general columns (PPL · KLD · top_p)</b></td></tr>

</tbody>

</table>

</div>

</div>

</div>

---

🚀 3. Usage

Quick start with Ollama

ollama run hf.co/pearsonkyle/tmax-27b-imatrix-MTP-GGUF:IQ2_M
# also: :Q2_K_S  ·  :IQ2_XS  ·  :Q2_K  ·  :IQ3_M  ·  :IQ4_XS  ·  :Q5_K_M

Building llama.cpp from source (GPU)

apt-get update && apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y
git clone https://github.com/ggml-org/llama.cpp
cmake llama.cpp -B llama.cpp/build -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=ON   # -DGGML_CUDA=OFF for CPU/Metal
cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-server
cp llama.cpp/build/bin/llama-* llama.cpp/

> MTP needs a recent llama.cpp--spec-type draft-mtp was merged in 2026-06. Build from current master.

Running the server with MTP speculative decoding

./llama-server \
    --model tmax-27b-IQ4_XS.gguf \
    --ctx-size 16384 --n-gpu-layers 999 \
    --spec-type draft-mtp --spec-draft-n-max 1 \
    --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 \
    --host 0.0.0.0 --port 1234

> Drop the two --spec-* flags to run without MTP. --spec-draft-n-max 1 is optimal (one nextn layer). If the model emits ` reasoning, also pass --chat-template-kwargs '{"enable_thinking": false}' so the answer lands in content`.

Querying via the OpenAI-compatible API

import json, urllib.request

def ask(content, max_tokens=256):
    body = {
        "messages": [{"role": "user", "content": content}],
        "max_tokens": max_tokens,
        # tmax may emit  reasoning. Set enable_thinking False
        # (or raise max_tokens) so the answer lands in "content".
        "chat_template_kwargs": {"enable_thinking": False},
    }
    req = urllib.request.Request("http://127.0.0.1:1234/v1/chat/completions",
                                 json.dumps(body).encode(),
                                 {"Content-Type": "application/json"})
    return json.loads(urllib.request.urlopen(req).read())["choices"][0]["message"]["content"]

print(ask("Write a Python function that reverses a linked list."))

---

🪪 4. License & attribution

Run pearsonkyle/tmax-27b-imatrix-MTP-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models