pearsonkyle/tmax-27b-imatrix-MTP-GGUF overview
<div style="font family: apple system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans serif; border: 1px solid bfdbfe; border radius: 16px; box shadow: 0 10px 15…
Runs locally from ~8.89 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| tmax-27b-IQ2_M.gguf | GGUF | IQ2_M | 9.74 GB | Download |
| tmax-27b-IQ2_XS.gguf | GGUF | IQ2_XS | 8.89 GB | Download |
| tmax-27b-IQ3_M.gguf | GGUF | IQ3_M | 12.14 GB | Download |
| tmax-27b-IQ4_XS.gguf | GGUF | IQ4_XS | 14.47 GB | Download |
| tmax-27b-Q2_K_S.gguf | GGUF | Q2_K_S | 9.96 GB | Download |
| tmax-27b-Q5_K_M.gguf | GGUF | Q5_K_M | 18.33 GB | Download |
Model Details
| Model ID | pearsonkyle/tmax-27b-imatrix-MTP-GGUF |
|---|---|
| Author | pearsonkyle |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | allenai/tmax-27b |
| Last modified | 2026-07-02T00:49:36.000Z |
Model README
---
library_name: gguf
base_model:
- allenai/tmax-27b
tags:
- gguf
- llama.cpp
- text-generation
- quantization
- quantized
- imatrix
- importance-matrix
- low-bit
- 2-bit
- 3-bit
- 4-bit
- iq2_xs
- iq2_m
- q2_k_s
- iq3_m
- iq4_xs
- q5_k_m
- 5-bit
- mtp
- multi-token-prediction
- speculative-decoding
- qwen3
- 27b
- tool-use
- function-calling
license: apache-2.0
language:
- en
pipeline_tag: text-generation
---
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #bfdbfe; border-radius: 16px; box-shadow: 0 10px 15px -3px rgba(0,0,0,0.05), 0 4px 6px -2px rgba(0,0,0,0.05); overflow: hidden; background: #ffffff; margin-bottom: 30px;">
<div style="background: linear-gradient(135deg, #1d4ed8 0%, #1e3a8a 100%); padding: 24px; color: white;">
<div style="display: flex; align-items: center; justify-content: space-between; flex-wrap: wrap; gap: 10px;">
<h1 style="margin: 0; font-size: 26px; font-weight: 800; display: flex; align-items: center; gap: 12px; color: white; border: none;">🧊 allenai/tmax-27b — imatrix GGUF (2 → 5-bit) + MTP</h1>
<span style="background: #f59e0b; color: #1c1917; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; text-transform: uppercase; letter-spacing: 0.5px;">imatrix + MTP</span>
</div>
</div>
<div style="display: flex; gap: 8px; flex-wrap: wrap; padding: 12px 24px; background: #f8fafc; border-bottom: 1px solid #e2e8f0;">
<span style="background: #dbeafe; color: #1e40af; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #bfdbfe;">📦 8.47 – 17.91 GiB</span>
<span style="background: #d1fae5; color: #065f46; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #a7f3d0;"> IQ2_XS → Q5_K_M · 7 quants</span>
<span style="background: #ede9fe; color: #5b21b6; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #ddd6fe;">⚡ MTP grafted (Q8) · 95.6% accept · @n=1</span>
<span style="background: #fce7f3; color: #9d174d; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #fbcfe8;">🏗️ Qwen3.6-27B derivative</span>
<span style="background: #fef3c7; color: #92400e; font-size: 11px; font-weight: 700; padding: 4px 10px; border-radius: 20px; border: 1px solid #fde68a;">🏅 KLD 0.0014 · top_p 95.1% (Q5_K_M)</span>
</div>
<div style="padding: 24px; display: flex; flex-direction: column; gap: 20px;">
<div style="background: #eff6ff; border-left: 5px solid #2563eb; padding: 16px; border-radius: 0 8px 8px 0;">
<h3 style="margin: 0 0 8px 0; font-size: 15px; color: #1e3a8a; font-weight: 700; display: flex; align-items: center; gap: 6px;"><span>🧊</span> What this is</h3>
<p style="margin: 0; font-size: 13px; color: #334155; line-height: 1.7;">A bunch of importance-matrix quantizations of <b>allenai/tmax-27b</b>, from aggressive 2-bit up to near-lossless 5-bit, each calibrated with a <b>hybrid importance matrix</b> from real agentic-coding usage logs + wiki text — and each shipping a <b>Multi-Token-Prediction (MTP) draft head bundled in at Q8_0</b> for built-in speculative decoding (<code>--spec-type draft-mtp</code>). The 2-bit row is benchmarked <b>head-to-head against Qwopus3.6</b> below; pick a higher-bit row when you have the VRAM.</p>
</div>
<div style="display: grid; grid-template-columns: repeat(auto-fit, minmax(200px, 1fr)); gap: 15px; margin-top: 10px;">
<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa; box-shadow: inset 0 2px 4px rgba(0,0,0,0.02);"><span style="font-weight: 700; color: #1e3a8a; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase; letter-spacing: 0.5px;">📉 ~2–6× smaller on disk</span><span style="font-size: 13px; color: #4b5563; line-height: 1.5;">8.47–17.91 GiB on disk (incl. the bundled MTP head) vs ~54 GiB for FP16. Tuned for English + Python agentic-coding workloads (see calibration scope below).</span></div>
<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa; box-shadow: inset 0 2px 4px rgba(0,0,0,0.02);"><span style="font-weight: 700; color: #1e3a8a; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase; letter-spacing: 0.5px;">⚡ MTP speculative decoding</span><span style="font-size: 13px; color: #4b5563; line-height: 1.5;">Grafted Qwopus3.6 nextn head at Q8_0, 95.6% draft acceptance at n=1 on IQ4_XS. Built-in — no separate draft model required.</span></div>
<div style="border: 1px solid #e2e8f0; padding: 14px; border-radius: 8px; background: #fafafa; box-shadow: inset 0 2px 4px rgba(0,0,0,0.02);"><span style="font-weight: 700; color: #1e3a8a; font-size: 12px; display: block; margin-bottom: 6px; text-transform: uppercase; letter-spacing: 0.5px;">🤖 SWE-rebench 70% pass</span><span style="font-size: 13px; color: #4b5563; line-height: 1.5;">IQ2_XS / Q2_K_S / IQ3_M / IQ4_XS all resolve 7/10 SWE-rebench instances — strong even at 2-bit. Full results in §1.</span></div>
</div>
</div>
</div>
🧰 1. Files & comparison
| | Q2_K (plain) | IQ2_XS | IQ2_M | Q2_K_S | IQ3_M | IQ4_XS | Q5_K_M |
|---|---|---|---|---|---|---|---|
| Technique | plain | hybrid imatrix | hybrid imatrix | hybrid imatrix | hybrid imatrix | hybrid imatrix | hybrid imatrix |
| File | Q2_K | IQ2_XS | IQ2_M | Q2_K_S | IQ3_M | IQ4_XS | Q5_K_M |
| Quality | ❌ | ⭐⭐ | ⭐ | ⭐⭐ | ⭐⭐⭐ | ⭐⭐⭐ | ⭐⭐⭐ |
| Size (GiB) | 9.98 | 8.47 | 9.32 | 9.54 | 11.72 | 14.05 | 17.91 |
| BPW | 3.186 | 2.704 | 2.976 | 3.048 | 3.742 | 4.486 | 5.720 |
| PPL (general) | 7.6005 | 25.5923 | 20.1178 | 18.1105 | 20.0037 | 13.5009 | 14.1304 |
| KLD med (general) | 0.1727 | 0.1345 | 0.0767 | 0.0825 | 0.0265 | 0.0059 | 0.0014 |
| top_p (general) | 73.03% | 72.89% | 78.21% | 78.34% | 83.72% | 91.50% | 95.04% |
Head-to-head vs Qwopus3.6-27B-Coder
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #bfdbfe; border-radius: 12px; overflow: hidden; background: #ffffff; margin: 16px 0;">
<div style="background: linear-gradient(135deg, #2563eb 0%, #1e3a8a 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px;">⚔️ IQ2_M head-to-head</div>
<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">
<p style="margin: 0 0 10px 0;">Both models are Qwen3.6-27B derivatives benched on the <b>identical</b> general eval corpus and the <b>same</b> llama.cpp build (lower KLD / higher top_p is better). Qwopus3.6 numbers are re-measured here, not quoted from its README, so the comparison is apples-to-apples.</p>
<table style="width:100%; border-collapse:collapse; font-size:12px; margin:0;">
<thead><tr style="background:#eff6ff;"><th style="padding:8px 10px; border:1px solid #bfdbfe; text-align:left; color:#1e3a8a;">Metric (IQ2_M)</th><th style="padding:8px 10px; border:1px solid #bfdbfe; text-align:left; color:#1e3a8a;">tmax-27b</th><th style="padding:8px 10px; border:1px solid #bfdbfe; text-align:left; color:#1e3a8a;">Qwopus3.6-27B-Coder</th></tr></thead>
<tbody>
<tr><td style="padding:8px 10px; border:1px solid #bfdbfe;"><b>KLD med (general)</b></td><td style="padding:8px 10px; border:1px solid #bfdbfe;">0.0767</td><td style="padding:8px 10px; border:1px solid #bfdbfe;">0.0535</td></tr>
<tr><td style="padding:8px 10px; border:1px solid #bfdbfe;"><b>top_p (general)</b></td><td style="padding:8px 10px; border:1px solid #bfdbfe;">78.21%</td><td style="padding:8px 10px; border:1px solid #bfdbfe;">83.23%</td></tr>
<tr><td style="padding:8px 10px; border:1px solid #bfdbfe;"><b>PPL (general)</b></td><td style="padding:8px 10px; border:1px solid #bfdbfe;">20.1178</td><td style="padding:8px 10px; border:1px solid #bfdbfe;">8.5961</td></tr>
</tbody>
</table>
</div>
</div>
Agentic coding across quants (SWE-rebench)
Every quant run as a coding agent (mini-swe-agent) over the **same 10 held-out
SWE-rebench instances**, one clean Docker container each. pass_rate = fraction whose
patch makes the gold FAIL_TO_PASS tests pass; patch_rate = fraction that produced a
non-empty diff. Token/step counts are per instance.
| Metric | Q2_K | IQ2_XS | IQ2_M | Q2_K_S | IQ3_M | IQ4_XS |
|---|---|---|---|---|---|---|
| pass_rate | 50% | 70% | 60% | 70% | 70% | 70% |
| patch_rate | 100% | 100% | 100% | 100% | 100% | 100% |
| resolved | 5/10 | 7/10 | 6/10 | 7/10 | 7/10 | 7/10 |
| tokens | 621,931 | 784,972 | 596,658 | 529,560 | 770,113 | 791,474 |
| steps | 38.7 | 49.8 | 40.9 | 37.1 | 47.5 | 48.3 |
| tool-err | 11% | 9% | 10% | 12% | 10% | 9% |
> Same 10-instance holdout as the head-to-head, no speculative decoding. At 2-bit
> the quants stay surprisingly capable agents; higher-bit rows trade size for headroom.
> Only one repetition is reported above and uncertainties on the pass rate can vary ~5-10%
> due to sampling variance.
>
> ⚠️ _These agentic numbers were measured on the pre-recalibration quants (see the
> "Recalibrated" note in §1). General-eval fidelity is unchanged by the recalibration, so
> they remain broadly indicative, but a fresh agentic run on the recalibrated GGUFs has not
> yet been done._
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #fde68a; border-radius: 12px; background: #fffbeb; padding: 16px; margin: 16px 0; color: #92400e; font-size: 13px; line-height: 1.7;">
<b>⚠️ Caveat.</b> The 2-bit rows (Q2_K / IQ2_*) are aggressive — strong for their size
but rougher on tmax than on Qwopus3.6 (see the head-to-head). Prefer <b>IQ3_M</b> or
<b>IQ4_XS</b> when you have the VRAM; reach for the 2-bit rows only when memory is the
binding constraint.
</div>
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; border: 1px solid #bfdbfe; border-radius: 12px; background: #eff6ff; padding: 16px; margin: 16px 0; color: #1e3a8a; font-size: 13px; line-height: 1.7;">
<b>📋 Calibration scope — English & Python, agentic coding.</b> The importance
matrix (and the windowed packing that shaped it) was calibrated on <b>real
agentic-coding sessions that are overwhelmingly English-language and
Python-centric</b> (Claude Code, opencode, qwen code). Expect <b>weaker fidelity on
other natural languages, non-Python ecosystems, and general-chat workloads</b>.
</div>
---
🔬 2. How they were made
<div style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, sans-serif; display: flex; flex-direction: column; gap: 16px; margin-bottom: 24px;">
<div style="border: 1px solid #bfdbfe; border-radius: 12px; overflow: hidden; background: #ffffff;">
<div style="background: linear-gradient(135deg, #2563eb 0%, #1e3a8a 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px;">🧮 2.1 Hybrid importance matrix</div>
<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">
<p style="margin: 0 0 10px 0;">At 2 bits the quantizer must decide <i>where</i> to spend its limited precision. An <b>importance matrix</b> measures, per input channel, how much each channel drives a layer's output on a calibration corpus, and tells <code>llama-quantize</code> to preserve the high-impact channels. This release blends activation energy <code>E[a²]</code> with weight-column energy <code>‖W[:, c]‖² · E[a²]</code>, collected at ctx=4096. Linear-attention / SSM tensors (Qwen3.6 is a hybrid architecture) pass through with raw <code>E[a²]</code>. The output is a standard GGUF with no runtime overhead.</p>
</div>
</div>
<div style="border: 1px solid #ddd6fe; border-radius: 12px; overflow: hidden; background: #ffffff;">
<div style="background: linear-gradient(135deg, #7c3aed 0%, #5b21b6 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px;">⚡ 2.2 Bundled MTP (grafted)</div>
<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">
<p style="margin: 0 0 10px 0;">tmax-27b's released checkpoint dropped the Qwen3.6 MTP draft head (its <code>mtp_num_hidden_layers</code> flag is vestigial). But tmax and <b>Jackrong/Qwopus3.6-27B-Coder</b> are <b>byte-for-byte architecture-identical</b> finetunes of the same Qwen3.6-27B base (hidden 5120, 64 layers, 24/4 GQA, head_dim 256, vocab 248320, identical hybrid attention layout), so Qwopus3.6's trained nextn head is dimensionally compatible. We <b>graft it onto tmax's trunk</b> as <code>blk.64</code> and keep it near-lossless at <b>Q8_0</b> in every quant; llama.cpp serves it as built-in speculative decoding (<code>--spec-type draft-mtp</code>).</p>
<p style="margin: 0 0 10px 0;">Despite being a <i>cross-finetune</i> graft, it drafts at high acceptance — measured on the IQ4_XS trunk over agentic-coding prompts (draft acceptance = fraction of drafted tokens the trunk confirms):</p>
<table style="width:100%; border-collapse:collapse; font-size:12px; margin:0;">
<thead><tr style="background:#f5f3ff;"><th style="padding:8px 10px; border:1px solid #ddd6fe; text-align:left; color:#5b21b6;"><code>--spec-draft-n-max</code></th><th style="padding:8px 10px; border:1px solid #ddd6fe; text-align:left; color:#5b21b6;">Draft acceptance</th></tr></thead>
<tbody>
<tr><td style="padding:8px 10px; border:1px solid #ddd6fe;"><b>1</b> (recommended)</td><td style="padding:8px 10px; border:1px solid #ddd6fe;"><b>95.6%</b></td></tr>
<tr><td style="padding:8px 10px; border:1px solid #ddd6fe;">2</td><td style="padding:8px 10px; border:1px solid #ddd6fe;">84.7%</td></tr>
<tr><td style="padding:8px 10px; border:1px solid #ddd6fe;">3</td><td style="padding:8px 10px; border:1px solid #ddd6fe;">74.6%</td></tr>
<tr><td style="padding:8px 10px; border:1px solid #ddd6fe;">4</td><td style="padding:8px 10px; border:1px solid #ddd6fe;">74.0%</td></tr>
</tbody>
</table>
<p style="margin: 10px 0 0 0;"><b><code>--spec-draft-n-max 1</code> is recommended</b> (one nextn layer; higher <code>n</code> drafts longer chains at falling acceptance). Speedup is GPU-bandwidth-dependent — spec decoding helps most on memory-bound CUDA GPUs; the win is smaller on unified-memory/Metal. Acceptance is also prompt-dependent (higher on predictable code, lower on open-ended text).</p>
</div>
</div>
<div style="border: 1px solid #bfdbfe; border-radius: 12px; overflow: hidden; background: #ffffff;">
<div style="background: linear-gradient(135deg, #2563eb 0%, #1e3a8a 100%); padding: 12px 16px; color: white; font-weight: 700; font-size: 14px;">📚 2.3 Calibration & evaluation data</div>
<div style="padding: 16px; font-size: 13px; color: #334155; line-height: 1.7;">
<p style="margin: 0 0 10px 0;">Calibration and every eval corpus are disjoint by construction — the tool-call eval is the held-out 10% of sessions, windowed exactly like calibration but never seen by it — so §1 measures generalization, not fit. All shipped under <code>calibration_data/</code>.</p>
<table style="width:100%; border-collapse:collapse; font-size:12px; margin:0;">
<thead><tr style="background:#eff6ff;"><th style="padding:8px 10px; border:1px solid #bfdbfe; text-align:left; color:#1e3a8a;">Corpus</th><th style="padding:8px 10px; border:1px solid #bfdbfe; text-align:left; color:#1e3a8a;">Source</th><th style="padding:8px 10px; border:1px solid #bfdbfe; text-align:left; color:#1e3a8a;">Used for</th></tr></thead>
<tbody>
<tr><td style="padding:8px 10px; border:1px solid #bfdbfe;"><b>Calibration</b></td><td style="padding:8px 10px; border:1px solid #bfdbfe;">~500k tokens of usage-log text (windowed) + all of <code>wiki.test.raw</code></td><td style="padding:8px 10px; border:1px solid #bfdbfe;">hybrid imatrix collection</td></tr>
<tr><td style="padding:8px 10px; border:1px solid #bfdbfe;"><b>Eval — tools</b> (in-distribution)</td><td style="padding:8px 10px; border:1px solid #bfdbfe;">held-out logtrain session slice (10%), windowed like calibration but disjoint from it</td><td style="padding:8px 10px; border:1px solid #bfdbfe;"><b>§1 tools columns (PPL · KLD · top_p)</b></td></tr>
<tr><td style="padding:8px 10px; border:1px solid #bfdbfe;"><b>Eval — general</b></td><td style="padding:8px 10px; border:1px solid #bfdbfe;"><code>combined_en_tiny</code> (broad English) from <code>eaddario/imatrix-calibration</code></td><td style="padding:8px 10px; border:1px solid #bfdbfe;"><b>§1 general columns (PPL · KLD · top_p)</b></td></tr>
</tbody>
</table>
</div>
</div>
</div>
---
🚀 3. Usage
Quick start with Ollama
ollama run hf.co/pearsonkyle/tmax-27b-imatrix-MTP-GGUF:IQ2_M
# also: :Q2_K_S · :IQ2_XS · :Q2_K · :IQ3_M · :IQ4_XS · :Q5_K_M
Building llama.cpp from source (GPU)
apt-get update && apt-get install pciutils build-essential cmake curl libcurl4-openssl-dev -y
git clone https://github.com/ggml-org/llama.cpp
cmake llama.cpp -B llama.cpp/build -DBUILD_SHARED_LIBS=OFF -DGGML_CUDA=ON # -DGGML_CUDA=OFF for CPU/Metal
cmake --build llama.cpp/build --config Release -j --clean-first --target llama-cli llama-server
cp llama.cpp/build/bin/llama-* llama.cpp/
> MTP needs a recent llama.cpp — --spec-type draft-mtp was merged in 2026-06. Build from current master.
Running the server with MTP speculative decoding
./llama-server \
--model tmax-27b-IQ4_XS.gguf \
--ctx-size 16384 --n-gpu-layers 999 \
--spec-type draft-mtp --spec-draft-n-max 1 \
--flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 \
--host 0.0.0.0 --port 1234
> Drop the two --spec-* flags to run without MTP. --spec-draft-n-max 1 is optimal (one nextn layer). If the model emits ` reasoning, also pass --chat-template-kwargs '{"enable_thinking": false}' so the answer lands in content`.
Querying via the OpenAI-compatible API
import json, urllib.request
def ask(content, max_tokens=256):
body = {
"messages": [{"role": "user", "content": content}],
"max_tokens": max_tokens,
# tmax may emit reasoning. Set enable_thinking False
# (or raise max_tokens) so the answer lands in "content".
"chat_template_kwargs": {"enable_thinking": False},
}
req = urllib.request.Request("http://127.0.0.1:1234/v1/chat/completions",
json.dumps(body).encode(),
{"Content-Type": "application/json"})
return json.loads(urllib.request.urlopen(req).read())["choices"][0]["message"]["content"]
print(ask("Write a Python function that reverses a linked list."))
---
🪪 4. License & attribution
- License: Apache-2.0 — inherited from the base model
allenai/tmax-27b. Ai2 asks that use follow its Responsible Use Guidelines. - Base weights:
allenai/tmax-27b(Qwen3.6-27B derivative; released checkpoint is the text trunk only — vision + MTP head dropped). - MTP draft head grafted from:
pearsonkyle/Qwopus3.6-27B-Coder-imatrix-2bit-MTP-GGUF(also a Qwen3.6-27B finetune; same head it ships natively). Also the head-to-head baseline. - Calibration + quantization performed locally with Quant-Tuner.
Run pearsonkyle/tmax-27b-imatrix-MTP-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models