GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

plunderstruck/Nex-N2-mini-ROCmFP4-GGUF overview

<div style="border:2px solid currentColor; font family:ui monospace,'SF Mono','Cascadia Mono',Consolas,'Liberation Mono',monospace;" <div style="border bottom:…

ggufrocmfp4qwen3.5nex-n2coderagenticmoeimatrixstrix-haloamdrocmvulkanenbase_model:nex-agi/Nex-N2-minibase_model:quantized:nex-agi/Nex-N2-minilicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~18.08 GB disk (24 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
3,373
Likes
2
Pipeline

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Nex-N2-mini-ROCmFP4-STRIX-embF16-imatrix-headQ6.ggufGGUFGGUF18.08 GBDownload

Model Details

Model IDplunderstruck/Nex-N2-mini-ROCmFP4-GGUF
Authorplunderstruck
Pipeline
Licenseapache-2.0
Base modelnex-agi/Nex-N2-mini
Last modified2026-06-21T04:43:26.000Z

Model README

---

base_model: nex-agi/Nex-N2-mini

license: apache-2.0

library_name: gguf

tags:

  • gguf
  • rocmfp4
  • qwen3.5
  • nex-n2
  • coder
  • agentic
  • moe
  • imatrix
  • strix-halo
  • amd
  • rocm
  • vulkan

language:

  • en

base_model_relation: quantized

---

<div style="border:2px solid currentColor; font-family:ui-monospace,'SF Mono','Cascadia Mono',Consolas,'Liberation Mono',monospace;">

<div style="border-bottom:1px solid currentColor; padding:6px 12px; font-size:11px; letter-spacing:3px; text-transform:uppercase; opacity:0.7; text-align:center;">PLUNDERSTRUCK // ROCmFP4 QUANTIZED MODEL // STRIX HALO · gfx1151</div>

<div style="padding:14px; display:flex; flex-wrap:wrap; align-items:center; justify-content:center; gap:18px;">

<pre style="margin:0; flex:0 0 auto; font-family:ui-monospace,'SF Mono','Cascadia Mono',Consolas,monospace; font-size:5px; line-height:1.1; letter-spacing:0;">

███████▆▄▁ ██████▎▐██████████████▖▀██████▙ ▟██████▘

██████████▆▃▁ ██████▎▐███████████████▙▝▜██████▖ ▗██████▛

█████████████▆ ██████▎▐████▀▀▀▀▀▀▀▀▀▀▀▀▀ ▀██████▙▟██████▘

██████▏▀▜█████ ██████▎▐████▃▃▃▃▃▃▃▃▃▃▃▃▃▖ ▔▜██████▘ ▀█▛

██████▏▗▂▔▀▜██ ██████▎▐█████████████████▍ ▝███▛▁▟▙

██████▏▐█▇▅▂▔▀ ██████▎▐█████████████████▍ ▔▀▘▃████▖

██████▏▐████▇▄▂██████▎▐████▀▀▀▀▀▀▀▀▀▀▀▀▀ ▗▟█▖▁▟██████▙▁

██████▏▝█████████████▎▐████▅▅▅▅▅▅▅▅▅▅▅▅▅ ▄██████▛▜██████▖

██████▏ ▝▜██████████▎▐███████████████▛▗▟██████▘ ▝██████▙▁

██████▏ ▀▜███████▎▐██████████████▘▄██████▛ ▜██████▖

</pre>

<div style="flex:0 1 auto; max-width:100%; text-align:center;">

<div style="font-size:23px; font-weight:800; letter-spacing:1px;">NEX-N2-MINI</div>

<div style="font-size:12.5px; letter-spacing:1px; opacity:0.8; margin-top:5px;"><span style="white-space:nowrap;">4-BIT ROCmFP4</span> · <span style="white-space:nowrap;">CODE-WEIGHTED IMATRIX</span> · <span style="white-space:nowrap;">HIGH-SPARSITY MoE (3B ACTIVE)</span> · <span style="white-space:nowrap;">AGENTIC CODER</span> · <span style="white-space:nowrap;">SINGLE AMD APU</span></div>

</div>

</div>

<table style="display:table; table-layout:fixed; width:100%; margin:0; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px;">

<tr>

<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">FORMAT</div><div style="font-weight:700;">ROCmFP4 4-BIT</div></td>

<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">PRECISION</div><div style="font-weight:700;">~4.5 BPW</div></td>

<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">SIZE</div><div style="font-weight:700;">18.4 GB</div></td>

<td style="border-top:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">CONTEXT</div><div style="font-weight:700;">131 K</div></td>

</tr>

<tr>

<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">ARCH</div><div style="font-weight:700;">qwen35moe</div></td>

<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">PARAMS</div><div style="font-weight:700;">35B / 3B ACTIVE</div></td>

<td style="border-top:1px solid currentColor; border-right:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">BACKEND</div><div style="font-weight:700;">VULKAN0</div></td>

<td style="border-top:1px solid currentColor; padding:8px 12px;"><div style="font-size:10px; letter-spacing:1px; opacity:0.6;">LICENSE</div><div style="font-weight:700;">APACHE-2.0</div></td>

</tr>

</table>

</div>

<div style="border:2px solid #dc2626; padding:10px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px; margin:14px 0;">

<b style="color:#dc2626; letter-spacing:1px;">⚠ REQUIRES THE ROCmFP4 FORK</b><br>

The custom <code>q4_0_rocmfp4</code> / <code>q4_0_rocmfp4_fast</code> tensor types <b>will not load in stock llama.cpp, LM Studio, or Ollama</b>. Build/run with <a href="https://github.com/charlie12345/ROCmFPX">charlie12345/ROCmFPX</a> · branch <code>mtp-rocmfp4-strix</code>.

</div>

<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:14px 0; opacity:0.85;">

<b>NOTE //</b> Ignore HuggingFace's auto-detected "F16" badge — its parser can't read ROCmFP4 and mislabels by the f16 embeddings. These are <b>~4.5 bpw 4-bit</b> files; pick by filename.

</div>

<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">01</span> · FILES</div>

<div style="overflow:hidden; border-radius:0;">

<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">

<thead><tr>

<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">File</th>

<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Size</th>

<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Output head</th>

<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Pick if</th>

</tr></thead>

<tbody>

<tr><td style="border:1px solid currentColor; padding:7px 10px;"><code>…-STRIX-embF16-imatrix-headQ6.gguf</code> ★</td><td style="border:1px solid currentColor; padding:7px 10px;">18.4 GB</td><td style="border:1px solid currentColor; padding:7px 10px;">Q6_K</td><td style="border:1px solid currentColor; padding:7px 10px;"><b>the one build</b> — best speed/quality balance: f16 embeddings + Q6 output head on the fast single-scale body</td></tr>

</tbody>

</table>

</div>

One file — the best speed/quality balance in ROCmFP4 for Strix Halo. It keeps the two quality levers that are actually felt — genuine f16 token embeddings (from BF16) and a Q6_K output head — on the fast single-scale q4_0_rocmfp4_fast body + the code-weighted imatrix (see §04). Not the leanest-fastest possible (a 4-bit output head squeezes out a few more tok/s, at a fidelity cost), and not the most faithful possible (see the base-model fidelity link in §04) — it's the point where speed and quality meet best. The Qwen (ChatML) chat template is baked into the GGUF — just pass --jinja.

<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">02</span> · QUICK START</div>

Run from the folder holding the .gguf:

env HSA_OVERRIDE_GFX_VERSION=11.5.1 GGML_HIP_ENABLE_UNIFIED_MEMORY=1 \
llama-server \
  -m Nex-N2-mini-ROCmFP4-STRIX-embF16-imatrix-headQ6.gguf \
  --alias nex-n2-mini \
  --host 0.0.0.0 \
  --port 8080 \
  -dev Vulkan0 \
  -ngl 999 \
  -fa on \
  -c 131072 \
  -b 2048 \
  -ub 256 \
  -t 16 \
  -tb 16 \
  -ctk f16 \
  -ctv f16 \
  -cpent 256 \
  -ctxcp 32 \
  --cache-reuse 256 \
  --cache-ram 65536 \
  --temp 0.6 \
  --top-p 0.95 \
  --top-k 20 \
  --min-p 0.0 \
  --jinja \
  --parallel 1 \
  --metrics \
  --no-mmap

<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:14px 0; opacity:0.85;">

<b>NOTE //</b> No <code>--spec-*</code> / <code>--spec-type draft-mtp</code> flags — Nex-N2-mini ships <b>without an MTP head</b> (non-speculative). At ~72 t/s it doesn't need speculative decoding to be quick.

</div>

<div style="overflow:hidden; border-radius:0;">

<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px;">

<thead><tr>

<th style="border:1px solid currentColor; padding:6px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px; width:40%;">Flag</th>

<th style="border:1px solid currentColor; padding:6px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Function</th>

</tr></thead>

<tbody>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>HSA_OVERRIDE_GFX_VERSION=11.5.1</code></td><td style="border:1px solid currentColor; padding:6px 10px;">treat the APU as gfx1151 (Strix Halo)</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>GGML_HIP_ENABLE_UNIFIED_MEMORY=1</code></td><td style="border:1px solid currentColor; padding:6px 10px;">allow use of the full 128 GB unified memory</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-dev Vulkan0</code></td><td style="border:1px solid currentColor; padding:6px 10px;">run on Vulkan — fastest backend for ROCmFP4 on Strix Halo</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-ngl 999 · -fa on</code></td><td style="border:1px solid currentColor; padding:6px 10px;">offload all layers · flash attention</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-c 131072</code></td><td style="border:1px solid currentColor; padding:6px 10px;">context length (128K)</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-b 2048 · -ub 256 · -t/-tb 16</code></td><td style="border:1px solid currentColor; padding:6px 10px;">prefill batch / micro-batch · CPU threads</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-ctk f16 · -ctv f16</code></td><td style="border:1px solid currentColor; padding:6px 10px;">f16 KV cache — how we run it; drop to <code>q8_0</code>/<code>q4_0</code> to use less memory</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>-cpent · -ctxcp · --cache-reuse · --cache-ram 65536</code></td><td style="border:1px solid currentColor; padding:6px 10px;">cross-turn KV checkpointing + 64 GB resident reuse cache</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--temp 0.6 --top-p 0.95 --top-k 20 --min-p 0.0</code></td><td style="border:1px solid currentColor; padding:6px 10px;">base-model recommended sampling</td></tr>

<tr><td style="border:1px solid currentColor; padding:6px 10px;"><code>--jinja --parallel 1 --metrics --no-mmap</code></td><td style="border:1px solid currentColor; padding:6px 10px;">apply baked ChatML template · single slot · metrics · weights in RAM</td></tr>

</tbody>

</table>

</div>

<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">03</span> · AGENTIC CODING / TOOLS</div>

Nex-N2-mini is an agentic / "thinking" coder — agentic tool-use trained. To get native tool calls, your client must use the qwen3_coder tool-call parser. Without it the model tends to narrate code instead of emitting structured tool calls.

<div style="overflow:hidden; border-radius:0;">

<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">

<tbody>

<tr><td style="border:1px solid currentColor; padding:8px 11px; width:30%;">CHAT TEMPLATE</td><td style="border:1px solid currentColor; padding:8px 11px;">Qwen (ChatML) — baked into the GGUF; pass <code>--jinja</code></td></tr>

<tr><td style="border:1px solid currentColor; padding:8px 11px;">TOOL-CALL PARSER</td><td style="border:1px solid currentColor; padding:8px 11px;"><code>qwen3_coder</code> — set in your client/runtime</td></tr>

<tr><td style="border:1px solid currentColor; padding:8px 11px;">SAMPLING</td><td style="border:1px solid currentColor; padding:8px 11px;">temp <code>0.6</code> · top-p <code>0.95</code> · top-k <code>20</code> (base-model recommended)</td></tr>

</tbody>

</table>

</div>

<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">04</span> · PERFORMANCE &amp; QUALITY</div>

<div style="overflow:hidden; border-radius:0;">

<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">

<tbody>

<tr><td style="border:1px solid currentColor; padding:8px 11px; width:42%;">DECODE · short-context</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;">~72 t/s (Vulkan / Ryzen AI Max+ 395)</td></tr>

<tr><td style="border:1px solid currentColor; padding:8px 11px;">SWE-BENCH VERIFIED · base model</td><td style="border:1px solid currentColor; padding:8px 11px; font-weight:700;">74.4</td></tr>

<tr><td style="border:1px solid currentColor; padding:8px 11px;">ACTIVE PARAMS</td><td style="border:1px solid currentColor; padding:8px 11px;">3B of 35B (high-sparsity MoE)</td></tr>

<tr><td style="border:1px solid currentColor; padding:8px 11px;">QUANTIZATION</td><td style="border:1px solid currentColor; padding:8px 11px;">fast single-scale body + f16 embeddings + Q6 head + code-weighted imatrix</td></tr>

</tbody>

</table>

</div>

This is the best speed/quality balance in ROCmFP4 — by design, not the absolute fastest. It keeps the two quality levers that are actually felt — genuine f16 token embeddings and a Q6_K output head — on the fast single-scale q4_0_rocmfp4_fast body. A leaner 4-bit-output-head build is a few tok/s faster but degrades fidelity you'll notice; an all-dual-scale body buys a KL improvement that sits inside the measurement noise while costing decode speed. The fast body + f16 embeddings + Q6 head is the point where those meet best.

How we landed on this recipe. We ran the full body-kernel / head-precision / dual-scale sweep — KL divergence vs the BF16 reference plus llama-bench decode — on the dense Qwen3.6-27B sibling, where the same q4_0_rocmfp4 levers apply. The frontier there was unambiguous: the all-dual-scale body and selective higher-precision tensors both traded decode speed for a KL gain inside the noise, so the fast body + f16 embeddings + Q6 head won the balance. We carry that conclusion to this MoE rather than re-running the whole sweep per model — see the 27B sweep for the numbers and the format-limit reasoning. (Directional internal measurements — reproduce before citing.)

<div style="border:1px solid currentColor; padding:8px 13px; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12px; margin:12px 0; opacity:0.9;">

<b>WANT MAXIMUM FIDELITY INSTEAD OF SPEED?</b> Grab a <b>Q6_K / Q8_0 GGUF of the base</b> from <a href="https://huggingface.co/nex-agi/Nex-N2-mini"><b>nex-agi/Nex-N2-mini</b></a> — those higher-bit GGUFs run on this same fork. We optimize for throughput in ROCmFP4; if you want the last bit of fidelity over speed, a higher-bit quant of the base is the one to grab.

</div>

The imatrix — code-weighted, and measured (it helps here). Quantized with an importance matrix from a code-weighted calibration mix (~2.6:1 code:general — eaddario code + Kalomaze groups_merged via froggeric/imatrix). Measured by KL-divergence + perplexity vs the true BF16 on a held-out code slice (disjoint from calibration):

<div style="overflow:hidden; border-radius:0;">

<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">

<thead><tr>

<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Metric (vs BF16, held-out code)</th>

<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">No-imatrix</th>

<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Imatrix</th>

<th style="border:1px solid currentColor; padding:7px 10px; text-align:left; text-transform:uppercase; font-size:10px; letter-spacing:1px;">Change</th>

</tr></thead>

<tbody>

<tr><td style="border:1px solid currentColor; padding:7px 10px;"><b>Perplexity</b></td><td style="border:1px solid currentColor; padding:7px 10px;">4.076</td><td style="border:1px solid currentColor; padding:7px 10px;"><b>4.013</b></td><td style="border:1px solid currentColor; padding:7px 10px;"><b>−1.5%</b> (recovers &gt;½ the 4-bit loss; ~3.3σ)</td></tr>

<tr><td style="border:1px solid currentColor; padding:7px 10px;"><b>Median KLD</b></td><td style="border:1px solid currentColor; padding:7px 10px;">0.0184</td><td style="border:1px solid currentColor; padding:7px 10px;"><b>0.0159</b></td><td style="border:1px solid currentColor; padding:7px 10px;"><b>−13%</b></td></tr>

<tr><td style="border:1px solid currentColor; padding:7px 10px;">RMS Δp</td><td style="border:1px solid currentColor; padding:7px 10px;">8.57%</td><td style="border:1px solid currentColor; padding:7px 10px;">8.00%</td><td style="border:1px solid currentColor; padding:7px 10px;">−7%</td></tr>

<tr><td style="border:1px solid currentColor; padding:7px 10px;"><b>Same top token as BF16</b></td><td style="border:1px solid currentColor; padding:7px 10px;">88.97%</td><td style="border:1px solid currentColor; padding:7px 10px;"><b>89.44%</b></td><td style="border:1px solid currentColor; padding:7px 10px;">+0.5 pp</td></tr>

</tbody>

</table>

</div>

For this model the imatrix is a clean win — better on every metric, including perplexity. (It's model-dependent — on the dense Qwopus-Coder the same recipe worsened code-PPL, so we shipped that one without imatrix. Always measure.)

<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">05</span> · BUILD (REPRODUCIBLE)</div>

# code-weighted imatrix on the BF16 (single pass)
llama-imatrix -m Nex-N2-mini-bf16.gguf -f code-weighted-calib.txt -o nexn2.imatrix -c 512 -ngl 999

# quant -> ROCmFP4 with the imatrix + genuine f16 embeddings
llama-quantize --token-embedding-type f16 --imatrix nexn2.imatrix \
  Nex-N2-mini-bf16.gguf \
  Nex-N2-mini-ROCmFP4-STRIX-embF16-imatrix.gguf  Q4_0_ROCMFP4_STRIX

# THE ONE BUILD (★): add the Q6_K output head on the fast single-scale body — best speed/quality balance (§04)
llama-quantize --token-embedding-type f16 --output-tensor-type q6_K --imatrix nexn2.imatrix \
  Nex-N2-mini-bf16.gguf \
  Nex-N2-mini-ROCmFP4-STRIX-embF16-imatrix-headQ6.gguf  Q4_0_ROCMFP4_STRIX

> Experimental research build for AMD Strix Halo — hardware/driver/prompt-sensitive, may not reproduce elsewhere. Not native FP4 tensor-core execution.

<div style="font-family:ui-monospace,'SF Mono',Consolas,monospace; font-weight:800; font-size:14px; letter-spacing:2px; text-transform:uppercase; border-bottom:2px solid currentColor; padding-bottom:5px; margin:26px 0 12px;"><span style="color:#ea580c;">06</span> · LINEAGE &amp; CREDITS</div>

<div style="overflow:hidden; border-radius:0;">

<table style="width:100%; border-collapse:collapse; border-radius:0; font-family:ui-monospace,'SF Mono',Consolas,monospace; font-size:12.5px;">

<tbody>

<tr><td style="border:1px solid currentColor; padding:8px 11px; width:26%;">BASE MODEL</td><td style="border:1px solid currentColor; padding:8px 11px;"><a href="https://huggingface.co/nex-agi/Nex-N2-mini">nex-agi/Nex-N2-mini</a> (Apache-2.0) · Qwen3.5-35B-A3B lineage (35B total / 3B active MoE)</td></tr>

<tr><td style="border:1px solid currentColor; padding:8px 11px;">CALIBRATION</td><td style="border:1px solid currentColor; padding:8px 11px;"><a href="https://huggingface.co/datasets/eaddario/imatrix-calibration">eaddario/imatrix-calibration</a> (code) + Kalomaze <code>groups_merged</code> via <a href="https://huggingface.co/datasets/froggeric/imatrix">froggeric/imatrix</a> (general)</td></tr>

<tr><td style="border:1px solid currentColor; padding:8px 11px;">FORMAT + RUNTIME</td><td style="border:1px solid currentColor; padding:8px 11px;"><a href="https://github.com/charlie12345/ROCmFPX">charlie12345/ROCmFPX</a> (based on llama.cpp, MIT)</td></tr>

</tbody>

</table>

</div>

Derivative quantization — verify the base model's license before redistribution / use.

Run plunderstruck/Nex-N2-mini-ROCmFP4-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models