GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Schackay3/Qwen3-Next-80B-A3B-Instruct-ternary-GGUF overview

Qwen3 Next 80B A3B Instruct — Ternary 1.58 bit expert GGUF To our knowledge, this is the first publicly released ternarized Mixture of Experts model. All prior…

ggufternary1.58-bitbitnetmoequantizationllama.cpptext-generationenbase_model:Qwen/Qwen3-Next-80B-A3B-Instructbase_model:quantized:Qwen/Qwen3-Next-80B-A3B-Instructlicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~27.26 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
1
Pipeline
text-generation
Author

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
qwen3-next-80b-a3b-instruct-ternary-Q2_K.ggufGGUFQ2_K27.26 GBDownload

Model Details

Model IDSchackay3/Qwen3-Next-80B-A3B-Instruct-ternary-GGUF
AuthorSchackay3
Pipelinetext-generation
Licenseapache-2.0
Base modelQwen/Qwen3-Next-80B-A3B-Instruct
Last modified2026-08-12T21:53:54.000Z

Model README

---

license: apache-2.0

base_model: Qwen/Qwen3-Next-80B-A3B-Instruct

base_model_relation: quantized

tags:

  • gguf
  • ternary
  • 1.58-bit
  • bitnet
  • moe
  • quantization
  • llama.cpp

language:

  • en

pipeline_tag: text-generation

---

Qwen3-Next-80B-A3B-Instruct — Ternary (1.58-bit expert) GGUF

**To our knowledge, this is the first publicly released ternarized

Mixture-of-Experts model.** All prior published ternary conversions

(PT-BitNet, PT2-LLM, Falcon3-1.58bit, Llama3-8B-1.58) target dense models.

Every one of the model's 24,576 expert weight matrices (512 experts x 48

layers, gate/up/down projections) was quantized to ternary values

{-1, 0, +1} with per-block scale and offset, using GPTQ-style error-

compensating reconstruction. The conversion ran for **34 hours on a single

NVIDIA Quadro RTX 4000 (8GB)** in a 2016-era dual-Xeon workstation.

What was quantized

| Component | Treatment |

|---|---|

| MoE experts (mlp.experts.*) — >90% of parameters | ternary (group-128, asymmetric) |

| Shared expert | ternary |

| Starved experts (<32 calibration tokens routed) | left bf16 (see below) |

| Router / gate | untouched bf16 |

| DeltaNet linear attention, full attention | untouched bf16 |

| Embeddings, norms, lm_head | untouched bf16 |

The trunk was deliberately left at full precision: it is a small fraction of

total parameters and by far the riskiest part to quantize aggressively.

Measured results

Reference machine: 2x Intel Xeon E5-2697 v4 (2016, no AVX-512, no AMX),

256GB DDR4-1866, Quadro RTX 4000 8GB. llama.cpp b10331,

flags -ngl 999 -ncmoe 999 -fa 1 -t 18 (attention on GPU, experts in system RAM).

Quality (wikitext2 perplexity, 2048 ctx)

| Model | PPL |

|---|---|

| bf16 original | 5.75 |

| this ternary conversion | 8.69 (+51%) |

| reference: naive rounding to ternary, no calibration | ~8,960x worse |

| reference: PT2-LLM published dense LLaMA-2-7B ternary | 11.56 |

Speed & size (vs the stock Q4_K_M release of the same model)

| | Q4_K_M | this model (Q2_K) | delta |

|---|---|---|---|

| generation | 16.74 tok/s | 23.74 tok/s | +42% |

| prompt processing | 58.50 tok/s | 107.08 tok/s | +83% |

| size | 45.08 GB | 27.25 GB | -40% |

Coherence (greedy decode)

Q: What are the three primary colors?
A: The three primary colors are red, blue

"The capital of France is" -> "Paris. The capital of the United States"
"def reverse_string(s):"   -> "\n    # Write your code here.\n"

Honest limitations

  • This is meaningfully worse than the original. A +51% perplexity increase

is real and you will likely notice it on hard reasoning and long-form code.

If you need maximum quality, use the standard Q4_K_M release.

  • Calibration was wikitext2 only (128 x 2048 tokens) — generic English prose

for a model trained on code, math, multilingual and agentic data. Domain-

matched calibration would very likely improve results.

  • SSR (Structural Similarity-based Reordering) was not enabled, nor was a

distillation teacher used. Both are known to help.

  • The GGUF is Q2_K (2.94 bpw), not a true 1.58-bit format. llama.cpp's

native ternary formats (TQ1_0/TQ2_0) are symmetric with 256-element blocks;

the GPTQ reconstruction produces asymmetric group-128 ternary (measured:

exactly 3 distinct values per 128-block, offsets averaging |mu|/alpha =

0.10-0.19). Q2_K stores a scale and min per sub-block and represents this

far more faithfully. A native asymmetric group-128 ternary kernel would

recover the remaining ~0.9 bpw and push speed toward ~35 tok/s.

  • MTP (multi-token prediction) head is not included (--no-mtp), so

speculative decoding is unavailable.

  • Single eval set. No MMLU / HumanEval / GSM8K numbers yet.

Novel finding: expert utilization varies with depth

Experts routed fewer than 32 calibration tokens were left in bf16 rather than

quantized on unreliable statistics. The count of such "starved" experts is not

uniform across the network:

| layers | starved experts (of 512) |

|---|---|

| 0-31 (early/middle) | 20-70 |

| 46-47 (late) | 114-131 |

Later layers concentrate routing on fewer experts. We have not seen this

measured elsewhere and it has direct implications for MoE quantization: a

uniform calibration budget under-serves deep layers.

Bug found in PT2-LLM

solve_closed_form_alpha_mu() in the upstream PT2-LLM quantizer has its

epsilon commented out, so denom reaches exactly 0 for all-zero ternary rows,

producing NaN that silently propagates through the entire weight matrix.

This is harmless on large dense matrices but fatal on small expert matrices:

484 of 512 experts failed before the fix, 0 after. Fixed by clamping the

denominator and applying the AGA fallback per-row instead of per-matrix.

Usage

llama-cli -m qwen3-ternary-Q2_K.gguf -ngl 999 --n-cpu-moe 999 -t 18 -fa on -cnv

Tune -t to your core count; on dual-socket systems, single-socket thread

counts often beat using all cores (NUMA contention). On this machine 18 threads

was fastest for generation, 24 for prompt processing.

Credits & license

Released under Apache 2.0, matching the base model.

Run Schackay3/Qwen3-Next-80B-A3B-Instruct-ternary-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models