GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

sakamakismile/Nex-N2.5-Max-GGUF overview

Nex N2.5 Max GGUF llama.cpp GGUF quantizations of nex agi/Nex N2.5 Max https://huggingface.co/nex agi/Nex N2.5 Max 1.6T parameter text only MoE, 384 routed exp…

llama.cppggufdeepseek_v4moenex-n2.5lna-labtext-generationenjazhbase_model:nex-agi/Nex-N2.5-Maxbase_model:quantized:nex-agi/Nex-N2.5-Maxlicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~18.14 GB disk (24 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
1
Pipeline
text-generation

Repository Files & Downloads

39 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Q3_K_M/Nex-N2.5-Max-Q3_K_M-00001-of-00018.ggufGGUFQ3_K_M38.25 GBDownload
Q3_K_M/Nex-N2.5-Max-Q3_K_M-00002-of-00018.ggufGGUFQ3_K_M41.91 GBDownload
Q3_K_M/Nex-N2.5-Max-Q3_K_M-00003-of-00018.ggufGGUFQ3_K_M37.61 GBDownload
Q3_K_M/Nex-N2.5-Max-Q3_K_M-00004-of-00018.ggufGGUFQ3_K_M41.91 GBDownload
Q3_K_M/Nex-N2.5-Max-Q3_K_M-00005-of-00018.ggufGGUFQ3_K_M37.61 GBDownload
Q3_K_M/Nex-N2.5-Max-Q3_K_M-00006-of-00018.ggufGGUFQ3_K_M41.91 GBDownload
Q3_K_M/Nex-N2.5-Max-Q3_K_M-00007-of-00018.ggufGGUFQ3_K_M37.61 GBDownload
Q3_K_M/Nex-N2.5-Max-Q3_K_M-00008-of-00018.ggufGGUFQ3_K_M41.91 GBDownload
Q3_K_M/Nex-N2.5-Max-Q3_K_M-00009-of-00018.ggufGGUFQ3_K_M37.61 GBDownload
Q3_K_M/Nex-N2.5-Max-Q3_K_M-00010-of-00018.ggufGGUFQ3_K_M41.91 GBDownload
Q3_K_M/Nex-N2.5-Max-Q3_K_M-00011-of-00018.ggufGGUFQ3_K_M37.61 GBDownload
Q3_K_M/Nex-N2.5-Max-Q3_K_M-00012-of-00018.ggufGGUFQ3_K_M41.91 GBDownload
Q3_K_M/Nex-N2.5-Max-Q3_K_M-00013-of-00018.ggufGGUFQ3_K_M37.61 GBDownload
Q3_K_M/Nex-N2.5-Max-Q3_K_M-00014-of-00018.ggufGGUFQ3_K_M41.91 GBDownload
Q3_K_M/Nex-N2.5-Max-Q3_K_M-00015-of-00018.ggufGGUFQ3_K_M37.61 GBDownload
Q3_K_M/Nex-N2.5-Max-Q3_K_M-00016-of-00018.ggufGGUFQ3_K_M41.91 GBDownload
Q3_K_M/Nex-N2.5-Max-Q3_K_M-00017-of-00018.ggufGGUFQ3_K_M37.61 GBDownload
Q3_K_M/Nex-N2.5-Max-Q3_K_M-00018-of-00018.ggufGGUFQ3_K_M22.60 GBDownload
Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00001-of-00021.ggufGGUFQ4K38.40 GBDownload
Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00002-of-00021.ggufGGUFQ4K41.03 GBDownload
Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00003-of-00021.ggufGGUFQ4K41.03 GBDownload
Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00004-of-00021.ggufGGUFQ4K41.03 GBDownload
Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00005-of-00021.ggufGGUFQ4K41.03 GBDownload
Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00006-of-00021.ggufGGUFQ4K41.03 GBDownload
Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00007-of-00021.ggufGGUFQ4K41.03 GBDownload
Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00008-of-00021.ggufGGUFQ4K41.03 GBDownload
Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00009-of-00021.ggufGGUFQ4K41.03 GBDownload
Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00010-of-00021.ggufGGUFQ4K41.03 GBDownload
Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00011-of-00021.ggufGGUFQ4K41.03 GBDownload
Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00012-of-00021.ggufGGUFQ4K41.03 GBDownload
Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00013-of-00021.ggufGGUFQ4K41.03 GBDownload
Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00014-of-00021.ggufGGUFQ4K41.03 GBDownload
Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00015-of-00021.ggufGGUFQ4K41.03 GBDownload
Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00016-of-00021.ggufGGUFQ4K41.03 GBDownload
Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00017-of-00021.ggufGGUFQ4K41.03 GBDownload
Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00018-of-00021.ggufGGUFQ4K41.03 GBDownload
Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00019-of-00021.ggufGGUFQ4K41.03 GBDownload
Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00020-of-00021.ggufGGUFQ4K41.03 GBDownload
Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00021-of-00021.ggufGGUFQ4K18.14 GBDownload

Model Details

Model IDsakamakismile/Nex-N2.5-Max-GGUF
Authorsakamakismile
Pipelinetext-generation
Licenseapache-2.0
Base modelnex-agi/Nex-N2.5-Max
Last modified2026-09-09T22:07:03.000Z

Model README

---

license: apache-2.0

base_model: nex-agi/Nex-N2.5-Max

base_model_relation: quantized

pipeline_tag: text-generation

library_name: llama.cpp

tags:

  • gguf
  • deepseek_v4
  • moe
  • nex-n2.5
  • lna-lab

language:

  • en
  • ja
  • zh

---

Nex-N2.5-Max GGUF (llama.cpp)

GGUF quantizations of nex-agi/Nex-N2.5-Max

(1.6T-parameter text-only MoE, 384 routed experts / top-6, 61 layers, 1M context, MTP head),

made at Lna-Lab on 2026-09-09 for pure-CPU inference on a 1 TiB DDR5 workstation.

Nex-N2.5-Max uses the DeepSeek-V4-Pro architecture unchanged (DeepseekV4ForCausalLM,

hidden 7168, MoE inter 3072, indexer 64 heads / top-1024, hyper-connections x4, 3 hash-routed layers).

Any llama.cpp build that runs DeepSeek-V4 runs these files.

Files

| folder | recipe | size | BPW | notes |

|---|---|---|---|---|

| Q4K-exps-Q8dense/ | routed experts Q4_K (pure), everything else (attention, shared expert, indexer, router, embeddings, output) Q8_0 | 898 GB (836 GiB) | 4.57 | fastest on CPU; needs ~880 GiB of free RAM |

| Q3_K_M/ | stock llama.cpp Q3_K_M mixture | 748 GB (697 GiB) | 3.81 | fits alongside other residents; slower on CPU (no repacked kernel for Q3_K) |

Each folder is a standard llama.cpp split (-00001-of-000NN.gguf); point -m at the first shard.

Measured (not estimated)

Threadripper PRO 9985WX (64 cores), 1 TiB DDR5-8ch, no GPU (-ngl 0 -t 64), llama.cpp master 465e49b, warm page cache, single stream:

| quant | prompt tok/s | generation tok/s |

|---|---|---|

| Q4K-exps-Q8dense | 27.8 | 5.35 |

| Q3_K_M | 25.6 | 3.04 |

Q3_K_M is slower although smaller: llama.cpp has an AVX-512 repacked kernel for Q4_K but not for Q3_K.

The repack also copies the Q4_K experts into anonymous memory (RssAnon ≈ 811 GiB for the Q4 file),

so mmap does not save you RAM with that file. Cold first run (page faults) is ~0.5 tok/s; warm the file first

(cat file > /dev/null in parallel chunks, ~50 s on a fast RAID) or just wait for the first generation.

Quality spot check (Japanese haiku with one-line commentary, thinking off): both quants produce clean Japanese

with sensible kigo; we could not tell them apart on that task. The model thinks in English when

enable_thinking is on (gpt-oss style), and answers in the user's language.

Running

llama-server -m Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00001-of-00021.gguf \
  -ngl 0 -t 64 -c 16384 -fa on --jinja --port 8030

The chat template is embedded (from the original repo) and supports chat_template_kwargs:

enable_thinking (default true) and thinking_mode (interleaved | full | drop).

Known issue (llama.cpp 465e49b): /v1/chat/completions can return HTTP 500

"The model produced output that does not match the expected peg-native format" after a complete answer.

The model is fine; the server-side output parser is not. Work-around: render with POST /apply-template

and generate with POST /completion.

How these were made

The original checkpoint stores routed experts as FP8 e4m3 with 128x128 block scales stored as F32

(expert_dtype: fp8, scale_fmt: ue8m0), whereas the official DeepSeek-V4 release packs experts as MXFP4.

llama.cpp's convert_hf_to_gguf.py (DeepSeek-V4 path) dequantizes the FP8 tensors, drops the scales, and then

tries to repack experts as MXFP4 -> KeyError: Missing routed expert tensors. It also reads F32 scales as e8m0 bytes.

llama.cpp-convert-fp8-experts.patch (against conversion/deepseek.py, master 465e49b) fixes both:

  1. float-typed block scales are used as values;
  2. FP8-dequantized experts are stacked (lazily, 384 per projection) and emitted as ordinary expert tensors

(they land as Q8_0 in the intermediate, per the converter's existing FP8 policy).

Pipeline: convert_hf_to_gguf.py --outtype bf16 (intermediate: experts Q8_0 + rest BF16, 1.70 TB, 2h18m) ->

llama-quantize --allow-requantize:

# Q4K-exps-Q8dense
llama-quantize --allow-requantize --pure --output-tensor-type q8_0 --token-embedding-type q8_0 \
  --tensor-type 'attn_.*=q8_0' --tensor-type 'ffn_.*_shexp=q8_0' --tensor-type 'indexer.*=q8_0' --tensor-type 'ffn_gate_inp=q8_0' \
  intermediate.gguf out.gguf Q4_K_M 48
# Q3_K_M
llama-quantize --allow-requantize intermediate.gguf out.gguf Q3_K_M 48

No importance matrix was used. A stock Q4_K_M was also baked (951 GB, 4.84 BPW) but is not published:

it does not fit in 1 TiB RAM next to anything else, and its attention weights are 4-bit anyway.

MTP layer is not included (the converter's main-model path skips mtp.*; those tensors are BF16 in the source).

Credits

  • Model: Nex-AGI, Apache-2.0. Architecture: DeepSeek.
  • Quantization and converter patch: YUKI (Claude Fable 5.1), researching together with Ken at Lna-Lab — @Tono_Ken3

Run sakamakismile/Nex-N2.5-Max-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models