sakamakismile/Nex-N2.5-Max-GGUF overview
Nex N2.5 Max GGUF llama.cpp GGUF quantizations of nex agi/Nex N2.5 Max https://huggingface.co/nex agi/Nex N2.5 Max 1.6T parameter text only MoE, 384 routed exp…
Runs locally from ~18.14 GB disk (24 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Q3_K_M/Nex-N2.5-Max-Q3_K_M-00001-of-00018.gguf | GGUF | Q3_K_M | 38.25 GB | Download |
| Q3_K_M/Nex-N2.5-Max-Q3_K_M-00002-of-00018.gguf | GGUF | Q3_K_M | 41.91 GB | Download |
| Q3_K_M/Nex-N2.5-Max-Q3_K_M-00003-of-00018.gguf | GGUF | Q3_K_M | 37.61 GB | Download |
| Q3_K_M/Nex-N2.5-Max-Q3_K_M-00004-of-00018.gguf | GGUF | Q3_K_M | 41.91 GB | Download |
| Q3_K_M/Nex-N2.5-Max-Q3_K_M-00005-of-00018.gguf | GGUF | Q3_K_M | 37.61 GB | Download |
| Q3_K_M/Nex-N2.5-Max-Q3_K_M-00006-of-00018.gguf | GGUF | Q3_K_M | 41.91 GB | Download |
| Q3_K_M/Nex-N2.5-Max-Q3_K_M-00007-of-00018.gguf | GGUF | Q3_K_M | 37.61 GB | Download |
| Q3_K_M/Nex-N2.5-Max-Q3_K_M-00008-of-00018.gguf | GGUF | Q3_K_M | 41.91 GB | Download |
| Q3_K_M/Nex-N2.5-Max-Q3_K_M-00009-of-00018.gguf | GGUF | Q3_K_M | 37.61 GB | Download |
| Q3_K_M/Nex-N2.5-Max-Q3_K_M-00010-of-00018.gguf | GGUF | Q3_K_M | 41.91 GB | Download |
| Q3_K_M/Nex-N2.5-Max-Q3_K_M-00011-of-00018.gguf | GGUF | Q3_K_M | 37.61 GB | Download |
| Q3_K_M/Nex-N2.5-Max-Q3_K_M-00012-of-00018.gguf | GGUF | Q3_K_M | 41.91 GB | Download |
| Q3_K_M/Nex-N2.5-Max-Q3_K_M-00013-of-00018.gguf | GGUF | Q3_K_M | 37.61 GB | Download |
| Q3_K_M/Nex-N2.5-Max-Q3_K_M-00014-of-00018.gguf | GGUF | Q3_K_M | 41.91 GB | Download |
| Q3_K_M/Nex-N2.5-Max-Q3_K_M-00015-of-00018.gguf | GGUF | Q3_K_M | 37.61 GB | Download |
| Q3_K_M/Nex-N2.5-Max-Q3_K_M-00016-of-00018.gguf | GGUF | Q3_K_M | 41.91 GB | Download |
| Q3_K_M/Nex-N2.5-Max-Q3_K_M-00017-of-00018.gguf | GGUF | Q3_K_M | 37.61 GB | Download |
| Q3_K_M/Nex-N2.5-Max-Q3_K_M-00018-of-00018.gguf | GGUF | Q3_K_M | 22.60 GB | Download |
| Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00001-of-00021.gguf | GGUF | Q4K | 38.40 GB | Download |
| Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00002-of-00021.gguf | GGUF | Q4K | 41.03 GB | Download |
| Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00003-of-00021.gguf | GGUF | Q4K | 41.03 GB | Download |
| Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00004-of-00021.gguf | GGUF | Q4K | 41.03 GB | Download |
| Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00005-of-00021.gguf | GGUF | Q4K | 41.03 GB | Download |
| Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00006-of-00021.gguf | GGUF | Q4K | 41.03 GB | Download |
| Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00007-of-00021.gguf | GGUF | Q4K | 41.03 GB | Download |
| Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00008-of-00021.gguf | GGUF | Q4K | 41.03 GB | Download |
| Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00009-of-00021.gguf | GGUF | Q4K | 41.03 GB | Download |
| Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00010-of-00021.gguf | GGUF | Q4K | 41.03 GB | Download |
| Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00011-of-00021.gguf | GGUF | Q4K | 41.03 GB | Download |
| Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00012-of-00021.gguf | GGUF | Q4K | 41.03 GB | Download |
| Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00013-of-00021.gguf | GGUF | Q4K | 41.03 GB | Download |
| Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00014-of-00021.gguf | GGUF | Q4K | 41.03 GB | Download |
| Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00015-of-00021.gguf | GGUF | Q4K | 41.03 GB | Download |
| Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00016-of-00021.gguf | GGUF | Q4K | 41.03 GB | Download |
| Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00017-of-00021.gguf | GGUF | Q4K | 41.03 GB | Download |
| Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00018-of-00021.gguf | GGUF | Q4K | 41.03 GB | Download |
| Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00019-of-00021.gguf | GGUF | Q4K | 41.03 GB | Download |
| Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00020-of-00021.gguf | GGUF | Q4K | 41.03 GB | Download |
| Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00021-of-00021.gguf | GGUF | Q4K | 18.14 GB | Download |
Model Details
| Model ID | sakamakismile/Nex-N2.5-Max-GGUF |
|---|---|
| Author | sakamakismile |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | nex-agi/Nex-N2.5-Max |
| Last modified | 2026-09-09T22:07:03.000Z |
Model README
---
license: apache-2.0
base_model: nex-agi/Nex-N2.5-Max
base_model_relation: quantized
pipeline_tag: text-generation
library_name: llama.cpp
tags:
- gguf
- deepseek_v4
- moe
- nex-n2.5
- lna-lab
language:
- en
- ja
- zh
---
Nex-N2.5-Max GGUF (llama.cpp)
GGUF quantizations of nex-agi/Nex-N2.5-Max
(1.6T-parameter text-only MoE, 384 routed experts / top-6, 61 layers, 1M context, MTP head),
made at Lna-Lab on 2026-09-09 for pure-CPU inference on a 1 TiB DDR5 workstation.
Nex-N2.5-Max uses the DeepSeek-V4-Pro architecture unchanged (DeepseekV4ForCausalLM,
hidden 7168, MoE inter 3072, indexer 64 heads / top-1024, hyper-connections x4, 3 hash-routed layers).
Any llama.cpp build that runs DeepSeek-V4 runs these files.
Files
| folder | recipe | size | BPW | notes |
|---|---|---|---|---|
| Q4K-exps-Q8dense/ | routed experts Q4_K (pure), everything else (attention, shared expert, indexer, router, embeddings, output) Q8_0 | 898 GB (836 GiB) | 4.57 | fastest on CPU; needs ~880 GiB of free RAM |
| Q3_K_M/ | stock llama.cpp Q3_K_M mixture | 748 GB (697 GiB) | 3.81 | fits alongside other residents; slower on CPU (no repacked kernel for Q3_K) |
Each folder is a standard llama.cpp split (-00001-of-000NN.gguf); point -m at the first shard.
Measured (not estimated)
Threadripper PRO 9985WX (64 cores), 1 TiB DDR5-8ch, no GPU (-ngl 0 -t 64), llama.cpp master 465e49b, warm page cache, single stream:
| quant | prompt tok/s | generation tok/s |
|---|---|---|
| Q4K-exps-Q8dense | 27.8 | 5.35 |
| Q3_K_M | 25.6 | 3.04 |
Q3_K_M is slower although smaller: llama.cpp has an AVX-512 repacked kernel for Q4_K but not for Q3_K.
The repack also copies the Q4_K experts into anonymous memory (RssAnon ≈ 811 GiB for the Q4 file),
so mmap does not save you RAM with that file. Cold first run (page faults) is ~0.5 tok/s; warm the file first
(cat file > /dev/null in parallel chunks, ~50 s on a fast RAID) or just wait for the first generation.
Quality spot check (Japanese haiku with one-line commentary, thinking off): both quants produce clean Japanese
with sensible kigo; we could not tell them apart on that task. The model thinks in English when
enable_thinking is on (gpt-oss style), and answers in the user's language.
Running
llama-server -m Q4K-exps-Q8dense/Nex-N2.5-Max-Q4K-exps-Q8dense-00001-of-00021.gguf \
-ngl 0 -t 64 -c 16384 -fa on --jinja --port 8030
The chat template is embedded (from the original repo) and supports chat_template_kwargs:
enable_thinking (default true) and thinking_mode (interleaved | full | drop).
Known issue (llama.cpp 465e49b): /v1/chat/completions can return HTTP 500
"The model produced output that does not match the expected peg-native format" after a complete answer.
The model is fine; the server-side output parser is not. Work-around: render with POST /apply-template
and generate with POST /completion.
How these were made
The original checkpoint stores routed experts as FP8 e4m3 with 128x128 block scales stored as F32
(expert_dtype: fp8, scale_fmt: ue8m0), whereas the official DeepSeek-V4 release packs experts as MXFP4.
llama.cpp's convert_hf_to_gguf.py (DeepSeek-V4 path) dequantizes the FP8 tensors, drops the scales, and then
tries to repack experts as MXFP4 -> KeyError: Missing routed expert tensors. It also reads F32 scales as e8m0 bytes.
llama.cpp-convert-fp8-experts.patch (against conversion/deepseek.py, master 465e49b) fixes both:
- float-typed block scales are used as values;
- FP8-dequantized experts are stacked (lazily, 384 per projection) and emitted as ordinary expert tensors
(they land as Q8_0 in the intermediate, per the converter's existing FP8 policy).
Pipeline: convert_hf_to_gguf.py --outtype bf16 (intermediate: experts Q8_0 + rest BF16, 1.70 TB, 2h18m) ->
llama-quantize --allow-requantize:
# Q4K-exps-Q8dense
llama-quantize --allow-requantize --pure --output-tensor-type q8_0 --token-embedding-type q8_0 \
--tensor-type 'attn_.*=q8_0' --tensor-type 'ffn_.*_shexp=q8_0' --tensor-type 'indexer.*=q8_0' --tensor-type 'ffn_gate_inp=q8_0' \
intermediate.gguf out.gguf Q4_K_M 48
# Q3_K_M
llama-quantize --allow-requantize intermediate.gguf out.gguf Q3_K_M 48
No importance matrix was used. A stock Q4_K_M was also baked (951 GB, 4.84 BPW) but is not published:
it does not fit in 1 TiB RAM next to anything else, and its attention weights are 4-bit anyway.
MTP layer is not included (the converter's main-model path skips mtp.*; those tensors are BF16 in the source).
Credits
- Model: Nex-AGI, Apache-2.0. Architecture: DeepSeek.
- Quantization and converter patch: YUKI (Claude Fable 5.1), researching together with Ken at Lna-Lab — @Tono_Ken3
Run sakamakismile/Nex-N2.5-Max-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models