GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

abenzerps/Nex-N2.5-mini-GGUF overview

Nex N2.5 mini GGUF GGUF files for Nex N2.5 mini https://huggingface.co/nex agi/Nex N2.5 mini , a long context agentic model for coding, tool use, computer use,…

ggufllama.cppqwen3.5qwen3.5-moelong-context262k-contextmultimodaltool-callingtext-generationconversationalenzhbase_model:nex-agi/Nex-N2.5-minibase_model:quantized:nex-agi/Nex-N2.5-minilicense:apache-2.0endpoints_compatibleregion:usimatrix

Runs locally from ~857.6 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
9
Pipeline
text-generation
Author

Repository Files & Downloads

18 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Nex-N2.5-mini-IQ1_M.ggufGGUFIQ1_M7.67 GBDownload
Nex-N2.5-mini-IQ1_S.ggufGGUFIQ1_S6.97 GBDownload
Nex-N2.5-mini-IQ2_M.ggufGGUFIQ2_M10.86 GBDownload
Nex-N2.5-mini-IQ2_XS.ggufGGUFIQ2_XS9.79 GBDownload
Nex-N2.5-mini-IQ2_XXS.ggufGGUFIQ2_XXS8.85 GBDownload
Nex-N2.5-mini-IQ3_XXS.ggufGGUFIQ3_XXS12.69 GBDownload
Nex-N2.5-mini-Q2_K.ggufGGUFQ2_K12.05 GBDownload
Nex-N2.5-mini-Q3_K_M.ggufGGUFQ3_K_M15.61 GBDownload
Nex-N2.5-mini-Q3_K_S.ggufGGUFQ3_K_S14.14 GBDownload
Nex-N2.5-mini-Q4_0.ggufGGUFQ4_018.36 GBDownload
Nex-N2.5-mini-Q4_K_M.ggufGGUFQ4_K_M19.71 GBDownload
Nex-N2.5-mini-Q4_K_S.ggufGGUFQ4_K_S18.52 GBDownload
Nex-N2.5-mini-Q5_K_M.ggufGGUFQ5_K_M23.03 GBDownload
Nex-N2.5-mini-Q5_K_S.ggufGGUFQ5_K_S22.33 GBDownload
Nex-N2.5-mini-Q6_K.ggufGGUFQ6_K26.56 GBDownload
Nex-N2.5-mini-Q8_0.ggufGGUFQ8_034.37 GBDownload
Nex-N2.5-mini-TQ1_0.ggufGGUFGGUF7.35 GBDownload
mmproj-Nex-N2.5-mini-F16.ggufGGUFF16857.6 MBDownload

Model Details

Model IDabenzerps/Nex-N2.5-mini-GGUF
Authorabenzerps
Pipelinetext-generation
Licenseapache-2.0
Base modelnex-agi/Nex-N2.5-mini
Last modified2026-09-08T22:02:25.000Z

Model README

---

license: apache-2.0

language:

- en

- zh

base_model: nex-agi/Nex-N2.5-mini

base_model_relation: quantized

pipeline_tag: text-generation

library_name: gguf

tags:

- gguf

- llama.cpp

- qwen3.5

- qwen3.5-moe

- long-context

- 262k-context

- multimodal

- tool-calling

---

Nex-N2.5-mini GGUF

GGUF files for Nex-N2.5-mini, a long-context agentic model for coding, tool use, computer use, and multimodal workloads. The source checkpoint supports a native context length of 262,144 tokens (256K).

Benchmarks

!Nex-N2.5 benchmark results

Benchmark results reported by Nex-AGI for the original Nex-N2.5 family. These figures are not measurements of this GGUF conversion.

GGUF files

| Quantization | File | Size (GB) |

| --- | --- | ---: |

| Q8_0 | Nex-N2.5-mini-Q8_0.gguf | 36.90 GB |

| Q6_K | Nex-N2.5-mini-Q6_K.gguf | 28.51 GB |

| Q5_K_M | Nex-N2.5-mini-Q5_K_M.gguf | 24.73 GB |

| Q5_K_S | Nex-N2.5-mini-Q5_K_S.gguf | 23.98 GB |

| Q4_K_M | Nex-N2.5-mini-Q4_K_M.gguf | 21.17 GB |

| Q4_K_S | Nex-N2.5-mini-Q4_K_S.gguf | 19.89 GB |

| Q4_0 | Nex-N2.5-mini-Q4_0.gguf | 19.72 GB |

| Q3_K_M | Nex-N2.5-mini-Q3_K_M.gguf | 16.76 GB |

| Q3_K_S | Nex-N2.5-mini-Q3_K_S.gguf | 15.18 GB |

| IQ3_XXS | Nex-N2.5-mini-IQ3_XXS.gguf | 13.62 GB |

| Q2_K | Nex-N2.5-mini-Q2_K.gguf | 12.94 GB |

| IQ2_M | Nex-N2.5-mini-IQ2_M.gguf | 11.66 GB |

| IQ2_XS | Nex-N2.5-mini-IQ2_XS.gguf | 10.51 GB |

| IQ2_XXS | Nex-N2.5-mini-IQ2_XXS.gguf | 9.50 GB |

| IQ1_M | Nex-N2.5-mini-IQ1_M.gguf | 8.24 GB |

| TQ1_0 | Nex-N2.5-mini-TQ1_0.gguf | 7.90 GB |

| IQ1_S | Nex-N2.5-mini-IQ1_S.gguf | 7.48 GB |

TQ1_0 is an experimental ternary quantization. IQ1_S and IQ1_M use importance-matrix quantization.

Multimodal projector

| File | Size | Description |

| --- | ---: | --- |

| mmproj-Nex-N2.5-mini-F16.gguf | 899 MB | F16 vision projector for runtimes with multimodal support |

The projector is optional for text-only use. Use it with a current llama.cpp build that supports the model's multimodal path.

Chat template

The GGUF embeds the upstream chat template. An external copy is provided as chat_template.jinja for runtimes that require a separate template file.

Usage

For text generation with llama.cpp:

~~~

llama-cli \

-m Nex-N2.5-mini-Q4_K_M.gguf \

-c 8192 --jinja \

--temp 0.7 --top-p 0.95 \

-p "Explain why reproducible builds matter."

~~~

For an OpenAI-compatible server:

~~~

llama-server \

-m Nex-N2.5-mini-Q4_K_M.gguf \

-c 8192 --jinja --host 0.0.0.0 --port 8080

~~~

Increase -c up to 262144 when sufficient memory is available. Tool-call behavior depends on the serving runtime and its parser integration; use the embedded template and verify tool calls in the target application.

Source and build

Run abenzerps/Nex-N2.5-mini-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models