abenzerps/Nex-N2.5-mini-GGUF overview
Nex N2.5 mini GGUF GGUF files for Nex N2.5 mini https://huggingface.co/nex agi/Nex N2.5 mini , a long context agentic model for coding, tool use, computer use,…
Runs locally from ~857.6 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Nex-N2.5-mini-IQ1_M.gguf | GGUF | IQ1_M | 7.67 GB | Download |
| Nex-N2.5-mini-IQ1_S.gguf | GGUF | IQ1_S | 6.97 GB | Download |
| Nex-N2.5-mini-IQ2_M.gguf | GGUF | IQ2_M | 10.86 GB | Download |
| Nex-N2.5-mini-IQ2_XS.gguf | GGUF | IQ2_XS | 9.79 GB | Download |
| Nex-N2.5-mini-IQ2_XXS.gguf | GGUF | IQ2_XXS | 8.85 GB | Download |
| Nex-N2.5-mini-IQ3_XXS.gguf | GGUF | IQ3_XXS | 12.69 GB | Download |
| Nex-N2.5-mini-Q2_K.gguf | GGUF | Q2_K | 12.05 GB | Download |
| Nex-N2.5-mini-Q3_K_M.gguf | GGUF | Q3_K_M | 15.61 GB | Download |
| Nex-N2.5-mini-Q3_K_S.gguf | GGUF | Q3_K_S | 14.14 GB | Download |
| Nex-N2.5-mini-Q4_0.gguf | GGUF | Q4_0 | 18.36 GB | Download |
| Nex-N2.5-mini-Q4_K_M.gguf | GGUF | Q4_K_M | 19.71 GB | Download |
| Nex-N2.5-mini-Q4_K_S.gguf | GGUF | Q4_K_S | 18.52 GB | Download |
| Nex-N2.5-mini-Q5_K_M.gguf | GGUF | Q5_K_M | 23.03 GB | Download |
| Nex-N2.5-mini-Q5_K_S.gguf | GGUF | Q5_K_S | 22.33 GB | Download |
| Nex-N2.5-mini-Q6_K.gguf | GGUF | Q6_K | 26.56 GB | Download |
| Nex-N2.5-mini-Q8_0.gguf | GGUF | Q8_0 | 34.37 GB | Download |
| Nex-N2.5-mini-TQ1_0.gguf | GGUF | GGUF | 7.35 GB | Download |
| mmproj-Nex-N2.5-mini-F16.gguf | GGUF | F16 | 857.6 MB | Download |
Model Details
| Model ID | abenzerps/Nex-N2.5-mini-GGUF |
|---|---|
| Author | abenzerps |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | nex-agi/Nex-N2.5-mini |
| Last modified | 2026-09-08T22:02:25.000Z |
Model README
---
license: apache-2.0
language:
- en
- zh
base_model: nex-agi/Nex-N2.5-mini
base_model_relation: quantized
pipeline_tag: text-generation
library_name: gguf
tags:
- gguf
- llama.cpp
- qwen3.5
- qwen3.5-moe
- long-context
- 262k-context
- multimodal
- tool-calling
---
Nex-N2.5-mini GGUF
GGUF files for Nex-N2.5-mini, a long-context agentic model for coding, tool use, computer use, and multimodal workloads. The source checkpoint supports a native context length of 262,144 tokens (256K).
Benchmarks
Benchmark results reported by Nex-AGI for the original Nex-N2.5 family. These figures are not measurements of this GGUF conversion.
GGUF files
| Quantization | File | Size (GB) |
| --- | --- | ---: |
| Q8_0 | Nex-N2.5-mini-Q8_0.gguf | 36.90 GB |
| Q6_K | Nex-N2.5-mini-Q6_K.gguf | 28.51 GB |
| Q5_K_M | Nex-N2.5-mini-Q5_K_M.gguf | 24.73 GB |
| Q5_K_S | Nex-N2.5-mini-Q5_K_S.gguf | 23.98 GB |
| Q4_K_M | Nex-N2.5-mini-Q4_K_M.gguf | 21.17 GB |
| Q4_K_S | Nex-N2.5-mini-Q4_K_S.gguf | 19.89 GB |
| Q4_0 | Nex-N2.5-mini-Q4_0.gguf | 19.72 GB |
| Q3_K_M | Nex-N2.5-mini-Q3_K_M.gguf | 16.76 GB |
| Q3_K_S | Nex-N2.5-mini-Q3_K_S.gguf | 15.18 GB |
| IQ3_XXS | Nex-N2.5-mini-IQ3_XXS.gguf | 13.62 GB |
| Q2_K | Nex-N2.5-mini-Q2_K.gguf | 12.94 GB |
| IQ2_M | Nex-N2.5-mini-IQ2_M.gguf | 11.66 GB |
| IQ2_XS | Nex-N2.5-mini-IQ2_XS.gguf | 10.51 GB |
| IQ2_XXS | Nex-N2.5-mini-IQ2_XXS.gguf | 9.50 GB |
| IQ1_M | Nex-N2.5-mini-IQ1_M.gguf | 8.24 GB |
| TQ1_0 | Nex-N2.5-mini-TQ1_0.gguf | 7.90 GB |
| IQ1_S | Nex-N2.5-mini-IQ1_S.gguf | 7.48 GB |
TQ1_0 is an experimental ternary quantization. IQ1_S and IQ1_M use importance-matrix quantization.
Multimodal projector
| File | Size | Description |
| --- | ---: | --- |
| mmproj-Nex-N2.5-mini-F16.gguf | 899 MB | F16 vision projector for runtimes with multimodal support |
The projector is optional for text-only use. Use it with a current llama.cpp build that supports the model's multimodal path.
Chat template
The GGUF embeds the upstream chat template. An external copy is provided as chat_template.jinja for runtimes that require a separate template file.
Usage
For text generation with llama.cpp:
~~~
llama-cli \
-m Nex-N2.5-mini-Q4_K_M.gguf \
-c 8192 --jinja \
--temp 0.7 --top-p 0.95 \
-p "Explain why reproducible builds matter."
~~~
For an OpenAI-compatible server:
~~~
llama-server \
-m Nex-N2.5-mini-Q4_K_M.gguf \
-c 8192 --jinja --host 0.0.0.0 --port 8080
~~~
Increase -c up to 262144 when sufficient memory is available. Tool-call behavior depends on the serving runtime and its parser integration; use the embedded template and verify tool calls in the target application.
Source and build
- Source model: nex-agi/Nex-N2.5-mini
- Source revision: 87420286149d9cce9bd46cd335ef9bda33c37c1b
- Conversion: ggml-org/llama.cpp commit f3f1a8f2760f28325a5ec20c05b171e5b7c83a29
- License: Apache-2.0
- Checksums: SHA256SUMS.txt
Run abenzerps/Nex-N2.5-mini-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models