abenzerps/Spark-X2.5-4B-GGUF overview
IMPORTANT Compatibility: These GGUF files require llama.cpp b10828 or later, which includes official support for the Spark X2.5 spark2 5 architecture. Applicat…
Runs locally from ~2.24 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | abenzerps/Spark-X2.5-4B-GGUF |
|---|---|
| Author | abenzerps |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | XHToken/Spark-X2.5-4B |
| Last modified | 2026-09-07T11:26:51.000Z |
Model README
---
base_model:
- XHToken/Spark-X2.5-4B
license: apache-2.0
library_name: gguf
pipeline_tag: text-generation
tags:
- gguf
- llama.cpp
- spark-x2.5
- long-context
- 1m-context
---
> [!IMPORTANT]
> Compatibility: These GGUF files require llama.cpp b10828 or later, which includes official support for the Spark-X2.5 (spark2_5) architecture. Applications with a bundled runtime must use an equivalent or newer build. llama.cpp support
Spark-X2.5-4B GGUF
GGUF quantizations of XHToken/Spark-X2.5-4B, a 4B general-purpose language model for reasoning, coding, tool use, and agentic workflows. Native context: 1,048,576 tokens (1M).
Benchmarks
!Spark-X2.5-4B benchmark results
Benchmark results reported by XHToken for Spark-X2.5-4B in thinking mode.
GGUF files
| Quantization | File | Size |
| --- | --- | ---: |
| Q4_0 | Spark-X2.5-4B-Q4_0.gguf | 2.41 GB |
| Q4_K_M | Spark-X2.5-4B-Q4_K_M.gguf | 2.60 GB |
| Q5_K_M | Spark-X2.5-4B-Q5_K_M.gguf | 2.98 GB |
| Q6_K | Spark-X2.5-4B-Q6_K.gguf | 3.38 GB |
| Q8_0 | Spark-X2.5-4B-Q8_0.gguf | 4.38 GB |
Includes the upstream chat_template.jinja. Checksums: SHA256SUMS.txt.
Usage
llama-cli -m Spark-X2.5-4B-Q4_K_M.gguf -c 131072 -cnv
Source
- Model: XHToken/Spark-X2.5-4B
- Revision:
ea14618d20e76b5b093d3ee20a5b9d733bb12410 - License: Apache-2.0
Run abenzerps/Spark-X2.5-4B-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models