XHToken/Spark-X2.5-1.7B-GGUF overview
Spark X2.5 1.7B GGUF NOTE This repository provides a BF16 GGUF conversion of Spark X2.5 1.7B. Spark X2.5 is a compact, general purpose language model for conve…
Runs locally from ~1.03 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | XHToken/Spark-X2.5-1.7B-GGUF |
|---|---|
| Author | XHToken |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | XHToken/Spark-X2.5-1.7B |
| Last modified | 2026-09-07T09:36:28.000Z |
Model README
---
license: apache-2.0
language:
- en
- zh
library_name: gguf
pipeline_tag: text-generation
base_model: XHToken/Spark-X2.5-1.7B
tags:
- gguf
- llama.cpp
- ollama
- lm-studio
- sparkx2_5
---
Spark-X2.5-1.7B-GGUF
> [!NOTE]
> This repository provides a BF16 GGUF conversion of Spark-X2.5-1.7B.
Spark-X2.5 is a compact, general-purpose language model for conversation, writing, translation, reasoning, coding, tool use, and agentic workflows. It uses a hybrid attention architecture, supports a native context length of up to 1M tokens, and covers more than 200 languages. For its architecture, training methods, benchmark results, fine-tuning, and citation, see the Spark-X2.5-1.7B.
Local Deployment
The GGUF file can be used for local inference with Ollama and LM Studio. Spark-X2.5 support is provided by XHToken/llama.cpp, so the Quick Starts below use this compatible implementation.
Ollama Quick Start
Build
git clone https://github.com/XHToken/llama.cpp.git llama.cpp-spark
git clone https://github.com/ollama/ollama.git ollama-spark
cd ollama-spark
export OLLAMA_LLAMA_CPP_SOURCE="$(cd ../llama.cpp-spark && pwd)"
cmake -S . -B build
cmake --build build --parallel 8
Import the GGUF
Replace the model path below with the absolute path to the downloaded GGUF file:
printf 'FROM /absolute/path/to/Spark-X2.5-1.7B.gguf\n' > ./Modelfile.spark
Create and Run
Start the Ollama server in the first terminal:
./ollama serve
Open a second terminal in the same ollama-spark directory:
./ollama create Spark-X2.5-1.7B -f ./Modelfile.spark
./ollama run Spark-X2.5-1.7B --think=false
--think=false disables thinking mode for faster, direct responses.
LM Studio Quick Start
Build the Compatible llama.cpp Runtime
git clone https://github.com/XHToken/llama.cpp.git llama.cpp-spark
cd llama.cpp-spark
cmake -S . -B build
cmake --build build --parallel 8
Configure LM Studio
- Close LM Studio.
- Back up the selected LM Studio runtime directory:
```text
<LM_STUDIO_HOME>/extensions/backends/<selected-runtime>/
```
- Copy the
llama.cpp-sparkbuild output into the selected runtime directory, replacing the existing runtime files. - Place
Spark-X2.5-1.7B.ggufin:
```text
<LM_STUDIO_HOME>/models/<org>/<name>/
```
Example runtime directory on Apple Silicon:
./build/bin/* -> ~/.lmstudio/extensions/backends/llama.cpp-mac-arm64-apple-metal-advsimd-<version>/
Run
Open LM Studio, select the model under My Models, click Load, and start a new chat.
You can also use the lms CLI:
lms ls
lms load <model>
lms chat <model>
License
Released under the Apache License 2.0.
Run XHToken/Spark-X2.5-1.7B-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models