Michionlion/Astrea-R8-Chat-9B-GGUF overview
Astrea R8 Chat 9B — Q8 0 GGUF This is an unofficial community Q8 0 GGUF conversion of Altworld/Astrea R8 Chat 9B https://huggingface.co/Altworld/Astrea R8 Chat…
Runs locally from ~8.87 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Astrea-R8-Chat-9B-Q8_0.gguf | GGUF | Q8_0 | 8.87 GB | Download |
Model Details
| Model ID | Michionlion/Astrea-R8-Chat-9B-GGUF |
|---|---|
| Author | Michionlion |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | Altworld/Astrea-R8-Chat-9B |
| Last modified | 2026-07-21T14:40:00.000Z |
Model README
---
base_model: Altworld/Astrea-R8-Chat-9B
license: apache-2.0
library_name: gguf
pipeline_tag: text-generation
tags:
- gguf
- llama.cpp
- qwen3.5
- chat
- creative-writing
- non-reasoning
---
Astrea R8 Chat 9B — Q8_0 GGUF
This is an unofficial community Q8_0 GGUF conversion of
for llama.cpp, with an explicit non-reasoning chat template.
File
| File | Quantization | Size |
|---|---:|---:|
| Astrea-R8-Chat-9B-Q8_0.gguf | Q8_0 | 9,527,501,280 bytes (8.87 GiB) |
Embedded non-reasoning template
The GGUF contains a hard non-reasoning Jinja template in
tokenizer.chat_template; no external template file is required.
Astrea's optional thinking mode was not reliable in local llama.cpp testing:
simple prompts could consume hundreds of tokens before emitting </think>, and often did not end reasoning at all, and simply responded as if reasoning was not enabled.
The bundled template therefore always places a closed, empty thinking block in
the prompt and does not expose an enable_thinking template variable. It also
omits hidden reasoning when replaying assistant messages into conversation
history. A standalone copy is included as chat_template.jinja for inspection.
llama.cpp
llama-server.exe `
--model Astrea-R8-Chat-9B-Q8_0.gguf `
--jinja `
--reasoning off `
--reasoning-format none `
--ctx-size 32768 `
--n-gpu-layers all `
--temp 0.8 `
--top-p 1.0 `
--top-k 0 `
--min-p 0.025 `
--repeat-penalty 1.08
The model metadata advertises a 262,144-token context window. Choose a context
size appropriate for your available VRAM/RAM. The command above starts at a
more conservative 32,768 tokens. I was able to easily run a much more ambitous setup with -ngl all --fit off -c 147456 -np 4 --kv-unified on a 16GB VRAM card (5070 Ti).
Conversion notes
- The original safetensors were converted to BF16 GGUF with llama.cpp's
convert_hf_to_gguf.py using --no-mtp. The downloaded checkpoint did not
contain the extra MTP-layer tensors declared by its configuration.
- BF16 was quantized with
llama-quantizeusingQ8_0. - llama.cpp's
gguf_new_metadata.pyembedded the hard non-reasoning template;
this metadata-only copy did not requantize tensors.
The tensor-only SHA-256 reported by llama-gguf-hash was identical before and
after the metadata rewrite:
20d213a0c5ee663ef6d02ffcff8d0b28cbff18b559c67aca5250cd5e6a22d624
The final whole-file checksums are in SHA256SUMS.
Validation
The final GGUF was loaded directly by llama-server without
--chat-template-file. Its exposed template matched the bundled standalone
Jinja, and a request that explicitly supplied enable_thinking=true still
returned normal content with no reasoning_content.
License and attribution
The source model is released under Apache-2.0. See LICENSE and NOTICE, and
refer to the source model card
for its intended use, evaluation results, and limitations.
Run Michionlion/Astrea-R8-Chat-9B-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models