Ttimms/Bible-Assistant-Qwen3.5-4B-v2-GGUF overview
Superseded 2026 09 05 . Bible Assistant Qwen3.5 4B v3.2 GGUF https://huggingface.co/Ttimms/Bible Assistant Qwen3.5 4B v3.2 GGUF is the current GGUF release sam…
Runs locally from ~2.34 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| bible-v2-4b-IQ4_XS-imat.gguf | GGUF | IQ4_XS | 2.34 GB | Download |
| bible-v2-4b-Q4_K_M-imat.gguf | GGUF | Q4_K_M | 2.52 GB | Download |
| bible-v2-4b-Q4_K_M.gguf | GGUF | Q4_K_M | 2.52 GB | Download |
| bible-v2-4b-Q5_K_M.gguf | GGUF | Q5_K_M | 2.86 GB | Download |
| bible-v2-4b-Q6_K.gguf | GGUF | Q6_K | 3.23 GB | Download |
| bible-v2-4b-Q8_0.gguf | GGUF | Q8_0 | 4.17 GB | Download |
| bible-v2-4b-f16.gguf | GGUF | F16 | 7.85 GB | Download |
Model Details
| Model ID | Ttimms/Bible-Assistant-Qwen3.5-4B-v2-GGUF |
|---|---|
| Author | Ttimms |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | Qwen/Qwen3.5-4B,Ttimms/Bible-Assistant-Qwen3.5-4B-v2 |
| Last modified | 2026-09-05T08:24:23.000Z |
Model README
---
license: apache-2.0
base_model:
- Qwen/Qwen3.5-4B
- Ttimms/Bible-Assistant-Qwen3.5-4B-v2
base_model_relation: quantized
tags:
- bible
- rag
- gguf
- llama.cpp
- qwen3.5
- text-generation
language:
- en
pipeline_tag: text-generation
quantized_by: Ttimms
license_link: https://huggingface.co/Qwen/Qwen3.5-4B/blob/main/LICENSE
---
> Superseded (2026-09-05). Bible-Assistant-Qwen3.5-4B-v3.2-GGUF
> is the current GGUF release -- same size class, better on every metric tested.
> This v2 quant set is kept for reproducibility, not as the recommended download.
Bible AI Assistant v2-4b — GGUF
GGUF quants of Ttimms/Bible-Assistant-Qwen3.5-4B-v2
— a Qwen3.5-4B SFT for retrieval-grounded Bible Q&A. See the base repo for the full
model card, training details, and the honest evaluation (this is an **interim
checkpoint**: strong on verbatim verse recall, weaker on open-ended thematic
answers; a v3 with teacher-distilled answers + GRPO is planned).
Architecture
graph TD
Base["Qwen/Qwen3.5-4B"]
SFT["bf16 LoRA SFT - 56k-example dataset"]
Merge["merge adapter -> bf16"]
Conv["convert_hf_to_gguf --no-mtp + llama-quantize (+imatrix)"]
ST["Bible-Assistant-Qwen3.5-4B-v2 (safetensors)"]
GG["...-v2-GGUF (Q4_K_M / Q5_K_M / Q6_K / Q8_0 / F16 +imat)"]
RAG["hybrid RAG: dense (nomic) + BM25 + RRF + bge-reranker-v2-m3"]
LLM["Ollama / llama.cpp"]
Base --> SFT --> Merge --> ST
Merge --> Conv --> GG
ST --> RAG --> LLM
GG --> LLM
Download
Grab one file, not the whole repo. -imat files use an importance matrix
(better quality at the same size).
| File | Quant | Size | Notes |
|---|---|--:|---|
| bible-v2-4b-Q4_K_M-imat.gguf | Q4_K_M + imatrix | 2.7 GB | recommended |
| bible-v2-4b-IQ4_XS-imat.gguf | IQ4_XS + imatrix | 2.5 GB | smallest usable |
| bible-v2-4b-Q5_K_M.gguf | Q5_K_M | 3.1 GB | |
| bible-v2-4b-Q6_K.gguf | Q6_K | 3.5 GB | |
| bible-v2-4b-Q8_0.gguf | Q8_0 | 4.5 GB | near-lossless |
| bible-v2-4b-f16.gguf | F16 | 8.4 GB | full precision |
Run it in
- llama.cpp —
llama-server -m <file>.gguf -ngl 99 - LM Studio (bundles a recent llama.cpp)
- koboldcpp
- Jan
- text-generation-webui
- Ollama — once its bundled llama.cpp includes this arch (see Requirements)
Requirements
Qwen3.5 is a hybrid architecture (qwen35 / Gated-DeltaNet + attention). You
need a recent llama.cpp — a build that includes the qwen35 hybrid arch
(commit 3173a56 or newer). Verified working with llama-server from a source
build.
- ✅ llama.cpp (current):
llama-server -m bible-v2-4b-Q4_K_M.gguf -ngl 99 - ✅ LM Studio (recent versions bundle a current llama.cpp)
- ⚠️ Ollama 0.33.x: the bundled llama.cpp is too old for the
qwen35arch
(check_tensor_dims: tensor 'blk.32.attn_norm.weight' not found). Use once
Ollama updates its runtime, or run llama.cpp directly.
Thinking mode
The Qwen3.5 chat template defaults to thinking on. This model was SFT'd
without <think> traces, so for a grounded RAG assistant you want it off:
- llama.cpp
/v1/chat/completions: pass
"chat_template_kwargs": {"enable_thinking": false}, or
- use a chat template that emits a closed empty
<think>\n\n</think>\n\nafter
<|im_start|>assistant\n (the project's deployment/pc/Modelfile does this).
Intended use
Retrieval-augmented Bible Q&A — the model expects retrieved verses in a Context:
block, then the question. It is not designed for context-free use, medical /
legal / financial advice, counselling (it redirects those to a pastor / crisis
line), or authoritative theological rulings.
License
Weights: Apache-2.0 (inherits from Qwen3.5-4B). Project code: MIT. Bible
translations: public domain. See the base repo for full attribution.
Run Ttimms/Bible-Assistant-Qwen3.5-4B-v2-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models