GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

schneewolflabs/B1.1-9B-GGUF overview

B1.1 9B — GGUF Q8 0 the quant every card number was measured on and vision mmproj for schneewolflabs/B1.1 9B https://huggingface.co/schneewolflabs/B1.1 9B — B1…

ggufagentstool-usereasoningqwen3.5enbase_model:schneewolflabs/B1.1-9Bbase_model:quantized:schneewolflabs/B1.1-9Blicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~879.0 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline

Repository Files & Downloads

2 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
B1.1-9B-Q8_0.ggufGGUFQ8_09.11 GBDownload
B1.1-9B-mmproj-f16.ggufGGUFF16879.0 MBDownload

Model Details

Model IDschneewolflabs/B1.1-9B-GGUF
Authorschneewolflabs
Pipeline
Licenseapache-2.0
Base modelschneewolflabs/B1.1-9B
Last modified2026-09-06T06:37:42.000Z

Model README

---

base_model: [schneewolflabs/B1.1-9B]

license: apache-2.0

language: [en]

tags: [gguf, agents, tool-use, reasoning, qwen3.5]

---

B1.1-9B — GGUF

Q8_0 (the quant every card number was measured on) and vision mmproj for

schneewolflabs/B1.1-9B — B1's

answer-after-thinking adapter merged at half strength, keeping most of the fix and most of B0's

persona.

llama-server -m B1.1-9B-Q8_0.gguf -ngl 99 -c 8192 --jinja -fa on -np 1 \
    --spec-type draft-mtp --spec-draft-n-max 4 \
    --mmproj B1.1-9B-mmproj-f16.gguf

Same architecture and layer count as B0-9B, so the

B0-9B-GGUF deployment guide — offload rules,

f16 KV cache, 6GB recipes — applies unchanged. With thinking on, pass tool definitions through

the native tools field and do not prefix /think; that request shape is where the

answer-after-thinking gain is largest.

Run schneewolflabs/B1.1-9B-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models