newbeeforever/notavox-qwen3.5-0.8b-gguf overview
Notavox Qwen3.5 0.8B GGUF This repository mirrors the Qwen3.5 0.8B Q4 K M.gguf quantization from unsloth/Qwen3.5 0.8B GGUF https://huggingface.co/unsloth/Qwen3…
Runs locally from ~507.8 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Qwen3.5-0.8B-Q4_K_M.gguf | GGUF | Q4_K_M | 507.8 MB | Download |
Model Details
| Model ID | newbeeforever/notavox-qwen3.5-0.8b-gguf |
|---|---|
| Author | newbeeforever |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | Qwen/Qwen3.5-0.8B |
| Last modified | 2026-07-28T13:24:18.000Z |
Model README
---
license: apache-2.0
base_model: Qwen/Qwen3.5-0.8B
library_name: llama.cpp
tags:
- gguf
- qwen3.5
- notavox
- on-device
- text-generation
---
Notavox Qwen3.5 0.8B GGUF
This repository mirrors the Qwen3.5-0.8B-Q4_K_M.gguf quantization from
for Notavox on-device, text-only voice-note summarization.
Included file
| File | Size | SHA-256 |
|---|---:|---|
| Qwen3.5-0.8B-Q4_K_M.gguf | 532,517,120 bytes | bd258782e35f7f458f8aced1adc053e6e92e89bc735ba3be89d38a06121dc517 |
Source revision: 6ab461498e2023f6e3c1baea90a8f0fe38ab64d0.
The visual projector is intentionally not included because Notavox uses this model for text-only transcript summarization. Recommended mobile runtime settings: 8,192-token context, non-thinking mode, and chunked summarization for long transcripts.
Run newbeeforever/notavox-qwen3.5-0.8b-gguf with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models