Model Intelligence Sheet
ibm-granite/granite-speech-4.1-2b-plus-GGUF overview
Granite Speech 4.1 2B Plus GGUF NOTE This repository contains models that have been converted to the GGUF format with various quantizations from an IBM Granite…
Runs locally from ~976.2 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
6 GGUF files detected
Direct downloads for local inference
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| granite-speech-4.1-2b-plus-Q4_K_M.gguf | GGUF | Q4_K_M | 976.2 MB | Download |
| granite-speech-4.1-2b-plus-Q5_K_M.gguf | GGUF | Q5_K_M | 1.10 GB | Download |
| granite-speech-4.1-2b-plus-Q6_K.gguf | GGUF | Q6_K | 1.25 GB | Download |
| granite-speech-4.1-2b-plus-Q8_0.gguf | GGUF | Q8_0 | 1.62 GB | Download |
| granite-speech-4.1-2b-plus-bf16.gguf | GGUF | BF16 | 3.04 GB | Download |
| mmproj-model-f16.gguf | GGUF | F16 | 1.09 GB | Download |
Model Details
| Model ID | ibm-granite/granite-speech-4.1-2b-plus-GGUF |
|---|---|
| Author | ibm-granite |
| Pipeline | — |
| License | apache-2.0 |
| Base model | ibm-granite/granite-speech-4.1-2b-plus |
| Last modified | 2026-06-30T22:40:17.000Z |
Model README
---
license: apache-2.0
library_name: llama.cpp
tags:
- language
- granite-4.1
- speech
- gguf
base_model:
- ibm-granite/granite-speech-4.1-2b-plus
---
Granite-Speech-4.1-2B-Plus (GGUF)
> [!NOTE]
> This repository contains models that have been converted to the GGUF format with various quantizations from an IBM Granite base model.
>
> Please reference the base model's full model card here:
> https://huggingface.co/ibm-granite/granite-speech-4.1-2b-plus
Requirements
- llama.cpp build: b9768
Run ibm-granite/granite-speech-4.1-2b-plus-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models