tinyopsec/TinyLlama-1.1B-Chat-v1.0-pruned2.4-GGUF overview
TinyLlama 1.1B Chat v1.0 pruned2.4 GGUF GGUF quantizations of RedHatAI/TinyLlama 1.1B Chat v1.0 pruned2.4 https://huggingface.co/RedHatAI/TinyLlama 1.1B Chat v…
Runs locally from ~412.1 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| model_f16.gguf | GGUF | F16 | 2.05 GB | Download |
| model_q2_k.gguf | GGUF | Q2_K | 412.1 MB | Download |
| model_q3_k_l.gguf | GGUF | Q3_K_L | 564.1 MB | Download |
| model_q3_k_m.gguf | GGUF | Q3_K_M | 523.0 MB | Download |
| model_q3_k_s.gguf | GGUF | Q3_K_S | 476.2 MB | Download |
| model_q4_k_m.gguf | GGUF | Q4_K_M | 636.9 MB | Download |
| model_q4_k_s.gguf | GGUF | Q4_K_S | 610.2 MB | Download |
| model_q5_k_m.gguf | GGUF | Q5_K_M | 745.8 MB | Download |
| model_q5_k_s.gguf | GGUF | Q5_K_S | 730.5 MB | Download |
| model_q6_k.gguf | GGUF | Q6_K | 861.6 MB | Download |
| model_q8_0.gguf | GGUF | Q8_0 | 1.09 GB | Download |
Model Details
| Model ID | tinyopsec/TinyLlama-1.1B-Chat-v1.0-pruned2.4-GGUF |
|---|---|
| Author | tinyopsec |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | RedHatAI/TinyLlama-1.1B-Chat-v1.0-pruned2.4 |
| Last modified | 2026-09-15T14:11:19.000Z |
Model README
---
license: apache-2.0
base_model: RedHatAI/TinyLlama-1.1B-Chat-v1.0-pruned2.4
tags:
- llama
- gguf
- sparse
- pruned
- quantized
- conversational
language:
- en
pipeline_tag: text-generation
library_name: gguf
---
TinyLlama-1.1B-Chat-v1.0-pruned2.4 GGUF
GGUF quantizations of RedHatAI/TinyLlama-1.1B-Chat-v1.0-pruned2.4, a 2:4 semi-structured sparse version of TinyLlama-1.1B-Chat-v1.0 pruned with SparseGPT via SparseML.
About the Original Model
TinyLlama-1.1B-Chat-v1.0-pruned2.4 is a compressed variant of TinyLlama 1.1B Chat, pruned to 2:4 sparsity (50% sparse) using the SparseGPT one-shot method. The pruning was performed by RedHat AI using SparseML on the open_platypus calibration dataset. The model retains conversational capability while being significantly more hardware-efficient.
Architecture: LlamaForCausalLM
Parameters: ~1.1B
Context length: 2048 tokens
Sparsity: 2:4 semi-structured (50%)
Quantization Files
| File | Bits | Approx Size | Use Case |
|------|------|-------------|----------|
| model_f16.gguf | 16 | ~2.2 GB | Maximum quality, reference |
| model_q8_0.gguf | 8 | ~1.2 GB | Best quality / size tradeoff |
| model_q6_k.gguf | 6 | ~0.9 GB | High quality, smaller |
| model_q5_k_m.gguf | 5 | ~0.8 GB | Balanced quality |
| model_q5_k_s.gguf | 5 | ~0.75 GB | Slightly smaller Q5 |
| model_q4_k_m.gguf | 4 | ~0.67 GB | Good quality, low RAM |
| model_q4_k_s.gguf | 4 | ~0.63 GB | Smaller Q4 |
| model_q3_k_l.gguf | 3 | ~0.55 GB | Low RAM, acceptable quality |
| model_q3_k_m.gguf | 3 | ~0.51 GB | Lower RAM |
| model_q3_k_s.gguf | 3 | ~0.47 GB | Minimum RAM Q3 |
| model_q2_k.gguf | 2 | ~0.38 GB | Extreme compression |
VRAM Requirements (Estimated)
| Quantization | VRAM |
|--------------|------|
| F16 | ~2.4 GB |
| Q8_0 | ~1.4 GB |
| Q6_K | ~1.1 GB |
| Q5_K_M | ~0.95 GB |
| Q4_K_M | ~0.85 GB |
| Q3_K_M | ~0.70 GB |
| Q2_K | ~0.55 GB |
Prompt Template
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
Inference
llama.cpp
./llama-cli -m model_q4_k_m.gguf -p "<|im_start|>user\nHello!<|im_end|>\n<|im_start|>assistant\n" -n 256
llama-cpp-python
from llama_cpp import Llama
llm = Llama(model_path="model_q4_k_m.gguf", n_ctx=2048)
output = llm.create_chat_completion(
messages=[{"role": "user", "content": "Hello!"}]
)
print(output["choices"][0]["message"]["content"])
LM Studio
- Open LM Studio
- Search for
tinyopsec/TinyLlama-1.1B-Chat-v1.0-pruned2.4-GGUF - Download the desired quantization
- Load and chat
Ollama
ollama run hf.co/tinyopsec/TinyLlama-1.1B-Chat-v1.0-pruned2.4-GGUF:Q4_K_M
Links
- Original model: RedHatAI/TinyLlama-1.1B-Chat-v1.0-pruned2.4
- Base model: TinyLlama/TinyLlama-1.1B-Chat-v1.0
- SparseGPT paper: arxiv:2301.00774
Run tinyopsec/TinyLlama-1.1B-Chat-v1.0-pruned2.4-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models