SaniaKhalid/tinyllama-trl-gguf overview
๐ฆ TinyLlama GGUF Quantized Model Model Description This is a GGUF quantized version of TinyLlama 1.1B parameters fine tuned using Unsloth and TRL with LoRA adโฆ
Runs locally from ~1.09 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| TinyLlama-Trl_q8.gguf | GGUF | Q8 | 1.09 GB | Download |
Model Details
| Model ID | SaniaKhalid/tinyllama-trl-gguf |
|---|---|
| Author | SaniaKhalid |
| Pipeline | โ |
| License | apache-2.0 |
| Base model | TinyLlama/TinyLlama-1.1B-Chat-v1.0 |
| Last modified | 2026-09-22T22:01:08.000Z |
Model README
---
language: en
license: apache-2.0
library_name: transformers
tags:
- tinyllama
- gguf
- quantized
- llama-cpp
- unsloth
- lora
- fine-tuned
base_model: TinyLlama/TinyLlama-1.1B-Chat-v1.0
model-index:
- name: tinyllama-trl-gguf
results: []
---
๐ฆ TinyLlama GGUF - Quantized Model
Model Description
This is a GGUF quantized version of TinyLlama (1.1B parameters) fine-tuned using Unsloth and TRL with LoRA adapters. The model has been optimized for efficient CPU inference with minimal memory footprint.
Key Features:
- GGUF Format: Optimized for llama.cpp and CPU inference
- Quantized: Reduced memory usage without significant quality loss
- Fine-tuned: Custom trained on specific dataset for improved performance
- Efficient: Runs on CPU, Raspberry Pi, and mobile devices
Model Details:
- Base Model: TinyLlama/TinyLlama-1.1B-Chat-v1.0
- Fine-tuning Method: LoRA (Low-Rank Adaptation) with Unsloth optimizations
- Format: GGUF (GGML Universal Format)
- Quantization: Q4_K_M (4-bit quantization)
- Parameters: 1.1 Billion
- Context Length: 2048 tokens
๐ ๏ธ Training Details
Base Model
| Parameter | Value |
|-----------|-------|
| Model | TinyLlama-1.1B-Chat-v1.0 |
| Architecture | Llama-based transformer |
| Parameters | 1.1 Billion |
| Context Length | 2048 tokens |
| Attention | Grouped-Query Attention (GQA) |
| Hidden Size | 2048 |
| Intermediate Size | 5632 |
| Number of Layers | 22 |
| Number of Heads | 32 |
| Head Dimension | 64 |
Fine-tuning Configuration
LoRA Parameters
LORA_R = 16 # Rank of LoRA matrices
LORA_ALPHA = 32 # Scaling factor (alpha/r = 2.0)
LORA_DROPOUT = 0.05 # Dropout for regularization
TARGET_MODULES = [ # Layers where LoRA is applied
"q_proj", # Query projection
"k_proj", # Key projection
"v_proj", # Value projection
"o_proj", # Output projection
"gate_proj", # Gate projection (MLP)
"up_proj", # Up projection (MLP)
"down_proj" # Down projection (MLP)
]
## ๐ Usage
### Option 1: Using llama-cpp-python (Recommended)
Install llama-cpp-python
pip install llama-cpp-python
from llama_cpp import Llama
Load the model
model_path = "arif-butt/tinyllama-trl-gguf" # or local path to .gguf file
llm = Llama(
model_path=model_path,
n_ctx=2048, # Context length
n_threads=4, # Number of CPU threads
n_gpu_layers=0, # Set >0 for GPU offloading
verbose=False,
)
Simple prompt
prompt = "Q: Name all the courses Arif butt teach?\nA:"
Generate response
output = llm(
prompt,
max_tokens=100, # Maximum tokens to generate
temperature=0.2, # Lower = more deterministic
top_p=0.95, # Nucleus sampling
repeat_penalty=1.1, # Penalize repetition
stop=["Q:", "\nQ:"], # Stop sequences
)
print(output["choices"][0]["text"])
โโ Simple Chat Interface for GGUF Model โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
from llama_cpp import Llama
import sys
class TinyLlamaChat:
def __init__(self, model_path="arif-butt/tinyllama-trl-gguf"):
"""Initialize the GGUF model"""
print("Loading model...")
self.llm = Llama(
model_path=model_path,
n_ctx=2048,
n_threads=4,
verbose=False,
)
print("โ Model loaded!")
def generate(self, prompt, max_tokens=100, temperature=0.7):
"""Generate response for a single prompt"""
output = self.llm(
prompt,
max_tokens=max_tokens,
temperature=temperature,
top_p=0.95,
repeat_penalty=1.1,
stop=["Q:", "\nQ:", "User:", "\nUser:", "Human:"],
)
return output["choices"][0]["text"]
def chat(self):
"""Interactive chat mode"""
print("\n๐ฌ Chat Mode (type 'quit' to exit)")
print("-" * 50)
while True:
user_input = input("\n๐ค You: ")
if user_input.lower() in ['quit', 'exit', 'q']:
break
prompt = f"Q: {user_input}\nA:"
response = self.generate(prompt, temperature=0.7)
print(f"๐ค Assistant: {response}")
Use the chat interface
if __name__ == "__main__":
chat = TinyLlamaChat()
chat.chat()
Run SaniaKhalid/tinyllama-trl-gguf with guIDE
Download guIDE โ the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face ยท Compare models