GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE โ†’
Model Intelligence Sheet

SaniaKhalid/tinyllama-trl-gguf overview

๐Ÿฆ™ TinyLlama GGUF Quantized Model Model Description This is a GGUF quantized version of TinyLlama 1.1B parameters fine tuned using Unsloth and TRL with LoRA adโ€ฆ

transformersgguftinyllamaquantizedllama-cppunslothlorafine-tunedenbase_model:TinyLlama/TinyLlama-1.1B-Chat-v1.0base_model:adapter:TinyLlama/TinyLlama-1.1B-Chat-v1.0license:apache-2.0endpoints_compatibleregion:usconversationalnot-for-all-audiences

Runs locally from ~1.09 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
โ€”

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
TinyLlama-Trl_q8.ggufGGUFQ81.09 GBDownload

Model Details

Model IDSaniaKhalid/tinyllama-trl-gguf
AuthorSaniaKhalid
Pipelineโ€”
Licenseapache-2.0
Base modelTinyLlama/TinyLlama-1.1B-Chat-v1.0
Last modified2026-09-22T22:01:08.000Z

Model README

---

language: en

license: apache-2.0

library_name: transformers

tags:

  • tinyllama
  • gguf
  • quantized
  • llama-cpp
  • unsloth
  • lora
  • fine-tuned

base_model: TinyLlama/TinyLlama-1.1B-Chat-v1.0

model-index:

  • name: tinyllama-trl-gguf

results: []

---

๐Ÿฆ™ TinyLlama GGUF - Quantized Model

Model Description

This is a GGUF quantized version of TinyLlama (1.1B parameters) fine-tuned using Unsloth and TRL with LoRA adapters. The model has been optimized for efficient CPU inference with minimal memory footprint.

Key Features:

  • GGUF Format: Optimized for llama.cpp and CPU inference
  • Quantized: Reduced memory usage without significant quality loss
  • Fine-tuned: Custom trained on specific dataset for improved performance
  • Efficient: Runs on CPU, Raspberry Pi, and mobile devices

Model Details:

  • Base Model: TinyLlama/TinyLlama-1.1B-Chat-v1.0
  • Fine-tuning Method: LoRA (Low-Rank Adaptation) with Unsloth optimizations
  • Format: GGUF (GGML Universal Format)
  • Quantization: Q4_K_M (4-bit quantization)
  • Parameters: 1.1 Billion
  • Context Length: 2048 tokens

๐Ÿ› ๏ธ Training Details

Base Model

| Parameter | Value |

|-----------|-------|

| Model | TinyLlama-1.1B-Chat-v1.0 |

| Architecture | Llama-based transformer |

| Parameters | 1.1 Billion |

| Context Length | 2048 tokens |

| Attention | Grouped-Query Attention (GQA) |

| Hidden Size | 2048 |

| Intermediate Size | 5632 |

| Number of Layers | 22 |

| Number of Heads | 32 |

| Head Dimension | 64 |

Fine-tuning Configuration

LoRA Parameters

LORA_R       = 16        # Rank of LoRA matrices
LORA_ALPHA   = 32        # Scaling factor (alpha/r = 2.0)
LORA_DROPOUT = 0.05      # Dropout for regularization
TARGET_MODULES = [       # Layers where LoRA is applied
    "q_proj",            # Query projection
    "k_proj",            # Key projection  
    "v_proj",            # Value projection
    "o_proj",            # Output projection
    "gate_proj",         # Gate projection (MLP)
    "up_proj",           # Up projection (MLP)
    "down_proj"          # Down projection (MLP)
]

## ๐Ÿš€ Usage

### Option 1: Using llama-cpp-python (Recommended)

Install llama-cpp-python

pip install llama-cpp-python

from llama_cpp import Llama

Load the model

model_path = "arif-butt/tinyllama-trl-gguf" # or local path to .gguf file

llm = Llama(

model_path=model_path,

n_ctx=2048, # Context length

n_threads=4, # Number of CPU threads

n_gpu_layers=0, # Set >0 for GPU offloading

verbose=False,

)

Simple prompt

prompt = "Q: Name all the courses Arif butt teach?\nA:"

Generate response

output = llm(

prompt,

max_tokens=100, # Maximum tokens to generate

temperature=0.2, # Lower = more deterministic

top_p=0.95, # Nucleus sampling

repeat_penalty=1.1, # Penalize repetition

stop=["Q:", "\nQ:"], # Stop sequences

)

print(output["choices"][0]["text"])

โ”€โ”€ Simple Chat Interface for GGUF Model โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€

from llama_cpp import Llama

import sys

class TinyLlamaChat:

def __init__(self, model_path="arif-butt/tinyllama-trl-gguf"):

"""Initialize the GGUF model"""

print("Loading model...")

self.llm = Llama(

model_path=model_path,

n_ctx=2048,

n_threads=4,

verbose=False,

)

print("โœ… Model loaded!")

def generate(self, prompt, max_tokens=100, temperature=0.7):

"""Generate response for a single prompt"""

output = self.llm(

prompt,

max_tokens=max_tokens,

temperature=temperature,

top_p=0.95,

repeat_penalty=1.1,

stop=["Q:", "\nQ:", "User:", "\nUser:", "Human:"],

)

return output["choices"][0]["text"]

def chat(self):

"""Interactive chat mode"""

print("\n๐Ÿ’ฌ Chat Mode (type 'quit' to exit)")

print("-" * 50)

while True:

user_input = input("\n๐Ÿ‘ค You: ")

if user_input.lower() in ['quit', 'exit', 'q']:

break

prompt = f"Q: {user_input}\nA:"

response = self.generate(prompt, temperature=0.7)

print(f"๐Ÿค– Assistant: {response}")

Use the chat interface

if __name__ == "__main__":

chat = TinyLlamaChat()

chat.chat()

Run SaniaKhalid/tinyllama-trl-gguf with guIDE

Download guIDE โ€” the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE โ†’ ยท Browse 524k+ models ยท Compare models

Source: Hugging Face ยท Compare models