GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

thesreedath/slm-125m-raft-rlaif-GGUF overview

SLM 125M RAFT RLAIF closed book + RAG, GGUF On device offline quantizations of thesreedath/slm 125m raft rlaif , our best 125M model with RAG standard Llama ar…

tfliteggufendpoints_compatibleregion:usconversational

Runs locally from ~75.6 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline

Repository Files & Downloads

3 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
slm-125m-raft-rlaif-Q4_K_M.ggufGGUFQ4_K_M75.6 MBDownload
slm-125m-raft-rlaif-Q8_0.ggufGGUFQ8_0128.1 MBDownload
slm-125m-raft-rlaif-f16.ggufGGUFF16240.6 MBDownload

Model Details

Model IDthesreedath/slm-125m-raft-rlaif-GGUF
Authorthesreedath
Pipeline
License
Base model
Last modified2026-07-24T11:52:41.000Z

Model README

SLM 125M - RAFT (RLAIF) closed-book + RAG, GGUF

On-device (offline) quantizations of thesreedath/slm-125m-raft-rlaif, our best 125M model with RAG (standard Llama architecture, 12 layers, 768 hidden, 16K vocab). Load with any llama.cpp runtime (Ollama, PocketPal AI, LM Studio, ChatterUI).

| file | size | use |

|---|---|---|

| Q4_K_M | ~79.2 MB | phone default |

| Q8_0 | ~134.3 MB | higher quality |

| F16 | ~252.3 MB | laptop quality ceiling |

For RAG, prepend retrieved passages as context blocks before the question.

Run thesreedath/slm-125m-raft-rlaif-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models