thesreedath/slm-125m-raft-rlaif-GGUF overview
SLM 125M RAFT RLAIF closed book + RAG, GGUF On device offline quantizations of thesreedath/slm 125m raft rlaif , our best 125M model with RAG standard Llama ar…
Runs locally from ~75.6 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | thesreedath/slm-125m-raft-rlaif-GGUF |
|---|---|
| Author | thesreedath |
| Pipeline | — |
| License | — |
| Base model | — |
| Last modified | 2026-07-24T11:52:41.000Z |
Model README
SLM 125M - RAFT (RLAIF) closed-book + RAG, GGUF
On-device (offline) quantizations of thesreedath/slm-125m-raft-rlaif, our best 125M model with RAG (standard Llama architecture, 12 layers, 768 hidden, 16K vocab). Load with any llama.cpp runtime (Ollama, PocketPal AI, LM Studio, ChatterUI).
| file | size | use |
|---|---|---|
| Q4_K_M | ~79.2 MB | phone default |
| Q8_0 | ~134.3 MB | higher quality |
| F16 | ~252.3 MB | laptop quality ceiling |
For RAG, prepend retrieved passages as context blocks before the question.
Run thesreedath/slm-125m-raft-rlaif-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models