thesreedath/slm-500m-raft-dpo-GGUF overview
SLM 500M RAFT RLAIF closed book + RAG, GGUF On device offline quantizations of thesreedath/slm 500m raft dpo , our best 500M model with RAG standard Llama arch…
Runs locally from ~325.9 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | thesreedath/slm-500m-raft-dpo-GGUF |
|---|---|
| Author | thesreedath |
| Pipeline | — |
| License | — |
| Base model | — |
| Last modified | 2026-07-24T11:14:12.000Z |
Model README
SLM 500M - RAFT (RLAIF) closed-book + RAG, GGUF
On-device (offline) quantizations of thesreedath/slm-500m-raft-dpo, our best 500M model with RAG (standard Llama architecture, 24 layers, 1280 hidden, 32K vocab). Load with any llama.cpp runtime (Ollama, PocketPal AI, LM Studio, ChatterUI).
| file | size | use |
|---|---|---|
| Q4_K_M | ~341.7 MB | phone default |
| Q8_0 | ~551.5 MB | higher quality |
| F16 | ~1036.9 MB | laptop quality ceiling |
For RAG, prepend retrieved passages as context blocks before the question.
Run thesreedath/slm-500m-raft-dpo-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models