thesreedath/gemma-2-2b-qa-raft-GGUF overview
Gemma 2 2B QA SFT + RAG RAFT , GGUF On device offline quantizations of thesreedath/gemma 2 2b qa raft , a LoRA fine tune of google/gemma 2 2b it for retrieval …
Runs locally from ~1.59 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | thesreedath/gemma-2-2b-qa-raft-GGUF |
|---|---|
| Author | thesreedath |
| Pipeline | — |
| License | — |
| Base model | — |
| Last modified | 2026-07-20T09:14:58.000Z |
Model README
Gemma 2 2B - QA-SFT + RAG (RAFT), GGUF
On-device (offline) quantizations of thesreedath/gemma-2-2b-qa-raft, a LoRA fine-tune of google/gemma-2-2b-it for retrieval-augmented QA. Load with any llama.cpp runtime (PocketPal AI, ChatterUI, MLC, Ollama).
| file | size | use |
|---|---|---|
| Q4_K_M | ~1.71 GB | recommended for phones |
| Q8_0 | ~2.78 GB | higher quality, more RAM |
Prompt format: prepend the retrieved context as <context id=i>...</context> blocks, then Question: ..., with the system line: You are a helpful assistant. Answer using ONLY the provided context...
Run thesreedath/gemma-2-2b-qa-raft-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models