GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

thesreedath/gemma-2-2b-qa-raft-GGUF overview

Gemma 2 2B QA SFT + RAG RAFT , GGUF On device offline quantizations of thesreedath/gemma 2 2b qa raft , a LoRA fine tune of google/gemma 2 2b it for retrieval …

onnxggufendpoints_compatibleregion:usconversational

Runs locally from ~1.59 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline

Repository Files & Downloads

2 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
gemma-2-2b-qa-raft-Q4_K_M.ggufGGUFQ4_K_M1.59 GBDownload
gemma-2-2b-qa-raft-Q8_0.ggufGGUFQ8_02.59 GBDownload

Model Details

Model IDthesreedath/gemma-2-2b-qa-raft-GGUF
Authorthesreedath
Pipeline
License
Base model
Last modified2026-07-20T09:14:58.000Z

Model README

Gemma 2 2B - QA-SFT + RAG (RAFT), GGUF

On-device (offline) quantizations of thesreedath/gemma-2-2b-qa-raft, a LoRA fine-tune of google/gemma-2-2b-it for retrieval-augmented QA. Load with any llama.cpp runtime (PocketPal AI, ChatterUI, MLC, Ollama).

| file | size | use |

|---|---|---|

| Q4_K_M | ~1.71 GB | recommended for phones |

| Q8_0 | ~2.78 GB | higher quality, more RAM |

Prompt format: prepend the retrieved context as <context id=i>...</context> blocks, then Question: ..., with the system line: You are a helpful assistant. Answer using ONLY the provided context...

Run thesreedath/gemma-2-2b-qa-raft-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models