GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Hob-forge/gpt-oss-20b-Q2_K-GGUF overview

GPT OSS 20B Q2 K GGUF 12GB VRAM Optimized Aggressively quantized version of OpenAI's GPT OSS 20B for 12GB VRAM GPUs with CPU offload. Why This Exists The offic…

ggufquantizedgpt-ossmoeq2_kllama.cppollamalm-studio16gbtext-generationbase_model:openai/gpt-oss-20bbase_model:quantized:openai/gpt-oss-20blicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~10.68 GB disk (12 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
891
Likes
1
Pipeline
text-generation
Author

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
gpt-oss-20b-Q2_K.ggufGGUFQ2_K10.68 GBDownload

Model Details

Model IDHob-forge/gpt-oss-20b-Q2_K-GGUF
AuthorHob-forge
Pipelinetext-generation
Licenseapache-2.0
Base modelopenai/gpt-oss-20b
Last modified2026-08-22T17:48:48.000Z

Model README

---

base_model: openai/gpt-oss-20b

base_model_relation: quantized

license: apache-2.0

pipeline_tag: text-generation

library_name: gguf

tags:

  • gguf
  • quantized
  • gpt-oss
  • moe
  • q2_k
  • llama.cpp
  • ollama
  • lm-studio
  • 16gb

---

GPT-OSS 20B - Q2_K GGUF (12GB VRAM Optimized)

Aggressively quantized version of OpenAI's GPT-OSS 20B for 12GB VRAM GPUs with CPU offload.

Why This Exists

The official GPT-OSS 20B requires 16GB VRAM. This Q2_K quantization runs comfortably on:

  • RTX 3080 (12GB)
  • RTX 4070 (12GB)
  • RTX 5070 (12GB)
  • Any 12GB+ GPU with CPU offload

Fast inference with GPU/CPU split - not just "it works" but actually usable for real tasks.

Quick Start

Ollama

# Download and run
ollama run Hob-forge/gpt-oss-20b-Q2_K-GGUF

llama.cpp

# With GPU offload (adjust layers based on your VRAM)
./llama-cli -m gpt-oss-20b-Q2_K.gguf -ngl 28 -c 4096

LM Studio

Just download and load - it will auto-detect optimal settings.

Model Details

| Property | Value |

|----------|-------|

| Parameters | 20.9B |

| Quantization | Q2_K |

| File Size | ~11GB |

| Context Length | 131,072 (use 4096-8192 for speed) |

| Architecture | GPT-OSS (MoE) |

Recommended Settings

num_gpu: 28        # Layers on GPU (adjust for your VRAM)
num_ctx: 4096      # Context window (increase if needed)
temperature: 0.5   # Good balance for most tasks

For 12GB VRAM, num_gpu: 28 leaves room for context. Reduce if you need larger context windows.

Performance Notes

  • Q2_K is aggressive quantization - expect some quality loss vs FP16
  • Still excellent for coding, reasoning, and general tasks
  • The speed/quality tradeoff is worth it for consumer hardware
  • Works great as a local coding assistant or agent backbone

Original Model

This is a quantized version of openai/gpt-oss-20b.

License

Apache 2.0 (same as original model)

Credits

  • Original model by OpenAI
  • Quantization for 12GB VRAM hardware

Run Hob-forge/gpt-oss-20b-Q2_K-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models