Adiuk/Llama-3.3-70B-Instruct-Q3_K_M-GGUF overview
Llama 3.3 70B Instruct Q3 K M GGUF Low bit GGUF quantizations of meta llama/Llama 3.3 70B Instruct https://huggingface.co/meta llama/Llama 3.3 70B Instruct , p…
Runs locally from ~31.91 GB disk (32 GB+ VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Llama-3.3-70B-Instruct-GGUF-Q3_K_M.gguf | GGUF | Q3_K_M | 31.91 GB | Download |
Model Details
Model README
---
license: llama3.3
base_model: meta-llama/Llama-3.3-70B-Instruct
tags: [gguf, quantized, aditurbo, llama.cpp]
---
Llama-3.3-70B-Instruct-Q3_K_M-GGUF
Low-bit GGUF quantizations of meta-llama/Llama-3.3-70B-Instruct,
produced with the AdiTurbo Engine (private llama.cpp fork, build 8661).
Quantized with a fresh 50-chunk wikitext-2 importance matrix. Intended for
offline/on-device inference (Eyla AIOS). See the AdiTurbo benchmark for
perplexity/throughput numbers.
Files
Llama-3.3-70B-Instruct-GGUF-Q3_K_M.gguf
This is a derivative quantization; the original model's license (llama3.3) applies.
Run Adiuk/Llama-3.3-70B-Instruct-Q3_K_M-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models