arelath/Meta-Llama-3-8B-Instruct-nanoquant-GGUF overview
Experiment 27: meta llama/Meta Llama 3 8B Instruct quality benchmark Status: completed Model: meta llama/Meta Llama 3 8B Instruct Revision: 8afb486c1db24fe5011…
Runs locally from ~2.33 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| meta-llama-3-8b-instruct-nanoquant.gguf | GGUF | GGUF | 2.33 GB | Download |
Model Details
| Model ID | arelath/Meta-Llama-3-8B-Instruct-nanoquant-GGUF |
|---|---|
| Author | arelath |
| Pipeline | text-generation |
| License | llama3 |
| Base model | meta-llama/Meta-Llama-3-8B-Instruct |
| Last modified | 2026-07-25T02:37:07.000Z |
Model README
---
base_model: meta-llama/Meta-Llama-3-8B-Instruct
language:
- en
license: llama3
pipeline_tag: text-generation
tags:
- gguf
- nanoquant
- quantized
base_model_relation: quantized
---
Experiment 27: meta-llama/Meta-Llama-3-8B-Instruct quality benchmark
- Status:
completed - Model:
meta-llama/Meta-Llama-3-8B-Instruct - Revision:
8afb486c1db24fe5011ec46dfbe5b5dccdb575c2 - Candidate run:
/workspace/NanoQuant/evidence/027/027-compress-and-benchmark-meta-llama-3-8b-instruct - Backend:
dense - Wall time: 141.53 seconds
completed means all evaluators returned finite metrics; it is not a BF16-quality acceptance gate.
Protocol
- WikiText-2: 64 samples × 128 tokens, batch 8
- WikiText token hash:
sha256:c7dc8810186996593a9a8a419db6bcae8e67e520c2412722425817fc5862a77d - Tasks: piqa, arc_easy, arc_challenge, hellaswag, winogrande, boolq; first 200 rows, batch 4
- Tokenizer hash:
sha256:8aa3159f5493d3660d5f6898b3b1a88d5b1626e4ac1f7dd5b60fed8f916080df
Quality results
| Benchmark | Metric | BF16 | NanoQuant | Delta | Ratio |
| --- | --- | ---: | ---: | ---: | ---: |
| WikiText-2 | perplexity ↓ | 24.956664 | 55.331093 | +30.374429 (+121.71%) | 2.2171x |
| piqa | acc_norm ↑ | 0.7700 | 0.6800 | -0.0900 | 0.8831x |
| arc_easy | acc_norm ↑ | 0.7550 | 0.4850 | -0.2700 | 0.6424x |
| arc_challenge | acc_norm ↑ | 0.5150 | 0.2900 | -0.2250 | 0.5631x |
| hellaswag | acc_norm ↑ | 0.6550 | 0.5500 | -0.1050 | 0.8397x |
| winogrande | acc ↑ | 0.7150 | 0.6350 | -0.0800 | 0.8881x |
| boolq | acc ↑ | 0.8300 | 0.7400 | -0.0900 | 0.8916x |
Runtime and memory
| Model | Elapsed seconds | Peak CUDA bytes | Peak host bytes |
| --- | ---: | ---: | ---: |
| BF16 | 29.73 | 18,194,890,752 | 33,399,840,768 |
| NanoQuant | 35.35 | 18,803,064,832 | 33,399,840,768 |
Provenance
- Experiment config hash:
sha256:801c828f06b50662d77573ed177f639142880d06b9d4f440261babe50ab94b96 - Launcher:
experiments/027-compress-and-benchmark-meta-llama-3-8b-instruct.py - Candidate identity:
{"config_hash":"sha256:044785fa22c02ddeb31c43cffde7172becdc19475a8e342b0a8344613a6ef2fb","model_hash":"sha256:d590dbc8c4a1851df6feb003b377003e7e4ededacc99ed35abb96844d236322d","plan_hash":"sha256-57b3f751fb9b821e18abad776ffc22f4f90d5b6d73378e90037728ca2b1cdcbb"} - Global tuning:
{"artifact_id":"sha256-68defbc90a8e75e2a59b51d40c054cff9bc393012039641fa3365b32787509e7","artifact_type":"global-tuning-result","schema_version":1}
Run arelath/Meta-Llama-3-8B-Instruct-nanoquant-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models