arelath/Llama-3.2-1B-Instruct-nanoquant-GGUF overview
Experiment 25: meta llama/Llama 3.2 1B Instruct quality benchmark Status: completed Model: meta llama/Llama 3.2 1B Instruct Revision: 9213176726f574b556790deb6…
Runs locally from ~392.4 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| llama-3-2-1b-instruct-nanoquant.gguf | GGUF | GGUF | 392.4 MB | Download |
Model Details
Model README
Experiment 25: meta-llama/Llama-3.2-1B-Instruct quality benchmark
- Status:
completed - Model:
meta-llama/Llama-3.2-1B-Instruct - Revision:
9213176726f574b556790deb65791e0c5aa438b6 - Candidate run:
D:\dev\research\NanoQuantRewrite\evidence\025\025-compress-and-benchmark-llama-3-2-1b-instruct - Backend:
dense - Wall time: 50.07 seconds
completed means all evaluators returned finite metrics; it is not a BF16-quality acceptance gate.
Protocol
- WikiText-2: 64 samples × 128 tokens, batch 8
- WikiText token hash:
sha256:c7dc8810186996593a9a8a419db6bcae8e67e520c2412722425817fc5862a77d - Tasks: piqa, arc_easy, arc_challenge, hellaswag, winogrande, boolq; first 200 rows, batch 4
- Tokenizer hash:
sha256:5409af4b5ead403c8c413b60287460703373a37222ce25ce929035b49b81719c
Quality results
| Benchmark | Metric | BF16 | NanoQuant | Delta | Ratio |
| --- | --- | ---: | ---: | ---: | ---: |
| WikiText-2 | perplexity ↓ | 36.856393 | 116.980145 | +80.123753 (+217.39%) | 3.1739x |
| piqa | acc_norm ↑ | 0.7350 | 0.6650 | -0.0700 | 0.9048x |
| arc_easy | acc_norm ↑ | 0.6200 | 0.4250 | -0.1950 | 0.6855x |
| arc_challenge | acc_norm ↑ | 0.3300 | 0.2950 | -0.0350 | 0.8939x |
| hellaswag | acc_norm ↑ | 0.6050 | 0.4550 | -0.1500 | 0.7521x |
| winogrande | acc ↑ | 0.6150 | 0.5550 | -0.0600 | 0.9024x |
| boolq | acc ↑ | 0.7500 | 0.6450 | -0.1050 | 0.8600x |
Runtime and memory
| Model | Elapsed seconds | Peak CUDA bytes | Peak host bytes |
| --- | ---: | ---: | ---: |
| BF16 | 19.98 | 5,920,260,096 | 4,135,366,656 |
| NanoQuant | 18.02 | 6,192,889,856 | 4,869,566,464 |
Provenance
- Experiment config hash:
sha256:5be8cce6ef0ec17fcd90ecf10c7971523799070d4c61eeda245207a1ac69b319 - Launcher:
experiments/025-compress-and-benchmark-llama-3-2-1b-instruct.py - Candidate identity:
{"config_hash":"sha256:a282bec0f20d082888d7322301dc95ef57f39a8b110317c9499aea8313acb3c4","model_hash":"sha256:0d4bbbbc32aa6ccf91325a258c47d5bf3839604149b6e081cb561c6a65e58f51","plan_hash":"sha256-3a9a46309740f9452a117627cb32f36f21f2fc9558d918aac3b601b5a0258da8"} - Global tuning:
{"artifact_id":"sha256-5c4902d094ed2d025b63e3f24ba3b9b0ccca3e5bc0541da64bc6a3def1ca0b4a","artifact_type":"global-tuning-result","schema_version":1}
Run arelath/Llama-3.2-1B-Instruct-nanoquant-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models