arelath/gemma-3-1b-it-nanoquant-GGUF overview
Experiment 17: google/gemma 3 1b it quality benchmark Status: completed Model: google/gemma 3 1b it Revision: dcc83ea841ab6100d6b47a070329e1ba4cf78752 Candidat…
Runs locally from ~398.0 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| gemma-3-1b-it-nanoquant.gguf | GGUF | GGUF | 398.0 MB | Download |
Model Details
Model README
Experiment 17: google/gemma-3-1b-it quality benchmark
- Status:
completed - Model:
google/gemma-3-1b-it - Revision:
dcc83ea841ab6100d6b47a070329e1ba4cf78752 - Candidate run:
D:\dev\research\NanoQuantRewrite\evidence\017\017-compress-and-benchmark-gemma-3-1b-it - Backend:
factorized - Wall time: 128.88 seconds
completed means all evaluators returned finite metrics; it is not a BF16-quality acceptance gate.
Protocol
- WikiText-2: 64 samples × 128 tokens, batch 8
- WikiText token hash:
sha256:ef19dc950344a837a1fd6e087c451ed9b26234408e85d0b0e3da4f6c7045ff27 - Tasks: piqa, arc_easy, arc_challenge, hellaswag, winogrande, boolq; first 200 rows, batch 4
- Tokenizer hash:
sha256:19317db471b30f6cfa877d781ecac1db28de6628e44e3751df0c44344444a811
Quality results
| Benchmark | Metric | BF16 | NanoQuant | Delta | Ratio |
| --- | --- | ---: | ---: | ---: | ---: |
| WikiText-2 | perplexity ↓ | 96.459609 | 274.912464 | +178.452856 (+185.00%) | 2.8500x |
| piqa | acc_norm ↑ | 0.7200 | 0.6100 | -0.1100 | 0.8472x |
| arc_easy | acc_norm ↑ | 0.6050 | 0.4000 | -0.2050 | 0.6612x |
| arc_challenge | acc_norm ↑ | 0.3950 | 0.2150 | -0.1800 | 0.5443x |
| hellaswag | acc_norm ↑ | 0.5850 | 0.4350 | -0.1500 | 0.7436x |
| winogrande | acc ↑ | 0.6250 | 0.4950 | -0.1300 | 0.7920x |
| boolq | acc ↑ | 0.8100 | 0.6150 | -0.1950 | 0.7593x |
Runtime and memory
| Model | Elapsed seconds | Peak CUDA bytes | Peak host bytes |
| --- | ---: | ---: | ---: |
| BF16 | 41.73 | 8,923,381,760 | 11,226,685,440 |
| NanoQuant | 76.10 | 8,940,158,976 | 11,226,685,440 |
Provenance
- Experiment config hash:
sha256:8fdbdbf81669a24f50146bca0ff2e696e0b438b67522a1cb413d127ff752e9c1 - Launcher:
experiments/017-compress-and-benchmark-gemma-3-1b-it.py - Candidate identity:
{"config_hash":"sha256:76a4f59a532cc3ae4cdac3654b92dd57007b1a9d22b8e08df622b522ab2f0e9f","model_hash":"sha256:32d5b5d041e98027bc7415107bc79b580f9cce407535b4e30134e8f8aed3b130","plan_hash":"sha256-1751998bfae1249ac1f32b104a44f1994b5403e491da71623b1aa5e4cfb49a15"} - Global tuning:
{"artifact_id":"sha256-b8a13d87903dc23188886a7c979d2d00444ec413feec5715124ec0b7173d5504","artifact_type":"global-tuning-result","schema_version":1}
Run arelath/gemma-3-1b-it-nanoquant-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models