GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

arelath/gemma-3-1b-it-nanoquant-GGUF overview

Experiment 17: google/gemma 3 1b it quality benchmark Status: completed Model: google/gemma 3 1b it Revision: dcc83ea841ab6100d6b47a070329e1ba4cf78752 Candidat…

ggufendpoints_compatibleregion:usconversational

Runs locally from ~398.0 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
Author

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
gemma-3-1b-it-nanoquant.ggufGGUFGGUF398.0 MBDownload

Model Details

Model IDarelath/gemma-3-1b-it-nanoquant-GGUF
Authorarelath
Pipeline
License
Base model
Last modified2026-07-19T07:11:42.000Z

Model README

Experiment 17: google/gemma-3-1b-it quality benchmark

  • Status: completed
  • Model: google/gemma-3-1b-it
  • Revision: dcc83ea841ab6100d6b47a070329e1ba4cf78752
  • Candidate run: D:\dev\research\NanoQuantRewrite\evidence\017\017-compress-and-benchmark-gemma-3-1b-it
  • Backend: factorized
  • Wall time: 128.88 seconds

completed means all evaluators returned finite metrics; it is not a BF16-quality acceptance gate.

Protocol

  • WikiText-2: 64 samples × 128 tokens, batch 8
  • WikiText token hash: sha256:ef19dc950344a837a1fd6e087c451ed9b26234408e85d0b0e3da4f6c7045ff27
  • Tasks: piqa, arc_easy, arc_challenge, hellaswag, winogrande, boolq; first 200 rows, batch 4
  • Tokenizer hash: sha256:19317db471b30f6cfa877d781ecac1db28de6628e44e3751df0c44344444a811

Quality results

| Benchmark | Metric | BF16 | NanoQuant | Delta | Ratio |

| --- | --- | ---: | ---: | ---: | ---: |

| WikiText-2 | perplexity ↓ | 96.459609 | 274.912464 | +178.452856 (+185.00%) | 2.8500x |

| piqa | acc_norm ↑ | 0.7200 | 0.6100 | -0.1100 | 0.8472x |

| arc_easy | acc_norm ↑ | 0.6050 | 0.4000 | -0.2050 | 0.6612x |

| arc_challenge | acc_norm ↑ | 0.3950 | 0.2150 | -0.1800 | 0.5443x |

| hellaswag | acc_norm ↑ | 0.5850 | 0.4350 | -0.1500 | 0.7436x |

| winogrande | acc ↑ | 0.6250 | 0.4950 | -0.1300 | 0.7920x |

| boolq | acc ↑ | 0.8100 | 0.6150 | -0.1950 | 0.7593x |

Runtime and memory

| Model | Elapsed seconds | Peak CUDA bytes | Peak host bytes |

| --- | ---: | ---: | ---: |

| BF16 | 41.73 | 8,923,381,760 | 11,226,685,440 |

| NanoQuant | 76.10 | 8,940,158,976 | 11,226,685,440 |

Provenance

  • Experiment config hash: sha256:8fdbdbf81669a24f50146bca0ff2e696e0b438b67522a1cb413d127ff752e9c1
  • Launcher: experiments/017-compress-and-benchmark-gemma-3-1b-it.py
  • Candidate identity: {"config_hash":"sha256:76a4f59a532cc3ae4cdac3654b92dd57007b1a9d22b8e08df622b522ab2f0e9f","model_hash":"sha256:32d5b5d041e98027bc7415107bc79b580f9cce407535b4e30134e8f8aed3b130","plan_hash":"sha256-1751998bfae1249ac1f32b104a44f1994b5403e491da71623b1aa5e4cfb49a15"}
  • Global tuning: {"artifact_id":"sha256-b8a13d87903dc23188886a7c979d2d00444ec413feec5715124ec0b7173d5504","artifact_type":"global-tuning-result","schema_version":1}

Run arelath/gemma-3-1b-it-nanoquant-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models