RemySkye/rwkv7-g1h-2.9b-i1-GGUF overview
rwkv7 g1h 2.9b i1 GGUF Model specific imatrix quants generated from rwkv7 g1h 2.9b 20260710 ctx10240 BF16.gguf https://huggingface.co/RemySkye/rwkv7 g1h 2.9b G…
Runs locally from ~4.2 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| rwkv7-g1h-2.9b-20260710-ctx10240.i1-IQ1_M.gguf | GGUF | IQ1_M | 905.1 MB | Download |
| rwkv7-g1h-2.9b-20260710-ctx10240.i1-IQ1_S.gguf | GGUF | IQ1_S | 848.9 MB | Download |
| rwkv7-g1h-2.9b-20260710-ctx10240.i1-IQ2_M.gguf | GGUF | IQ2_M | 1.14 GB | Download |
| rwkv7-g1h-2.9b-20260710-ctx10240.i1-IQ2_S.gguf | GGUF | IQ2_S | 1.06 GB | Download |
| rwkv7-g1h-2.9b-20260710-ctx10240.i1-IQ2_XS.gguf | GGUF | IQ2_XS | 1.05 GB | Download |
| rwkv7-g1h-2.9b-20260710-ctx10240.i1-IQ2_XXS.gguf | GGUF | IQ2_XXS | 998.9 MB | Download |
| rwkv7-g1h-2.9b-20260710-ctx10240.i1-IQ3_S.gguf | GGUF | IQ3_S | 1.41 GB | Download |
| rwkv7-g1h-2.9b-20260710-ctx10240.i1-IQ3_XXS.gguf | GGUF | IQ3_XXS | 1.28 GB | Download |
| rwkv7-g1h-2.9b-20260710-ctx10240.i1-IQ4_NL.gguf | GGUF | IQ4_NL | 1.75 GB | Download |
| rwkv7-g1h-2.9b-20260710-ctx10240.i1-IQ4_XS.gguf | GGUF | IQ4_XS | 1.67 GB | Download |
| rwkv7-g1h-2.9b-20260710-ctx10240.i1-Q2_K.gguf | GGUF | Q2_K | 1.16 GB | Download |
| rwkv7-g1h-2.9b-20260710-ctx10240.i1-Q3_K_L.gguf | GGUF | Q3_K_L | 1.78 GB | Download |
| rwkv7-g1h-2.9b-20260710-ctx10240.i1-Q3_K_M.gguf | GGUF | Q3_K_M | 1.64 GB | Download |
| rwkv7-g1h-2.9b-20260710-ctx10240.i1-Q3_K_S.gguf | GGUF | Q3_K_S | 1.41 GB | Download |
| rwkv7-g1h-2.9b-20260710-ctx10240.i1-Q4_0.gguf | GGUF | Q4_0 | 1.75 GB | Download |
| rwkv7-g1h-2.9b-20260710-ctx10240.i1-Q4_1.gguf | GGUF | Q4_1 | 1.90 GB | Download |
| rwkv7-g1h-2.9b-20260710-ctx10240.i1-Q4_K_M.gguf | GGUF | Q4_K_M | 1.91 GB | Download |
| rwkv7-g1h-2.9b-20260710-ctx10240.i1-Q4_K_S.gguf | GGUF | Q4_K_S | 1.75 GB | Download |
| rwkv7-g1h-2.9b-20260710-ctx10240.i1-Q5_0.gguf | GGUF | Q5_0 | 2.06 GB | Download |
| rwkv7-g1h-2.9b-20260710-ctx10240.i1-Q5_1.gguf | GGUF | Q5_1 | 2.22 GB | Download |
| rwkv7-g1h-2.9b-20260710-ctx10240.i1-Q5_K_M.gguf | GGUF | Q5_K_M | 2.15 GB | Download |
| rwkv7-g1h-2.9b-20260710-ctx10240.i1-Q5_K_S.gguf | GGUF | Q5_K_S | 2.06 GB | Download |
| rwkv7-g1h-2.9b-20260710-ctx10240.i1-Q6_K.gguf | GGUF | Q6_K | 2.39 GB | Download |
| rwkv7-g1h-2.9b-20260710-ctx10240.i1-imatrix.gguf | GGUF | GGUF | 4.2 MB | Download |
Model Details
| Model ID | RemySkye/rwkv7-g1h-2.9b-i1-GGUF |
|---|---|
| Author | RemySkye |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | BlinkDL/rwkv7-g1 |
| Last modified | 2026-07-26T06:50:32.000Z |
Model README
---
license: apache-2.0
base_model: BlinkDL/rwkv7-g1
library_name: llama.cpp
pipeline_tag: text-generation
tags:
- rwkv
- rwkv7
- gguf
- quantized
- imatrix
---
rwkv7-g1h-2.9b i1 GGUF
Model-specific imatrix quants generated from rwkv7-g1h-2.9b-20260710-ctx10240-BF16.gguf. The BF16 master remains in the linked static-GGUF repository and is not duplicated here.
Calibration
- Dataset:
lemon07r/pile-calibration-v5 - Dataset revision:
ea863bb930b9959dd78095c165aa1376d14e698b - llama.cpp revision:
c92e806d1c81091c9035edce99c35374da1b465e - Context per chunk:
1024tokens - Chunks:
512 - Approximate evaluated tokens:
524288 - Output-weight collection: disabled, following llama.cpp's default recommendation
- Raw JSONL SHA-256:
54d1f2bd6a80cc75a72e7f8a62e23438e23f06287aab3855c5a7826127fae7be - Deterministically curated corpus SHA-256:
8c7de8adad55f0be5e00f394b7c7efca7c6d13fdd4a2f476bc08c0491fe98027
The source dataset is diverse and duplicate-free, but contains a few book-length outliers. Before calibration, corrupt/severely repetitive records are removed, long records are capped using four separated excerpts, records are deterministically shuffled, and rare scripts are lightly interleaved into the early calibration window.
Files
| Quant | PPL | BF16 retained |
| ------------ | ---------: | ------------: |
| i1-Q6_K | 6.8764 | 99.62% |
| i1-Q5_1 | 6.952 | 98.54% |
| i1-Q5_K_M | 6.9247 | 98.93% |
| i1-Q5_K_S | 6.9551 | 98.49% |
| i1-Q5_0 | 7.179 | 95.42% |
| i1-Q4_1 | 23.2667 | 29.44% |
| i1-Q4_K_M | 6.9945 | 97.94% |
| i1-Q4_K_S | 16.5402 | 41.42% |
| i1-Q4_0 | 231.5514 | 2.96% |
| i1-IQ4_NL | 8.8656 | 77.27% |
| i1-IQ4_XS | 9.0385 | 75.79% |
| i1-Q3_K_L | 7.2029 | 95.10% |
| i1-Q3_K_M | 7.231 | 94.74% |
| i1-IQ3_S | 554.0051 | 1.24% |
| i1-Q3_K_S | 8229.9054 | 0.0832% |
| i1-Q2_K | 12573.2758 | 0.0545% |
| i1-IQ3_XXS | 475.9699 | 1.44% |
| i1-IQ2_M | 1073.0608 | 0.638% |
| i1-IQ2_S | 1242.044 | 0.552% |
| i1-IQ2_XS | 42504.01 | 0.0161% |
| i1-IQ2_XXS | 49975.9054 | 0.0137% |
| i1-IQ1_M | 22106.1854 | 0.0310% |
| i1-IQ1_S | 23790.2745 | 0.0288% |
RWKV-aware mixed quantizations
The Q3_K_M, Q3_K_L, Q4_K_M, and Q5_K_M files use custom RWKV-aware recipes with explicit tensor assignments. Higher precision is used for the token embeddings and selected value, time-mix output, and channel-mix tensors where it is expected to preserve the most quality.
Earlier automated files with these names were removed after verification showed that llama.cpp's generic mixed recipes did not recognize RWKV's time_mix_ and channel_mix_ tensor roles. Because of that, the automated M and L variants had collapsed to the same effective layouts as the retained S variants.
These replacement files have genuinely different tensor layouts, providing additional size and quality choices between the existing S variants and the larger quantizations.
The files in this i1 repository use the same RWKV-aware tensor layouts together with this model's existing importance matrix.
Run RemySkye/rwkv7-g1h-2.9b-i1-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models