RemySkye/rwkv7-g1h-13.3b-GGUF overview
rwkv7 g1h 13.3b 20260710 ctx10240 GGUF GGUF conversions of rwkv7 g1h 13.3b 20260710 ctx10240.pth https://huggingface.co/BlinkDL/rwkv7 g1/blob/6d5762253b343eec6…
Runs locally from ~5.07 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| rwkv7-g1h-13.3b-20260710-ctx10240-BF16.gguf | GGUF | BF16 | 24.91 GB | Download |
| rwkv7-g1h-13.3b-20260710-ctx10240-IQ4_XS.gguf | GGUF | IQ4_XS | 7.45 GB | Download |
| rwkv7-g1h-13.3b-20260710-ctx10240-Q2_K.gguf | GGUF | Q2_K | 5.07 GB | Download |
| rwkv7-g1h-13.3b-20260710-ctx10240-Q3_K_L.gguf | GGUF | Q3_K_L | 7.83 GB | Download |
| rwkv7-g1h-13.3b-20260710-ctx10240-Q3_K_M.gguf | GGUF | Q3_K_M | 7.14 GB | Download |
| rwkv7-g1h-13.3b-20260710-ctx10240-Q3_K_S.gguf | GGUF | Q3_K_S | 6.26 GB | Download |
| rwkv7-g1h-13.3b-20260710-ctx10240-Q4_0.gguf | GGUF | Q4_0 | 7.81 GB | Download |
| rwkv7-g1h-13.3b-20260710-ctx10240-Q4_1.gguf | GGUF | Q4_1 | 8.54 GB | Download |
| rwkv7-g1h-13.3b-20260710-ctx10240-Q4_K_M.gguf | GGUF | Q4_K_M | 8.48 GB | Download |
| rwkv7-g1h-13.3b-20260710-ctx10240-Q4_K_S.gguf | GGUF | Q4_K_S | 7.81 GB | Download |
| rwkv7-g1h-13.3b-20260710-ctx10240-Q5_0.gguf | GGUF | Q5_0 | 9.27 GB | Download |
| rwkv7-g1h-13.3b-20260710-ctx10240-Q5_1.gguf | GGUF | Q5_1 | 10.00 GB | Download |
| rwkv7-g1h-13.3b-20260710-ctx10240-Q5_K_M.gguf | GGUF | Q5_K_M | 9.62 GB | Download |
| rwkv7-g1h-13.3b-20260710-ctx10240-Q5_K_S.gguf | GGUF | Q5_K_S | 9.27 GB | Download |
| rwkv7-g1h-13.3b-20260710-ctx10240-Q6_K.gguf | GGUF | Q6_K | 10.83 GB | Download |
| rwkv7-g1h-13.3b-20260710-ctx10240-Q8_0.gguf | GGUF | Q8_0 | 13.72 GB | Download |
Model Details
| Model ID | RemySkye/rwkv7-g1h-13.3b-GGUF |
|---|---|
| Author | RemySkye |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | BlinkDL/rwkv7-g1 |
| Last modified | 2026-07-26T06:46:54.000Z |
Model README
---
license: apache-2.0
base_model: BlinkDL/rwkv7-g1
library_name: llama.cpp
pipeline_tag: text-generation
tags:
- rwkv
- rwkv7
- gguf
- quantized
---
rwkv7-g1h-13.3b-20260710-ctx10240 GGUF
GGUF conversions of rwkv7-g1h-13.3b-20260710-ctx10240.pth. This repository is one model only; the other parameter sizes are published in separate repositories.
Conversion
The source checkpoint is BF16. It was converted directly from .pth to a BF16 GGUF with RWKV's converter, then the standard release ladder below was produced from that BF16 GGUF using the pinned llama.cpp quantizer. No safetensors staging file and no importance matrix were used.
- Source revision:
6d5762253b343eec6cfbf5ed62f872f30a4cd89c - rwkv-mobile revision:
ebfb744281c31a07aad5606ec7473f79f837e92a - llama.cpp revision:
c92e806d1c81091c9035edce99c35374da1b465e - Context marker in source filename:
ctx10240
Files
BF16master GGUFQ2_KQ3_K_SIQ4_XSQ4_K_SQ5_K_SQ6_KQ8_0Q4_0Q4_1Q5_0Q5_1
| Quant | PPL | BF16 retained |
| -------- | -------: | ------------: |
| BF16 | 5.363 | 100.00% |
| Q8_0 | 5.367 | 99.93% |
| Q6_K | 5.381 | 99.67% |
| Q5_1 | 5.4362 | 98.65% |
| Q5_K_M | 5.4277 | 98.81% |
| Q5_K_S | 5.4329 | 98.71% |
| Q5_0 | 5.4536 | 98.34% |
| Q4_1 | 5.6231 | 95.37% |
| Q4_K_M | 5.5459 | 96.70% |
| Q4_K_S | 5.5857 | 96.01% |
| Q4_0 | 5.6735 | 94.53% |
| IQ4_XS | 5.5776 | 96.15% |
| Q3_K_L | 5.7207 | 93.75% |
| Q3_K_M | 5.8051 | 92.38% |
| Q3_K_S | 6.2314 | 86.06% |
| Q2_K | 271.9763 | 1.97% |
RWKV-aware mixed quantizations
The Q3_K_M, Q3_K_L, Q4_K_M, and Q5_K_M files use custom RWKV-aware recipes with explicit tensor assignments. Higher precision is used for the token embeddings and selected value, time-mix output, and channel-mix tensors where it is expected to preserve the most quality.
Earlier automated files with these names were removed after verification showed that llama.cpp's generic mixed recipes did not recognize RWKV's time_mix_ and channel_mix_ tensor roles. Because of that, the automated M and L variants had collapsed to the same effective layouts as the retained S variants.
These replacement files have genuinely different tensor layouts, providing additional size and quality choices between the existing S variants and the larger quantizations.
Prompting
Follow the prompt guide in the upstream model card. Avoid a trailing space at the end of the input.
Run RemySkye/rwkv7-g1h-13.3b-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models