Soulfate24/ReaderLM-v2-ASHQ1-Remix-GGUF overview
ReaderLM v2 ASHQ1 Remix This is a GGUF quantized version of the original model. 📈 Release Benchmarks wiki.test.raw, symmetric FA auto reference | Model | Size…
Runs locally from ~2.0 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| imatrix.gguf | GGUF | GGUF | 2.0 MB | Download |
| model-ASHQ1-Compact-33pc.gguf | GGUF | GGUF | 979.7 MB | Download |
| model-ASHQ1-Fidelity-48pc.gguf | GGUF | GGUF | 1.39 GB | Download |
| model-ASHQ1-Mini-30pc.gguf | GGUF | GGUF | 890.7 MB | Download |
| model-ASHQ1-Nano-27pc.gguf | GGUF | GGUF | 802.5 MB | Download |
| model-ASHQ1-Pico-24pc.gguf | GGUF | GGUF | 763.0 MB | Download |
| model-ASHQ1-Precision-42pc.gguf | GGUF | GGUF | 1.24 GB | Download |
| model-ASHQ1-Quality-36pc.gguf | GGUF | GGUF | 1.05 GB | Download |
Model Details
| Model ID | Soulfate24/ReaderLM-v2-ASHQ1-Remix-GGUF |
|---|---|
| Author | Soulfate24 |
| Pipeline | text-generation |
| License | cc-by-nc-4.0 |
| Base model | jinaai/ReaderLM-v2 |
| Last modified | 2026-09-10T10:52:40.000Z |
Model README
---
pipeline_tag: text-generation
language:
- multilingual
inference: false
license: cc-by-nc-4.0
library_name: transformers
base_model:
- jinaai/ReaderLM-v2
tags:
- quantization
- gguf
- ashq1
- imatrix
---
ReaderLM-v2 - ASHQ1-Remix
This is a GGUF quantized version of the original model.
📈 Release Benchmarks (wiki.test.raw, symmetric FA-auto reference)
| Model | Size | PPL | KLD | RMS Δp | top-p | Speed |
| :--- | ---: | ---: | ---: | ---: | ---: | ---: |
| Q8_0 (stock) | 1570 MiB | 15.9804 | 0.0022 | 1.15% | 97.6% | 4655 t/s |
| Fidelity-48pc | 1422 MiB | 15.9874 | 0.0044 | 1.58% | 96.4% | 4267 t/s |
| Precision-42pc 🥈 | 1268 MiB | 15.9426† | 0.0067 | 1.98% | 95.6% | 3939 t/s |
| Q6_K-imx (stock) | 1214 MiB | 15.9578† | 0.0074 | 2.08% | 95.2% | 4267 t/s |
| Quality-36pc ⭐ | 1073 MiB | 15.9779† | 0.0223 | 3.71% | 92.2% | 4267 t/s |
| Q5_K_M-imx (stock) | 1073 MiB | 15.9779† | 0.0223 | 3.71% | 92.2% | 4267 t/s |
| Compact-33pc | 980 MiB | 15.9972 | 0.0504 | 5.53% | 88.6% | 3939 t/s |
| Mini-30pc | 891 MiB | 16.0512 | 0.0751 | 6.61% | 86.3% | 4267 t/s |
| IQ4_XS-imx (stock) | 854 MiB | 16.0340 | 0.0815 | 6.96% | 86.0% | 4267 t/s |
| Nano-27pc | 803 MiB | 16.4019 | 0.1434 | 9.29% | 81.5% | 3939 t/s |
| Pico-24pc ✗ | 763 MiB | 17.0117 | 0.2145 | 11.49% | 77.5% | 3939 t/s |
| IQ3_M-imx (stock) | 741 MiB | 17.2831 | 0.2318 | 11.77% | 77.1% | 3939 t/s |
ℹ️ About ASHQ1-Remix Suite
Activation-aware GGUF quantization whose every ratio, floor, and cap traces to a measured experiment. Plain-BF16-native first; AutoRound lineage supported with explicit saturation bounds. Full seven-tier ladder validated across six model families.
🔗 Link: https://huggingface.co/Soulfate24/AutoRound-ASHQ1-Remix_Double-Quantization_Suite
Run Soulfate24/ReaderLM-v2-ASHQ1-Remix-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models