Model Intelligence Sheet
shannonkun/rugpt3xl-1.3b-sft-gguf overview
Выход за пределы 128 токенов резко снижает качество, так как в этом gguf не активирован механизм sparse attention. Исправление требует принятие патча и пересоз…
Runs locally from ~1.3 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
7 GGUF files detected
Direct downloads for local inference
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Sft_Full_20260610_050605-1.4B-F16.gguf | GGUF | F16 | 2.65 GB | Download |
| rugpt3xl_1.3B_sft_0.1-BF16.gguf | GGUF | BF16 | 2.65 GB | Download |
| rugpt3xl_1.3B_sft_0.1-IQ2_M.gguf | GGUF | IQ2_M | 547.6 MB | Download |
| rugpt3xl_1.3B_sft_0.1-IQ4_XS.gguf | GGUF | IQ4_XS | 766.3 MB | Download |
| rugpt3xl_1.3B_sft_0.1-Q5_K_M.gguf | GGUF | Q5_K_M | 1006.3 MB | Download |
| rugpt3xl_1.3B_sft_0.1-Q8_0.gguf | GGUF | Q8_0 | 1.42 GB | Download |
| rugpt3xl_imatrix.gguf | GGUF | GGUF | 1.3 MB | Download |
Model Details
| Model ID | shannonkun/rugpt3xl-1.3b-sft-gguf |
|---|---|
| Author | shannonkun |
| Pipeline | — |
| License | mit |
| Base model | shannonkun/rugpt3xl-1.3b-sft |
| Last modified | 2026-07-23T06:07:38.000Z |
Model README
---
license: mit
language:
- ru
base_model:
- shannonkun/rugpt3xl-1.3b-sft
datasets:
- IlyaGusev/saiga_preferences
---
Выход за пределы 128 токенов резко снижает качество, так как в этом gguf не активирован механизм sparse attention. Исправление требует принятие патча и пересоздание gguf: https://github.com/ggml-org/llama.cpp/pull/21161
Запускать лучше с шаблоном чата:
Run shannonkun/rugpt3xl-1.3b-sft-gguf with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models