tepirale author hub
Model Q2 0.gguf is completely degraded by the quantization. Model Q4 1.gguf shows no quantization problems. MTP active GPU RTX 3090 | 24 GB VRAM Q4 1 use: 18.8 GB VRAM Q5 K S use: 20 GB VRAM MTP active GPU RTX 6000 | 48 GB VRAM max tok | compl tok | time s | tok/s medido | tok/s…
Models
6
Downloads
2,351
Run models locally with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.