yosoyalguien/Meta-Llama-3.1-8B-Instruct-GGUF-TQ4_1S overview
yosoyalguien/Meta Llama 3.1 8B Instruct GGUF TQ4 1S TurboQuant TQ4 1S quantization of bartowski/Meta Llama 3.1 8B Instruct GGUF https://huggingface.co/bartowsk…
Runs locally from ~4.78 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Meta-Llama-3.1-8B-Instruct-tq4_1s.gguf | GGUF | GGUF | 4.78 GB | Download |
Model Details
| Model ID | yosoyalguien/Meta-Llama-3.1-8B-Instruct-GGUF-TQ4_1S |
|---|---|
| Author | yosoyalguien |
| Pipeline | — |
| License | other |
| Base model | bartowski/Meta-Llama-3.1-8B-Instruct-GGUF |
| Last modified | 2026-07-09T22:19:30.000Z |
Model README
---
base_model:
- bartowski/Meta-Llama-3.1-8B-Instruct-GGUF
license: other
tags:
- turboquant
- gguf
- quantization
- tq4_1s
---
yosoyalguien/Meta-Llama-3.1-8B-Instruct-GGUF-TQ4_1S
TurboQuant (TQ4_1S) quantization of bartowski/Meta-Llama-3.1-8B-Instruct-GGUF.
Details
| Field | Value |
|---|---|
| Parent model | bartowski/Meta-Llama-3.1-8B-Instruct-GGUF |
| Quantization type | TQ4_1S |
| Quantization tool | turboquant-plus-tqp-v0.2.0 |
| File | Meta-Llama-3.1-8B-Instruct-tq4_1s.gguf |
Usage
Use with llama.cpp (TurboQuant fork) or any GGUF-compatible runtime that supports the TQ4_1S type.
llama-server -m Meta-Llama-3.1-8B-Instruct-tq4_1s.gguf --port 8080
Disclaimer
This model was quantized using TurboQuant, an experimental KV cache compression and quantization method. Quality may differ from the parent model. Refer to the parent model for licensing and usage terms.
Run yosoyalguien/Meta-Llama-3.1-8B-Instruct-GGUF-TQ4_1S with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models