SH4P3S/apertus-4b-r2-Q5_K_M.gguf overview
apertus 4b r2 Q5 K M Round 2 QLoRA tool use fine tune of swiss ai/Apertus v1.1 4B Instruct https://huggingface.co/swiss ai/Apertus v1.1 4B Instruct , quantized…
Runs locally from ~2.59 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| apertus-4b-r2-Q5_K_M.gguf | GGUF | Q5_K_M | 2.59 GB | Download |
Model Details
Model README
---
license: apache-2.0
base_model: swiss-ai/Apertus-v1.1-4B-Instruct
tags:
- gguf
- llama.cpp
- tool-use
- apertus
---
apertus-4b-r2-Q5_K_M
Round-2 QLoRA tool-use fine-tune of swiss-ai/Apertus-v1.1-4B-Instruct, quantized to Q5_K_M for llama.cpp (converted from f16, never quant→quant).
Round-2 changes vs round 1: direct weather/fetch tool selection (no forced chaining), notes answers stay grounded in note hits, date anchoring (Current date: <Weekday>, <YYYY-MM-DD> as the last system-prompt line — the runtime must inject it), and a higher share of clarifying questions.
Intended to run with its paired system prompt and tools.gbnf grammar (lazy-triggered on <tool_call>); sampling temp 0.8, top_p 0.9, repeat_penalty 1.0, context 4096.
- File:
apertus-4b-r2-Q5_K_M.gguf - Size: 2,777,042,176 bytes
- SHA-256:
5219675afd2da7108c56f43284945b39e76e1e1d5aae1aca85a630c98d664794
Run SH4P3S/apertus-4b-r2-Q5_K_M.gguf with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models