realoperator42/trinity-nano-stheno-GGUF overview
trinity nano stheno — GGUF realoperator42/trinity nano stheno lora attention only LoRA, r=16, alpha=32 merged into arcee ai/Trinity Nano Preview https://huggin…
Runs locally from ~3.50 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | realoperator42/trinity-nano-stheno-GGUF |
|---|---|
| Author | realoperator42 |
| Pipeline | — |
| License | apache-2.0 |
| Base model | arcee-ai/Trinity-Nano-Preview |
| Last modified | 2026-08-08T22:10:02.000Z |
Model README
---
license: apache-2.0
base_model: arcee-ai/Trinity-Nano-Preview
tags: [gguf, llama.cpp, chatml, roleplay]
---
trinity-nano-stheno — GGUF
realoperator42/trinity-nano-stheno-lora (attention-only LoRA, r=16, alpha=32) merged into
arcee-ai/Trinity-Nano-Preview and converted to GGUF.
Architecture is afmoe, so you need a llama.cpp build that includes AFMOE support.
The embedded chat template is the adapter's corrected ChatML, with the training-only
{% generation %} markers removed so --jinja parses cleanly.
llama-cli -m trinity-nano-stheno-Q4_K_M.gguf --jinja -cnv
Files
| File | Size |
| --- | --- |
| trinity-nano-stheno-Q4_K_M.gguf | 3.8 GB |
| trinity-nano-stheno-Q8_0.gguf | 6.5 GB |
Run realoperator42/trinity-nano-stheno-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models