GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

tepirale/Ornith-Agents-A1-3.6-35B-A3B-GGUF-v2 overview

Se hizo una extraccion de MTP de un modelo 35b, formateo pesos. experimento 20 creo que elimine el template de los modelos el injerto MTP del otro modelo me gu…

ggufmtpqwen3.5moedare-tiesornithbase_model:tepirale/Ornith-Agents-A1-3.6-35B-A3B-dare_tiesbase_model:quantized:tepirale/Ornith-Agents-A1-3.6-35B-A3B-dare_tiesendpoints_compatibleregion:usconversational

Runs locally from ~866.8 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
5
Likes
0
Pipeline
Author

Repository Files & Downloads

6 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Ornith-Agents-A1-3.6-35B-A3B-dare_ties-Q4_K_M.ggufGGUFQ4_K_M20.55 GBDownload
Ornith-Agents-A1-3.6-35B-A3B-dare_ties-Q5_K_M.ggufGGUFQ5_K_M23.87 GBDownload
Ornith-Agents-A1-3.6-35B-A3B-dare_ties-Q6_K.ggufGGUFQ6_K27.39 GBDownload
Ornith-Agents-A1-3.6-35B-A3B-dare_ties-Q8_0.ggufGGUFQ8_035.21 GBDownload
Ornith-Agents-A1-3.6-35B-A3B-dare_ties-f16.ggufGGUFF1664.61 GBDownload
Ornith-Agents-A1-3.6-35B-A3B-dare_ties-mtp-sidecar.ggufGGUFGGUF866.8 MBDownload

Model Details

Model IDtepirale/Ornith-Agents-A1-3.6-35B-A3B-GGUF-v2
Authortepirale
Pipeline
License
Base modeltepirale/Ornith-Agents-A1-3.6-35B-A3B-dare_ties
Last modified2026-07-12T07:49:36.000Z

Model README

---

base_model: tepirale/Ornith-Agents-A1-3.6-35B-A3B-dare_ties

tags: [gguf, mtp, qwen3.5, moe, dare-ties, ornith]

---

Se hizo una extraccion de MTP de un modelo 35b, formateo pesos.

experimento 20

creo que elimine el template de los modelos

el injerto MTP del otro modelo me gusto con este modelo , me falta recrear este modelo para agregar el template

en algun momento lo borrare

wget -O /content/chat_template.jinja tepirale/Ornith-Agents-A1-3.6-35B-A3B-dare_ties

Uso (llama.cpp con MTP)

llama-server -m Ornith-Agents-A1-3.6-35B-A3B-dare_ties-Q4_K_M.gguf \
  --spec-type draft-mtp --spec-draft-n-max 2 \
  -ngl 99 -fa on -c 65536 --jinja \
  --cache-type-k q8_0 --cache-type-v q8_0
/content/llama.cpp/build-cuda/bin/llama-server \
  -m /content/work/gguf/Ornith-Agents-A1-3.6-35B-A3B-dare_ties-Q4_K_M.gguf \
  --chat-template-file /content/chat_template.jinja \
  --spec-type draft-mtp \
  --spec-draft-n-max 3 \
  --spec-draft-n-min 1 \
  --host 0.0.0.0 --port 8080 \
  --n-gpu-layers 999 \
  --ctx-size 20000 \
  --flash-attn on \
  --cont-batching \
  --parallel 4 \
  --cache-type-k q8_0 \
  --cache-type-v q8_0 \
  --metrics

Quants

| Quant | Tamaño |

|-------|--------|

| Q4_K_M | 20.5 GiB |

| Q5_K_M | 23.9 GiB |

| Q6_K | 27.4 GiB |

| Q8_0 | 35.2 GiB |

GPU 3090 24GB | uso 22GB libre ~2GB | modelo Q4_K_M.gguf

max_tok | compl_tok | time(s) | tok/s medido | tok/s nativo | draft acc% | finish

---------------------------------------------------------------------------------------

512 | 512 | 3.11 | 164.64 | 185.81154644562835 | 58.6% | length

1024 | 1024 | 6.38 | 160.6 | 168.1821675659376 | 49.6% | length

1536 | 1536 | 9.23 | 166.42 | 174.37983333556605 | 52.5% | length

2048 | 2048 | 11.82 | 173.27 | 180.66270396669316 | 56.3% | length

2560 | 2560 | 15.35 | 166.77 | 173.19459183649997 | 52.7% | length

3072 | 3072 | 18.23 | 168.52 | 174.07312295115594 | 53.1% | length

3584 | 3584 | 21.77 | 164.61 | 170.3211904082557 | 51.6% | length

4096 | 4096 | 24.66 | 166.11 | 171.41060296776368 | 52.3% | length

4608 | 4608 | 26.66 | 172.83 | 178.31869156802577 | 56.8% | length

5120 | 5120 | 30.05 | 170.39 | 175.3896167291549 | 54.9% | length

5632 | 5632 | 34.0 | 165.65 | 170.00963609304574 | 52.2% | length

6144 | 6144 | 37.16 | 165.34 | 170.1135424832823 | 52.4% | length

6656 | 6656 | 40.8 | 163.14 | 171.59782245044528 | 53.3% | length

7168 | 7168 | 47.57 | 150.68 | 153.01051061387676 | 54.1% | length

Run tepirale/Ornith-Agents-A1-3.6-35B-A3B-GGUF-v2 with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models