GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

VertexResearch/Vertex-0.6-35M-Instruct-GGUF overview

Vertex 0.6 35M Instruct — GGUF GGUF quantizations of VertexResearch/Vertex 0.6 35M Instruct https://huggingface.co/VertexResearch/Vertex 0.6 35M Instruct , a ≈…

llama.cppggufquantizedchatvertexqwen3text-generationenbase_model:VertexResearch/Vertex-0.6-35M-Instructbase_model:quantized:VertexResearch/Vertex-0.6-35M-Instructlicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~24.4 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
11
Likes
0
Pipeline
text-generation

Repository Files & Downloads

11 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Vertex-0.6-35M-Instruct-BF16.ggufGGUFBF1665.8 MBDownload
Vertex-0.6-35M-Instruct-Q2_K.ggufGGUFQ2_K24.4 MBDownload
Vertex-0.6-35M-Instruct-Q3_K_M.ggufGGUFQ3_K_M25.4 MBDownload
Vertex-0.6-35M-Instruct-Q3_K_S.ggufGGUFQ3_K_S24.4 MBDownload
Vertex-0.6-35M-Instruct-Q4_0.ggufGGUFQ4_025.2 MBDownload
Vertex-0.6-35M-Instruct-Q4_1.ggufGGUFQ4_126.5 MBDownload
Vertex-0.6-35M-Instruct-Q4_K_M.ggufGGUFQ4_K_M27.9 MBDownload
Vertex-0.6-35M-Instruct-Q4_K_S.ggufGGUFQ4_K_S27.1 MBDownload
Vertex-0.6-35M-Instruct-Q5_K_M.ggufGGUFQ5_K_M29.1 MBDownload
Vertex-0.6-35M-Instruct-Q5_K_S.ggufGGUFQ5_K_S28.7 MBDownload
Vertex-0.6-35M-Instruct-Q8_0.ggufGGUFQ8_035.5 MBDownload

Model Details

Model IDVertexResearch/Vertex-0.6-35M-Instruct-GGUF
AuthorVertexResearch
Pipelinetext-generation
Licenseapache-2.0
Base modelVertexResearch/Vertex-0.6-35M-Instruct
Last modified2026-08-26T12:41:59.000Z

Model README

---

license: apache-2.0

language:

  • en

base_model: VertexResearch/Vertex-0.6-35M-Instruct

pipeline_tag: text-generation

library_name: llama.cpp

tags:

  • gguf
  • quantized
  • chat
  • vertex
  • qwen3

---

Vertex-0.6-35M-Instruct — GGUF

GGUF quantizations of VertexResearch/Vertex-0.6-35M-Instruct,

a ≈34M-parameter Qwen3-architecture chat model. The ChatML chat template

and <|im_end|> stop token are embedded in the GGUF metadata, so LM Studio,

llama.cpp, and Ollama chat with it correctly out of the box — no manual

template setup needed.

Files

| Quant | Size | Notes |

|---|---|---|

| BF16 | 69 MB | Full precision |

| Q8_0 | 37 MB | Recommended — at this model size there is little reason to go lower |

| Q5_K_M | 31 MB | |

| Q5_K_S | 30 MB | |

| Q4_K_M | 29 MB | |

| Q4_K_S | 28 MB | |

| Q4_1 | 28 MB | Legacy |

| Q4_0 | 26 MB | Legacy |

| Q3_K_M | 27 MB | Quality loss noticeable on a model this small |

| Q3_K_S | 26 MB | Quality loss noticeable on a model this small |

| Q2_K | 26 MB | Not recommended at 34M params |

Note: the model's hidden size (384) is not a multiple of 256, so k-quants

fall back to legacy formats for some tensors — the sub-Q4 files save less

space than usual and Q8_0 remains the best pick.

Quick start

  • LM Studio: search for this repo, pick a quant (Q8_0 recommended),

download, and chat — the template is auto-detected.

  • llama.cpp: llama-cli -m Vertex-0.6-35M-Instruct-Q8_0.gguf
  • Ollama: ollama run hf.co/VertexResearch/Vertex-0.6-35M-Instruct-GGUF:Q8_0

Limitations

These models are not the most coherent yet and need more tuning: expect rambling, repetition, and inconsistent answers, especially over longer generations.

34M parameters: simple conversational ability only — expect weak factual

reliability, arithmetic, and instruction-following on complex rewrites.

English-centric, 1024-token context, no safety tuning.

Run VertexResearch/Vertex-0.6-35M-Instruct-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models