VertexResearch/Vertex-0.6-35M-Base-GGUF overview
Vertex 0.6 35M Base — GGUF GGUF quantizations of VertexResearch/Vertex 0.6 35M Base https://huggingface.co/VertexResearch/Vertex 0.6 35M Base , a ≈34M paramete…
Runs locally from ~24.4 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Vertex-0.6-35M-Base-BF16.gguf | GGUF | BF16 | 65.8 MB | Download |
| Vertex-0.6-35M-Base-Q2_K.gguf | GGUF | Q2_K | 24.4 MB | Download |
| Vertex-0.6-35M-Base-Q3_K_M.gguf | GGUF | Q3_K_M | 25.4 MB | Download |
| Vertex-0.6-35M-Base-Q3_K_S.gguf | GGUF | Q3_K_S | 24.4 MB | Download |
| Vertex-0.6-35M-Base-Q4_0.gguf | GGUF | Q4_0 | 25.2 MB | Download |
| Vertex-0.6-35M-Base-Q4_1.gguf | GGUF | Q4_1 | 26.4 MB | Download |
| Vertex-0.6-35M-Base-Q4_K_M.gguf | GGUF | Q4_K_M | 27.8 MB | Download |
| Vertex-0.6-35M-Base-Q4_K_S.gguf | GGUF | Q4_K_S | 27.1 MB | Download |
| Vertex-0.6-35M-Base-Q5_K_M.gguf | GGUF | Q5_K_M | 29.1 MB | Download |
| Vertex-0.6-35M-Base-Q5_K_S.gguf | GGUF | Q5_K_S | 28.7 MB | Download |
| Vertex-0.6-35M-Base-Q8_0.gguf | GGUF | Q8_0 | 35.5 MB | Download |
Model Details
| Model ID | VertexResearch/Vertex-0.6-35M-Base-GGUF |
|---|---|
| Author | VertexResearch |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | VertexResearch/Vertex-0.6-35M-Base |
| Last modified | 2026-08-26T12:41:50.000Z |
Model README
---
license: apache-2.0
language:
- en
base_model: VertexResearch/Vertex-0.6-35M-Base
pipeline_tag: text-generation
library_name: llama.cpp
tags:
- gguf
- quantized
- vertex
- qwen3
---
Vertex-0.6-35M-Base — GGUF
GGUF quantizations of VertexResearch/Vertex-0.6-35M-Base,
a ≈34M-parameter Qwen3-architecture base model. Works out of the box in
LM Studio, llama.cpp, and Ollama (Qwen3 is natively supported).
This is a raw base model — it continues text and has no chat template.
For chatting, use the instruct variant:
VertexResearch/Vertex-0.6-35M-Instruct.
Files
| Quant | Size | Notes |
|---|---|---|
| BF16 | 69 MB | Full precision |
| Q8_0 | 37 MB | Recommended — at this model size there is little reason to go lower |
| Q5_K_M | 31 MB | |
| Q5_K_S | 30 MB | |
| Q4_K_M | 29 MB | |
| Q4_K_S | 28 MB | |
| Q4_1 | 28 MB | Legacy |
| Q4_0 | 26 MB | Legacy |
| Q3_K_M | 27 MB | Quality loss noticeable on a model this small |
| Q3_K_S | 26 MB | Quality loss noticeable on a model this small |
| Q2_K | 26 MB | Not recommended at 34M params |
Note: the model's hidden size (384) is not a multiple of 256, so k-quants
fall back to legacy formats for some tensors — the sub-Q4 files save less
space than usual and Q8_0 remains the best pick.
Quick start
- LM Studio: search for this repo, pick a quant, download, and load.
Use it in completion/base mode (no chat template).
- llama.cpp:
llama-completion -m Vertex-0.6-35M-Base-Q8_0.gguf -p "Your prompt" -n 100 - Ollama:
ollama run hf.co/VertexResearch/Vertex-0.6-35M-Base-GGUF:Q8_0
Limitations
These models are not the most coherent yet and need more tuning: expect rambling, repetition, and inconsistent answers, especially over longer generations.
Run VertexResearch/Vertex-0.6-35M-Base-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models