North-ML1/Forge-1-Mini-GGUF overview
Forge 1 Mini GGUF GGUF builds for North ML1/Forge 1 Mini . Use the embedded ChatML template and stop on <|im end| . Modern llama.cpp conversation mode: bash ll…
Runs locally from ~2.9 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Forge-1-Mini-F16.gguf | GGUF | F16 | 9.9 MB | Download |
| Forge-1-Mini-Q2_K.gguf | GGUF | Q2_K | 2.9 MB | Download |
| Forge-1-Mini-Q3_K_M.gguf | GGUF | Q3_K_M | 3.1 MB | Download |
| Forge-1-Mini-Q4_K_M.gguf | GGUF | Q4_K_M | 3.8 MB | Download |
| Forge-1-Mini-Q8_0.gguf | GGUF | Q8_0 | 5.3 MB | Download |
| Forge-1-Mini-TQ1_0.gguf | GGUF | GGUF | 2.9 MB | Download |
Model Details
Model README
---
license: mit
language:
- en
tags:
- gguf
- llama
- llama.cpp
- causal-lm
- north-ml
- forge
base_model: North-ML1/Forge-1-Mini
---
Forge 1 Mini GGUF
GGUF builds for North-ML1/Forge-1-Mini.
Use the embedded ChatML template and stop on <|im_end|>.
Modern llama.cpp conversation mode:
llama-cli -m Forge-1-Mini-Q4_K_M.gguf -cnv -p "Hi" -st -n 64 --temp 0
The GGUF metadata explicitly sets tokenizer.ggml.eot_token_id=5, where token 5 is <|im_end|>.
from llama_cpp import Llama
llm = Llama(model_path="Forge-1-Mini-Q4_K_M.gguf", n_ctx=512)
out = llm.create_chat_completion(
messages=[{"role": "user", "content": "What is 2 + 2?"}],
max_tokens=96,
temperature=0.0,
stop=["<|im_end|>"],
)
print(out["choices"][0]["message"]["content"].strip())
Expected output:
4
Files
| File | Quantization | Size |
|---|---:|---:|
| Forge-1-Mini-F16.gguf | F16 | 9.9 MB |
| Forge-1-Mini-Q8_0.gguf | Q8_0 | 5.3 MB |
| Forge-1-Mini-Q4_K_M.gguf | Q4_K_M | 3.8 MB |
| Forge-1-Mini-Q3_K_M.gguf | Q3_K_M | 3.1 MB |
| Forge-1-Mini-Q2_K.gguf | Q2_K | 2.9 MB |
| Forge-1-Mini-TQ1_0.gguf | TQ1_0 | 2.9 MB |
Verification
All listed GGUF files were generated with llama.cpp llama-quantize and passed a llama-cpp-python smoke test using llama.cpp tokenization and greedy sampling:
Who are you? -> I am Forge-1-Mini, a tiny local assistant created by Arthur / North ML.
Hi -> Hi! I am Forge-1-Mini. How can I help?
What is 2 + 2? -> 4
Write a Python function that adds two numbers. -> def add(a, b): return a + b
Who is Jesus? -> Christians believe Jesus Christ is the eternal Son of God...
How should I treat someone I disagree with? -> Treat the person with dignity...
Note: this model has a 192-wide hidden dimension. Some K-quant and TQ tensors fall back to compatible GGML tensor types because those formats require 256-column divisibility. The files are valid GGUF outputs from llama.cpp and were tested after quantization.
Run North-ML1/Forge-1-Mini-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models