GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

North-ML1/Forge-1-Mini-GGUF overview

Forge 1 Mini GGUF GGUF builds for North ML1/Forge 1 Mini . Use the embedded ChatML template and stop on <|im end| . Modern llama.cpp conversation mode: bash ll…

ggufllamallama.cppcausal-lmnorth-mlforgeenbase_model:North-ML1/Forge-1-Minibase_model:quantized:North-ML1/Forge-1-Minilicense:mitendpoints_compatibleregion:usconversational

Runs locally from ~2.9 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
Author

Repository Files & Downloads

6 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Forge-1-Mini-F16.ggufGGUFF169.9 MBDownload
Forge-1-Mini-Q2_K.ggufGGUFQ2_K2.9 MBDownload
Forge-1-Mini-Q3_K_M.ggufGGUFQ3_K_M3.1 MBDownload
Forge-1-Mini-Q4_K_M.ggufGGUFQ4_K_M3.8 MBDownload
Forge-1-Mini-Q8_0.ggufGGUFQ8_05.3 MBDownload
Forge-1-Mini-TQ1_0.ggufGGUFGGUF2.9 MBDownload

Model Details

Model IDNorth-ML1/Forge-1-Mini-GGUF
AuthorNorth-ML1
Pipeline
Licensemit
Base modelNorth-ML1/Forge-1-Mini
Last modified2026-06-22T00:58:59.000Z

Model README

---

license: mit

language:

- en

tags:

  • gguf
  • llama
  • llama.cpp
  • causal-lm
  • north-ml
  • forge

base_model: North-ML1/Forge-1-Mini

---

Forge 1 Mini GGUF

GGUF builds for North-ML1/Forge-1-Mini.

Use the embedded ChatML template and stop on <|im_end|>.

Modern llama.cpp conversation mode:

llama-cli -m Forge-1-Mini-Q4_K_M.gguf -cnv -p "Hi" -st -n 64 --temp 0

The GGUF metadata explicitly sets tokenizer.ggml.eot_token_id=5, where token 5 is <|im_end|>.

from llama_cpp import Llama

llm = Llama(model_path="Forge-1-Mini-Q4_K_M.gguf", n_ctx=512)
out = llm.create_chat_completion(
    messages=[{"role": "user", "content": "What is 2 + 2?"}],
    max_tokens=96,
    temperature=0.0,
    stop=["<|im_end|>"],
)
print(out["choices"][0]["message"]["content"].strip())

Expected output:

4

Files

| File | Quantization | Size |

|---|---:|---:|

| Forge-1-Mini-F16.gguf | F16 | 9.9 MB |

| Forge-1-Mini-Q8_0.gguf | Q8_0 | 5.3 MB |

| Forge-1-Mini-Q4_K_M.gguf | Q4_K_M | 3.8 MB |

| Forge-1-Mini-Q3_K_M.gguf | Q3_K_M | 3.1 MB |

| Forge-1-Mini-Q2_K.gguf | Q2_K | 2.9 MB |

| Forge-1-Mini-TQ1_0.gguf | TQ1_0 | 2.9 MB |

Verification

All listed GGUF files were generated with llama.cpp llama-quantize and passed a llama-cpp-python smoke test using llama.cpp tokenization and greedy sampling:

Who are you? -> I am Forge-1-Mini, a tiny local assistant created by Arthur / North ML.
Hi -> Hi! I am Forge-1-Mini. How can I help?
What is 2 + 2? -> 4
Write a Python function that adds two numbers. -> def add(a, b): return a + b
Who is Jesus? -> Christians believe Jesus Christ is the eternal Son of God...
How should I treat someone I disagree with? -> Treat the person with dignity...

Note: this model has a 192-wide hidden dimension. Some K-quant and TQ tensors fall back to compatible GGML tensor types because those formats require 256-column divisibility. The files are valid GGUF outputs from llama.cpp and were tested after quantization.

Run North-ML1/Forge-1-Mini-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models