ContextReq/Pebble-10M-GGUF overview
Pebble 10M GGUF GGUF conversions of basically ai/Pebble 10M https://huggingface.co/basically ai/Pebble 10M Apache 2.0 . IMPORTANT: patched llama.cpp required P…
Runs locally from ~7.2 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | ContextReq/Pebble-10M-GGUF |
|---|---|
| Author | ContextReq |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | basically-ai/Pebble-10M |
| Last modified | 2026-09-04T01:45:32.000Z |
Model README
---
license: apache-2.0
pipeline_tag: text-generation
tags:
- gguf
- llama.cpp
- pebble
- mamba2
- hybrid
- small-language-model
base_model: basically-ai/Pebble-10M
---
Pebble-10M-GGUF
GGUF conversions of basically-ai/Pebble-10M (Apache 2.0).
IMPORTANT: patched llama.cpp required
Pebble uses a custom hybrid Mamba2 + attention architecture. These GGUFs carry
general.architecture = "pebble", which upstream llama.cpp refuses to load.
Everything needed to run them lives in the support repo:
rootendpoint/basicallyai_llama.cpp_support
llama.cpp-pebble.patch- adds thepebblearchitecture to llama.cpp
(applies cleanly against upstream commit 0eadefe)
basicallyai_to_gguf.py- standalone converter (numpy + safetensors only)numpy_reference.py- independent reference implementation used to verify correctness
Apply the patch, rebuild llama.cpp, then:
llama-cli -m pebble-10m-f16.gguf -p "The capital of France" -n 64
Files
| Quant | Size | Type |
|-------|------|------|
| f16 | 20.7 MB | F16 |
| q8_0 | 11.1 MB | mostly Q8_0 |
| q4_k_m | 7.5 MB | mostly Q4_K_M |
Verification
Outputs were cross-checked token-by-token against an independent pure-numpy
reference implementation over multiple prompts (CPU and CUDA backends, and
quantized KV cache). Identical greedy sequences up to genuine argmax ties.
Model
- 10M parameters, hidden 384, 8 layers (6 Mamba2 + 2 attention), ctx 512, vocab 2048
- A research-scale model: expect toy-level output quality.
CPU support (no GPU required)
Pure-CPU support for these models (no mamba-ssm, no CUDA) lives in the
basicallyai_cpu_support repository:
https://github.com/rootendpoint/basicallyai_cpu_support
It runs the original HF checkpoints in pure PyTorch on plain CPU
(~118 tok/s for 10M, ~65 tok/s for 25M on a Ryzen 5 2600X), verified
token-identical against this GGUF pipeline and an independent numpy oracle.
Run ContextReq/Pebble-10M-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models