Asher-1/sam3-gguf overview
633;E;echo "=== 3. hparams as seen by the C++ loader"\x3b timeout 60 ./build vis/diff tokenizer /tmp/hf verify/sam3 q4 0.gguf /tmp/tok diff 2 &1 | grep E "me…
Runs locally from ~22.5 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| sam2.1_hiera_base_plus_f16.gguf | GGUF | F16 | 155.7 MB | Download |
| sam2.1_hiera_base_plus_f32.gguf | GGUF | F32 | 308.5 MB | Download |
| sam2.1_hiera_base_plus_q4_0.gguf | GGUF | Q4_0 | 45.8 MB | Download |
| sam2.1_hiera_base_plus_q4_1.gguf | GGUF | Q4_1 | 50.5 MB | Download |
| sam2.1_hiera_base_plus_q8_0.gguf | GGUF | Q8_0 | 83.7 MB | Download |
| sam2.1_hiera_large_f16.gguf | GGUF | F16 | 430.0 MB | Download |
| sam2.1_hiera_large_f32.gguf | GGUF | F32 | 856.3 MB | Download |
| sam2.1_hiera_large_q4_0.gguf | GGUF | Q4_0 | 124.0 MB | Download |
| sam2.1_hiera_large_q4_1.gguf | GGUF | Q4_1 | 137.2 MB | Download |
| sam2.1_hiera_large_q8_0.gguf | GGUF | Q8_0 | 230.1 MB | Download |
| sam2.1_hiera_small_f16.gguf | GGUF | F16 | 89.2 MB | Download |
| sam2.1_hiera_small_f32.gguf | GGUF | F32 | 175.7 MB | Download |
| sam2.1_hiera_small_q4_0.gguf | GGUF | Q4_0 | 26.4 MB | Download |
| sam2.1_hiera_small_q4_1.gguf | GGUF | Q4_1 | 29.1 MB | Download |
| sam2.1_hiera_small_q8_0.gguf | GGUF | Q8_0 | 47.9 MB | Download |
| sam2.1_hiera_tiny_f16.gguf | GGUF | F16 | 75.6 MB | Download |
| sam2.1_hiera_tiny_f32.gguf | GGUF | F32 | 148.7 MB | Download |
| sam2.1_hiera_tiny_q4_0.gguf | GGUF | Q4_0 | 22.5 MB | Download |
| sam2.1_hiera_tiny_q4_1.gguf | GGUF | Q4_1 | 24.8 MB | Download |
| sam2.1_hiera_tiny_q8_0.gguf | GGUF | Q8_0 | 40.6 MB | Download |
| sam2_hiera_base_plus_f16.gguf | GGUF | F16 | 155.7 MB | Download |
| sam2_hiera_base_plus_f32.gguf | GGUF | F32 | 308.4 MB | Download |
| sam2_hiera_base_plus_q4_0.gguf | GGUF | Q4_0 | 45.8 MB | Download |
| sam2_hiera_base_plus_q4_1.gguf | GGUF | Q4_1 | 50.5 MB | Download |
| sam2_hiera_base_plus_q8_0.gguf | GGUF | Q8_0 | 83.7 MB | Download |
| sam2_hiera_large_f16.gguf | GGUF | F16 | 430.0 MB | Download |
| sam2_hiera_tiny_f16.gguf | GGUF | F16 | 75.6 MB | Download |
| sam2_hiera_tiny_f32.gguf | GGUF | F32 | 148.6 MB | Download |
| sam2_hiera_tiny_q4_0.gguf | GGUF | Q4_0 | 22.5 MB | Download |
| sam2_hiera_tiny_q4_1.gguf | GGUF | Q4_1 | 24.8 MB | Download |
| sam2_hiera_tiny_q8_0.gguf | GGUF | Q8_0 | 40.6 MB | Download |
| sam3-f16.gguf | GGUF | F16 | 1.71 GB | Download |
| sam3-f32.gguf | GGUF | F32 | 3.21 GB | Download |
| sam3-q4_0.gguf | GGUF | Q4_0 | 714.0 MB | Download |
| sam3-q4_1.gguf | GGUF | Q4_1 | 720.7 MB | Download |
| sam3-q8_0.gguf | GGUF | Q8_0 | 1.02 GB | Download |
| sam3-visual-f16.gguf | GGUF | F16 | 901.7 MB | Download |
| sam3-visual-q4_0.gguf | GGUF | Q4_0 | 275.8 MB | Download |
| sam3-visual-q4_1.gguf | GGUF | Q4_1 | 302.9 MB | Download |
| sam3-visual-q8_0.gguf | GGUF | Q8_0 | 493.1 MB | Download |
Model Details
Model README
]633;E;echo "=== 3. hparams as seen by the C++ loader"\x3b timeout 60 ./build-vis/diff_tokenizer /tmp/hf_verify/sam3-q4_0.gguf /tmp/tok_diff 2>&1 | grep -E "mem_selection|max_cond|max_obj" ;a7ba6c53-87c6-427f-92bf-6a5106a2fee1]633;C# github: https://github.com/Asher-1/sam3-ggml
Model Zoo — models/
This directory holds ready-to-run GGUF models for sam3.cpp.
Each file name encodes three things:
<family>_<backbone/size>_<precision>.gguf
| Part | Meaning |
|------|---------|
| sam3 / sam3-visual | SAM 3 — ViT-32 backbone + text encoder + DETR detector (850M params) |
| sam2 / sam2.1 | SAM 2 / SAM 2.1 — Meta's Hiera-backbone segmentation models (visual only) |
| tiny / small / base_plus / large | Backbone size (39M / 46M / 81M / 224M params) |
| f32 / f16 / q8_0 / q4_1 / q4_0 | Weight precision (see Precision guide) |
> Architecture lineage: this directory covers 2 architectures —
> SAM 3 (sam3-, sam3-visual-) and the SAM 2 family
> (sam2, sam2.1). There are no SAM 1 checkpoints (SAM 1 / ViT-B/L/H is a separate
> architecture not shipped by this project). The SAM 2 family is visual-only
> (points/box + tracking); SAM 3 full adds text-prompted detection (PCS):
> type "cat" and get every cat in the image.
Download
All checkpoints listed in this README are published as ready-to-run GGUF
files on Hugging Face — no PyTorch weights or conversion step needed:
> https://huggingface.co/Asher-1/sam3-gguf
> Tracker/tokenizer alignment patch (2026-09, updated on HF): the sam3-*
> checkpoints (f32/f16/q8_0/q4_1/q4_0) on the HF repo now carry the repaired
> BPE merge table (all 48894 official rows — the old converter silently
> dropped the 6 #-first merges) and the SAM3 tracker-alignment hparams
> (max_cond_frames_in_attn / use_memory_selection / mf_threshold_x100).
> Locally cached copies downloaded before 2026-09-13 can be repaired in place
> with the idempotent scripts/fix_gguf_merges.py (a dry-run prints "already
> patched; nothing to do" when current). sam2 / sam3-visual- files are
> unaffected (no tokenizer / SAM2-compatible defaults).
The repo mirrors this directory 1:1 (40 files, ~14 GB total).
# All 40 models into models/
huggingface-cli download Asher-1/sam3-gguf --local-dir models
# Or just one file
huggingface-cli download Asher-1/sam3-gguf sam3-f16.gguf --local-dir models
# Or plain curl (single-file direct URL)
curl -L -o models/sam3-f16.gguf \
https://huggingface.co/Asher-1/sam3-gguf/resolve/main/sam3-f16.gguf
(huggingface-cli comes from pip install huggingface_hub.)
Quick pick
| You want… | Pick |
|-----------|------|
| Text-prompted detection ("type cat, get every cat") | sam3-f16.gguf (1.8 GB) or sam3-q8_0.gguf (1.1 GB) |
| Best visual quality-to-speed balance on GPU | sam2.1_hiera_base_plus_f16.gguf (156 MB) |
| Fastest interactive point/box segmentation on any device | sam2.1_hiera_tiny_q4_0.gguf (23 MB) |
| Best segmentation quality | sam2.1_hiera_large_f16.gguf (431 MB) or _q8_0 (231 MB) |
| Debugging / numerical reference (never for deployment) | sam2.1_hiera_tiny_f32.gguf |
Model files
Sizes below are the actual .gguf files in this directory. Latency is a
single-image PVS run (encode + segment) at 1008×1008 on RTX 3060 CUDA,
point (315,250) on tests/cat.jpg. The current SAM 3 F16 result uses
sam3_encode_image_pvs(), 2 warmups and 7 timed runs (p50); the remaining
rows are the earlier all-model snapshot. score = mask IoU confidence.
SAM 3 (850M params — ViT-32 backbone + text encoder + DETR decoder)
| File | Size | Load | Encode | Segment | Total | score |
|------|------|-----:|-------:|--------:|------:|------:|
| sam3-f32.gguf | 3.3 GB | 3.3 s | 4.2 s | 0.21 s | 7.7 s | 0.953 |
| sam3-f16.gguf | 1.8 GB | 0.81 s | 0.566 s | 0.032 s | 1.41 s | 0.953 |
| sam3-q8_0.gguf | 1.1 GB | 1.4 s | 3.5 s | 0.20 s | 5.1 s | 0.953 |
| sam3-q4_1.gguf | 730 MB | 1.1 s | 3.6 s | 0.21 s | 4.9 s | 0.937 |
| sam3-q4_0.gguf | 707 MB | 1.5 s | 3.6 s | 0.25 s | 5.3 s | 0.915 |
Best for: text-prompted detection (PCS) + point/box segmentation (PVS) +
video tracking in one model. The full SAM 3 is the only family here that
supports text prompts; the visual path matches sam3-visual exactly.
SAM 3 Visual (no text encoder — PVS + tracking only)
| File | Size | Load | Encode | Segment | Total | score |
|------|------|-----:|-------:|--------:|------:|------:|
| sam3-visual-f16.gguf | 902 MB | 1.5 s | 2.3 s | 0.22 s | 4.0 s | 0.952 |
| sam3-visual-q8_0.gguf | 494 MB | 0.7 s | 2.2 s | 0.22 s | 3.1 s | 0.953 |
| sam3-visual-q4_1.gguf | 303 MB | 0.7 s | 2.2 s | 0.21 s | 3.1 s | 0.937 |
| sam3-visual-q4_0.gguf | 276 MB | 0.6 s | 2.2 s | 0.23 s | 3.0 s | 0.915 |
Best for: SAM 3-quality segmentation without the text encoder — half the
size and ~40% faster than full SAM 3. Same PVS + tracking capabilities as
sam2.1_hiera_base_plus but with the stronger SAM 3 backbone.
SAM 2 (Hiera backbone, visual only)
| File | Size | Load | Encode | Segment | Total | score |
|------|------|-----:|-------:|--------:|------:|------:|
| sam2_hiera_tiny_f16.gguf | 76 MB | 0.79 s | 1.43 s | 0.37 s | 2.6 s | 0.959 |
| sam2_hiera_tiny_f32.gguf | 149 MB | 0.90 s | 0.98 s | 0.20 s | 2.1 s | 0.959 |
| sam2_hiera_tiny_q8_0.gguf | 41 MB | 0.39 s | 0.82 s | 0.16 s | 1.4 s | 0.959 |
| sam2_hiera_tiny_q4_1.gguf | 25 MB | 0.35 s | 0.81 s | 0.16 s | 1.3 s | 0.930 |
| sam2_hiera_tiny_q4_0.gguf | 23 MB | 0.43 s | 0.83 s | 0.17 s | 1.4 s | 0.933 |
| sam2_hiera_base_plus_f16.gguf | 156 MB | 0.82 s | 1.22 s | 0.20 s | 2.2 s | 0.957 |
| sam2_hiera_base_plus_f32.gguf | 309 MB | 0.97 s | 1.22 s | 0.17 s | 2.4 s | 0.957 |
| sam2_hiera_base_plus_q8_0.gguf | 84 MB | 0.93 s | 1.57 s | 0.28 s | 2.8 s | 0.955 |
| sam2_hiera_base_plus_q4_1.gguf | 51 MB | 0.73 s | 1.16 s | 0.18 s | 2.1 s | 0.954 |
| sam2_hiera_base_plus_q4_0.gguf | 46 MB | 0.65 s | 1.19 s | 0.22 s | 2.1 s | 0.952 |
| sam2_hiera_large_f16.gguf | 430 MB | 1.42 s | 1.47 s | 0.25 s | 3.1 s | 0.909 |
SAM 2.1 (improved SAM 2, same Hiera architecture)
| File | Size | Load | Encode | Segment | Total | score |
|------|------|-----:|-------:|--------:|------:|------:|
| sam2.1_hiera_tiny_f32.gguf | 149 MB | 0.78 s | 1.11 s | 0.21 s | 2.1 s | 0.943 |
| sam2.1_hiera_tiny_f16.gguf | 76 MB | 0.60 s | 1.03 s | 0.23 s | 1.9 s | 0.943 |
| sam2.1_hiera_tiny_q8_0.gguf | 41 MB | 0.62 s | 1.11 s | 0.24 s | 2.0 s | 0.945 |
| sam2.1_hiera_tiny_q4_1.gguf | 25 MB | 0.62 s | 0.95 s | 0.19 s | 1.8 s | 0.956 |
| sam2.1_hiera_tiny_q4_0.gguf | 23 MB | 0.73 s | 1.13 s | 0.25 s | 2.1 s | 0.927 |
| sam2.1_hiera_small_f32.gguf | 176 MB | 0.69 s | 0.88 s | 0.18 s | 1.8 s | 0.945 |
| sam2.1_hiera_small_f16.gguf | 90 MB | 0.75 s | 0.99 s | 0.19 s | 1.9 s | 0.945 |
| sam2.1_hiera_small_q8_0.gguf | 48 MB | 0.61 s | 0.96 s | 0.19 s | 1.8 s | 0.944 |
| sam2.1_hiera_small_q4_1.gguf | 30 MB | 0.72 s | 1.27 s | 0.19 s | 2.2 s | 0.947 |
| sam2.1_hiera_small_q4_0.gguf | 27 MB | 0.62 s | 1.11 s | 0.19 s | 1.9 s | 0.949 |
| sam2.1_hiera_base_plus_f32.gguf | 309 MB | 0.99 s | 1.25 s | 0.21 s | 2.5 s | 0.953 |
| sam2.1_hiera_base_plus_f16.gguf | 156 MB | 0.71 s | 1.09 s | 0.18 s | 2.0 s | 0.953 |
| sam2.1_hiera_base_plus_q8_0.gguf | 84 MB | 0.79 s | 1.40 s | 0.21 s | 2.4 s | 0.954 |
| sam2.1_hiera_base_plus_q4_1.gguf | 51 MB | 0.74 s | 1.44 s | 0.24 s | 2.4 s | 0.944 |
| sam2.1_hiera_base_plus_q4_0.gguf | 46 MB | 0.79 s | 1.34 s | 0.23 s | 2.4 s | 0.936 |
| sam2.1_hiera_large_f32.gguf | 857 MB | 1.42 s | 1.57 s | 0.18 s | 3.2 s | 0.940 |
| sam2.1_hiera_large_f16.gguf | 431 MB | 1.00 s | 1.29 s | 0.17 s | 2.5 s | 0.940 |
| sam2.1_hiera_large_q8_0.gguf | 231 MB | 0.76 s | 1.44 s | 0.19 s | 2.4 s | 0.938 |
| sam2.1_hiera_large_q4_1.gguf | 138 MB | 0.75 s | 1.68 s | 0.23 s | 2.7 s | 0.928 |
| sam2.1_hiera_large_q4_0.gguf | 124 MB | 0.77 s | 1.49 s | 0.23 s | 2.5 s | 0.900 |
Charts
- Latency chart — 40 locally benchmarked checkpoints (RTX 3060 CUDA)
- Effect grid — 40 locally benchmarked checkpoints on cat.jpg
Precision guide
| Precision | Relative size | Quality | Use |
|-----------|---------------|---------|-----|
| f32 | 1.0× | reference | Debugging, numerical checks only — never deploy |
| f16 | 0.5× | ≈ f32 | Recommended default — near-lossless, half the size |
| q8_0 | 0.25× | very close to f16 | Big models (large/sam3) when f16 is too big |
| q4_1 | ~0.14× | good (retains scale + offset) | Aggressive size cuts with better fidelity than q4_0 |
| q4_0 | ~0.13× | acceptable for interactive use | Smallest files; quality gap is visible on thin structures |
Size selection guide
| Need | SAM 3 | SAM 3 Visual | base_plus | tiny |
|------|-------|--------------|-----------|------|
| Text prompts (PCS) | Yes | - | - | - |
| PVS + tracking | Yes | Yes | Yes | Yes |
| Encode latency (RTX 3060) | 0.566 s (F16 PVS) | snapshot: 2.2 s | ~1.1–1.6 s | ~0.8–1.1 s |
| Size (f16) | 1.8 GB | 902 MB | 156 MB | 76 MB |
- SAM 2 vs SAM 2.1: prefer 2.1 for new projects (better training data and
tracking; same architecture, same speed, same sizes).
- Video tracking: tiny is the practical choice for interactive playback on
CPU; larger backbones work well on GPU.
- Point/box (PVS) + tracking work on every model here; **text-prompted
detection (PCS) requires a SAM 3 checkpoint* (the sam3- files above).
Run Asher-1/sam3-gguf with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models