GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Hanno-Labs/dinghy-law-0.6b-v1-gguf overview

dinghy law 0.6b v1 — GGUF GGUF / llama.cpp build of Hanno Labs/dinghy law 0.6b v1 https://huggingface.co/Hanno Labs/dinghy law 0.6b v1 , a compact 0.6B legal t…

ggufllama.cppembeddingslegalsentence-transformersbase_model:Hanno-Labs/dinghy-law-0.6b-v1base_model:quantized:Hanno-Labs/dinghy-law-0.6b-v1license:apache-2.0endpoints_compatibleregion:us

Runs locally from ~471.9 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline

Repository Files & Downloads

3 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
dinghy-law-0.6b-v1-Q6_K.ggufGGUFQ6_K471.9 MBDownload
dinghy-law-0.6b-v1-Q8_0.ggufGGUFQ8_0609.5 MBDownload
dinghy-law-0.6b-v1-f16.ggufGGUFF161.12 GBDownload

Model Details

Model IDHanno-Labs/dinghy-law-0.6b-v1-gguf
AuthorHanno-Labs
Pipeline
Licenseapache-2.0
Base modelHanno-Labs/dinghy-law-0.6b-v1
Last modified2026-07-14T14:29:10.000Z

Model README

---

license: apache-2.0

base_model: Hanno-Labs/dinghy-law-0.6b-v1

tags:

- gguf

- llama.cpp

- embeddings

- legal

- sentence-transformers

library_name: gguf

---

dinghy-law-0.6b-v1 — GGUF

GGUF / llama.cpp build of Hanno-Labs/dinghy-law-0.6b-v1, a compact (0.6B) legal text-embedding model

(contrastive fine-tune of Qwen/Qwen3-Embedding-0.6B). Source revision 8f63ca78b621eb9c242429ea2472389d23f8902f.

What this is (backbone-only)

The sentence-transformers pipeline is Transformer -> last-token Pooling -> Dense -> L2 Normalize. llama.cpp carries

the backbone: Transformer -> last-token pool -> L2 norm. The 2_Dense head is dropped — it is a

near-orthogonal rotation (normalized ortho-defect 0.00103, unit row-norms, ~zero bias), and cosine similarity is

invariant under rotation, so dropping it does not change retrieval ranking. The **f16 GGUF reproduces the HF backbone

embedding at cosine 1.00000** on legal probes.

Files

  • dinghy-law-0.6b-v1-f16.gguf — 1.20 GB
  • dinghy-law-0.6b-v1-Q8_0.gguf — 0.64 GB
  • dinghy-law-0.6b-v1-Q6_K.gguf — 0.49 GB

Recommended: Q8_0 (near-lossless, best default). Q6_K for a smaller footprint with negligible quality loss.

f16 as the full-precision reference.

Fidelity — cosine(HF backbone, GGUF) on legal probes

| file | mean cos | min cos |

|---|---|---|

| dinghy-law-0.6b-v1-f16.gguf | 1.00000 | 1.00000 |

| dinghy-law-0.6b-v1-Q8_0.gguf | 0.99911 | 0.99896 |

| dinghy-law-0.6b-v1-Q6_K.gguf | 0.99313 | 0.99009 |

Worst min-cos across quants: 0.99009 (PASS).

Usage (llama.cpp)

Embeddings use last-token pooling and L2 normalization:

llama-embedding -m dinghy-law-0.6b-v1-Q8_0.gguf \
  -p "Retrieve statutes governing landlord obligations for habitability." \
  --pooling last --embd-normalize 2

For retrieval, format the query with the Qwen3-Embedding instruction prefix (documents are embedded raw):

Instruct: {task_instruction}\nQuery: {query}

e.g. task instruction Identify the most relevant statutes for the given situation. for statute retrieval.

Provenance

Quantized from the source safetensors with llama.cpp/convert_hf_to_gguf.py (f16) + llama-quantize (Q8_0, Q6_K).

MTEB(Law, v1) Mean(Task) nDCG@10 of the full model = 65.83. License: apache-2.0.

Run Hanno-Labs/dinghy-law-0.6b-v1-gguf with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models