pbhappliedsystems/qwen-2.5-3b-instruct-gguf-q4-k-m overview
Quantized, converted, and evaluated by PBH Applied Systems, LLC — Applied AI/ML Consulting · LLM Optimization & Deployment · Quantized AI Infrastructure 🔬 This repository is part of a production-oriented evaluation series. Every model published under pbhappliedsystems has been independently evaluated using quant_eval v7.21 — a proprietary behavioral evaluation harness developed by PBH Applied Systems. Scores measure real agent-adjacent task performance across structured output, tool dispatch, multi-turn state retention, and multi-step planning families — not perplexity or benchmark leaderboard proxies. ⚖️ License Notice: This model is governed by the Qwen Research License, which permits non-commercial use only. Commercial use requires a separate license from Alibaba Cloud. See the License section for full details. ---
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| qwen-2.5-3B-instruct-gguf-Q4-K-M.gguf | GGUF | — | 1.80 GB | Download |
Model Details Live
Metadata Inspector
Normalized metadata (stored in metadata_json)
{
"metadata": {},
"card_data": {
"language": [
"en",
"zh",
"fr",
"es",
"pt",
"de",
"it",
"ru",
"ja",
"ko",
"ar"
],
"license": "other",
"license_name": "qwen-research",
"license_link": "https://huggingface.co/Qwen/Qwen2.5-3B-Instruct/blob/main/LICENSE",
"base_model": "Qwen/Qwen2.5-3B-Instruct",
"tags": [
"gguf",
"quantized",
"q4_k_m",
"qwen",
"qwen2.5",
"instruct",
"llama-cpp",
"edge-deployment",
"structured-output",
"pbh-applied-systems",
"quant-eval"
],
"frontmatter": {
"language": [
"en",
"zh",
"fr",
"es",
"pt",
"de",
"it",
"ru",
"ja",
"ko",
"ar"
],
"license": "other",
"license_name": "qwen-research",
"license_link": "https://huggingface.co/Qwen/Qwen2.5-3B-Instruct/blob/main/LICENSE",
"base_model": "Qwen/Qwen2.5-3B-Instruct",
"tags": [
"gguf",
"quantized",
"q4_k_m",
"qwen",
"qwen2.5",
"instruct",
"llama-cpp",
"edge-deployment",
"structured-output",
"pbh-applied-systems",
"quant-eval"
]
},
"hero_image_url": "",
"summary": "**Quantized, converted, and evaluated by PBH Applied Systems, LLC** — Applied AI/ML Consulting · LLM Optimization & Deployment · Quantized AI Infrastructure > 🔬 **This repository is part of a production-oriented evaluation series.** Every model published under pbhappliedsystems has been independently evaluated using **quant_eval v7.21** — a proprietary behavioral evaluation harness developed by PBH Applied Systems. Scores measure real agent-adjacent task performance across structured output, tool dispatch, multi-turn state retention, and multi-step planning families — not perplexity or benchmark leaderboard proxies. > ⚖️ **License Notice:** This model is governed by the **Qwen Research License**, which permits **non-commercial use only**. Commercial use requires a separate license from Alibaba Cloud. See the License section for full details. ---",
"quick_links": [],
"benchmark_table_html": "",
"readme_markdown": "---\nlanguage:\n - en\n - zh\n - fr\n - es\n - pt\n - de\n - it\n - ru\n - ja\n - ko\n - ar\nlicense: other\nlicense_name: qwen-research\nlicense_link: https://huggingface.co/Qwen/Qwen2.5-3B-Instruct/blob/main/LICENSE\nbase_model: Qwen/Qwen2.5-3B-Instruct\ntags:\n - gguf\n - quantized\n - q4_k_m\n - qwen\n - qwen2.5\n - instruct\n - llama-cpp\n - edge-deployment\n - structured-output\n - pbh-applied-systems\n - quant-eval\n---\n\n# Qwen2.5-3B-Instruct · GGUF Q4\\_K\\_M\n\n**Quantized, converted, and evaluated by [PBH Applied Systems, LLC](https://pbhappliedsystems.com)**\n— Applied AI/ML Consulting · LLM Optimization & Deployment · Quantized AI Infrastructure\n\n> 🔬 **This repository is part of a production-oriented evaluation series.** Every model published under [`pbhappliedsystems`](https://huggingface.co/pbhappliedsystems) has been independently evaluated using **quant_eval v7.21** — a proprietary behavioral evaluation harness developed by PBH Applied Systems. Scores measure real agent-adjacent task performance across structured output, tool dispatch, multi-turn state retention, and multi-step planning families — not perplexity or benchmark leaderboard proxies.\n\n> ⚖️ **License Notice:** This model is governed by the **Qwen Research License**, which permits **non-commercial use only**. Commercial use requires a separate license from Alibaba Cloud. See the [License](#license) section for full details.\n\n---\n\n## Model Description\n\nThis repository contains the **4-bit quantized (Q4\\_K\\_M)** GGUF of [`Qwen/Qwen2.5-3B-Instruct`](https://huggingface.co/Qwen/Qwen2.5-3B-Instruct), a 3-billion parameter instruction-tuned model from Alibaba Cloud (September 2024 release). Qwen2.5-3B-Instruct is the smallest model in the PBH Applied Systems evaluated series, and at Q4\\_K\\_M precision it represents the most hardware-accessible deployment option — requiring as little as 4 GB VRAM and fitting on consumer edge hardware.\n\nThe full-precision F16 baseline is published separately at [`pbhappliedsystems/qwen-2.5-3B-instruct-gguf-F16`](https://huggingface.co/pbhappliedsystems/qwen-2.5-3B-instruct-gguf-F16).\n\n### Key Characteristics\n\n- **Parameters:** 3B\n- **Format:** GGUF Q4\\_K\\_M\n- **File size:** 1.93 GB\n- **SHA256:** `9ab3bc9beaddaec3700d5cc754b52e1501a3fd172bc7fc3ee3eb8e1d388ee043`\n- **Minimum VRAM (GPU inference):** ~4 GB\n- **Recommended GPU tier:** Any CUDA-capable GPU · CPU inference viable\n- **Context window:** 32,768 tokens\n- **Inference speed (eval hardware):** avg **0.390 sec/case** on RTX 4090 — fastest Q4\\_K\\_M in the evaluated series\n- **License:** Qwen Research License (non-commercial)\n\n---\n\n## PBH Applied Systems Evaluation — quant\\_eval v7.21\n\n> **Evaluation conducted by PBH Applied Systems, LLC using quant_eval v7.21**\n> Run ID: `20260221_041137` · Fixtures: `golden_oracle_fixtures_v7_21` (SHA256: `6d71a0b9147c...`) · Seed: 42\n> Hardware: NVIDIA RTX 4090 · Total rows evaluated: 84 (42 F16 · 42 Q4\\_K\\_M)\n\n### Aggregate Scores (Q4\\_K\\_M)\n\nScores are normalized to [0.0 – 1.0]. Higher is better.\n\n| Dimension | Score | Series Context |\n|---|---:|---|\n| Task Completion | 0.4905 | Expected at 3B scale |\n| Reasoning | 0.3704 | Limited by parameter count |\n| Coherence | 0.9074 | Strong at 3B |\n| Instruction Following | 0.6599 | Moderate |\n| **Avg inference time** | **0.390 sec/case** | **Fastest in series** |\n\n### Per-Family Pass Rates\n\n#### F16 Baseline (`full_weight_transformers`)\n\n| Family | N | Pass Rate | Avg Secs | Notes |\n|---|---:|---:|---:|---|\n| json\\_multistep | 5 | 0.000 | 3.704 | checks\\_consistent\\_ok bottleneck |\n| stateful\\_followup | 2 | **1.000** | 0.525 | Both turns exact match |\n| toolcall\\_only | 2 | 0.000 | 0.405 | tool\\_name\\_ok=1, args\\_ok=0 |\n| mixed\\_brief\\_json | 2 | 0.000 | 0.695 | JSON valid; ANSWER line absent — see note |\n| toolcall | 2 | **1.000** | 0.745 | Stage-1 passes; final\\_mismatch — see note |\n| json | 4 | n/a | 2.072 | bucket\\_score avg = 10.000 |\n| fuzz | 20 | n/a | 1.437 | bucket\\_score avg = 7.500 — 15/20 pass |\n| mcq | 5 | n/a | 0.022 | Empty raw output on all 5 |\n\n#### Q4\\_K\\_M (`quantized_llama_cpp`)\n\n| Family | N | Pass Rate | Δ vs F16 | Avg Secs | Notes |\n|---|---:|---:|---:|---:|---|\n| json\\_multistep | 5 | **0.200** | **+0.200** | 0.985 | ms\\_easy\\_01 recovers |\n| stateful\\_followup | 2 | **1.000** | 0.000 | 0.125 | Perfect retention |\n| toolcall\\_only | 2 | 0.000 | 0.000 | 0.100 | Wrong arg key names |\n| mixed\\_brief\\_json | 2 | **1.000** | **+1.000** | 0.155 | Full recovery at Q4\\_K\\_M |\n| toolcall | 2 | **1.000** | 0.000 | 0.215 | Stage-1 passes; final\\_mismatch — see note |\n| json | 4 | n/a | — | 0.510 | bucket\\_score avg = 10.000 |\n| fuzz | 20 | n/a | — | 0.408 | bucket\\_score avg = 7.500 — same 15/20 |\n| mcq | 5 | n/a | — | 0.010 | 3/5 pass at Q4\\_K\\_M |\n\n---\n\n## Key Findings\n\n### Finding 1: The Speed Story — Fastest Model in the Evaluated Series\n\nAt **0.390 sec/case average**, this Q4\\_K\\_M variant is the fastest model evaluated in the PBH Applied Systems series. Individual family timings:\n\n| Family | Q4\\_K\\_M Avg Secs |\n|---|---:|\n| mcq | **0.010** |\n| stateful\\_followup | 0.125 |\n| toolcall\\_only | 0.100 |\n| mixed\\_brief\\_json | 0.155 |\n| toolcall | 0.215 |\n| fuzz | 0.408 |\n| json | 0.510 |\n| json\\_multistep | 0.985 |\n\nMCQ responses at 10 milliseconds. Stateful follow-up at 125 milliseconds. Even the most complex planning cases complete in under one second. At 1.93 GB, this model runs on hardware where no other model in this series fits — including systems without dedicated GPUs.\n\n**This speed profile is the primary deployment argument for Qwen2.5-3B Q4\\_K\\_M.** It is not the most capable model evaluated. It is the most accessible model evaluated, with a capability ceiling that is clearly defined by the evaluation data below.\n\n### Finding 2: checks\\_consistent\\_ok = 0.200 — The 3B Ceiling\n\nThe single most consistent failure signal across both runners and all json\\_multistep cases is `checks_consistent_ok`. Both runners achieve the same rate (0.200) — 1/5 cases pass. This is the signal that measures whether the model's intermediate reasoning steps are internally self-consistent.\n\n**This is a parameter-count limitation, not a quantization artifact.** The same 5 fuzz cases fail on both runners (fuzz_0001, _0002, _0003, _0010, _0012) at the same bucket scores, confirming the failure boundary is set by the model's capacity, not by the precision level. A 3B model that can produce valid JSON schemas reliably but cannot maintain consistent multi-step reasoning chains is behaving exactly as expected at this scale.\n\n| Signal | F16 Rate | Q4\\_K\\_M Rate |\n|---|---:|---:|\n| schema\\_ok | 0.800 | **1.000** |\n| checks\\_consistent\\_ok | 0.200 | 0.200 |\n| stop\\_semantics\\_ok | 0.200 | 0.400 |\n| oracle\\_equiv\\_ok | 0.200 | 0.400 |\n\nQ4\\_K\\_M actually improves over F16 on `schema_ok` (0.800 → 1.000), `stop_semantics_ok` (0.200 → 0.400), and `oracle_equiv_ok` (0.200 → 0.400). The `checks_consistent_ok` signal, which is the deepest reasoning consistency test, holds at 0.200 regardless of precision.\n\n### Finding 3: F16 Role-Token Contamination\n\nThe F16 evaluation reveals a specific issue with how the HuggingFace Transformers runner handles this model's chat template. Raw outputs for several families show truncated role token prefixes:\n\n- toolcall cases: `ician {\"tool_name\": \"add\", ...}` — the role token \"technician\" is truncated to \"ician\" and appears in the response body\n- stateful cases: `ician {\"counter\": 2}` — same prefix contamination\n- mixed_brief_json: `user ANSWER: 13 {...}` — the \"user\" role token appears as literal text\n\nThe JSON content in these cases is correct, but the role prefix causes extraction failures. This explains why `toolcall` shows `final_mismatch` at F16 (the stage-1 JSON is valid but the post-response text includes role tokens), and why `mixed_brief_json` fails `answer_line_ok` at F16 despite having `json_parse_ok=1` and `schema_ok=1`.\n\n**This is a runner configuration issue, not a model quality issue.** The Q4_K_M llama.cpp runner does not exhibit this behavior.\n\n### Finding 4: F16 ms\\_easy\\_01 — Wrong Output Type\n\n`ms_easy_01` at F16 is the only case in the evaluated series where a model returns an entirely wrong output type for json_multistep. The F16 model produces:\n\n```json\n[{\"shelf\": \"A\", \"item\": \"P\", \"can_place\": 1}]\n```\n\nAn array, where the task requires a specific schema object with `plan`, `checks`, and `final` keys. This causes `schema_ok=0` — the only F16 json_multistep case with a schema failure. At Q4\\_K\\_M, `ms_easy_01` recovers: `schema_ok=1`, `oracle_equiv_ok=1`, `checks_consistent_ok=1` — a clean pass.\n\n### Finding 5: toolcall\\_only — Schema Vocabulary Errors\n\nBoth `toolcall_only` cases at Q4\\_K\\_M show a consistent vocabulary mismatch:\n\n| Case | Q4\\_K\\_M Raw Output | Expected |\n|---|---|---|\n| toolonly\\_01 | `{\"tool\": \"add\", \"operands\": [5, 10]}` | `{\"tool_name\": \"add\", \"args\": {\"a\": 5, \"b\": 10}}` |\n| toolonly\\_02 | `{\"tool\": \"add\", \"operands\": [25, 75]}` | `{\"tool_name\": \"add\", \"args\": {\"a\": 25, \"b\": 75}}` |\n\nThe model uses `\"tool\"` instead of `\"tool_name\"` and `\"operands\"` (array) instead of `\"args\"` (object with named keys). Tool name recognition passes (`tool_name_ok=1`) because the extractor finds \"add\" in the payload. Argument extraction fails because `\"operands\": [5, 10]` is not a valid schema for `{\"a\": integer, \"b\": integer}`.\n\nThis is a schema vocabulary issue at 3B scale — the model knows the tool concept but not the exact field name schema. A system prompt that explicitly states the required key names (`tool_name`, `args.a`, `args.b`) would likely resolve this for basic tools.\n\n### Finding 6: toolcall — Correct Arithmetic, EOS Contamination\n\nBoth `toolcall` cases at Q4\\_K\\_M pass stage-1 (tool dispatch valid) but produce `final_mismatch`:\n\n| Case | Raw Output | Expected |\n|---|---|---|\n| tool\\_01 | `{...add(2,3)...}<\\|im_end\\|> 5<\\|im_end\\|>` | `5` |\n| tool\\_02 | `{...add(10,-4)...}<\\|im_end\\|> 6<\\|im_end\\|>` | `6` |\n\nThe arithmetic is correct. The EOS token (`<|im_end|>`) appended to the answer string causes the string comparison to fail. A `re.sub(r'<\\|im_end\\|>', '', raw)` strip resolves it. See the Usage section for the validated extraction pattern.\n\n### Finding 7: MCQ — A-Bias Pattern\n\n| Case | Q4\\_K\\_M Result | Raw |\n|---|---|---|\n| mcq\\_01 | ✅ PASS | `B` |\n| mcq\\_02 | ❌ FAIL | `A` (wrong) |\n| mcq\\_03 | ✅ PASS | `C` |\n| mcq\\_04 | ✅ PASS | `B` |\n| mcq\\_05 | ❌ FAIL | `A` (wrong) |\n\nThe two failures both produce `A`. The same A-bias pattern observed in Mistral-Nemo appears here at 3B scale. At F16, all 5 MCQ cases produce empty raw output (`raw=''`), suggesting the F16 runner cannot extract the choice from the model's response format at this parameter count.\n\n---\n\n## Signal-Level Diagnostics\n\n### Q4\\_K\\_M — json\\_multistep\n\n| Signal | F16 Rate | Q4\\_K\\_M Rate | Delta |\n|---|---:|---:|---:|\n| schema\\_ok | 0.800 | **1.000** | +0.200 |\n| checks\\_consistent\\_ok | 0.200 | 0.200 | 0.000 |\n| stop\\_semantics\\_ok | 0.200 | 0.400 | +0.200 |\n| oracle\\_equiv\\_ok | 0.200 | 0.400 | +0.200 |\n\n### Q4\\_K\\_M — stateful\\_followup\n\n| Signal | Rate |\n|---|---:|\n| turn1\\_parse\\_ok | 1.000 |\n| turn2\\_parse\\_ok | 1.000 |\n| turn1\\_exact\\_match | 1.000 |\n| turn2\\_exact\\_match | 1.000 |\n\n### Q4\\_K\\_M — toolcall\\_only\n\n| Signal | Rate |\n|---|---:|\n| tool\\_name\\_ok | 1.000 |\n| args\\_ok | 0.000 |\n\n### Q4\\_K\\_M — mixed\\_brief\\_json\n\n| Signal | Rate |\n|---|---:|\n| answer\\_line\\_ok | 1.000 |\n| json\\_parse\\_ok | 1.000 |\n| schema\\_ok | 1.000 |\n\n---\n\n## Recommended Use Cases\n\n### ✅ Deploy with Confidence (Q4\\_K\\_M, Non-Commercial)\n\n- **Stateful multi-turn agents** — Perfect two-turn retention (1.000) at 0.125 sec/case. The fastest stateful performance in the evaluated series.\n- **Structured JSON outputs (single-step)** — `json` bucket\\_score 10.000, fuzz bucket\\_score 7.500. Fast and reliable for constraint-adherent single-step placements.\n- **Hybrid brief + JSON responses** — `mixed_brief_json` passes at 1.000 in 0.155 sec. Clean ANSWER line + valid JSON.\n- **Edge and resource-constrained deployment** — 1.93 GB at Q4\\_K\\_M. Runs on 4 GB VRAM, embedded GPUs, and CPU-only environments where no other evaluated model fits.\n- **High-throughput batch processing** — At 0.390 sec/case average, this model can process thousands of structured inference tasks per hour on modest hardware.\n\n### ⚠️ Use with Guardrails (Q4\\_K\\_M)\n\n- **Scaffolded tool-calling** — `toolcall` stage-1 passes at 1.000 but add EOS stripping before final answer extraction. Arithmetic results are correct.\n- **Bare tool-call dispatch** — `toolcall_only` fails on args schema (`\"operands\"` vs `\"args\"`). Provide explicit schema examples in the system prompt or add a normalization layer.\n- **Multi-step planning (easy difficulty only)** — ms\\_easy\\_01 passes. Higher difficulties fail consistently on `checks_consistent_ok`. Use with external validation.\n- **MCQ applications** — 3/5 pass but A-bias exists. Add chain-of-thought prompting or validation for MCQ pipelines.\n\n### ❌ Not Recommended (Q4\\_K\\_M)\n\n- **Medium-to-hard multi-step planning** — ms\\_med and ms\\_hard cases fail on internal consistency. The 3B parameter count is the limiting factor, not the quantization.\n- **Any commercial use without a Qwen commercial license** — The Qwen Research License restricts use to non-commercial research and evaluation purposes. Contact Alibaba Cloud for commercial licensing.\n\n---\n\n## Hardware Requirements\n\n| Configuration | VRAM Required | Notes |\n|---|---|---|\n| Q4\\_K\\_M (this repo) · GPU | ~4 GB | Fits on embedded/mobile GPUs |\n| Q4\\_K\\_M · CPU only | 4 GB RAM | Viable; slower inference |\n| Q4\\_K\\_M · 32K context | ~8 GB | Full context window |\n| F16 (companion repo) · GPU | ~8 GB | 6.18 GB model + overhead |\n\n---\n\n## Usage\n\n### Installation\n\n```bash\npip install llama-cpp-python huggingface_hub\n```\n\nFor GPU acceleration (CUDA):\n\n```bash\nCMAKE_ARGS=\"-DGGML_CUDA=on\" pip install llama-cpp-python --force-reinstall --no-cache-dir\n```\n\n### Python — llama-cpp-python\n\n```python\nfrom huggingface_hub import hf_hub_download\nfrom llama_cpp import Llama\n\nmodel_path = hf_hub_download(\n repo_id=\"pbhappliedsystems/qwen-2.5-3B-instruct-gguf-Q4-K-M\",\n filename=\"qwen-2.5-3B-instruct-gguf-Q4-K-M.gguf\"\n)\n\nllm = Llama(\n model_path=model_path,\n n_ctx=4096,\n n_gpu_layers=-1, # -1 for full GPU; set 0 for CPU-only\n verbose=False,\n)\n\nresponse = llm.create_chat_completion(\n messages=[\n {\n \"role\": \"system\",\n \"content\": \"You are a precise assistant. Return structured JSON when asked. Use exactly the key names specified.\"\n },\n {\n \"role\": \"user\",\n \"content\": \"Summarize the following and return JSON with keys: summary, sentiment, action_items.\"\n }\n ],\n temperature=0.15,\n max_tokens=512,\n)\n\nprint(response[\"choices\"][0][\"message\"][\"content\"])\n```\n\nFor CPU-only deployment (no GPU required):\n\n```python\nllm = Llama(\n model_path=model_path,\n n_ctx=2048,\n n_gpu_layers=0, # CPU-only\n n_threads=8, # Tune to your CPU core count\n verbose=False,\n)\n```\n\nFor tool-calling with EOS stripping (addresses toolcall final_mismatch finding):\n\n```python\nimport json, re\nfrom huggingface_hub import hf_hub_download\nfrom llama_cpp import Llama\n\nmodel_path = hf_hub_download(\n repo_id=\"pbhappliedsystems/qwen-2.5-3B-instruct-gguf-Q4-K-M\",\n filename=\"qwen-2.5-3B-instruct-gguf-Q4-K-M.gguf\"\n)\n\nllm = Llama(model_path=model_path, n_ctx=2048, n_gpu_layers=-1, verbose=False)\n\ndef call_tool_with_cleanup(prompt: str) -> dict:\n \"\"\"\n Tool dispatch with EOS stripping.\n quant_eval v7.21: toolcall stage-1 pass=1.000; final_mismatch due to <|im_end|> suffix.\n Arithmetic is correct — strip EOS before downstream processing.\n \"\"\"\n response = llm.create_chat_completion(\n messages=[\n {\n \"role\": \"system\",\n \"content\": (\n \"You are a tool-calling assistant. \"\n \"When calling a tool, output JSON as: \"\n '{\"tool_name\": \"<name>\", \"args\": {\"a\": <n>, \"b\": <n>}}\\n'\n \"Then on the next line, output the result as a plain number.\"\n )\n },\n {\"role\": \"user\", \"content\": prompt}\n ],\n temperature=0.0,\n max_tokens=128,\n )\n raw = response[\"choices\"][0][\"message\"][\"content\"]\n # Strip EOS tokens before processing\n clean = re.sub(r'<\\|im_end\\|>', '', raw).strip()\n return {\"raw\": raw, \"clean\": clean}\n\nresult = call_tool_with_cleanup(\"Use the add tool to compute 5 plus 10.\")\nprint(result[\"clean\"])\n```\n\n### CLI — llama-cli\n\n```bash\nllama-cli \\\n --model qwen-2.5-3B-instruct-gguf-Q4-K-M.gguf \\\n --chat-template qwen2 \\\n --system-prompt \"You are a precise assistant. Use exactly the key names specified in any JSON schema.\" \\\n --prompt \"Return a JSON object with keys: summary, risk_level, action_items.\" \\\n --n-predict 512 \\\n --ctx-size 4096 \\\n --n-gpu-layers -1 \\\n --temp 0.15\n```\n\nFor server deployment:\n\n```bash\nllama-server \\\n --model qwen-2.5-3B-instruct-gguf-Q4-K-M.gguf \\\n --chat-template qwen2 \\\n --ctx-size 4096 \\\n --n-gpu-layers -1 \\\n --port 8080 \\\n --host 0.0.0.0\n```\n\nQuery via the OpenAI-compatible API:\n\n```python\nfrom openai import OpenAI\nimport re\n\nclient = OpenAI(base_url=\"http://localhost:8080/v1\", api_key=\"not-required\")\n\nresponse = client.chat.completions.create(\n model=\"qwen-2.5-3B-instruct-gguf-Q4-K-M\",\n messages=[{\"role\": \"user\", \"content\": \"Your prompt here\"}],\n temperature=0.15,\n)\n# Strip EOS tokens from output\nclean = re.sub(r'<\\|im_end\\|>', '', response.choices[0].message.content).strip()\nprint(clean)\n```\n\n---\n\n## Evaluation Artifacts\n\nThe full per-case evaluation CSV (`comparison_results_v7_21_Qwen2.5_3B_Instruct_20260221_041137.csv`) and `rollup.json` are published in this repository for independent verification.\n\n---\n\n## Artifact Provenance\n\n| Artifact | Format | Size | SHA256 |\n|---|---|---|---|\n| `qwen-2.5-3B-instruct-gguf-Q4-K-M.gguf` | GGUF Q4\\_K\\_M | 1.93 GB | `9ab3bc9beaddaec3700d5cc754b52e1501a3fd172bc7fc3ee3eb8e1d388ee043` |\n| F16 *(companion repo)* | GGUF F16 | 6.18 GB | `65a0239fc9f9a40e2d4f79ae5e158cad423c2476fe089c744d5e6a6ff6fc9330` |\n\nBoth artifacts were produced from `Qwen/Qwen2.5-3B-Instruct` using a custom-built llama.cpp conversion and quantization pipeline developed by PBH Applied Systems.\n\n---\n\n## Evaluation Methodology\n\n**quant_eval v7.21** is a proprietary behavioral evaluation harness developed by PBH Applied Systems.\n\n**Fixture set:** `golden_oracle_fixtures_v7_21` (SHA256: `6d71a0b9147c079371b02a94f3c149eb78a6adc03dc16ff6833b964fbf4174f0`)\n\n| Family | Description | Pass Signals |\n|---|---|---|\n| `fuzz` | Property-based regression; structured placement correctness | schema\\_ok, constraints\\_ok |\n| `json` | Single-step structured JSON with constraint rules | schema\\_ok, constraints\\_ok |\n| `json_multistep` | Multi-step planning with self-check and oracle verification | schema\\_ok, checks\\_consistent\\_ok, stop\\_semantics\\_ok, oracle\\_equiv\\_ok |\n| `mcq` | Multiple-choice extraction | choice\\_ok |\n| `stateful_followup` | Two-turn state tracking; turn-2 correct given turn-1 | turn1/2\\_parse\\_ok, turn1/2\\_exact\\_match |\n| `mixed_brief_json` | Hybrid: natural language answer + valid JSON block | answer\\_line\\_ok, json\\_parse\\_ok, schema\\_ok |\n| `toolcall` | Tool call embedded in response; parse + schema validation | stage1\\_tool\\_parse\\_ok, stage1\\_tool\\_schema\\_ok |\n| `toolcall_only` | Bare schema-only tool call; strict tool name + args check | tool\\_name\\_ok, args\\_ok |\n\n**Evaluation hardware:** NVIDIA RTX 4090 (24 GB VRAM)\n**Evaluation date:** February 21, 2026\n**quant_eval seed:** 42\n\n---\n\n## About PBH Applied Systems\n\n[**PBH Applied Systems, LLC**](https://pbhappliedsystems.com) is an Oklahoma City–based applied machine learning and AI systems company specializing in production-grade model evaluation, quantization pipelines, agentic AI infrastructure, and scalable AI-driven application development. The organization operates with a strong emphasis on engineering rigor, reproducibility, and real-world deployment constraints — particularly in environments where performance, cost efficiency, and reliability must be balanced against available hardware and budget.\n\n### Founder — Patrick Hill, M.S.\n\nPBH Applied Systems was founded by **Patrick Hill**, a Data Scientist and AI/ML Engineer with 10+ years of experience delivering advanced analytics, predictive modeling, and decision-support solutions across high-stakes operational environments. Patrick holds a **Master of Science in Software Engineering with concentrations in Artificial Intelligence and Machine Learning** (GPA: 4.0) and a B.S. in Business Finance.\n\n**Technical expertise spans:**\n\n- **Languages & Data:** Python, SQL, Linux, Pandas, NumPy, scikit-learn\n- **ML & Modeling:** Supervised and unsupervised learning, neural networks, NLP, transformers, regression, classification, forecasting, and feature engineering\n- **AI/ML Frameworks:** PyTorch, TensorFlow/Keras, HuggingFace Transformers, GGUF, llama.cpp, BitsAndBytes, PEFT, QLoRA\n- **Deployment & MLOps:** Flask APIs, Docker, CI/CD pipelines, REST endpoints, streaming inference, version control\n- **Data Platforms:** Jupyter, Databricks, Power BI, Matplotlib\n- **Quantization:** GGUF conversion, Q4\\_K\\_M / Q5\\_K\\_M / Q8\\_0 strategies, adapter-per-model evaluation architecture\n\n### Published Author\n\nPatrick is the author of **[Applied Machine Learning: Concepts, Tools, and Case Studies](https://a.co/d/05qat7Xz)** — a 1,200+ page practitioner-oriented textbook adopted as **required reading for CSC 373 – Machine Learning at the University of Advancing Technology**.\n\n### Core Service Areas\n\n**1. LLM Optimization & Deployment** · **2. AI Evaluation Frameworks** · **3. Agentic AI Infrastructure** · **4. Scalable AI Application Development** · **5. ML Pipeline Design & Analytics** · **6. Model & Agent Cataloging**\n\n---\n\n## 📞 Work With PBH Applied Systems\n\nThe Qwen2.5-3B Q4\\_K\\_M evaluation tells a clear story about what to expect at the 3B scale: sub-second inference, reliable structured outputs, solid stateful retention, and a well-defined ceiling on multi-step reasoning consistency. Every finding is verifiable — download the CSV and check the `checks_consistent_ok` signal yourself.\n\n👉 **[Book a Scoping Call](https://pbhappliedsystems.com)** — Discuss model selection, edge deployment strategy, or evaluation needs directly with Patrick.\n\n👉 **[Request an Evaluation Report](https://pbhappliedsystems.com)** — Full quant_eval behavioral audit for your target model(s). Engagements from $2,500.\n\n### Connect\n\n| | |\n|---|---|\n| 🌐 **Website** | [pbhappliedsystems.com](https://pbhappliedsystems.com) |\n| 📧 **Email** | [patrick@pbhappliedsystems.com](mailto:patrick@pbhappliedsystems.com) |\n| 💼 **LinkedIn** | [PBH Applied Systems, LLC](https://www.linkedin.com/company/pbh-applied-systems-llc) |\n| ▶️ **YouTube** | [@pbhappliedsystems](https://www.youtube.com/@pbhappliedsystems) |\n| 📸 **Instagram** | [@pbhappliedsystems](https://www.instagram.com/pbhappliedsystems) |\n| 👍 **Facebook** | [pbhappliedsystems](https://www.facebook.com/pbhappliedsystems) |\n\n---\n\n## License\n\nThis GGUF repository is governed by the **Qwen Research License Agreement**, inherited from the base model [`Qwen/Qwen2.5-3B-Instruct`](https://huggingface.co/Qwen/Qwen2.5-3B-Instruct).\n\n**Key terms:**\n\n- **Non-commercial use only** — The license grants a non-exclusive, royalty-free license for **research or evaluation purposes only**\n- **Commercial use requires a separate license** — Contact Alibaba Cloud to request a commercial license\n- **Redistribution** — Permitted with attribution and license copy included\n- **Attribution requirement** — If used to train or improve another AI model, display \"Built with Qwen\" or \"Improved using Qwen\" in documentation\n- **Governing law** — People's Republic of China; exclusive jurisdiction in Hangzhou City People's Courts\n\nFull license text: [Qwen Research License Agreement](https://huggingface.co/Qwen/Qwen2.5-3B-Instruct/blob/main/LICENSE)\n\nThe quant_eval evaluation methodology, fixture set, and scoring framework are proprietary to PBH Applied Systems, LLC and are not included in this repository.\n\n---\n\n*GGUF conversion, quantization, and behavioral evaluation performed by [PBH Applied Systems, LLC](https://pbhappliedsystems.com) · quant_eval v7.21 · Run ID: `20260221_041137`*\n",
"related_quantizations": []
},
"tags": [
"gguf",
"quantized",
"q4_k_m",
"qwen",
"qwen2.5",
"instruct",
"llama-cpp",
"edge-deployment",
"structured-output",
"pbh-applied-systems",
"quant-eval",
"en",
"zh",
"fr",
"es",
"pt",
"de",
"it",
"ru",
"ja",
"ko",
"ar",
"base_model:Qwen/Qwen2.5-3B-Instruct",
"base_model:quantized:Qwen/Qwen2.5-3B-Instruct",
"license:other",
"endpoints_compatible",
"region:us",
"conversational"
],
"likes": 0,
"downloads": 530,
"gated": false,
"private": false,
"last_modified": "2026-04-15T07:03:48.000Z",
"created_at": "2026-04-12T06:01:39.000Z",
"pipeline_tag": "",
"library_name": ""
}
Source payload excerpt (from Hugging Face API)
{
"_id": "69db35430a7d72741cb3ae28",
"id": "pbhappliedsystems/qwen-2.5-3B-instruct-gguf-Q4-K-M",
"modelId": "pbhappliedsystems/qwen-2.5-3B-instruct-gguf-Q4-K-M",
"sha": "c239150fd8f97060737146a7a16de78472bb6d77",
"createdAt": "2026-04-12T06:01:39.000Z",
"lastModified": "2026-04-15T07:03:48.000Z",
"author": "pbhappliedsystems",
"downloads": 530,
"likes": 0,
"gated": false,
"private": false,
"pipeline_tag": "",
"library_name": "",
"siblings_count": 7
}