GraySoft
Projects Models About FAQ Contact Download guIDE →
Model Intelligence Sheet

pbhappliedsystems/qwen-2.5-3b-instruct-gguf-q4-k-m overview

Quantized, converted, and evaluated by PBH Applied Systems, LLC — Applied AI/ML Consulting · LLM Optimization & Deployment · Quantized AI Infrastructure 🔬 This repository is part of a production-oriented evaluation series. Every model published under pbhappliedsystems has been independently evaluated using quant_eval v7.21 — a proprietary behavioral evaluation harness developed by PBH Applied Systems. Scores measure real agent-adjacent task performance across structured output, tool dispatch, multi-turn state retention, and multi-step planning families — not perplexity or benchmark leaderboard proxies. ⚖️ License Notice: This model is governed by the Qwen Research License, which permits non-commercial use only. Commercial use requires a separate license from Alibaba Cloud. See the License section for full details. ---

ggufquantizedq4_k_mqwenqwen2.5instructllama-cppedge-deploymentstructured-outputpbh-applied-systemsquant-evalenzhfresptdeitrujakoarbase_model:Qwen/Qwen2.5-3B-Instructbase_model:quantized:Qwen/Qwen2.5-3B-Instructlicense:otherendpoints_compatibleregion:usconversational
pbhappliedsystems/qwen-2.5-3b-instruct-gguf-q4-k-m visual
Downloads
530
Likes
0
Pipeline
Library
Visibility
Public
Access
Open

Repository Files & Downloads

1 files detected
Direct downloads for all repository files
FileTypeQuantizationSizeLink
qwen-2.5-3B-instruct-gguf-Q4-K-M.gguf GGUF 1.80 GB Download

Model Details Live

Model Slug
pbhappliedsystems/qwen-2.5-3b-instruct-gguf-q4-k-m
Author
pbhappliedsystems
Pipeline Task
Library
Created
2026-04-12
Last Modified
2026-04-15
Gated
No
Private
No
HF SHA
c239150fd8f97060737146a7a16de78472bb6d77
License
other
Language
en, zh, fr, es, pt, de, it, ru, ja, ko, ar
Base Model
Qwen/Qwen2.5-3B-Instruct

Metadata Inspector

Normalized metadata (stored in metadata_json)
{
  "metadata": {},
  "card_data": {
    "language": [
      "en",
      "zh",
      "fr",
      "es",
      "pt",
      "de",
      "it",
      "ru",
      "ja",
      "ko",
      "ar"
    ],
    "license": "other",
    "license_name": "qwen-research",
    "license_link": "https://huggingface.co/Qwen/Qwen2.5-3B-Instruct/blob/main/LICENSE",
    "base_model": "Qwen/Qwen2.5-3B-Instruct",
    "tags": [
      "gguf",
      "quantized",
      "q4_k_m",
      "qwen",
      "qwen2.5",
      "instruct",
      "llama-cpp",
      "edge-deployment",
      "structured-output",
      "pbh-applied-systems",
      "quant-eval"
    ],
    "frontmatter": {
      "language": [
        "en",
        "zh",
        "fr",
        "es",
        "pt",
        "de",
        "it",
        "ru",
        "ja",
        "ko",
        "ar"
      ],
      "license": "other",
      "license_name": "qwen-research",
      "license_link": "https://huggingface.co/Qwen/Qwen2.5-3B-Instruct/blob/main/LICENSE",
      "base_model": "Qwen/Qwen2.5-3B-Instruct",
      "tags": [
        "gguf",
        "quantized",
        "q4_k_m",
        "qwen",
        "qwen2.5",
        "instruct",
        "llama-cpp",
        "edge-deployment",
        "structured-output",
        "pbh-applied-systems",
        "quant-eval"
      ]
    },
    "hero_image_url": "",
    "summary": "**Quantized, converted, and evaluated by PBH Applied Systems, LLC** — Applied AI/ML Consulting · LLM Optimization & Deployment · Quantized AI Infrastructure > 🔬 **This repository is part of a production-oriented evaluation series.** Every model published under pbhappliedsystems has been independently evaluated using **quant_eval v7.21** — a proprietary behavioral evaluation harness developed by PBH Applied Systems. Scores measure real agent-adjacent task performance across structured output, tool dispatch, multi-turn state retention, and multi-step planning families — not perplexity or benchmark leaderboard proxies. > ⚖️ **License Notice:** This model is governed by the **Qwen Research License**, which permits **non-commercial use only**. Commercial use requires a separate license from Alibaba Cloud. See the License section for full details. ---",
    "quick_links": [],
    "benchmark_table_html": "",
    "readme_markdown": "---\nlanguage:\n  - en\n  - zh\n  - fr\n  - es\n  - pt\n  - de\n  - it\n  - ru\n  - ja\n  - ko\n  - ar\nlicense: other\nlicense_name: qwen-research\nlicense_link: https://huggingface.co/Qwen/Qwen2.5-3B-Instruct/blob/main/LICENSE\nbase_model: Qwen/Qwen2.5-3B-Instruct\ntags:\n  - gguf\n  - quantized\n  - q4_k_m\n  - qwen\n  - qwen2.5\n  - instruct\n  - llama-cpp\n  - edge-deployment\n  - structured-output\n  - pbh-applied-systems\n  - quant-eval\n---\n\n# Qwen2.5-3B-Instruct · GGUF Q4\\_K\\_M\n\n**Quantized, converted, and evaluated by [PBH Applied Systems, LLC](https://pbhappliedsystems.com)**\n— Applied AI/ML Consulting · LLM Optimization & Deployment · Quantized AI Infrastructure\n\n> 🔬 **This repository is part of a production-oriented evaluation series.** Every model published under [`pbhappliedsystems`](https://huggingface.co/pbhappliedsystems) has been independently evaluated using **quant_eval v7.21** — a proprietary behavioral evaluation harness developed by PBH Applied Systems. Scores measure real agent-adjacent task performance across structured output, tool dispatch, multi-turn state retention, and multi-step planning families — not perplexity or benchmark leaderboard proxies.\n\n> ⚖️ **License Notice:** This model is governed by the **Qwen Research License**, which permits **non-commercial use only**. Commercial use requires a separate license from Alibaba Cloud. See the [License](#license) section for full details.\n\n---\n\n## Model Description\n\nThis repository contains the **4-bit quantized (Q4\\_K\\_M)** GGUF of [`Qwen/Qwen2.5-3B-Instruct`](https://huggingface.co/Qwen/Qwen2.5-3B-Instruct), a 3-billion parameter instruction-tuned model from Alibaba Cloud (September 2024 release). Qwen2.5-3B-Instruct is the smallest model in the PBH Applied Systems evaluated series, and at Q4\\_K\\_M precision it represents the most hardware-accessible deployment option — requiring as little as 4 GB VRAM and fitting on consumer edge hardware.\n\nThe full-precision F16 baseline is published separately at [`pbhappliedsystems/qwen-2.5-3B-instruct-gguf-F16`](https://huggingface.co/pbhappliedsystems/qwen-2.5-3B-instruct-gguf-F16).\n\n### Key Characteristics\n\n- **Parameters:** 3B\n- **Format:** GGUF Q4\\_K\\_M\n- **File size:** 1.93 GB\n- **SHA256:** `9ab3bc9beaddaec3700d5cc754b52e1501a3fd172bc7fc3ee3eb8e1d388ee043`\n- **Minimum VRAM (GPU inference):** ~4 GB\n- **Recommended GPU tier:** Any CUDA-capable GPU · CPU inference viable\n- **Context window:** 32,768 tokens\n- **Inference speed (eval hardware):** avg **0.390 sec/case** on RTX 4090 — fastest Q4\\_K\\_M in the evaluated series\n- **License:** Qwen Research License (non-commercial)\n\n---\n\n## PBH Applied Systems Evaluation — quant\\_eval v7.21\n\n> **Evaluation conducted by PBH Applied Systems, LLC using quant_eval v7.21**\n> Run ID: `20260221_041137` · Fixtures: `golden_oracle_fixtures_v7_21` (SHA256: `6d71a0b9147c...`) · Seed: 42\n> Hardware: NVIDIA RTX 4090 · Total rows evaluated: 84 (42 F16 · 42 Q4\\_K\\_M)\n\n### Aggregate Scores (Q4\\_K\\_M)\n\nScores are normalized to [0.0 – 1.0]. Higher is better.\n\n| Dimension | Score | Series Context |\n|---|---:|---|\n| Task Completion | 0.4905 | Expected at 3B scale |\n| Reasoning | 0.3704 | Limited by parameter count |\n| Coherence | 0.9074 | Strong at 3B |\n| Instruction Following | 0.6599 | Moderate |\n| **Avg inference time** | **0.390 sec/case** | **Fastest in series** |\n\n### Per-Family Pass Rates\n\n#### F16 Baseline (`full_weight_transformers`)\n\n| Family | N | Pass Rate | Avg Secs | Notes |\n|---|---:|---:|---:|---|\n| json\\_multistep | 5 | 0.000 | 3.704 | checks\\_consistent\\_ok bottleneck |\n| stateful\\_followup | 2 | **1.000** | 0.525 | Both turns exact match |\n| toolcall\\_only | 2 | 0.000 | 0.405 | tool\\_name\\_ok=1, args\\_ok=0 |\n| mixed\\_brief\\_json | 2 | 0.000 | 0.695 | JSON valid; ANSWER line absent — see note |\n| toolcall | 2 | **1.000** | 0.745 | Stage-1 passes; final\\_mismatch — see note |\n| json | 4 | n/a | 2.072 | bucket\\_score avg = 10.000 |\n| fuzz | 20 | n/a | 1.437 | bucket\\_score avg = 7.500 — 15/20 pass |\n| mcq | 5 | n/a | 0.022 | Empty raw output on all 5 |\n\n#### Q4\\_K\\_M (`quantized_llama_cpp`)\n\n| Family | N | Pass Rate | Δ vs F16 | Avg Secs | Notes |\n|---|---:|---:|---:|---:|---|\n| json\\_multistep | 5 | **0.200** | **+0.200** | 0.985 | ms\\_easy\\_01 recovers |\n| stateful\\_followup | 2 | **1.000** | 0.000 | 0.125 | Perfect retention |\n| toolcall\\_only | 2 | 0.000 | 0.000 | 0.100 | Wrong arg key names |\n| mixed\\_brief\\_json | 2 | **1.000** | **+1.000** | 0.155 | Full recovery at Q4\\_K\\_M |\n| toolcall | 2 | **1.000** | 0.000 | 0.215 | Stage-1 passes; final\\_mismatch — see note |\n| json | 4 | n/a | — | 0.510 | bucket\\_score avg = 10.000 |\n| fuzz | 20 | n/a | — | 0.408 | bucket\\_score avg = 7.500 — same 15/20 |\n| mcq | 5 | n/a | — | 0.010 | 3/5 pass at Q4\\_K\\_M |\n\n---\n\n## Key Findings\n\n### Finding 1: The Speed Story — Fastest Model in the Evaluated Series\n\nAt **0.390 sec/case average**, this Q4\\_K\\_M variant is the fastest model evaluated in the PBH Applied Systems series. Individual family timings:\n\n| Family | Q4\\_K\\_M Avg Secs |\n|---|---:|\n| mcq | **0.010** |\n| stateful\\_followup | 0.125 |\n| toolcall\\_only | 0.100 |\n| mixed\\_brief\\_json | 0.155 |\n| toolcall | 0.215 |\n| fuzz | 0.408 |\n| json | 0.510 |\n| json\\_multistep | 0.985 |\n\nMCQ responses at 10 milliseconds. Stateful follow-up at 125 milliseconds. Even the most complex planning cases complete in under one second. At 1.93 GB, this model runs on hardware where no other model in this series fits — including systems without dedicated GPUs.\n\n**This speed profile is the primary deployment argument for Qwen2.5-3B Q4\\_K\\_M.** It is not the most capable model evaluated. It is the most accessible model evaluated, with a capability ceiling that is clearly defined by the evaluation data below.\n\n### Finding 2: checks\\_consistent\\_ok = 0.200 — The 3B Ceiling\n\nThe single most consistent failure signal across both runners and all json\\_multistep cases is `checks_consistent_ok`. Both runners achieve the same rate (0.200) — 1/5 cases pass. This is the signal that measures whether the model's intermediate reasoning steps are internally self-consistent.\n\n**This is a parameter-count limitation, not a quantization artifact.** The same 5 fuzz cases fail on both runners (fuzz_0001, _0002, _0003, _0010, _0012) at the same bucket scores, confirming the failure boundary is set by the model's capacity, not by the precision level. A 3B model that can produce valid JSON schemas reliably but cannot maintain consistent multi-step reasoning chains is behaving exactly as expected at this scale.\n\n| Signal | F16 Rate | Q4\\_K\\_M Rate |\n|---|---:|---:|\n| schema\\_ok | 0.800 | **1.000** |\n| checks\\_consistent\\_ok | 0.200 | 0.200 |\n| stop\\_semantics\\_ok | 0.200 | 0.400 |\n| oracle\\_equiv\\_ok | 0.200 | 0.400 |\n\nQ4\\_K\\_M actually improves over F16 on `schema_ok` (0.800 → 1.000), `stop_semantics_ok` (0.200 → 0.400), and `oracle_equiv_ok` (0.200 → 0.400). The `checks_consistent_ok` signal, which is the deepest reasoning consistency test, holds at 0.200 regardless of precision.\n\n### Finding 3: F16 Role-Token Contamination\n\nThe F16 evaluation reveals a specific issue with how the HuggingFace Transformers runner handles this model's chat template. Raw outputs for several families show truncated role token prefixes:\n\n- toolcall cases: `ician {\"tool_name\": \"add\", ...}` — the role token \"technician\" is truncated to \"ician\" and appears in the response body\n- stateful cases: `ician {\"counter\": 2}` — same prefix contamination\n- mixed_brief_json: `user ANSWER: 13 {...}` — the \"user\" role token appears as literal text\n\nThe JSON content in these cases is correct, but the role prefix causes extraction failures. This explains why `toolcall` shows `final_mismatch` at F16 (the stage-1 JSON is valid but the post-response text includes role tokens), and why `mixed_brief_json` fails `answer_line_ok` at F16 despite having `json_parse_ok=1` and `schema_ok=1`.\n\n**This is a runner configuration issue, not a model quality issue.** The Q4_K_M llama.cpp runner does not exhibit this behavior.\n\n### Finding 4: F16 ms\\_easy\\_01 — Wrong Output Type\n\n`ms_easy_01` at F16 is the only case in the evaluated series where a model returns an entirely wrong output type for json_multistep. The F16 model produces:\n\n```json\n[{\"shelf\": \"A\", \"item\": \"P\", \"can_place\": 1}]\n```\n\nAn array, where the task requires a specific schema object with `plan`, `checks`, and `final` keys. This causes `schema_ok=0` — the only F16 json_multistep case with a schema failure. At Q4\\_K\\_M, `ms_easy_01` recovers: `schema_ok=1`, `oracle_equiv_ok=1`, `checks_consistent_ok=1` — a clean pass.\n\n### Finding 5: toolcall\\_only — Schema Vocabulary Errors\n\nBoth `toolcall_only` cases at Q4\\_K\\_M show a consistent vocabulary mismatch:\n\n| Case | Q4\\_K\\_M Raw Output | Expected |\n|---|---|---|\n| toolonly\\_01 | `{\"tool\": \"add\", \"operands\": [5, 10]}` | `{\"tool_name\": \"add\", \"args\": {\"a\": 5, \"b\": 10}}` |\n| toolonly\\_02 | `{\"tool\": \"add\", \"operands\": [25, 75]}` | `{\"tool_name\": \"add\", \"args\": {\"a\": 25, \"b\": 75}}` |\n\nThe model uses `\"tool\"` instead of `\"tool_name\"` and `\"operands\"` (array) instead of `\"args\"` (object with named keys). Tool name recognition passes (`tool_name_ok=1`) because the extractor finds \"add\" in the payload. Argument extraction fails because `\"operands\": [5, 10]` is not a valid schema for `{\"a\": integer, \"b\": integer}`.\n\nThis is a schema vocabulary issue at 3B scale — the model knows the tool concept but not the exact field name schema. A system prompt that explicitly states the required key names (`tool_name`, `args.a`, `args.b`) would likely resolve this for basic tools.\n\n### Finding 6: toolcall — Correct Arithmetic, EOS Contamination\n\nBoth `toolcall` cases at Q4\\_K\\_M pass stage-1 (tool dispatch valid) but produce `final_mismatch`:\n\n| Case | Raw Output | Expected |\n|---|---|---|\n| tool\\_01 | `{...add(2,3)...}<\\|im_end\\|>  5<\\|im_end\\|>` | `5` |\n| tool\\_02 | `{...add(10,-4)...}<\\|im_end\\|>  6<\\|im_end\\|>` | `6` |\n\nThe arithmetic is correct. The EOS token (`<|im_end|>`) appended to the answer string causes the string comparison to fail. A `re.sub(r'<\\|im_end\\|>', '', raw)` strip resolves it. See the Usage section for the validated extraction pattern.\n\n### Finding 7: MCQ — A-Bias Pattern\n\n| Case | Q4\\_K\\_M Result | Raw |\n|---|---|---|\n| mcq\\_01 | ✅ PASS | `B` |\n| mcq\\_02 | ❌ FAIL | `A` (wrong) |\n| mcq\\_03 | ✅ PASS | `C` |\n| mcq\\_04 | ✅ PASS | `B` |\n| mcq\\_05 | ❌ FAIL | `A` (wrong) |\n\nThe two failures both produce `A`. The same A-bias pattern observed in Mistral-Nemo appears here at 3B scale. At F16, all 5 MCQ cases produce empty raw output (`raw=''`), suggesting the F16 runner cannot extract the choice from the model's response format at this parameter count.\n\n---\n\n## Signal-Level Diagnostics\n\n### Q4\\_K\\_M — json\\_multistep\n\n| Signal | F16 Rate | Q4\\_K\\_M Rate | Delta |\n|---|---:|---:|---:|\n| schema\\_ok | 0.800 | **1.000** | +0.200 |\n| checks\\_consistent\\_ok | 0.200 | 0.200 | 0.000 |\n| stop\\_semantics\\_ok | 0.200 | 0.400 | +0.200 |\n| oracle\\_equiv\\_ok | 0.200 | 0.400 | +0.200 |\n\n### Q4\\_K\\_M — stateful\\_followup\n\n| Signal | Rate |\n|---|---:|\n| turn1\\_parse\\_ok | 1.000 |\n| turn2\\_parse\\_ok | 1.000 |\n| turn1\\_exact\\_match | 1.000 |\n| turn2\\_exact\\_match | 1.000 |\n\n### Q4\\_K\\_M — toolcall\\_only\n\n| Signal | Rate |\n|---|---:|\n| tool\\_name\\_ok | 1.000 |\n| args\\_ok | 0.000 |\n\n### Q4\\_K\\_M — mixed\\_brief\\_json\n\n| Signal | Rate |\n|---|---:|\n| answer\\_line\\_ok | 1.000 |\n| json\\_parse\\_ok | 1.000 |\n| schema\\_ok | 1.000 |\n\n---\n\n## Recommended Use Cases\n\n### ✅ Deploy with Confidence (Q4\\_K\\_M, Non-Commercial)\n\n- **Stateful multi-turn agents** — Perfect two-turn retention (1.000) at 0.125 sec/case. The fastest stateful performance in the evaluated series.\n- **Structured JSON outputs (single-step)** — `json` bucket\\_score 10.000, fuzz bucket\\_score 7.500. Fast and reliable for constraint-adherent single-step placements.\n- **Hybrid brief + JSON responses** — `mixed_brief_json` passes at 1.000 in 0.155 sec. Clean ANSWER line + valid JSON.\n- **Edge and resource-constrained deployment** — 1.93 GB at Q4\\_K\\_M. Runs on 4 GB VRAM, embedded GPUs, and CPU-only environments where no other evaluated model fits.\n- **High-throughput batch processing** — At 0.390 sec/case average, this model can process thousands of structured inference tasks per hour on modest hardware.\n\n### ⚠️ Use with Guardrails (Q4\\_K\\_M)\n\n- **Scaffolded tool-calling** — `toolcall` stage-1 passes at 1.000 but add EOS stripping before final answer extraction. Arithmetic results are correct.\n- **Bare tool-call dispatch** — `toolcall_only` fails on args schema (`\"operands\"` vs `\"args\"`). Provide explicit schema examples in the system prompt or add a normalization layer.\n- **Multi-step planning (easy difficulty only)** — ms\\_easy\\_01 passes. Higher difficulties fail consistently on `checks_consistent_ok`. Use with external validation.\n- **MCQ applications** — 3/5 pass but A-bias exists. Add chain-of-thought prompting or validation for MCQ pipelines.\n\n### ❌ Not Recommended (Q4\\_K\\_M)\n\n- **Medium-to-hard multi-step planning** — ms\\_med and ms\\_hard cases fail on internal consistency. The 3B parameter count is the limiting factor, not the quantization.\n- **Any commercial use without a Qwen commercial license** — The Qwen Research License restricts use to non-commercial research and evaluation purposes. Contact Alibaba Cloud for commercial licensing.\n\n---\n\n## Hardware Requirements\n\n| Configuration | VRAM Required | Notes |\n|---|---|---|\n| Q4\\_K\\_M (this repo) · GPU | ~4 GB | Fits on embedded/mobile GPUs |\n| Q4\\_K\\_M · CPU only | 4 GB RAM | Viable; slower inference |\n| Q4\\_K\\_M · 32K context | ~8 GB | Full context window |\n| F16 (companion repo) · GPU | ~8 GB | 6.18 GB model + overhead |\n\n---\n\n## Usage\n\n### Installation\n\n```bash\npip install llama-cpp-python huggingface_hub\n```\n\nFor GPU acceleration (CUDA):\n\n```bash\nCMAKE_ARGS=\"-DGGML_CUDA=on\" pip install llama-cpp-python --force-reinstall --no-cache-dir\n```\n\n### Python — llama-cpp-python\n\n```python\nfrom huggingface_hub import hf_hub_download\nfrom llama_cpp import Llama\n\nmodel_path = hf_hub_download(\n    repo_id=\"pbhappliedsystems/qwen-2.5-3B-instruct-gguf-Q4-K-M\",\n    filename=\"qwen-2.5-3B-instruct-gguf-Q4-K-M.gguf\"\n)\n\nllm = Llama(\n    model_path=model_path,\n    n_ctx=4096,\n    n_gpu_layers=-1,  # -1 for full GPU; set 0 for CPU-only\n    verbose=False,\n)\n\nresponse = llm.create_chat_completion(\n    messages=[\n        {\n            \"role\": \"system\",\n            \"content\": \"You are a precise assistant. Return structured JSON when asked. Use exactly the key names specified.\"\n        },\n        {\n            \"role\": \"user\",\n            \"content\": \"Summarize the following and return JSON with keys: summary, sentiment, action_items.\"\n        }\n    ],\n    temperature=0.15,\n    max_tokens=512,\n)\n\nprint(response[\"choices\"][0][\"message\"][\"content\"])\n```\n\nFor CPU-only deployment (no GPU required):\n\n```python\nllm = Llama(\n    model_path=model_path,\n    n_ctx=2048,\n    n_gpu_layers=0,    # CPU-only\n    n_threads=8,       # Tune to your CPU core count\n    verbose=False,\n)\n```\n\nFor tool-calling with EOS stripping (addresses toolcall final_mismatch finding):\n\n```python\nimport json, re\nfrom huggingface_hub import hf_hub_download\nfrom llama_cpp import Llama\n\nmodel_path = hf_hub_download(\n    repo_id=\"pbhappliedsystems/qwen-2.5-3B-instruct-gguf-Q4-K-M\",\n    filename=\"qwen-2.5-3B-instruct-gguf-Q4-K-M.gguf\"\n)\n\nllm = Llama(model_path=model_path, n_ctx=2048, n_gpu_layers=-1, verbose=False)\n\ndef call_tool_with_cleanup(prompt: str) -> dict:\n    \"\"\"\n    Tool dispatch with EOS stripping.\n    quant_eval v7.21: toolcall stage-1 pass=1.000; final_mismatch due to <|im_end|> suffix.\n    Arithmetic is correct — strip EOS before downstream processing.\n    \"\"\"\n    response = llm.create_chat_completion(\n        messages=[\n            {\n                \"role\": \"system\",\n                \"content\": (\n                    \"You are a tool-calling assistant. \"\n                    \"When calling a tool, output JSON as: \"\n                    '{\"tool_name\": \"<name>\", \"args\": {\"a\": <n>, \"b\": <n>}}\\n'\n                    \"Then on the next line, output the result as a plain number.\"\n                )\n            },\n            {\"role\": \"user\", \"content\": prompt}\n        ],\n        temperature=0.0,\n        max_tokens=128,\n    )\n    raw = response[\"choices\"][0][\"message\"][\"content\"]\n    # Strip EOS tokens before processing\n    clean = re.sub(r'<\\|im_end\\|>', '', raw).strip()\n    return {\"raw\": raw, \"clean\": clean}\n\nresult = call_tool_with_cleanup(\"Use the add tool to compute 5 plus 10.\")\nprint(result[\"clean\"])\n```\n\n### CLI — llama-cli\n\n```bash\nllama-cli \\\n  --model qwen-2.5-3B-instruct-gguf-Q4-K-M.gguf \\\n  --chat-template qwen2 \\\n  --system-prompt \"You are a precise assistant. Use exactly the key names specified in any JSON schema.\" \\\n  --prompt \"Return a JSON object with keys: summary, risk_level, action_items.\" \\\n  --n-predict 512 \\\n  --ctx-size 4096 \\\n  --n-gpu-layers -1 \\\n  --temp 0.15\n```\n\nFor server deployment:\n\n```bash\nllama-server \\\n  --model qwen-2.5-3B-instruct-gguf-Q4-K-M.gguf \\\n  --chat-template qwen2 \\\n  --ctx-size 4096 \\\n  --n-gpu-layers -1 \\\n  --port 8080 \\\n  --host 0.0.0.0\n```\n\nQuery via the OpenAI-compatible API:\n\n```python\nfrom openai import OpenAI\nimport re\n\nclient = OpenAI(base_url=\"http://localhost:8080/v1\", api_key=\"not-required\")\n\nresponse = client.chat.completions.create(\n    model=\"qwen-2.5-3B-instruct-gguf-Q4-K-M\",\n    messages=[{\"role\": \"user\", \"content\": \"Your prompt here\"}],\n    temperature=0.15,\n)\n# Strip EOS tokens from output\nclean = re.sub(r'<\\|im_end\\|>', '', response.choices[0].message.content).strip()\nprint(clean)\n```\n\n---\n\n## Evaluation Artifacts\n\nThe full per-case evaluation CSV (`comparison_results_v7_21_Qwen2.5_3B_Instruct_20260221_041137.csv`) and `rollup.json` are published in this repository for independent verification.\n\n---\n\n## Artifact Provenance\n\n| Artifact | Format | Size | SHA256 |\n|---|---|---|---|\n| `qwen-2.5-3B-instruct-gguf-Q4-K-M.gguf` | GGUF Q4\\_K\\_M | 1.93 GB | `9ab3bc9beaddaec3700d5cc754b52e1501a3fd172bc7fc3ee3eb8e1d388ee043` |\n| F16 *(companion repo)* | GGUF F16 | 6.18 GB | `65a0239fc9f9a40e2d4f79ae5e158cad423c2476fe089c744d5e6a6ff6fc9330` |\n\nBoth artifacts were produced from `Qwen/Qwen2.5-3B-Instruct` using a custom-built llama.cpp conversion and quantization pipeline developed by PBH Applied Systems.\n\n---\n\n## Evaluation Methodology\n\n**quant_eval v7.21** is a proprietary behavioral evaluation harness developed by PBH Applied Systems.\n\n**Fixture set:** `golden_oracle_fixtures_v7_21` (SHA256: `6d71a0b9147c079371b02a94f3c149eb78a6adc03dc16ff6833b964fbf4174f0`)\n\n| Family | Description | Pass Signals |\n|---|---|---|\n| `fuzz` | Property-based regression; structured placement correctness | schema\\_ok, constraints\\_ok |\n| `json` | Single-step structured JSON with constraint rules | schema\\_ok, constraints\\_ok |\n| `json_multistep` | Multi-step planning with self-check and oracle verification | schema\\_ok, checks\\_consistent\\_ok, stop\\_semantics\\_ok, oracle\\_equiv\\_ok |\n| `mcq` | Multiple-choice extraction | choice\\_ok |\n| `stateful_followup` | Two-turn state tracking; turn-2 correct given turn-1 | turn1/2\\_parse\\_ok, turn1/2\\_exact\\_match |\n| `mixed_brief_json` | Hybrid: natural language answer + valid JSON block | answer\\_line\\_ok, json\\_parse\\_ok, schema\\_ok |\n| `toolcall` | Tool call embedded in response; parse + schema validation | stage1\\_tool\\_parse\\_ok, stage1\\_tool\\_schema\\_ok |\n| `toolcall_only` | Bare schema-only tool call; strict tool name + args check | tool\\_name\\_ok, args\\_ok |\n\n**Evaluation hardware:** NVIDIA RTX 4090 (24 GB VRAM)\n**Evaluation date:** February 21, 2026\n**quant_eval seed:** 42\n\n---\n\n## About PBH Applied Systems\n\n[**PBH Applied Systems, LLC**](https://pbhappliedsystems.com) is an Oklahoma City–based applied machine learning and AI systems company specializing in production-grade model evaluation, quantization pipelines, agentic AI infrastructure, and scalable AI-driven application development. The organization operates with a strong emphasis on engineering rigor, reproducibility, and real-world deployment constraints — particularly in environments where performance, cost efficiency, and reliability must be balanced against available hardware and budget.\n\n### Founder — Patrick Hill, M.S.\n\nPBH Applied Systems was founded by **Patrick Hill**, a Data Scientist and AI/ML Engineer with 10+ years of experience delivering advanced analytics, predictive modeling, and decision-support solutions across high-stakes operational environments. Patrick holds a **Master of Science in Software Engineering with concentrations in Artificial Intelligence and Machine Learning** (GPA: 4.0) and a B.S. in Business Finance.\n\n**Technical expertise spans:**\n\n- **Languages & Data:** Python, SQL, Linux, Pandas, NumPy, scikit-learn\n- **ML & Modeling:** Supervised and unsupervised learning, neural networks, NLP, transformers, regression, classification, forecasting, and feature engineering\n- **AI/ML Frameworks:** PyTorch, TensorFlow/Keras, HuggingFace Transformers, GGUF, llama.cpp, BitsAndBytes, PEFT, QLoRA\n- **Deployment & MLOps:** Flask APIs, Docker, CI/CD pipelines, REST endpoints, streaming inference, version control\n- **Data Platforms:** Jupyter, Databricks, Power BI, Matplotlib\n- **Quantization:** GGUF conversion, Q4\\_K\\_M / Q5\\_K\\_M / Q8\\_0 strategies, adapter-per-model evaluation architecture\n\n### Published Author\n\nPatrick is the author of **[Applied Machine Learning: Concepts, Tools, and Case Studies](https://a.co/d/05qat7Xz)** — a 1,200+ page practitioner-oriented textbook adopted as **required reading for CSC 373 – Machine Learning at the University of Advancing Technology**.\n\n### Core Service Areas\n\n**1. LLM Optimization & Deployment** · **2. AI Evaluation Frameworks** · **3. Agentic AI Infrastructure** · **4. Scalable AI Application Development** · **5. ML Pipeline Design & Analytics** · **6. Model & Agent Cataloging**\n\n---\n\n## 📞 Work With PBH Applied Systems\n\nThe Qwen2.5-3B Q4\\_K\\_M evaluation tells a clear story about what to expect at the 3B scale: sub-second inference, reliable structured outputs, solid stateful retention, and a well-defined ceiling on multi-step reasoning consistency. Every finding is verifiable — download the CSV and check the `checks_consistent_ok` signal yourself.\n\n👉 **[Book a Scoping Call](https://pbhappliedsystems.com)** — Discuss model selection, edge deployment strategy, or evaluation needs directly with Patrick.\n\n👉 **[Request an Evaluation Report](https://pbhappliedsystems.com)** — Full quant_eval behavioral audit for your target model(s). Engagements from $2,500.\n\n### Connect\n\n| | |\n|---|---|\n| 🌐 **Website** | [pbhappliedsystems.com](https://pbhappliedsystems.com) |\n| 📧 **Email** | [patrick@pbhappliedsystems.com](mailto:patrick@pbhappliedsystems.com) |\n| 💼 **LinkedIn** | [PBH Applied Systems, LLC](https://www.linkedin.com/company/pbh-applied-systems-llc) |\n| ▶️ **YouTube** | [@pbhappliedsystems](https://www.youtube.com/@pbhappliedsystems) |\n| 📸 **Instagram** | [@pbhappliedsystems](https://www.instagram.com/pbhappliedsystems) |\n| 👍 **Facebook** | [pbhappliedsystems](https://www.facebook.com/pbhappliedsystems) |\n\n---\n\n## License\n\nThis GGUF repository is governed by the **Qwen Research License Agreement**, inherited from the base model [`Qwen/Qwen2.5-3B-Instruct`](https://huggingface.co/Qwen/Qwen2.5-3B-Instruct).\n\n**Key terms:**\n\n- **Non-commercial use only** — The license grants a non-exclusive, royalty-free license for **research or evaluation purposes only**\n- **Commercial use requires a separate license** — Contact Alibaba Cloud to request a commercial license\n- **Redistribution** — Permitted with attribution and license copy included\n- **Attribution requirement** — If used to train or improve another AI model, display \"Built with Qwen\" or \"Improved using Qwen\" in documentation\n- **Governing law** — People's Republic of China; exclusive jurisdiction in Hangzhou City People's Courts\n\nFull license text: [Qwen Research License Agreement](https://huggingface.co/Qwen/Qwen2.5-3B-Instruct/blob/main/LICENSE)\n\nThe quant_eval evaluation methodology, fixture set, and scoring framework are proprietary to PBH Applied Systems, LLC and are not included in this repository.\n\n---\n\n*GGUF conversion, quantization, and behavioral evaluation performed by [PBH Applied Systems, LLC](https://pbhappliedsystems.com) · quant_eval v7.21 · Run ID: `20260221_041137`*\n",
    "related_quantizations": []
  },
  "tags": [
    "gguf",
    "quantized",
    "q4_k_m",
    "qwen",
    "qwen2.5",
    "instruct",
    "llama-cpp",
    "edge-deployment",
    "structured-output",
    "pbh-applied-systems",
    "quant-eval",
    "en",
    "zh",
    "fr",
    "es",
    "pt",
    "de",
    "it",
    "ru",
    "ja",
    "ko",
    "ar",
    "base_model:Qwen/Qwen2.5-3B-Instruct",
    "base_model:quantized:Qwen/Qwen2.5-3B-Instruct",
    "license:other",
    "endpoints_compatible",
    "region:us",
    "conversational"
  ],
  "likes": 0,
  "downloads": 530,
  "gated": false,
  "private": false,
  "last_modified": "2026-04-15T07:03:48.000Z",
  "created_at": "2026-04-12T06:01:39.000Z",
  "pipeline_tag": "",
  "library_name": ""
}
Source payload excerpt (from Hugging Face API)
{
  "_id": "69db35430a7d72741cb3ae28",
  "id": "pbhappliedsystems/qwen-2.5-3B-instruct-gguf-Q4-K-M",
  "modelId": "pbhappliedsystems/qwen-2.5-3B-instruct-gguf-Q4-K-M",
  "sha": "c239150fd8f97060737146a7a16de78472bb6d77",
  "createdAt": "2026-04-12T06:01:39.000Z",
  "lastModified": "2026-04-15T07:03:48.000Z",
  "author": "pbhappliedsystems",
  "downloads": 530,
  "likes": 0,
  "gated": false,
  "private": false,
  "pipeline_tag": "",
  "library_name": "",
  "siblings_count": 7
}