GraySoft
Projects Models About FAQ Contact Download guIDE →
Model Intelligence Sheet

pbhappliedsystems/ministral-3-14b-instruct-2512-gguf-q4-k-m overview

Quantized, converted, and evaluated by PBH Applied Systems, LLC — Applied AI/ML Consulting · LLM Optimization & Deployment · Quantized AI Infrastructure 🔬 This repository is part of a production-oriented evaluation series. Every model published under pbhappliedsystems has been independently evaluated using quant_eval v7.21 — a proprietary behavioral evaluation harness developed by PBH Applied Systems. Scores measure real agent-adjacent task performance across structured output, tool dispatch, multi-turn state retention, and multi-step planning families — not perplexity or benchmark leaderboard proxies. ---

ggufquantizedq4_k_mmistralinstructllama-cppagentictool-callingstructured-outputpbh-applied-systemsquant-evalenbase_model:mistralai/Ministral-3-14B-Instruct-2512base_model:quantized:mistralai/Ministral-3-14B-Instruct-2512license:apache-2.0endpoints_compatibleregion:usconversational
pbhappliedsystems/ministral-3-14b-instruct-2512-gguf-q4-k-m visual
Downloads
1,030
Likes
0
Pipeline
Library
Visibility
Public
Access
Open

Repository Files & Downloads

1 files detected
Direct downloads for all repository files
FileTypeQuantizationSizeLink
ministral-3-14b-instruct-2512-gguf-Q4-K-M.gguf GGUF 7.67 GB Download

Model Details Live

Model Slug
pbhappliedsystems/ministral-3-14b-instruct-2512-gguf-q4-k-m
Author
pbhappliedsystems
Pipeline Task
Library
Created
2026-04-11
Last Modified
2026-04-15
Gated
No
Private
No
HF SHA
8cd64684d68cab4cb8bce35c4ea7070f1e43f375
License
apache-2.0
Language
en
Base Model
mistralai/Ministral-3-14B-Instruct-2512

Metadata Inspector

Normalized metadata (stored in metadata_json)
{
  "metadata": {},
  "card_data": {
    "language": [
      "en"
    ],
    "license": "apache-2.0",
    "base_model": "mistralai/Ministral-3-14B-Instruct-2512",
    "tags": [
      "gguf",
      "quantized",
      "q4_k_m",
      "mistral",
      "instruct",
      "llama-cpp",
      "agentic",
      "tool-calling",
      "structured-output",
      "pbh-applied-systems",
      "quant-eval"
    ],
    "frontmatter": {
      "language": [
        "en"
      ],
      "license": "apache-2.0",
      "base_model": "mistralai/Ministral-3-14B-Instruct-2512",
      "tags": [
        "gguf",
        "quantized",
        "q4_k_m",
        "mistral",
        "instruct",
        "llama-cpp",
        "agentic",
        "tool-calling",
        "structured-output",
        "pbh-applied-systems",
        "quant-eval"
      ]
    },
    "hero_image_url": "",
    "summary": "**Quantized, converted, and evaluated by PBH Applied Systems, LLC** — Applied AI/ML Consulting · LLM Optimization & Deployment · Quantized AI Infrastructure > 🔬 **This repository is part of a production-oriented evaluation series.** Every model published under pbhappliedsystems has been independently evaluated using **quant_eval v7.21** — a proprietary behavioral evaluation harness developed by PBH Applied Systems. Scores measure real agent-adjacent task performance across structured output, tool dispatch, multi-turn state retention, and multi-step planning families — not perplexity or benchmark leaderboard proxies. ---",
    "quick_links": [],
    "benchmark_table_html": "",
    "readme_markdown": "---\nlanguage:\n  - en\nlicense: apache-2.0\nbase_model: mistralai/Ministral-3-14B-Instruct-2512\ntags:\n  - gguf\n  - quantized\n  - q4_k_m\n  - mistral\n  - instruct\n  - llama-cpp\n  - agentic\n  - tool-calling\n  - structured-output\n  - pbh-applied-systems\n  - quant-eval\n---\n\n# Ministral-3-14B-Instruct-2512 · GGUF Q4\\_K\\_M\n\n**Quantized, converted, and evaluated by [PBH Applied Systems, LLC](https://pbhappliedsystems.com)**\n— Applied AI/ML Consulting · LLM Optimization & Deployment · Quantized AI Infrastructure\n\n> 🔬 **This repository is part of a production-oriented evaluation series.** Every model published under [`pbhappliedsystems`](https://huggingface.co/pbhappliedsystems) has been independently evaluated using **quant_eval v7.21** — a proprietary behavioral evaluation harness developed by PBH Applied Systems. Scores measure real agent-adjacent task performance across structured output, tool dispatch, multi-turn state retention, and multi-step planning families — not perplexity or benchmark leaderboard proxies.\n\n---\n\n## Model Description\n\nThis repository contains the **4-bit quantized (Q4\\_K\\_M)** GGUF of [`mistralai/Ministral-3-14B-Instruct-2512`](https://huggingface.co/mistralai/Ministral-3-14B-Instruct-2512), a 14-billion parameter instruction-tuned model from Mistral AI (December 2025 release).\n\nThe Q4\\_K\\_M format applies 4-bit quantization with K-quant medium precision, targeting a balance of inference speed and output fidelity suitable for deployment on consumer and professional GPU hardware. The full-precision F16 baseline is published separately at [`pbhappliedsystems/ministral-3-14b-instruct-2512-gguf-F16`](https://huggingface.co/pbhappliedsystems/ministral-3-14b-instruct-2512-gguf-F16).\n\n### Key Characteristics\n\n- **Parameters:** 14B\n- **Format:** GGUF Q4\\_K\\_M\n- **File size:** 8.24 GB\n- **SHA256:** `a23910514ee512aa28db8dddd390c26a73b9c318dcdec374ae02d722d9658749`\n- **Minimum VRAM (GPU inference):** ~10–12 GB (T4 class or better)\n- **Recommended GPU tier:** NVIDIA T4 (16 GB) · RTX 3080/4080 · A10G\n- **Context window:** 32,768 tokens (per base model specification)\n- **Inference speed (eval hardware):** avg **3.77 sec/case** on RTX 4090\n\n---\n\n## PBH Applied Systems Evaluation — quant\\_eval v7.21\n\n> **Evaluation conducted by PBH Applied Systems, LLC using quant_eval v7.21**\n> Run ID: `20260209_170235` · Fixtures: `golden_oracle_fixtures_v7_21` (SHA256: `6d71a0b9147c...`) · Seed: 42\n> Hardware: NVIDIA RTX 4090 · Total rows evaluated: 84 (42 F16 · 42 Q4\\_K\\_M)\n\n### Aggregate Scores (Q4\\_K\\_M)\n\nScores are normalized to [0.0 – 1.0]. Higher is better.\n\n| Dimension | Score |\n|---|---:|\n| Task Completion | 0.6809 |\n| Reasoning | 0.9148 |\n| Coherence | 0.9259 |\n| Instruction Following | 0.9689 |\n| **Avg inference time** | **3.77 sec/case** |\n\n### Per-Family Pass Rates\n\nThe evaluation runs 8 task families. Pass rate is a conjunction of all gating signals for that family — a strict, production-oriented measure. Families marked `n/a` use bucket scoring rather than binary pass/fail.\n\n#### F16 Baseline (`full_weight_transformers`)\n\n| Family | N | Pass Rate | Notes |\n|---|---:|---:|---|\n| json\\_multistep | 5 | 0.600 | Tier-1 gating: schema, checks, stop semantics, oracle equivalence |\n| stateful\\_followup | 2 | **1.000** | Both turns parse and match expected state |\n| toolcall\\_only | 2 | **1.000** | Tool name + args schema correct |\n| mixed\\_brief\\_json | 2 | **1.000** | Answer line + JSON schema correct |\n| toolcall | 2 | **1.000** | Tool parse + schema valid |\n| json | 4 | n/a | bucket\\_score avg = 10.000 |\n| fuzz | 20 | n/a | bucket\\_score avg = 10.000 |\n| mcq | 5 | n/a | bucket\\_score avg = 1.000 |\n\n#### Q4\\_K\\_M (`quantized_llama_cpp`)\n\n| Family | N | Pass Rate | Δ vs F16 | Notes |\n|---|---:|---:|---:|---|\n| json\\_multistep | 5 | 0.600 | 0.000 | No degradation at Tier-1 |\n| stateful\\_followup | 2 | **1.000** | 0.000 | Perfect retention |\n| toolcall\\_only | 2 | **0.000** | **−1.000** | ⚠️ Full degradation — see below |\n| mixed\\_brief\\_json | 2 | **1.000** | 0.000 | No degradation |\n| toolcall | 2 | **1.000** | 0.000 | No degradation |\n| json | 4 | n/a | — | bucket\\_score avg = 10.000 |\n| fuzz | 20 | n/a | — | bucket\\_score avg = 10.000 |\n| mcq | 5 | n/a | — | bucket\\_score avg = 1.000 |\n\n### ⚠️ Quantization Degradation Finding — toolcall\\_only\n\n**`toolcall_only` degraded from 1.000 (F16) to 0.000 (Q4\\_K\\_M).** Both test cases failed on `tool_name_ok` and `args_ok` simultaneously. This means the quantized model, when asked to emit a bare tool-call JSON with no surrounding prose or chain-of-thought scaffold, lost the ability to produce a schema-valid payload entirely.\n\n**Practical implication:** Do not deploy this Q4\\_K\\_M variant in pipelines where the model is expected to emit raw tool-call JSON without an external schema enforcement layer or retry loop. The `toolcall` family (tool call embedded in a broader response) remained at 1.000, indicating this degradation is specific to the strict schema-only output format.\n\n**This is exactly the kind of signal that pre-deployment quantization evaluation is designed to surface.** The F16 baseline passes cleanly; the Q4\\_K\\_M variant does not. Without independent evaluation of both formats before deployment, this failure mode reaches production silently.\n\n### Signal-Level Diagnostics (Q4\\_K\\_M)\n\n#### json\\_multistep\n\n| Signal | Rate | Tier |\n|---|---:|---|\n| schema\\_ok | 1.000 | Tier-1 (gating) |\n| checks\\_consistent\\_ok | 0.800 | Tier-1 (gating) |\n| stop\\_semantics\\_ok | 1.000 | Tier-1 (gating) |\n| oracle\\_equiv\\_ok | 0.600 | Tier-1 (gating) |\n| final\\_consistent\\_ok | 0.000 | Tier-2 (tracked, non-gating) |\n| final\\_match\\_reported | 0.000 | Tier-2 (tracked, non-gating) |\n\n> **Note on Tier-2:** `final_consistent_ok` and `final_match_reported` are not gating signals. Most deployed agentic systems compute and validate state externally; the model is not expected to report its own final state with exact fidelity. Tier-1 oracle equivalence (0.600) is the production-relevant signal here.\n\n#### stateful\\_followup\n\n| Signal | Rate | Tier |\n|---|---:|---|\n| turn1\\_parse\\_ok | 1.000 | Tier-1 |\n| turn2\\_parse\\_ok | 1.000 | Tier-1 |\n| turn1\\_exact\\_match | 1.000 | Tier-1 |\n| turn2\\_exact\\_match | 1.000 | Tier-1 |\n\n#### toolcall\\_only\n\n| Signal | Rate | Tier |\n|---|---:|---|\n| tool\\_name\\_ok | 0.000 | Tier-1 |\n| args\\_ok | 0.000 | Tier-1 |\n\n#### mixed\\_brief\\_json\n\n| Signal | Rate | Tier |\n|---|---:|---|\n| answer\\_line\\_ok | 1.000 | Tier-1 |\n| json\\_parse\\_ok | 1.000 | Tier-1 |\n| schema\\_ok | 1.000 | Tier-1 |\n\n---\n\n## Recommended Use Cases\n\nDerived from `catalog_recommendation.json` (quant_eval v7.21, run `20260209_170235`).\n\n### ✅ Deploy with Confidence (Q4\\_K\\_M)\n\n- **Stateful multi-turn agents** — Two-turn state retention is perfect (1.000). Suitable for conversational agents where turn-2 depends on turn-1 parsed state.\n- **Structured JSON outputs (single-step)** — bucket\\_score avg of 10.000 on both `json` and `fuzz` families indicates consistently valid structured outputs.\n- **Hybrid brief + JSON responses** — `mixed_brief_json` passes at 1.000; combining a natural language answer line with a JSON payload is reliable.\n- **Tool-calling with response scaffolding** — `toolcall` (tool call embedded within a broader response) passes at 1.000. Use with a response template or instruction scaffold.\n- **JSON multi-step with external validation loop** — Pass rate of 0.600 is below the conservative PASS threshold but workable when an external planner or repair loop verifies each step.\n\n### ⚠️ Use with Guardrails (Q4\\_K\\_M)\n\n- **Bare tool-call dispatch (schema-only output)** — `toolcall_only` failed completely (0.000). A schema enforcement layer, retry policy, or output parser is required for reliable tool dispatch without surrounding prose.\n\n### ❌ Not Recommended (Q4\\_K\\_M)\n\n- **Unassisted multi-step planning** — Where planning correctness must hold without external verification or oracle validation.\n\n---\n\n## Hardware Requirements\n\n| Configuration | VRAM Required | Recommended GPU |\n|---|---|---|\n| Q4\\_K\\_M (this repo) · GPU only | ~10–12 GB | T4 16 GB · RTX 3080/4080 · A10G |\n| Q4\\_K\\_M · CPU offload fallback | 8 GB VRAM + 4 GB RAM | Any CUDA-capable GPU |\n| F16 baseline (companion repo) | ~30 GB | A100 40 GB · RTX 4090 · 2× A10G |\n\n---\n\n## Usage\n\n### Installation\n\n```bash\npip install llama-cpp-python huggingface_hub\n```\n\nFor GPU acceleration (CUDA):\n\n```bash\nCMAKE_ARGS=\"-DGGML_CUDA=on\" pip install llama-cpp-python --force-reinstall --no-cache-dir\n```\n\n### Python — llama-cpp-python\n\n```python\nfrom huggingface_hub import hf_hub_download\nfrom llama_cpp import Llama\n\n# Download directly from HuggingFace Hub\nmodel_path = hf_hub_download(\n    repo_id=\"pbhappliedsystems/ministral-3-14b-instruct-2512-gguf-Q4-K-M\",\n    filename=\"ministral-3-14b-instruct-2512-gguf-Q4-K-M.gguf\"\n)\n\nllm = Llama(\n    model_path=model_path,\n    n_ctx=8192,          # context window; increase up to 32768 per model spec\n    n_gpu_layers=-1,     # -1 offloads all layers to GPU\n    verbose=False,\n)\n\nresponse = llm.create_chat_completion(\n    messages=[\n        {\n            \"role\": \"system\",\n            \"content\": \"You are a helpful, concise assistant. Respond in structured JSON when asked.\"\n        },\n        {\n            \"role\": \"user\",\n            \"content\": \"Summarize the following contract clause and flag any obligations: ...\"\n        }\n    ],\n    temperature=0.15,\n    max_tokens=1024,\n)\n\nprint(response[\"choices\"][0][\"message\"][\"content\"])\n```\n\nFor tool-calling use cases, enforce output schema externally (see evaluation findings above):\n\n```python\nimport json\nfrom huggingface_hub import hf_hub_download\nfrom llama_cpp import Llama\n\nmodel_path = hf_hub_download(\n    repo_id=\"pbhappliedsystems/ministral-3-14b-instruct-2512-gguf-Q4-K-M\",\n    filename=\"ministral-3-14b-instruct-2512-gguf-Q4-K-M.gguf\"\n)\n\nllm = Llama(\n    model_path=model_path,\n    n_ctx=4096,\n    n_gpu_layers=-1,\n    verbose=False,\n)\n\ndef call_with_tool_enforcement(prompt: str, retries: int = 3) -> dict:\n    \"\"\"\n    Wrap tool-call dispatch with schema enforcement and retry.\n    Required for Q4_K_M: toolcall_only pass rate = 0.000 (see eval above).\n    \"\"\"\n    for attempt in range(retries):\n        response = llm.create_chat_completion(\n            messages=[\n                {\"role\": \"system\", \"content\": \"Respond only with a valid JSON tool call.\"},\n                {\"role\": \"user\", \"content\": prompt}\n            ],\n            temperature=0.0,\n            max_tokens=256,\n        )\n        raw = response[\"choices\"][0][\"message\"][\"content\"].strip()\n        try:\n            parsed = json.loads(raw)\n            assert \"tool_name\" in parsed and \"args\" in parsed\n            return parsed\n        except (json.JSONDecodeError, AssertionError):\n            if attempt == retries - 1:\n                raise ValueError(f\"Tool call failed after {retries} attempts. Raw: {raw}\")\n\nresult = call_with_tool_enforcement(\"Place item P on shelf A.\")\n```\n\n### CLI — llama-cli\n\n```bash\n# One-shot prompt\nllama-cli \\\n  --model ministral-3-14b-instruct-2512-gguf-Q4-K-M.gguf \\\n  --chat-template mistral \\\n  --system-prompt \"You are a helpful assistant.\" \\\n  --prompt \"Summarize the following and return a JSON object with keys: summary, risk_level, action_items.\" \\\n  --n-predict 512 \\\n  --ctx-size 8192 \\\n  --n-gpu-layers -1 \\\n  --temp 0.15\n```\n\nFor server deployment (OpenAI-compatible endpoint):\n\n```bash\nllama-server \\\n  --model ministral-3-14b-instruct-2512-gguf-Q4-K-M.gguf \\\n  --chat-template mistral \\\n  --ctx-size 8192 \\\n  --n-gpu-layers -1 \\\n  --port 8080 \\\n  --host 0.0.0.0\n```\n\nThen query via the OpenAI-compatible API:\n\n```python\nfrom openai import OpenAI\n\nclient = OpenAI(base_url=\"http://localhost:8080/v1\", api_key=\"not-required\")\n\nresponse = client.chat.completions.create(\n    model=\"ministral-3-14b-instruct-2512-gguf-Q4-K-M\",\n    messages=[{\"role\": \"user\", \"content\": \"Your prompt here\"}],\n    temperature=0.15,\n)\nprint(response.choices[0].message.content)\n```\n\n---\n\n## Artifact Provenance\n\n| Artifact | Format | Size | SHA256 |\n|---|---|---|---|\n| `ministral-3-14b-instruct-2512-gguf-Q4-K-M.gguf` | GGUF Q4\\_K\\_M | 8.24 GB | `a23910514ee512aa28db8dddd390c26a73b9c318dcdec374ae02d722d9658749` |\n| F16 *(companion repo)* | GGUF F16 | 27.0 GB | `74ea113134173d29f8daba097457500e831eace3741de002846b3ab89781fd52` |\n\nBoth artifacts were produced from `mistralai/Ministral-3-14B-Instruct-2512` using a custom-built llama.cpp conversion and quantization pipeline developed by PBH Applied Systems. Conversion and quantization were performed on the full HuggingFace snapshot without modification to model weights prior to conversion.\n\n---\n\n## Evaluation Methodology\n\n**quant_eval v7.21** is a proprietary behavioral evaluation harness developed by PBH Applied Systems. It evaluates both the full-precision (F16) and quantized variants of a model against an identical fixture set, enabling direct comparison of capability retention across quantization levels.\n\n**Fixture set:** `golden_oracle_fixtures_v7_21` (SHA256: `6d71a0b9147c079371b02a94f3c149eb78a6adc03dc16ff6833b964fbf4174f0`)\n\n**Task families evaluated:**\n\n| Family | Description | Pass Signals |\n|---|---|---|\n| `fuzz` | Property-based regression; structured placement correctness | schema\\_ok, constraints\\_ok |\n| `json` | Single-step structured JSON with constraint rules | schema\\_ok, constraints\\_ok |\n| `json_multistep` | Multi-step planning with self-check and oracle verification | schema\\_ok, checks\\_consistent\\_ok, stop\\_semantics\\_ok, oracle\\_equiv\\_ok |\n| `mcq` | Multiple-choice extraction | choice\\_ok |\n| `stateful_followup` | Two-turn state tracking; turn-2 correct given turn-1 | turn1/2\\_parse\\_ok, turn1/2\\_exact\\_match |\n| `mixed_brief_json` | Hybrid: natural language answer + valid JSON block | answer\\_line\\_ok, json\\_parse\\_ok, schema\\_ok |\n| `toolcall` | Tool call embedded in response; parse + schema validation | stage1\\_tool\\_parse\\_ok, stage1\\_tool\\_schema\\_ok |\n| `toolcall_only` | Bare schema-only tool call; strict tool name + args check | tool\\_name\\_ok, args\\_ok |\n\nScores are conservative conjunctions — a case passes only when **all** gating signals succeed. This is intentional: partial success in a deployed agent is often indistinguishable from failure at the system level.\n\n**Evaluation hardware:** NVIDIA RTX 4090 (24 GB VRAM)\n**Evaluation date:** February 9, 2026\n**quant_eval seed:** 42\n\n---\n\n## About PBH Applied Systems\n\n[**PBH Applied Systems, LLC**](https://pbhappliedsystems.com) is an Oklahoma City–based applied machine learning and AI systems company specializing in production-grade model evaluation, quantization pipelines, agentic AI infrastructure, and scalable AI-driven application development. The organization operates with a strong emphasis on engineering rigor, reproducibility, and real-world deployment constraints — particularly in environments where performance, cost efficiency, and reliability must be balanced against available hardware and budget.\n\n### Founder — Patrick Hill, M.S.\n\nPBH Applied Systems was founded by **Patrick Hill**, a Data Scientist and AI/ML Engineer with 10+ years of experience delivering advanced analytics, predictive modeling, and decision-support solutions across high-stakes operational environments. Patrick holds a **Master of Science in Software Engineering with concentrations in Artificial Intelligence and Machine Learning** (GPA: 4.0) and a B.S. in Business Finance.\n\n**Technical expertise spans:**\n\n- **Languages & Data:** Python, SQL, Linux, Pandas, NumPy, scikit-learn\n- **ML & Modeling:** Supervised and unsupervised learning, neural networks, NLP, transformers, regression, classification, forecasting, and feature engineering\n- **AI/ML Frameworks:** PyTorch, TensorFlow/Keras, HuggingFace Transformers, GGUF, llama.cpp, BitsAndBytes, PEFT, QLoRA\n- **Deployment & MLOps:** Flask APIs, Docker, CI/CD pipelines, REST endpoints, streaming inference, version control\n- **Data Platforms:** Jupyter, Databricks, Power BI, Matplotlib\n- **Quantization:** GGUF conversion, Q4\\_K\\_M / Q5\\_K\\_M / Q8\\_0 strategies, adapter-per-model evaluation architecture\n\n### Published Author\n\nPatrick is the author of **[Applied Machine Learning: Concepts, Tools, and Case Studies](https://a.co/d/05qat7Xz)** — a 1,200+ page practitioner-oriented textbook covering statistical modeling, supervised and unsupervised learning, neural networks, NLP, and real-world decision-support case studies. The text has been adopted as **required reading for CSC 373 – Machine Learning at the University of Advancing Technology**, and reflects the same philosophy applied across all PBH systems: prioritize practical correctness over theoretical novelty, favor interpretable and reliable solutions, and introduce complexity only when justified by data and deployment constraints.\n\n### Core Service Areas\n\n**1. LLM Optimization & Deployment**\nEnd-to-end conversion of full-weight HuggingFace models to production-ready GGUF format, with quantization strategies matched to target hardware and latency requirements. Custom-built llama.cpp pipelines with adapter-per-model architecture ensuring strict separation of concerns and universal cross-model compatibility.\n\n**2. AI Evaluation Frameworks**\nProprietary behavioral evaluation via quant_eval — multi-run, timestamped pipelines producing structured artifacts, SHA256-verified manifests, per-family pass rates, F16 vs. quantized delta analysis, and deployment-ready recommendations. Evaluation batteries cover structured JSON output, multi-step reasoning, tool-calling fidelity, MCQ benchmarking, and fuzz/regression testing.\n\n**3. Agentic AI Infrastructure**\nDesign and deployment of agent-oriented architectures using LlamaIndex ReAct agents, Flask orchestration layers, and serverless GPU inference. Full pipeline from model selection through quantization, evaluation, and production serving — including lead capture flows, budget controls, and API gateway integration.\n\n**4. Scalable AI Application Development**\nProduction-grade multimodal AI applications integrating quantized LLMs, Whisper (speech-to-text), and BLIP (vision) via modular Flask APIs with Dockerized deployment and streaming-style responses. Advanced time-series forecasting systems featuring custom lightweight attention mechanisms, ensemble meta-learning, Bayesian hyperparameter optimization with resource-aware OOM backoff, and FinBERT sentiment fusion for hybrid structured/unstructured data pipelines.\n\n**5. ML Pipeline Design & Analytics**\nEnd-to-end data and model pipelines engineered for decision-support and operational forecasting. Encompasses feature engineering, leak-free forward-chaining cross-validation, KPI dashboard development, and analytical governance procedures designed for reproducibility at scale. Proven track record of translating complex model outputs into actionable insights for senior stakeholders across large-scale operational datasets.\n\n**6. Model & Agent Cataloging**\nStructured model catalog publishing with reproducible artifacts, standardized reporting, and clear performance tradeoff documentation — enabling engineering teams to make informed deployment decisions without re-running evaluations from scratch.\n\n### Engineering Principles\n\n- **Reproducibility first** — Every run produces structured artifacts, versioned manifests, and comparable outputs\n- **Universality as a requirement** — Systems work across models without custom rewrites per deployment\n- **No silent behavior changes** — Evaluation logic, prompts, and workflows are locked and versioned\n- **GPU utilization is non-negotiable** — All pipelines are designed to fully leverage available hardware\n- **Separation of IP and operations** — Core intellectual property is maintained independently of client deliverables\n\n---\n\n## 📞 Work With PBH Applied Systems\n\nThe `toolcall_only` degradation documented in this card — **1.000 (F16) → 0.000 (Q4\\_K\\_M)** — is a representative example of what rigorous pre-deployment evaluation surfaces. This finding is invisible to perplexity scores, benchmark leaderboards, and casual manual testing. It only appears when you run the model against structured behavioral tasks under production-equivalent conditions.\n\n**If you are selecting, deploying, or building on quantized open-weight models, evaluation like this belongs in your deployment process.**\n\n👉 **[Book a Scoping Call](https://pbhappliedsystems.com)** — Discuss your model selection, quantization strategy, or deployment architecture directly with Patrick.\n\n👉 **[Request an Evaluation Report](https://pbhappliedsystems.com)** — A full quant_eval behavioral audit for your target model(s): per-family pass rates, F16 vs. quantized delta analysis, failure cluster diagnostics, and a deployment recommendation. Engagements from $2,500.\n\n### Connect\n\n| | |\n|---|---|\n| 🌐 **Website** | [pbhappliedsystems.com](https://pbhappliedsystems.com) |\n| 📧 **Email** | [patrick@pbhappliedsystems.com](mailto:patrick@pbhappliedsystems.com) |\n| 💼 **LinkedIn** | [PBH Applied Systems, LLC](https://www.linkedin.com/company/pbh-applied-systems-llc) |\n| ▶️ **YouTube** | [@pbhappliedsystems](https://www.youtube.com/@pbhappliedsystems) |\n| 📸 **Instagram** | [@pbhappliedsystems](https://www.instagram.com/pbhappliedsystems) |\n| 👍 **Facebook** | [pbhappliedsystems](https://www.facebook.com/pbhappliedsystems) |\n\n---\n\n## License\n\nThis GGUF repository inherits the license of the base model:\n**Apache 2.0** — [`mistralai/Ministral-3-14B-Instruct-2512`](https://huggingface.co/mistralai/Ministral-3-14B-Instruct-2512)\n\nThe quant_eval evaluation methodology, fixture set, and scoring framework are proprietary to PBH Applied Systems, LLC and are not included in this repository.\n\n---\n\n*GGUF conversion, quantization, and behavioral evaluation performed by [PBH Applied Systems, LLC](https://pbhappliedsystems.com) · quant_eval v7.21 · Run ID: `20260209_170235`*\n",
    "related_quantizations": []
  },
  "tags": [
    "gguf",
    "quantized",
    "q4_k_m",
    "mistral",
    "instruct",
    "llama-cpp",
    "agentic",
    "tool-calling",
    "structured-output",
    "pbh-applied-systems",
    "quant-eval",
    "en",
    "base_model:mistralai/Ministral-3-14B-Instruct-2512",
    "base_model:quantized:mistralai/Ministral-3-14B-Instruct-2512",
    "license:apache-2.0",
    "endpoints_compatible",
    "region:us",
    "conversational"
  ],
  "likes": 0,
  "downloads": 1030,
  "gated": false,
  "private": false,
  "last_modified": "2026-04-15T06:58:22.000Z",
  "created_at": "2026-04-11T06:33:53.000Z",
  "pipeline_tag": "",
  "library_name": ""
}
Source payload excerpt (from Hugging Face API)
{
  "_id": "69d9eb512f3fedd8e2b5f678",
  "id": "pbhappliedsystems/ministral-3-14b-instruct-2512-gguf-Q4-K-M",
  "modelId": "pbhappliedsystems/ministral-3-14b-instruct-2512-gguf-Q4-K-M",
  "sha": "8cd64684d68cab4cb8bce35c4ea7070f1e43f375",
  "createdAt": "2026-04-11T06:33:53.000Z",
  "lastModified": "2026-04-15T06:58:22.000Z",
  "author": "pbhappliedsystems",
  "downloads": 1030,
  "likes": 0,
  "gated": false,
  "private": false,
  "pipeline_tag": "",
  "library_name": "",
  "siblings_count": 7
}