GraySoft
Projects Models About FAQ Contact Download guIDE →
Model Intelligence Sheet

pbhappliedsystems/ministral-3-14b-reasoning-2512-gguf-f16 overview

Converted and evaluated by PBH Applied Systems, LLC — Applied AI/ML Consulting · LLM Optimization & Deployment · Quantized AI Infrastructure 🔬 This repository is part of a production-oriented evaluation series. Every model published under pbhappliedsystems has been independently evaluated using quanteval v7.21 — a proprietary behavioral evaluation harness developed by PBH Applied Systems. Scores measure real agent-adjacent task performance across structured output, tool dispatch, multi-turn state retention, and multi-step planning families — not perplexity or benchmark leaderboard proxies. 📌 This is the full-precision F16 baseline repository. The evaluated Q4\K\M deployment variant is published at pbhappliedsystems/ministral-3-14b-reasoning-2512-gguf-Q4-K-M. That card documents the complete F16 vs. Q4\K\_M comparison — including the 55.7× inference speedup and the hard planning case capability regression — in full detail. 🧠 This is a reasoning-architecture model. Ministral-3-14B-Reasoning-2512 generates extended internal chain-of-thought traces before producing its final answer. At full F16 precision, this reasoning capability is fully expressed — and measurable. The evaluation data below documents exactly where it helps, where it creates extraction friction, and why it matters for deployment decisions. ---

gguff16full-precisionmistralreasoningchain-of-thoughtllama-cppagenticstructured-outputpbh-applied-systemsquant-evalbaselineenbase_model:mistralai/Ministral-3-14B-Reasoning-2512base_model:quantized:mistralai/Ministral-3-14B-Reasoning-2512license:apache-2.0endpoints_compatibleregion:usconversational
pbhappliedsystems/ministral-3-14b-reasoning-2512-gguf-f16 visual
Downloads
151
Likes
0
Pipeline
Library
Visibility
Public
Access
Open

Repository Files & Downloads

1 files detected
Direct downloads for all repository files
FileTypeQuantizationSizeLink
ministral-3-14b-reasoning-2512-gguf-F16.gguf GGUF F16 25.17 GB Download

Model Details Live

Model Slug
pbhappliedsystems/ministral-3-14b-reasoning-2512-gguf-f16
Author
pbhappliedsystems
Pipeline Task
Library
Created
2026-04-11
Last Modified
2026-04-15
Gated
No
Private
No
HF SHA
40cc4b4407e6ebb72f660fb1746a6c4944f42d3e
License
apache-2.0
Language
en
Base Model
mistralai/Ministral-3-14B-Reasoning-2512

Metadata Inspector

Normalized metadata (stored in metadata_json)
{
  "metadata": {},
  "card_data": {
    "language": [
      "en"
    ],
    "license": "apache-2.0",
    "base_model": "mistralai/Ministral-3-14B-Reasoning-2512",
    "tags": [
      "gguf",
      "f16",
      "full-precision",
      "mistral",
      "reasoning",
      "chain-of-thought",
      "llama-cpp",
      "agentic",
      "structured-output",
      "pbh-applied-systems",
      "quant-eval",
      "baseline"
    ],
    "frontmatter": {
      "language": [
        "en"
      ],
      "license": "apache-2.0",
      "base_model": "mistralai/Ministral-3-14B-Reasoning-2512",
      "tags": [
        "gguf",
        "f16",
        "full-precision",
        "mistral",
        "reasoning",
        "chain-of-thought",
        "llama-cpp",
        "agentic",
        "structured-output",
        "pbh-applied-systems",
        "quant-eval",
        "baseline"
      ]
    },
    "hero_image_url": "",
    "summary": "**Converted and evaluated by PBH Applied Systems, LLC** — Applied AI/ML Consulting · LLM Optimization & Deployment · Quantized AI Infrastructure > 🔬 **This repository is part of a production-oriented evaluation series.** Every model published under pbhappliedsystems has been independently evaluated using **quant_eval v7.21** — a proprietary behavioral evaluation harness developed by PBH Applied Systems. Scores measure real agent-adjacent task performance across structured output, tool dispatch, multi-turn state retention, and multi-step planning families — not perplexity or benchmark leaderboard proxies. > 📌 **This is the full-precision F16 baseline repository.** The evaluated Q4\\_K\\_M deployment variant is published at pbhappliedsystems/ministral-3-14b-reasoning-2512-gguf-Q4-K-M. That card documents the complete F16 vs. Q4\\_K\\_M comparison — including the **55.7× inference speedup** and the hard planning case capability regression — in full detail. > 🧠 **This is a reasoning-architecture model.** Ministral-3-14B-Reasoning-2512 generates extended internal chain-of-thought traces before producing its final answer. At full F16 precision, this reasoning capability is fully expressed — and measurable. The evaluation data below documents exactly where it helps, where it creates extraction friction, and why it matters for deployment decisions. ---",
    "quick_links": [],
    "benchmark_table_html": "",
    "readme_markdown": "---\nlanguage:\n  - en\nlicense: apache-2.0\nbase_model: mistralai/Ministral-3-14B-Reasoning-2512\ntags:\n  - gguf\n  - f16\n  - full-precision\n  - mistral\n  - reasoning\n  - chain-of-thought\n  - llama-cpp\n  - agentic\n  - structured-output\n  - pbh-applied-systems\n  - quant-eval\n  - baseline\n---\n\n# Ministral-3-14B-Reasoning-2512 · GGUF F16\n\n**Converted and evaluated by [PBH Applied Systems, LLC](https://pbhappliedsystems.com)**\n— Applied AI/ML Consulting · LLM Optimization & Deployment · Quantized AI Infrastructure\n\n> 🔬 **This repository is part of a production-oriented evaluation series.** Every model published under [`pbhappliedsystems`](https://huggingface.co/pbhappliedsystems) has been independently evaluated using **quant_eval v7.21** — a proprietary behavioral evaluation harness developed by PBH Applied Systems. Scores measure real agent-adjacent task performance across structured output, tool dispatch, multi-turn state retention, and multi-step planning families — not perplexity or benchmark leaderboard proxies.\n\n> 📌 **This is the full-precision F16 baseline repository.** The evaluated Q4\\_K\\_M deployment variant is published at [`pbhappliedsystems/ministral-3-14b-reasoning-2512-gguf-Q4-K-M`](https://huggingface.co/pbhappliedsystems/ministral-3-14b-reasoning-2512-gguf-Q4-K-M). That card documents the complete F16 vs. Q4\\_K\\_M comparison — including the **55.7× inference speedup** and the hard planning case capability regression — in full detail.\n\n> 🧠 **This is a reasoning-architecture model.** Ministral-3-14B-Reasoning-2512 generates extended internal chain-of-thought traces before producing its final answer. At full F16 precision, this reasoning capability is fully expressed — and measurable. The evaluation data below documents exactly where it helps, where it creates extraction friction, and why it matters for deployment decisions.\n\n---\n\n## Model Description\n\nThis repository contains the **full-precision F16 GGUF** of [`mistralai/Ministral-3-14B-Reasoning-2512`](https://huggingface.co/mistralai/Ministral-3-14B-Reasoning-2512), a 14-billion parameter reasoning-tuned model from Mistral AI (December 2025 release).\n\nThe F16 format preserves all original float16 weights without quantization, enabling the model's full chain-of-thought capability. In the PBH Applied Systems evaluation pipeline, this F16 run (`20260209_212510`) served as the baseline cache generation pass — producing the `full_weight_cache.json` used as the reference anchor for the subsequent Q4\\_K\\_M comparison run (`20260209_233252`). This two-pass architecture ensures that F16 and Q4\\_K\\_M results are measured against the identical fixture set under controlled, reproducible conditions.\n\nFor most production deployments, the Q4\\_K\\_M variant is the appropriate choice due to its dramatically lower VRAM requirement and 55.7× faster inference. The F16 is the right choice when the full chain-of-thought is required — for hard planning tasks, deliberative reasoning workflows, or any use case where the thinking trace itself is part of the deliverable.\n\n### Key Characteristics\n\n- **Parameters:** 14B\n- **Architecture:** Reasoning (extended chain-of-thought at full precision)\n- **Format:** GGUF F16 (full precision)\n- **File size:** 27.0 GB\n- **SHA256:** `7645d01deed3415326c7c2bf8b58280e234021a91b4b3ade52b4735c976ad221`\n- **Minimum VRAM (GPU inference):** ~30 GB\n- **Recommended GPU tier:** A100 40 GB · RTX 4090 (24 GB, with offload) · 2× A10G\n- **Context window:** 32,768 tokens (per base model specification)\n- **Inference speed (eval hardware):** avg **65.67 sec/case** on RTX 4090\n\n> **Why 65.67 sec/case?** This is the cost of full reasoning at F16. The model generates extended chain-of-thought traces before each response. The `json` family averaged **105.82 sec/case**; the fuzz suite peaked at **176.47 sec** on a single case. This is not a performance problem — it is the reasoning capability operating as designed at full precision. For comparison, the Q4\\_K\\_M variant averages **1.18 sec/case** with the chain-of-thought substantially compressed.\n\n---\n\n## PBH Applied Systems Evaluation — quant\\_eval v7.21\n\n> **Evaluation conducted by PBH Applied Systems, LLC using quant_eval v7.21**\n> Run ID: `20260209_212510` · Fixtures: `golden_oracle_fixtures_v7_21` (SHA256: `6d71a0b9147c...`) · Seed: 42\n> Hardware: NVIDIA RTX 4090 · Runner: `full_weight_transformers` (F16 only) · Total rows: 42\n\n**Note on aggregate scores:** The normalized aggregate dimensions (task completion, reasoning, coherence, instruction following) are computed from the combined F16 + Q4\\_K\\_M comparison run and are reported on the Q4\\_K\\_M card. This F16 card reports per-family pass rates from the `full_weight_transformers` runner — the authoritative signal for understanding full-precision capability.\n\n### Per-Family Pass Rates — F16 (`full_weight_transformers`)\n\n| Family | N | Pass Rate | Avg Secs | Min | Max | Notes |\n|---|---:|---:|---:|---:|---:|---|\n| json\\_multistep | 5 | 0.600 | 99.64 | 78.26 | 110.73 | 3 pass — 2× oracle\\_equiv\\_ok failure |\n| stateful\\_followup | 2 | **1.000** | 15.93 | 13.83 | 18.03 | Both turns exact match |\n| toolcall\\_only | 2 | 0.000 | 14.41 | 13.34 | 15.47 | Architectural — see note |\n| mixed\\_brief\\_json | 2 | **1.000** | 12.73 | 12.36 | 13.11 | Answer line + JSON schema correct |\n| toolcall | 2 | **1.000** | 22.90 | 22.87 | 22.92 | Tool parse + schema valid |\n| json | 4 | n/a | 105.82 | 104.77 | 107.02 | bucket\\_score avg = 10.000 |\n| fuzz | 20 | n/a | 84.22 | 34.74 | 176.47 | bucket\\_score avg = 10.000 |\n| mcq | 5 | n/a | 4.02 | 3.97 | 4.12 | bucket\\_score avg = 0.800 |\n\n### json\\_multistep — Case-Level Breakdown\n\n| Case | Difficulty | Result | Secs | Failure Signal |\n|---|---|---|---:|---|\n| ms\\_easy\\_01 | Easy | ❌ FAIL | 78.26 | oracle\\_equiv\\_ok=0 |\n| ms\\_easy\\_02 | Easy | ❌ FAIL | 92.18 | oracle\\_equiv\\_ok=0 |\n| ms\\_med\\_01 | Medium | ✅ PASS | 108.47 | — |\n| ms\\_med\\_02 | Medium | ✅ PASS | 108.54 | — |\n| ms\\_hard\\_01 | Hard | ✅ **PASS** | 110.73 | — |\n\n**The F16 passes the hard case.** `ms_hard_01` — the most demanding planning case in the fixture set — is solved correctly by the F16 model in 110.73 seconds of deliberation. The Q4\\_K\\_M variant fails this same case (checks\\_consistent\\_ok=0, oracle\\_equiv\\_ok=0) in 3.39 seconds. This is the clearest empirical argument for deploying F16 when planning correctness on harder cases is a requirement.\n\nThe easy case failures (`ms_easy_01`, `ms_easy_02`) are a pattern specific to the Reasoning architecture: the model's extended deliberation on simple cases sometimes leads it to over-engineer the solution, producing an oracle\\_equiv\\_ok failure where the Instruct variant succeeds without the reasoning overhead. The chain-of-thought is a double-edged signal — it helps on hard cases and occasionally over-complicates easy ones.\n\n### ⚠️ toolcall\\_only — Architectural Limitation (Both Runners)\n\n`toolcall_only` fails at 0.000 on the F16 runner, and also at 0.000 on the Q4\\_K\\_M runner. **This is not a quantization finding — it is an architectural characteristic of the Reasoning model** that applies at both precision levels.\n\n| Signal | F16 Rate | Q4\\_K\\_M Rate |\n|---|---:|---:|\n| tool\\_name\\_ok | 0.500 | 1.000 |\n| args\\_ok | 0.000 | 0.000 |\n\n**Case-level detail:**\n- `toolonly_01` (15.47s): tool\\_name\\_ok=1, args\\_ok=0, schema\\_ok=0 — tool is identified but args extraction fails amid reasoning output\n- `toolonly_02` (13.34s): tool\\_name\\_ok=0, args\\_ok=0, schema\\_ok=0 — the reasoning trace fully interferes with tool name extraction\n\nThe chain-of-thought output structure is incompatible with bare schema-only tool dispatch at full precision. The `toolcall` family (tool call embedded within a broader response) passes at 1.000 — the reasoning traces are absorbed naturally into the response structure when there is surrounding prose scaffolding.\n\n**Practical implication:** Use response-scaffolded tool calling at both precision levels. See the Q4\\_K\\_M card for the `call_tool_scaffolded()` implementation pattern.\n\n### MCQ — Backtick Wrapping Artifact\n\n`mcq_02` fails with `detail=invalid_choice raw='```'` at 3.97 seconds. The F16 model emits a markdown code fence delimiter as its extracted response — the reasoning trace wraps the answer in a code block before the choice letter, and the extractor captures the fence instead of the letter.\n\nAll other MCQ cases pass correctly at 3.97–4.12 sec each. This is a single-case formatting artifact, not a general MCQ capability failure. A post-processing step that strips markdown fences before choice extraction would resolve it.\n\n| Case | Secs | Result | Detail |\n|---|---:|---|---|\n| mcq\\_01 | 3.97 | ✅ PASS | ok |\n| mcq\\_02 | 3.97 | ❌ FAIL | `invalid_choice raw='```'` |\n| mcq\\_03 | 3.97 | ✅ PASS | ok |\n| mcq\\_04 | 4.12 | ✅ PASS | ok |\n| mcq\\_05 | 4.07 | ✅ PASS | ok |\n\n### Signal-Level Diagnostics (F16)\n\n#### json\\_multistep\n\n| Signal | Rate | Tier |\n|---|---:|---|\n| schema\\_ok | 1.000 | Tier-1 (gating) |\n| checks\\_consistent\\_ok | 1.000 | Tier-1 (gating) |\n| stop\\_semantics\\_ok | 1.000 | Tier-1 (gating) |\n| oracle\\_equiv\\_ok | 0.600 | Tier-1 (gating) |\n| final\\_consistent\\_ok | 0.000 | Tier-2 (tracked, non-gating) |\n| final\\_match\\_reported | 0.000 | Tier-2 (tracked, non-gating) |\n\n> **checks\\_consistent\\_ok = 1.000 at F16** is a meaningful signal. The Q4\\_K\\_M variant drops this to 0.800. At full precision, the model's internal reasoning maintains self-consistent intermediate check states across all five cases — it is only oracle equivalence (final placement correctness) that fails on the two easy cases.\n\n#### stateful\\_followup\n\n| Signal | Rate | Tier |\n|---|---:|---|\n| turn1\\_parse\\_ok | 1.000 | Tier-1 |\n| turn2\\_parse\\_ok | 1.000 | Tier-1 |\n| turn1\\_exact\\_match | 1.000 | Tier-1 |\n| turn2\\_exact\\_match | 1.000 | Tier-1 |\n\n#### toolcall\\_only\n\n| Signal | Rate | Tier |\n|---|---:|---|\n| tool\\_name\\_ok | 0.500 | Tier-1 (gating) |\n| args\\_ok | 0.000 | Tier-1 (gating) |\n\n#### mixed\\_brief\\_json\n\n| Signal | Rate | Tier |\n|---|---:|---|\n| answer\\_line\\_ok | 1.000 | Tier-1 |\n| json\\_parse\\_ok | 1.000 | Tier-1 |\n| schema\\_ok | 1.000 | Tier-1 |\n\n---\n\n## When to Deploy F16 vs. Q4\\_K\\_M\n\n| Criterion | F16 (this repo) | Q4\\_K\\_M |\n|---|---|---|\n| VRAM available | ~30 GB | ~10–12 GB |\n| Latency acceptable | > 60 sec/response | < 2 sec/response |\n| Hard planning correctness | ✅ ms\\_hard\\_01 passes | ❌ ms\\_hard\\_01 fails |\n| Chain-of-thought traces needed | ✅ Full deliberation | ❌ Compressed/suppressed |\n| Bare tool dispatch | ❌ Architectural limit | ❌ Architectural limit |\n| Scaffolded tool dispatch | ✅ 1.000 | ✅ 1.000 |\n| Stateful multi-turn | ✅ 1.000 | ✅ 1.000 |\n| Throughput priority | ❌ Low | ✅ High |\n| Deployment hardware | A100 · 2× A10G | T4 · RTX 3080/4080 |\n\n---\n\n## Recommended Use Cases — F16\n\n### ✅ Deploy with Confidence\n\n- **Hard multi-step planning** — The only variant in this pair that passes `ms_hard_01`. Use F16 when planning horizon difficulty is unpredictable or when hard cases must be handled correctly.\n- **Deliberative reasoning workflows** — Use cases where the chain-of-thought trace itself is the deliverable: audit trails, decision documentation, explainable AI pipelines.\n- **Stateful multi-turn agents** — Perfect two-turn state retention (1.000). The reasoning architecture maintains context across turns reliably at full precision.\n- **Hybrid brief + JSON outputs** — `mixed_brief_json` passes at 1.000. The reasoning traces do not interfere with mixed-format output at full precision.\n- **Scaffolded tool-calling** — `toolcall` passes at 1.000. Embed tool dispatch within a broader response structure.\n- **Structured single-step JSON** — `json` and `fuzz` both achieve bucket\\_score avg of 10.000, with perfect constraint adherence despite the 105-second deliberation cost.\n\n### ⚠️ Use with Extraction Normalization\n\n- **MCQ and single-choice outputs** — Add a post-processing step to strip markdown fences before choice extraction. One of five cases fails due to backtick wrapping; the underlying answer is present but embedded in a code block.\n- **Multi-step planning at easy difficulty** — The easy cases (ms\\_easy\\_01, ms\\_easy\\_02) fail at F16 due to reasoning over-engineering. Use with an external validator or prefer Q4\\_K\\_M for easy-difficulty tasks.\n\n### ❌ Not Recommended\n\n- **Bare tool-call dispatch (schema-only output)** — `toolcall_only` fails at 0.000. This is an architectural characteristic, not a precision issue. Use scaffolded tool calling instead.\n- **High-throughput pipelines** — At 65.67 sec/case average and up to 176 seconds for fuzz cases, F16 is not suitable for latency-sensitive or batch-at-scale workloads.\n\n---\n\n## Hardware Requirements\n\n| Configuration | VRAM Required | Recommended GPU |\n|---|---|---|\n| F16 (this repo) · full GPU offload | ~30 GB | A100 40 GB · 2× A10G · RTX 4090 (partial) |\n| F16 · mixed CPU/GPU offload | 16–24 GB VRAM + 16 GB RAM | RTX 3090/4090 with `n_gpu_layers` tuning |\n| Q4\\_K\\_M (companion repo) | ~10–12 GB | T4 16 GB · RTX 3080/4080 · A10G |\n\n---\n\n## Usage\n\n### Installation\n\n```bash\npip install llama-cpp-python huggingface_hub\n```\n\nFor GPU acceleration (CUDA):\n\n```bash\nCMAKE_ARGS=\"-DGGML_CUDA=on\" pip install llama-cpp-python --force-reinstall --no-cache-dir\n```\n\n### Python — llama-cpp-python\n\n```python\nfrom huggingface_hub import hf_hub_download\nfrom llama_cpp import Llama\n\n# Note: 27 GB download — ensure sufficient disk space and ~30 GB VRAM\nmodel_path = hf_hub_download(\n    repo_id=\"pbhappliedsystems/ministral-3-14b-reasoning-2512-gguf-F16\",\n    filename=\"ministral-3-14b-reasoning-2512-gguf-F16.gguf\"\n)\n\nllm = Llama(\n    model_path=model_path,\n    n_ctx=16384,      # Increase for long reasoning traces; up to 32768 per model spec\n    n_gpu_layers=-1,  # -1 offloads all layers to GPU; reduce if VRAM < 30 GB\n    verbose=False,\n)\n\nresponse = llm.create_chat_completion(\n    messages=[\n        {\n            \"role\": \"system\",\n            \"content\": (\n                \"You are a precise reasoning assistant. Think through each problem step-by-step \"\n                \"before giving your final answer.\"\n            )\n        },\n        {\n            \"role\": \"user\",\n            \"content\": \"Analyze the following contract for logical inconsistencies and conflicting obligations. Return a structured JSON summary with keys: inconsistencies, obligations, risk_level, recommendation.\"\n        }\n    ],\n    temperature=0.7,\n    max_tokens=4096,  # Reasoning traces at F16 can be long — allocate generously\n)\n\nprint(response[\"choices\"][0][\"message\"][\"content\"])\n```\n\nFor partial GPU offload when VRAM is between 16–24 GB:\n\n```python\nllm = Llama(\n    model_path=model_path,\n    n_ctx=8192,\n    n_gpu_layers=20,   # Tune based on VRAM; remainder offloads to CPU\n    verbose=True,      # Enable to monitor layer offload and memory usage\n)\n```\n\nFor MCQ with extraction normalization (addresses the backtick wrapping artifact noted above):\n\n```python\nimport re\nfrom huggingface_hub import hf_hub_download\nfrom llama_cpp import Llama\n\nmodel_path = hf_hub_download(\n    repo_id=\"pbhappliedsystems/ministral-3-14b-reasoning-2512-gguf-F16\",\n    filename=\"ministral-3-14b-reasoning-2512-gguf-F16.gguf\"\n)\n\nllm = Llama(model_path=model_path, n_ctx=4096, n_gpu_layers=-1, verbose=False)\n\ndef extract_mcq_choice(raw: str, valid_choices: list = [\"A\", \"B\", \"C\", \"D\"]) -> str:\n    \"\"\"\n    Extract MCQ choice from reasoning model output.\n    Strips markdown fences and reasoning traces before extraction.\n    Addresses quant_eval v7.21 finding: mcq_02 fails with invalid_choice raw='```'\n    \"\"\"\n    # Strip markdown fences\n    clean = re.sub(r'```[^\\n]*\\n?', '', raw).strip()\n    # Find last occurrence of a valid choice letter (reasoning may mention choices mid-trace)\n    for char in reversed(clean.split()):\n        candidate = char.strip('.,;:()\\'\"').upper()\n        if candidate in valid_choices:\n            return candidate\n    raise ValueError(f\"No valid choice found in: {raw[:200]}\")\n\nresponse = llm.create_chat_completion(\n    messages=[\n        {\"role\": \"system\", \"content\": \"Answer with only the letter of the correct choice.\"},\n        {\"role\": \"user\", \"content\": \"Which of the following is a primary color? A) Green B) Orange C) Blue D) Purple\"}\n    ],\n    temperature=0.0,\n    max_tokens=512,\n)\nraw = response[\"choices\"][0][\"message\"][\"content\"]\nchoice = extract_mcq_choice(raw)\nprint(f\"Extracted choice: {choice}\")\n```\n\n### CLI — llama-cli\n\n```bash\n# Deliberative reasoning prompt (expect 30–120 second response time at full precision)\nllama-cli \\\n  --model ministral-3-14b-reasoning-2512-gguf-F16.gguf \\\n  --chat-template mistral \\\n  --system-prompt \"You are a precise reasoning assistant. Think step-by-step before responding.\" \\\n  --prompt \"Analyze the following multi-step decision problem and return a JSON plan with checks at each step.\" \\\n  --n-predict 4096 \\\n  --ctx-size 16384 \\\n  --n-gpu-layers -1 \\\n  --temp 0.15\n```\n\nFor server deployment:\n\n```bash\nllama-server \\\n  --model ministral-3-14b-reasoning-2512-gguf-F16.gguf \\\n  --chat-template mistral \\\n  --ctx-size 16384 \\\n  --n-gpu-layers -1 \\\n  --port 8080 \\\n  --host 0.0.0.0\n```\n\nQuery via the OpenAI-compatible API:\n\n```python\nfrom openai import OpenAI\n\nclient = OpenAI(base_url=\"http://localhost:8080/v1\", api_key=\"not-required\")\n\nresponse = client.chat.completions.create(\n    model=\"ministral-3-14b-reasoning-2512-gguf-F16\",\n    messages=[{\"role\": \"user\", \"content\": \"Your prompt here\"}],\n    temperature=0.7,\n    timeout=300,  # F16 reasoning responses may take 60–180 seconds\n)\nprint(response.choices[0].message.content)\n```\n\n---\n\n## Artifact Provenance\n\n| Artifact | Format | Size | SHA256 |\n|---|---|---|---|\n| `ministral-3-14b-reasoning-2512-gguf-F16.gguf` | GGUF F16 | 27.0 GB | `7645d01deed3415326c7c2bf8b58280e234021a91b4b3ade52b4735c976ad221` |\n| Q4\\_K\\_M *(companion repo)* | GGUF Q4\\_K\\_M | 8.24 GB | `e7171d96748ddc948fd6d9edb3d1c6e3f9ba6b855ff964aee98519788da330c2` |\n\nThe F16 GGUF was converted from the `mistralai/Ministral-3-14B-Reasoning-2512` HuggingFace snapshot using a custom-built llama.cpp conversion pipeline developed by PBH Applied Systems, without modification to model weights prior to conversion.\n\n**Two-pass evaluation architecture:** The F16 evaluation run (`20260209_212510`) operated in cache-generation mode (`skip_quant=true`), producing the `full_weight_cache.json` that served as the reference baseline for the subsequent Q4\\_K\\_M comparison run (`20260209_233252`). This architecture ensures that F16 and Q4\\_K\\_M results are measured against the identical fixture set under reproducible, controlled conditions, and that the F16 results cached during the first pass are the exact same responses compared against in the second pass.\n\n---\n\n## Evaluation Methodology\n\n**quant_eval v7.21** is a proprietary behavioral evaluation harness developed by PBH Applied Systems. The two-run architecture evaluates the full-precision (F16) model first, caches its results, then evaluates the quantized variant against the same fixture set — enabling exact apples-to-apples comparison.\n\n**Fixture set:** `golden_oracle_fixtures_v7_21` (SHA256: `6d71a0b9147c079371b02a94f3c149eb78a6adc03dc16ff6833b964fbf4174f0`)\n\n| Family | Description | Pass Signals |\n|---|---|---|\n| `fuzz` | Property-based regression; structured placement correctness | schema\\_ok, constraints\\_ok |\n| `json` | Single-step structured JSON with constraint rules | schema\\_ok, constraints\\_ok |\n| `json_multistep` | Multi-step planning with self-check and oracle verification | schema\\_ok, checks\\_consistent\\_ok, stop\\_semantics\\_ok, oracle\\_equiv\\_ok |\n| `mcq` | Multiple-choice extraction | choice\\_ok |\n| `stateful_followup` | Two-turn state tracking; turn-2 correct given turn-1 | turn1/2\\_parse\\_ok, turn1/2\\_exact\\_match |\n| `mixed_brief_json` | Hybrid: natural language answer + valid JSON block | answer\\_line\\_ok, json\\_parse\\_ok, schema\\_ok |\n| `toolcall` | Tool call embedded in response; parse + schema validation | stage1\\_tool\\_parse\\_ok, stage1\\_tool\\_schema\\_ok |\n| `toolcall_only` | Bare schema-only tool call; strict tool name + args check | tool\\_name\\_ok, args\\_ok |\n\n**Evaluation hardware:** NVIDIA RTX 4090 (24 GB VRAM)\n**F16 evaluation date:** February 9, 2026\n**quant_eval seed:** 42\n\n---\n\n## About PBH Applied Systems\n\n[**PBH Applied Systems, LLC**](https://pbhappliedsystems.com) is an Oklahoma City–based applied machine learning and AI systems company specializing in production-grade model evaluation, quantization pipelines, agentic AI infrastructure, and scalable AI-driven application development. The organization operates with a strong emphasis on engineering rigor, reproducibility, and real-world deployment constraints — particularly in environments where performance, cost efficiency, and reliability must be balanced against available hardware and budget.\n\n### Founder — Patrick Hill, M.S.\n\nPBH Applied Systems was founded by **Patrick Hill**, a Data Scientist and AI/ML Engineer with 10+ years of experience delivering advanced analytics, predictive modeling, and decision-support solutions across high-stakes operational environments. Patrick holds a **Master of Science in Software Engineering with concentrations in Artificial Intelligence and Machine Learning** (GPA: 4.0) and a B.S. in Business Finance.\n\n**Technical expertise spans:**\n\n- **Languages & Data:** Python, SQL, Linux, Pandas, NumPy, scikit-learn\n- **ML & Modeling:** Supervised and unsupervised learning, neural networks, NLP, transformers, regression, classification, forecasting, and feature engineering\n- **AI/ML Frameworks:** PyTorch, TensorFlow/Keras, HuggingFace Transformers, GGUF, llama.cpp, BitsAndBytes, PEFT, QLoRA\n- **Deployment & MLOps:** Flask APIs, Docker, CI/CD pipelines, REST endpoints, streaming inference, version control\n- **Data Platforms:** Jupyter, Databricks, Power BI, Matplotlib\n- **Quantization:** GGUF conversion, Q4\\_K\\_M / Q5\\_K\\_M / Q8\\_0 strategies, adapter-per-model evaluation architecture\n\n### Published Author\n\nPatrick is the author of **[Applied Machine Learning: Concepts, Tools, and Case Studies](https://a.co/d/05qat7Xz)** — a 1,200+ page practitioner-oriented textbook adopted as **required reading for CSC 373 – Machine Learning at the University of Advancing Technology**.\n\n### Core Service Areas\n\n**1. LLM Optimization & Deployment** — End-to-end GGUF conversion and quantization with custom llama.cpp pipelines and adapter-per-model architecture.\n\n**2. AI Evaluation Frameworks** — Proprietary behavioral evaluation via quant_eval: per-family pass rates, F16 vs. quantized delta analysis, failure cluster diagnostics, deployment recommendations.\n\n**3. Agentic AI Infrastructure** — LlamaIndex ReAct agents, Flask orchestration, serverless GPU inference, full pipeline from model selection to production serving.\n\n**4. Scalable AI Application Development** — Multimodal applications (quantized LLMs + Whisper + BLIP), Dockerized Flask APIs, advanced time-series forecasting with custom attention mechanisms, Bayesian hyperparameter optimization, and FinBERT sentiment fusion.\n\n**5. ML Pipeline Design & Analytics** — Feature engineering, forward-chaining cross-validation, KPI dashboards, analytical governance at scale.\n\n**6. Model & Agent Cataloging** — Structured catalog publishing with reproducible artifacts and clear performance tradeoff documentation.\n\n---\n\n## 📞 Work With PBH Applied Systems\n\nThis F16 card documents the full reasoning capability of Ministral-3-14B-Reasoning-2512 at maximum precision. The [Q4\\_K\\_M companion card](https://huggingface.co/pbhappliedsystems/ministral-3-14b-reasoning-2512-gguf-Q4-K-M) shows what changes when you quantize — including which capabilities survive the 55.7× speedup and which do not. **The hard planning case that takes 110 seconds to solve correctly at F16 fails in 3 seconds at Q4\\_K\\_M. That difference is only visible when you evaluate both.**\n\n👉 **[Book a Scoping Call](https://pbhappliedsystems.com)** — Discuss your reasoning model deployment strategy, quantization tradeoffs, or agentic architecture directly with Patrick.\n\n👉 **[Request an Evaluation Report](https://pbhappliedsystems.com)** — A full quant_eval behavioral audit: per-family pass rates, F16 vs. quantized delta analysis, failure cluster diagnostics, and a deployment recommendation. Engagements from $2,500.\n\n### Connect\n\n| | |\n|---|---|\n| 🌐 **Website** | [pbhappliedsystems.com](https://pbhappliedsystems.com) |\n| 📧 **Email** | [patrick@pbhappliedsystems.com](mailto:patrick@pbhappliedsystems.com) |\n| 💼 **LinkedIn** | [PBH Applied Systems, LLC](https://www.linkedin.com/company/pbh-applied-systems-llc) |\n| ▶️ **YouTube** | [@pbhappliedsystems](https://www.youtube.com/@pbhappliedsystems) |\n| 📸 **Instagram** | [@pbhappliedsystems](https://www.instagram.com/pbhappliedsystems) |\n| 👍 **Facebook** | [pbhappliedsystems](https://www.facebook.com/pbhappliedsystems) |\n\n---\n\n## License\n\nThis GGUF repository inherits the license of the base model:\n**Apache 2.0** — [`mistralai/Ministral-3-14B-Reasoning-2512`](https://huggingface.co/mistralai/Ministral-3-14B-Reasoning-2512)\n\nThe quant_eval evaluation methodology, fixture set, and scoring framework are proprietary to PBH Applied Systems, LLC and are not included in this repository.\n\n---\n\n*GGUF conversion and behavioral evaluation performed by [PBH Applied Systems, LLC](https://pbhappliedsystems.com) · quant_eval v7.21 · F16 Run ID: `20260209_212510`*\n",
    "related_quantizations": []
  },
  "tags": [
    "gguf",
    "f16",
    "full-precision",
    "mistral",
    "reasoning",
    "chain-of-thought",
    "llama-cpp",
    "agentic",
    "structured-output",
    "pbh-applied-systems",
    "quant-eval",
    "baseline",
    "en",
    "base_model:mistralai/Ministral-3-14B-Reasoning-2512",
    "base_model:quantized:mistralai/Ministral-3-14B-Reasoning-2512",
    "license:apache-2.0",
    "endpoints_compatible",
    "region:us",
    "conversational"
  ],
  "likes": 0,
  "downloads": 151,
  "gated": false,
  "private": false,
  "last_modified": "2026-04-15T06:59:11.000Z",
  "created_at": "2026-04-11T07:01:13.000Z",
  "pipeline_tag": "",
  "library_name": ""
}
Source payload excerpt (from Hugging Face API)
{
  "_id": "69d9f1b9a20155247ffbffd0",
  "id": "pbhappliedsystems/ministral-3-14b-reasoning-2512-gguf-F16",
  "modelId": "pbhappliedsystems/ministral-3-14b-reasoning-2512-gguf-F16",
  "sha": "40cc4b4407e6ebb72f660fb1746a6c4944f42d3e",
  "createdAt": "2026-04-11T07:01:13.000Z",
  "lastModified": "2026-04-15T06:59:11.000Z",
  "author": "pbhappliedsystems",
  "downloads": 151,
  "likes": 0,
  "gated": false,
  "private": false,
  "pipeline_tag": "",
  "library_name": "",
  "siblings_count": 7
}