GraySoft
Projects Models About FAQ Contact Download guIDE →
Model Intelligence Sheet

pbhappliedsystems/ministral-3-14b-instruct-2512-gguf-f16 overview

Converted and evaluated by PBH Applied Systems, LLC — Applied AI/ML Consulting · LLM Optimization & Deployment · Quantized AI Infrastructure 🔬 This repository is part of a production-oriented evaluation series. Every model published under pbhappliedsystems has been independently evaluated using quanteval v7.21 — a proprietary behavioral evaluation harness developed by PBH Applied Systems. Scores measure real agent-adjacent task performance across structured output, tool dispatch, multi-turn state retention, and multi-step planning families — not perplexity or benchmark leaderboard proxies. 📌 This is the full-precision F16 baseline repository. The evaluated and deployment-ready Q4\K\M variant is published at pbhappliedsystems/ministral-3-14b-instruct-2512-gguf-Q4-K-M. That card documents the full F16 vs. Q4\K\M comparison, including a complete quantization degradation finding on toolcallonly (1.000 → 0.000). ---

gguff16full-precisionmistralinstructllama-cppagentictool-callingstructured-outputpbh-applied-systemsquant-evalbaselineenbase_model:mistralai/Ministral-3-14B-Instruct-2512base_model:quantized:mistralai/Ministral-3-14B-Instruct-2512license:apache-2.0endpoints_compatibleregion:usconversational
pbhappliedsystems/ministral-3-14b-instruct-2512-gguf-f16 visual
Downloads
167
Likes
0
Pipeline
Library
Visibility
Public
Access
Open

Repository Files & Downloads

1 files detected
Direct downloads for all repository files
FileTypeQuantizationSizeLink
ministral-3-14b-instruct-2512-gguf-F16.gguf GGUF F16 25.17 GB Download

Model Details Live

Model Slug
pbhappliedsystems/ministral-3-14b-instruct-2512-gguf-f16
Author
pbhappliedsystems
Pipeline Task
Library
Created
2026-04-11
Last Modified
2026-04-15
Gated
No
Private
No
HF SHA
2876a290261b0cdd90f9fc1af90f08c1d12abd7f
License
apache-2.0
Language
en
Base Model
mistralai/Ministral-3-14B-Instruct-2512

Metadata Inspector

Normalized metadata (stored in metadata_json)
{
  "metadata": {},
  "card_data": {
    "language": [
      "en"
    ],
    "license": "apache-2.0",
    "base_model": "mistralai/Ministral-3-14B-Instruct-2512",
    "tags": [
      "gguf",
      "f16",
      "full-precision",
      "mistral",
      "instruct",
      "llama-cpp",
      "agentic",
      "tool-calling",
      "structured-output",
      "pbh-applied-systems",
      "quant-eval",
      "baseline"
    ],
    "frontmatter": {
      "language": [
        "en"
      ],
      "license": "apache-2.0",
      "base_model": "mistralai/Ministral-3-14B-Instruct-2512",
      "tags": [
        "gguf",
        "f16",
        "full-precision",
        "mistral",
        "instruct",
        "llama-cpp",
        "agentic",
        "tool-calling",
        "structured-output",
        "pbh-applied-systems",
        "quant-eval",
        "baseline"
      ]
    },
    "hero_image_url": "",
    "summary": "**Converted and evaluated by PBH Applied Systems, LLC** — Applied AI/ML Consulting · LLM Optimization & Deployment · Quantized AI Infrastructure > 🔬 **This repository is part of a production-oriented evaluation series.** Every model published under pbhappliedsystems has been independently evaluated using **quant_eval v7.21** — a proprietary behavioral evaluation harness developed by PBH Applied Systems. Scores measure real agent-adjacent task performance across structured output, tool dispatch, multi-turn state retention, and multi-step planning families — not perplexity or benchmark leaderboard proxies. > 📌 **This is the full-precision F16 baseline repository.** The evaluated and deployment-ready Q4\\_K\\_M variant is published at pbhappliedsystems/ministral-3-14b-instruct-2512-gguf-Q4-K-M. That card documents the full F16 vs. Q4\\_K\\_M comparison, including a complete quantization degradation finding on toolcall_only (1.000 → 0.000). ---",
    "quick_links": [],
    "benchmark_table_html": "",
    "readme_markdown": "---\nlanguage:\n  - en\nlicense: apache-2.0\nbase_model: mistralai/Ministral-3-14B-Instruct-2512\ntags:\n  - gguf\n  - f16\n  - full-precision\n  - mistral\n  - instruct\n  - llama-cpp\n  - agentic\n  - tool-calling\n  - structured-output\n  - pbh-applied-systems\n  - quant-eval\n  - baseline\n---\n\n# Ministral-3-14B-Instruct-2512 · GGUF F16\n\n**Converted and evaluated by [PBH Applied Systems, LLC](https://pbhappliedsystems.com)**\n— Applied AI/ML Consulting · LLM Optimization & Deployment · Quantized AI Infrastructure\n\n> 🔬 **This repository is part of a production-oriented evaluation series.** Every model published under [`pbhappliedsystems`](https://huggingface.co/pbhappliedsystems) has been independently evaluated using **quant_eval v7.21** — a proprietary behavioral evaluation harness developed by PBH Applied Systems. Scores measure real agent-adjacent task performance across structured output, tool dispatch, multi-turn state retention, and multi-step planning families — not perplexity or benchmark leaderboard proxies.\n\n> 📌 **This is the full-precision F16 baseline repository.** The evaluated and deployment-ready Q4\\_K\\_M variant is published at [`pbhappliedsystems/ministral-3-14b-instruct-2512-gguf-Q4-K-M`](https://huggingface.co/pbhappliedsystems/ministral-3-14b-instruct-2512-gguf-Q4-K-M). That card documents the full F16 vs. Q4\\_K\\_M comparison, including a complete quantization degradation finding on `toolcall_only` (1.000 → 0.000).\n\n---\n\n## Model Description\n\nThis repository contains the **full-precision F16 GGUF** of [`mistralai/Ministral-3-14B-Instruct-2512`](https://huggingface.co/mistralai/Ministral-3-14B-Instruct-2512), a 14-billion parameter instruction-tuned model from Mistral AI (December 2025 release).\n\nThe F16 format preserves the original float16 weights without quantization. It serves two purposes in the PBH Applied Systems evaluation pipeline: as the **reference baseline** against which Q4\\_K\\_M capability retention is measured, and as a high-fidelity inference option for deployments where VRAM is not a constraint and maximum output quality is required.\n\nFor most production deployments, the Q4\\_K\\_M variant is the appropriate choice. The F16 is the ground truth from which quantization tradeoffs are measured.\n\n### Key Characteristics\n\n- **Parameters:** 14B\n- **Format:** GGUF F16 (full precision)\n- **File size:** 27.0 GB\n- **SHA256:** `74ea113134173d29f8daba097457500e831eace3741de002846b3ab89781fd52`\n- **Minimum VRAM (GPU inference):** ~30 GB\n- **Recommended GPU tier:** A100 40 GB · RTX 4090 (24 GB, with offload) · 2× A10G\n- **Context window:** 32,768 tokens (per base model specification)\n- **Inference speed (eval hardware):** avg **6.37 sec/case** on RTX 4090\n\n> **Inference note:** F16 runs at 6.37 sec/case avg vs. 3.77 sec/case for Q4\\_K\\_M — approximately 69% slower. The json\\_multistep family in particular averaged **17.79 sec/case** at full precision due to multi-step generation length.\n\n---\n\n## PBH Applied Systems Evaluation — quant\\_eval v7.21\n\n> **Evaluation conducted by PBH Applied Systems, LLC using quant_eval v7.21**\n> Run ID: `20260206_213615` · Fixtures: `golden_oracle_fixtures_v7_21` (SHA256: `6d71a0b9147c...`) · Seed: 42\n> Hardware: NVIDIA RTX 4090 · Runner: `full_weight_transformers` · Total F16 rows: 42\n\n**Note on aggregate scores:** The normalized aggregate dimensions (task completion, reasoning, coherence, instruction following) are reported for the Q4\\_K\\_M variant only, as they are computed from the combined comparison run. F16 evaluation is reported at the per-family pass rate level, which is the authoritative signal for deployment decisions.\n\n### Per-Family Pass Rates — F16 (`full_weight_transformers`)\n\n| Family | N | Pass Rate | Avg Secs | Notes |\n|---|---:|---:|---:|---|\n| json\\_multistep | 5 | 0.600 | 17.79 | 3 pass, 2 fail — see case breakdown below |\n| stateful\\_followup | 2 | **1.000** | 3.08 | Both turns parse and match expected state |\n| toolcall\\_only | 2 | **1.000**\\* | 3.81 | Gating passed; schema wrapper non-compliance — see note |\n| mixed\\_brief\\_json | 2 | **1.000** | 2.42 | Answer line + JSON schema correct |\n| toolcall | 2 | **1.000** | 3.54 | Tool parse + schema valid |\n| json | 4 | n/a | 8.01 | bucket\\_score avg = 10.000 |\n| fuzz | 20 | n/a | 5.97 | bucket\\_score avg = 10.000 |\n| mcq | 5 | n/a | 0.32 | bucket\\_score avg = 1.000 |\n\n### json\\_multistep — Case-Level Breakdown\n\n| Case | Difficulty | Result | Failure Signal |\n|---|---|---|---|\n| ms\\_easy\\_01 | Easy | ✅ PASS | — |\n| ms\\_easy\\_02 | Easy | ❌ FAIL | oracle\\_equiv\\_ok=0 |\n| ms\\_med\\_01 | Medium | ✅ PASS | — |\n| ms\\_med\\_02 | Medium | ✅ PASS | — |\n| ms\\_hard\\_01 | Hard | ❌ FAIL | checks\\_consistent\\_ok=0 + oracle\\_equiv\\_ok=0 |\n\nThe model handles easy and medium planning cases reliably. `ms_easy_02` failure is a plan divergence (model chose `[A, B]` where oracle expected `[A, A]`). `ms_hard_01` failure involves an inconsistent intermediate check alongside an incorrect final placement — the harder the planning horizon, the less reliable the output without an external validator.\n\n### ⚠️ toolcall\\_only — Schema Wrapper Non-Compliance (F16)\n\n**Pass rate: 1.000 (gating) — but `schema_ok=0` on both cases.**\n\nThe F16 model correctly identifies the tool and extracts valid arguments (`tool_name_ok=1`, `args_ok=1`), satisfying the gating condition for a pass. However, both cases emit a non-standard outer wrapper key (`\"tool\"`) instead of the expected `\"tool_name\"`, triggering `schema_ok=0` and `detail=schema_error`.\n\nThis is a **schema discipline issue**, not a capability failure. The model knows what tool to call and what arguments to supply — it simply wraps them in a non-conformant key. This is qualitatively distinct from the Q4\\_K\\_M variant, which fails completely on `tool_name_ok=0` and `args_ok=0` simultaneously.\n\n| Signal | F16 Rate | Q4\\_K\\_M Rate | Interpretation |\n|---|---:|---:|---|\n| tool\\_name\\_ok | 1.000 | 0.000 | F16 identifies tool correctly |\n| args\\_ok | 1.000 | 0.000 | F16 extracts valid args |\n| schema\\_ok | 0.000 | 0.000 | Both fail outer wrapper schema |\n| **Gating pass rate** | **1.000** | **0.000** | F16 passes; Q4\\_K\\_M fails entirely |\n\n**Practical implication for F16 deployment:** A strict schema validation layer will catch the wrapper key mismatch. Add a normalization step that maps `\"tool\"` → `\"tool_name\"` in the response parser, or instruct-tune the system prompt to enforce the exact key format. The underlying capability is intact; the schema discipline requires enforcement.\n\n### Signal-Level Diagnostics (F16)\n\n#### json\\_multistep\n\n| Signal | Rate | Tier |\n|---|---:|---|\n| schema\\_ok | 1.000 | Tier-1 (gating) |\n| checks\\_consistent\\_ok | 0.800 | Tier-1 (gating) |\n| stop\\_semantics\\_ok | 1.000 | Tier-1 (gating) |\n| oracle\\_equiv\\_ok | 0.600 | Tier-1 (gating) |\n| final\\_consistent\\_ok | 0.000 | Tier-2 (tracked, non-gating) |\n| final\\_match\\_reported | 0.000 | Tier-2 (tracked, non-gating) |\n\n> **Note on Tier-2:** `final_consistent_ok` and `final_match_reported` are not gating signals. Most deployed agentic systems compute and validate state externally. Tier-1 oracle equivalence (0.600) is the production-relevant signal.\n\n#### stateful\\_followup\n\n| Signal | Rate | Tier |\n|---|---:|---|\n| turn1\\_parse\\_ok | 1.000 | Tier-1 |\n| turn2\\_parse\\_ok | 1.000 | Tier-1 |\n| turn1\\_exact\\_match | 1.000 | Tier-1 |\n| turn2\\_exact\\_match | 1.000 | Tier-1 |\n\n#### toolcall\\_only\n\n| Signal | Rate | Tier |\n|---|---:|---|\n| tool\\_name\\_ok | 1.000 | Tier-1 (gating) |\n| args\\_ok | 1.000 | Tier-1 (gating) |\n| schema\\_ok | 0.000 | Non-gating (tracked) |\n\n#### mixed\\_brief\\_json\n\n| Signal | Rate | Tier |\n|---|---:|---|\n| answer\\_line\\_ok | 1.000 | Tier-1 |\n| json\\_parse\\_ok | 1.000 | Tier-1 |\n| schema\\_ok | 1.000 | Tier-1 |\n\n---\n\n## Recommended Use Cases — F16\n\n### ✅ Deploy with Confidence (F16)\n\n- **Stateful multi-turn agents** — Perfect two-turn state retention (1.000). Both turns parse and match expected state exactly.\n- **Structured JSON outputs (single-step)** — bucket\\_score avg of 10.000 on both `json` and `fuzz`; consistently valid structured outputs.\n- **Hybrid brief + JSON responses** — `mixed_brief_json` passes at 1.000.\n- **Tool-calling with response scaffolding** — `toolcall` passes at 1.000. Tool call embedded in a broader response is fully reliable.\n- **Tool-only dispatch with schema normalization** — `toolcall_only` passes gating at 1.000. Add a wrapper key normalization step for strict schema compliance.\n- **JSON multi-step with external validation loop** — 0.600 pass rate; workable with an external planner or repair loop.\n\n### ⚠️ Use with Guardrails (F16)\n\n- **Strict bare tool-call dispatch** — Schema wrapper non-compliance (`schema_ok=0`) requires a normalization layer for systems enforcing exact JSON key format.\n- **Hard multi-step planning without validation** — `ms_hard_01` fails at both check consistency and oracle equivalence.\n\n### ❌ Not Recommended (F16)\n\n- **Unassisted multi-step planning** — Where planning correctness must hold without external verification or oracle validation, particularly at medium-to-hard difficulty.\n\n---\n\n## Hardware Requirements\n\n| Configuration | VRAM Required | Recommended GPU |\n|---|---|---|\n| F16 (this repo) · full GPU offload | ~30 GB | A100 40 GB · 2× A10G · RTX 4090 (partial offload) |\n| F16 · mixed CPU/GPU offload | 16–24 GB VRAM + 16 GB RAM | RTX 3090/4090 with `n_gpu_layers` tuning |\n| Q4\\_K\\_M (companion repo) | ~10–12 GB | T4 16 GB · RTX 3080/4080 · A10G |\n\nFor most production use cases, the Q4\\_K\\_M variant at ~10–12 GB VRAM and 3.77 sec/case is the appropriate deployment target. The F16 is recommended when maximum output fidelity is required and hardware constraints allow, or when using the model as the reference baseline in a quantization evaluation pipeline.\n\n---\n\n## Usage\n\n### Installation\n\n```bash\npip install llama-cpp-python huggingface_hub\n```\n\nFor GPU acceleration (CUDA):\n\n```bash\nCMAKE_ARGS=\"-DGGML_CUDA=on\" pip install llama-cpp-python --force-reinstall --no-cache-dir\n```\n\n### Python — llama-cpp-python\n\n```python\nfrom huggingface_hub import hf_hub_download\nfrom llama_cpp import Llama\n\n# Download F16 GGUF directly from HuggingFace Hub\n# Note: 27 GB download — ensure sufficient disk space and ~30 GB VRAM\nmodel_path = hf_hub_download(\n    repo_id=\"pbhappliedsystems/ministral-3-14b-instruct-2512-gguf-F16\",\n    filename=\"ministral-3-14b-instruct-2512-gguf-F16.gguf\"\n)\n\nllm = Llama(\n    model_path=model_path,\n    n_ctx=8192,          # context window; increase up to 32768 per model spec\n    n_gpu_layers=-1,     # -1 offloads all layers to GPU; reduce if VRAM < 30 GB\n    verbose=False,\n)\n\nresponse = llm.create_chat_completion(\n    messages=[\n        {\n            \"role\": \"system\",\n            \"content\": \"You are a helpful, concise assistant. Respond in structured JSON when asked.\"\n        },\n        {\n            \"role\": \"user\",\n            \"content\": \"Summarize the following contract clause and flag any obligations: ...\"\n        }\n    ],\n    temperature=0.15,\n    max_tokens=1024,\n)\n\nprint(response[\"choices\"][0][\"message\"][\"content\"])\n```\n\nFor partial GPU offload when VRAM is between 16–24 GB:\n\n```python\nllm = Llama(\n    model_path=model_path,\n    n_ctx=4096,\n    n_gpu_layers=20,   # Tune based on available VRAM; remainder runs on CPU\n    verbose=True,      # Enable to monitor layer offload and memory usage\n)\n```\n\nFor tool-calling with schema normalization (addressing the wrapper non-compliance noted above):\n\n```python\nimport json, re\nfrom huggingface_hub import hf_hub_download\nfrom llama_cpp import Llama\n\nmodel_path = hf_hub_download(\n    repo_id=\"pbhappliedsystems/ministral-3-14b-instruct-2512-gguf-F16\",\n    filename=\"ministral-3-14b-instruct-2512-gguf-F16.gguf\"\n)\n\nllm = Llama(model_path=model_path, n_ctx=4096, n_gpu_layers=-1, verbose=False)\n\ndef normalize_tool_wrapper(raw: str) -> dict:\n    \"\"\"\n    Normalize F16 schema wrapper non-compliance.\n    Maps non-standard 'tool' key -> 'tool_name' before validation.\n    See quant_eval v7.21 toolcall_only finding: schema_ok=0 on both F16 cases.\n    \"\"\"\n    # Extract JSON block from markdown fences if present\n    match = re.search(r'```(?:json)?\\s*([\\s\\S]*?)```', raw)\n    payload = match.group(1).strip() if match else raw.strip()\n    parsed = json.loads(payload)\n    # Normalize wrapper key\n    if \"tool\" in parsed and \"tool_name\" not in parsed:\n        parsed[\"tool_name\"] = parsed.pop(\"tool\")\n    assert \"tool_name\" in parsed and \"args\" in parsed\n    return parsed\n\nresponse = llm.create_chat_completion(\n    messages=[\n        {\"role\": \"system\", \"content\": \"Respond only with a valid JSON tool call.\"},\n        {\"role\": \"user\", \"content\": \"Add 5 and 10.\"}\n    ],\n    temperature=0.0,\n    max_tokens=256,\n)\nraw = response[\"choices\"][0][\"message\"][\"content\"]\nresult = normalize_tool_wrapper(raw)\nprint(result)\n```\n\n### CLI — llama-cli\n\n```bash\n# One-shot prompt (ensure sufficient VRAM before running)\nllama-cli \\\n  --model ministral-3-14b-instruct-2512-gguf-F16.gguf \\\n  --chat-template mistral \\\n  --system-prompt \"You are a helpful assistant.\" \\\n  --prompt \"Summarize the following and return a JSON object with keys: summary, risk_level, action_items.\" \\\n  --n-predict 512 \\\n  --ctx-size 8192 \\\n  --n-gpu-layers -1 \\\n  --temp 0.15\n```\n\nFor server deployment (OpenAI-compatible endpoint):\n\n```bash\nllama-server \\\n  --model ministral-3-14b-instruct-2512-gguf-F16.gguf \\\n  --chat-template mistral \\\n  --ctx-size 8192 \\\n  --n-gpu-layers -1 \\\n  --port 8080 \\\n  --host 0.0.0.0\n```\n\nQuery via the OpenAI-compatible API:\n\n```python\nfrom openai import OpenAI\n\nclient = OpenAI(base_url=\"http://localhost:8080/v1\", api_key=\"not-required\")\n\nresponse = client.chat.completions.create(\n    model=\"ministral-3-14b-instruct-2512-gguf-F16\",\n    messages=[{\"role\": \"user\", \"content\": \"Your prompt here\"}],\n    temperature=0.15,\n)\nprint(response.choices[0].message.content)\n```\n\n---\n\n## Artifact Provenance\n\n| Artifact | Format | Size | SHA256 |\n|---|---|---|---|\n| `ministral-3-14b-instruct-2512-gguf-F16.gguf` | GGUF F16 | 27.0 GB | `74ea113134173d29f8daba097457500e831eace3741de002846b3ab89781fd52` |\n| Q4\\_K\\_M *(companion repo)* | GGUF Q4\\_K\\_M | 8.24 GB | `a23910514ee512aa28db8dddd390c26a73b9c318dcdec374ae02d722d9658749` |\n\nBoth artifacts were produced from `mistralai/Ministral-3-14B-Instruct-2512` using a custom-built llama.cpp conversion pipeline developed by PBH Applied Systems. Conversion was performed on the full HuggingFace snapshot without modification to model weights prior to conversion.\n\n---\n\n## Evaluation Methodology\n\n**quant_eval v7.21** is a proprietary behavioral evaluation harness developed by PBH Applied Systems. The F16 evaluation run (`20260206_213615`) produces the `full_weight_cache.json` used as the reference baseline in the subsequent Q4\\_K\\_M comparison run (`20260209_170235`). This two-run architecture — F16 first, Q4\\_K\\_M second against the cached F16 results — enables exact apples-to-apples comparison of capability retention across quantization levels on an identical fixture set.\n\n**Fixture set:** `golden_oracle_fixtures_v7_21` (SHA256: `6d71a0b9147c079371b02a94f3c149eb78a6adc03dc16ff6833b964fbf4174f0`)\n\n**Task families evaluated:**\n\n| Family | Description | Pass Signals |\n|---|---|---|\n| `fuzz` | Property-based regression; structured placement correctness | schema\\_ok, constraints\\_ok |\n| `json` | Single-step structured JSON with constraint rules | schema\\_ok, constraints\\_ok |\n| `json_multistep` | Multi-step planning with self-check and oracle verification | schema\\_ok, checks\\_consistent\\_ok, stop\\_semantics\\_ok, oracle\\_equiv\\_ok |\n| `mcq` | Multiple-choice extraction | choice\\_ok |\n| `stateful_followup` | Two-turn state tracking; turn-2 correct given turn-1 | turn1/2\\_parse\\_ok, turn1/2\\_exact\\_match |\n| `mixed_brief_json` | Hybrid: natural language answer + valid JSON block | answer\\_line\\_ok, json\\_parse\\_ok, schema\\_ok |\n| `toolcall` | Tool call embedded in response; parse + schema validation | stage1\\_tool\\_parse\\_ok, stage1\\_tool\\_schema\\_ok |\n| `toolcall_only` | Bare schema-only tool call; strict tool name + args check | tool\\_name\\_ok, args\\_ok |\n\nScores are conservative conjunctions — a case passes only when **all** gating signals succeed.\n\n**Evaluation hardware:** NVIDIA RTX 4090 (24 GB VRAM)\n**F16 evaluation date:** February 6, 2026\n**quant_eval seed:** 42\n\n---\n\n## About PBH Applied Systems\n\n[**PBH Applied Systems, LLC**](https://pbhappliedsystems.com) is an Oklahoma City–based applied machine learning and AI systems company specializing in production-grade model evaluation, quantization pipelines, agentic AI infrastructure, and scalable AI-driven application development. The organization operates with a strong emphasis on engineering rigor, reproducibility, and real-world deployment constraints — particularly in environments where performance, cost efficiency, and reliability must be balanced against available hardware and budget.\n\n### Founder — Patrick Hill, M.S.\n\nPBH Applied Systems was founded by **Patrick Hill**, a Data Scientist and AI/ML Engineer with 10+ years of experience delivering advanced analytics, predictive modeling, and decision-support solutions across high-stakes operational environments. Patrick holds a **Master of Science in Software Engineering with concentrations in Artificial Intelligence and Machine Learning** (GPA: 4.0) and a B.S. in Business Finance.\n\n**Technical expertise spans:**\n\n- **Languages & Data:** Python, SQL, Linux, Pandas, NumPy, scikit-learn\n- **ML & Modeling:** Supervised and unsupervised learning, neural networks, NLP, transformers, regression, classification, forecasting, and feature engineering\n- **AI/ML Frameworks:** PyTorch, TensorFlow/Keras, HuggingFace Transformers, GGUF, llama.cpp, BitsAndBytes, PEFT, QLoRA\n- **Deployment & MLOps:** Flask APIs, Docker, CI/CD pipelines, REST endpoints, streaming inference, version control\n- **Data Platforms:** Jupyter, Databricks, Power BI, Matplotlib\n- **Quantization:** GGUF conversion, Q4\\_K\\_M / Q5\\_K\\_M / Q8\\_0 strategies, adapter-per-model evaluation architecture\n\n### Published Author\n\nPatrick is the author of **[Applied Machine Learning: Concepts, Tools, and Case Studies](https://a.co/d/05qat7Xz)** — a 1,200+ page practitioner-oriented textbook covering statistical modeling, supervised and unsupervised learning, neural networks, NLP, and real-world decision-support case studies. The text has been adopted as **required reading for CSC 373 – Machine Learning at the University of Advancing Technology**, and reflects the same philosophy applied across all PBH systems: prioritize practical correctness over theoretical novelty, favor interpretable and reliable solutions, and introduce complexity only when justified by data and deployment constraints.\n\n### Core Service Areas\n\n**1. LLM Optimization & Deployment**\nEnd-to-end conversion of full-weight HuggingFace models to production-ready GGUF format, with quantization strategies matched to target hardware and latency requirements. Custom-built llama.cpp pipelines with adapter-per-model architecture ensuring strict separation of concerns and universal cross-model compatibility.\n\n**2. AI Evaluation Frameworks**\nProprietary behavioral evaluation via quant_eval — multi-run, timestamped pipelines producing structured artifacts, SHA256-verified manifests, per-family pass rates, F16 vs. quantized delta analysis, and deployment-ready recommendations. Evaluation batteries cover structured JSON output, multi-step reasoning, tool-calling fidelity, MCQ benchmarking, and fuzz/regression testing.\n\n**3. Agentic AI Infrastructure**\nDesign and deployment of agent-oriented architectures using LlamaIndex ReAct agents, Flask orchestration layers, and serverless GPU inference. Full pipeline from model selection through quantization, evaluation, and production serving — including lead capture flows, budget controls, and API gateway integration.\n\n**4. Scalable AI Application Development**\nProduction-grade multimodal AI applications integrating quantized LLMs, Whisper (speech-to-text), and BLIP (vision) via modular Flask APIs with Dockerized deployment and streaming-style responses. Advanced time-series forecasting systems featuring custom lightweight attention mechanisms, ensemble meta-learning, Bayesian hyperparameter optimization with resource-aware OOM backoff, and FinBERT sentiment fusion for hybrid structured/unstructured data pipelines.\n\n**5. ML Pipeline Design & Analytics**\nEnd-to-end data and model pipelines engineered for decision-support and operational forecasting. Encompasses feature engineering, leak-free forward-chaining cross-validation, KPI dashboard development, and analytical governance procedures designed for reproducibility at scale. Proven track record of translating complex model outputs into actionable insights for senior stakeholders across large-scale operational datasets.\n\n**6. Model & Agent Cataloging**\nStructured model catalog publishing with reproducible artifacts, standardized reporting, and clear performance tradeoff documentation — enabling engineering teams to make informed deployment decisions without re-running evaluations from scratch.\n\n### Engineering Principles\n\n- **Reproducibility first** — Every run produces structured artifacts, versioned manifests, and comparable outputs\n- **Universality as a requirement** — Systems work across models without custom rewrites per deployment\n- **No silent behavior changes** — Evaluation logic, prompts, and workflows are locked and versioned\n- **GPU utilization is non-negotiable** — All pipelines are designed to fully leverage available hardware\n- **Separation of IP and operations** — Core intellectual property is maintained independently of client deliverables\n\n---\n\n## 📞 Work With PBH Applied Systems\n\nThis F16 card documents what the model can do at full precision. The [Q4\\_K\\_M companion card](https://huggingface.co/pbhappliedsystems/ministral-3-14b-instruct-2512-gguf-Q4-K-M) documents what degrades when you quantize — including a complete `toolcall_only` failure (1.000 → 0.000) that is invisible without running both evaluations. **Without evaluating both formats against the same fixture set before deployment, you are making a deployment decision without the data to support it.**\n\n👉 **[Book a Scoping Call](https://pbhappliedsystems.com)** — Discuss your model selection, quantization strategy, or deployment architecture directly with Patrick.\n\n👉 **[Request an Evaluation Report](https://pbhappliedsystems.com)** — A full quant_eval behavioral audit for your target model(s): per-family pass rates, F16 vs. quantized delta analysis, failure cluster diagnostics, and a deployment recommendation. Engagements from $2,500.\n\n### Connect\n\n| | |\n|---|---|\n| 🌐 **Website** | [pbhappliedsystems.com](https://pbhappliedsystems.com) |\n| 📧 **Email** | [patrick@pbhappliedsystems.com](mailto:patrick@pbhappliedsystems.com) |\n| 💼 **LinkedIn** | [PBH Applied Systems, LLC](https://www.linkedin.com/company/pbh-applied-systems-llc) |\n| ▶️ **YouTube** | [@pbhappliedsystems](https://www.youtube.com/@pbhappliedsystems) |\n| 📸 **Instagram** | [@pbhappliedsystems](https://www.instagram.com/pbhappliedsystems) |\n| 👍 **Facebook** | [pbhappliedsystems](https://www.facebook.com/pbhappliedsystems) |\n\n---\n\n## License\n\nThis GGUF repository inherits the license of the base model:\n**Apache 2.0** — [`mistralai/Ministral-3-14B-Instruct-2512`](https://huggingface.co/mistralai/Ministral-3-14B-Instruct-2512)\n\nThe quant_eval evaluation methodology, fixture set, and scoring framework are proprietary to PBH Applied Systems, LLC and are not included in this repository.\n\n---\n\n*GGUF conversion and behavioral evaluation performed by [PBH Applied Systems, LLC](https://pbhappliedsystems.com) · quant_eval v7.21 · F16 Run ID: `20260206_213615`*\n",
    "related_quantizations": []
  },
  "tags": [
    "gguf",
    "f16",
    "full-precision",
    "mistral",
    "instruct",
    "llama-cpp",
    "agentic",
    "tool-calling",
    "structured-output",
    "pbh-applied-systems",
    "quant-eval",
    "baseline",
    "en",
    "base_model:mistralai/Ministral-3-14B-Instruct-2512",
    "base_model:quantized:mistralai/Ministral-3-14B-Instruct-2512",
    "license:apache-2.0",
    "endpoints_compatible",
    "region:us",
    "conversational"
  ],
  "likes": 0,
  "downloads": 167,
  "gated": false,
  "private": false,
  "last_modified": "2026-04-15T06:57:42.000Z",
  "created_at": "2026-04-11T05:12:22.000Z",
  "pipeline_tag": "",
  "library_name": ""
}
Source payload excerpt (from Hugging Face API)
{
  "_id": "69d9d836ca95b5d9945bb0df",
  "id": "pbhappliedsystems/ministral-3-14b-instruct-2512-gguf-F16",
  "modelId": "pbhappliedsystems/ministral-3-14b-instruct-2512-gguf-F16",
  "sha": "2876a290261b0cdd90f9fc1af90f08c1d12abd7f",
  "createdAt": "2026-04-11T05:12:22.000Z",
  "lastModified": "2026-04-15T06:57:42.000Z",
  "author": "pbhappliedsystems",
  "downloads": 167,
  "likes": 0,
  "gated": false,
  "private": false,
  "pipeline_tag": "",
  "library_name": "",
  "siblings_count": 7
}