pbhappliedsystems/mistral-nemo-instruct-2407-gguf-f16 overview
Converted and evaluated by PBH Applied Systems, LLC — Applied AI/ML Consulting · LLM Optimization & Deployment · Quantized AI Infrastructure 🔬 This repository is part of a production-oriented evaluation series. Every model published under pbhappliedsystems has been independently evaluated using quanteval v7.21 — a proprietary behavioral evaluation harness developed by PBH Applied Systems. Scores measure real agent-adjacent task performance across structured output, tool dispatch, multi-turn state retention, and multi-step planning families — not perplexity or benchmark leaderboard proxies. 📌 This is the full-precision F16 baseline repository. The evaluated Q4\K\M deployment variant is published at pbhappliedsystems/mistral-nemo-instruct-2407-gguf-Q4-K-M. That card documents the full F16 vs. Q4\K\M comparison — including the json\multistep degradation (0.600 → 0.400), a complete ms\hard\01 breakdown at Q4\K\M, and a tool\02 final\mismatch finding that does not appear at F16. ---
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| mistral-nemo-instruct-2407-gguf-F16.gguf | GGUF | F16 | 22.82 GB | Download |
Model Details Live
Metadata Inspector
Normalized metadata (stored in metadata_json)
{
"metadata": {},
"card_data": {
"language": [
"en",
"fr",
"de",
"es",
"it",
"pt",
"ru",
"zh",
"ja"
],
"license": "apache-2.0",
"base_model": "mistralai/Mistral-Nemo-Instruct-2407",
"tags": [
"gguf",
"f16",
"full-precision",
"mistral",
"nemo",
"instruct",
"llama-cpp",
"agentic",
"tool-calling",
"structured-output",
"multilingual",
"pbh-applied-systems",
"quant-eval",
"baseline"
],
"frontmatter": {
"language": [
"en",
"fr",
"de",
"es",
"it",
"pt",
"ru",
"zh",
"ja"
],
"license": "apache-2.0",
"base_model": "mistralai/Mistral-Nemo-Instruct-2407",
"tags": [
"gguf",
"f16",
"full-precision",
"mistral",
"nemo",
"instruct",
"llama-cpp",
"agentic",
"tool-calling",
"structured-output",
"multilingual",
"pbh-applied-systems",
"quant-eval",
"baseline"
]
},
"hero_image_url": "",
"summary": "**Converted and evaluated by PBH Applied Systems, LLC** — Applied AI/ML Consulting · LLM Optimization & Deployment · Quantized AI Infrastructure > 🔬 **This repository is part of a production-oriented evaluation series.** Every model published under pbhappliedsystems has been independently evaluated using **quant_eval v7.21** — a proprietary behavioral evaluation harness developed by PBH Applied Systems. Scores measure real agent-adjacent task performance across structured output, tool dispatch, multi-turn state retention, and multi-step planning families — not perplexity or benchmark leaderboard proxies. > 📌 **This is the full-precision F16 baseline repository.** The evaluated Q4\\_K\\_M deployment variant is published at pbhappliedsystems/mistral-nemo-instruct-2407-gguf-Q4-K-M. That card documents the full F16 vs. Q4\\_K\\_M comparison — including the json\\_multistep degradation (0.600 → 0.400), a complete ms\\_hard\\_01 breakdown at Q4\\_K\\_M, and a tool\\_02 final\\_mismatch finding that does not appear at F16. ---",
"quick_links": [],
"benchmark_table_html": "",
"readme_markdown": "---\nlanguage:\n - en\n - fr\n - de\n - es\n - it\n - pt\n - ru\n - zh\n - ja\nlicense: apache-2.0\nbase_model: mistralai/Mistral-Nemo-Instruct-2407\ntags:\n - gguf\n - f16\n - full-precision\n - mistral\n - nemo\n - instruct\n - llama-cpp\n - agentic\n - tool-calling\n - structured-output\n - multilingual\n - pbh-applied-systems\n - quant-eval\n - baseline\n---\n\n# Mistral-Nemo-Instruct-2407 · GGUF F16\n\n**Converted and evaluated by [PBH Applied Systems, LLC](https://pbhappliedsystems.com)**\n— Applied AI/ML Consulting · LLM Optimization & Deployment · Quantized AI Infrastructure\n\n> 🔬 **This repository is part of a production-oriented evaluation series.** Every model published under [`pbhappliedsystems`](https://huggingface.co/pbhappliedsystems) has been independently evaluated using **quant_eval v7.21** — a proprietary behavioral evaluation harness developed by PBH Applied Systems. Scores measure real agent-adjacent task performance across structured output, tool dispatch, multi-turn state retention, and multi-step planning families — not perplexity or benchmark leaderboard proxies.\n\n> 📌 **This is the full-precision F16 baseline repository.** The evaluated Q4\\_K\\_M deployment variant is published at [`pbhappliedsystems/mistral-nemo-instruct-2407-gguf-Q4-K-M`](https://huggingface.co/pbhappliedsystems/mistral-nemo-instruct-2407-gguf-Q4-K-M). That card documents the full F16 vs. Q4\\_K\\_M comparison — including the json\\_multistep degradation (0.600 → 0.400), a complete ms\\_hard\\_01 breakdown at Q4\\_K\\_M, and a tool\\_02 final\\_mismatch finding that does not appear at F16.\n\n---\n\n## Model Description\n\nThis repository contains the **full-precision F16 GGUF** of [`mistralai/Mistral-Nemo-Instruct-2407`](https://huggingface.co/mistralai/Mistral-Nemo-Instruct-2407), a 12-billion parameter instruction-tuned model developed by Mistral AI in collaboration with NVIDIA (July 2024 release). Mistral-Nemo features the Tekken tokenizer and a **128,000-token context window** — the largest context in the PBH Applied Systems evaluated series outside of the Qwen2.5-14B-Instruct-1M, which supports a 1 million token window.\n\nThe F16 format preserves all original float16 weights without quantization. In the PBH Applied Systems evaluation pipeline, this F16 run (`20260211_014604`) served as the baseline cache generation pass — producing the `full_weight_cache.json` used as the reference anchor for the subsequent Q4\\_K\\_M comparison run (`20260211_022944`).\n\nFor most production deployments, the Q4\\_K\\_M variant is the appropriate choice. The F16 is the right choice when maximum output fidelity, clean tool-execution pipelines, and the full 128K context window at maximum precision are required.\n\n### Key Characteristics\n\n- **Parameters:** 12B\n- **Format:** GGUF F16 (full precision)\n- **File size:** 24.5 GB\n- **SHA256:** `cc7b8c5c3f129ad32aee562017b5e56f1284ee7d92445be292c644a42b3c9556`\n- **Context window:** 128,000 tokens (Tekken tokenizer)\n- **Minimum VRAM (GPU inference):** ~26 GB\n- **Recommended GPU tier:** A100 40 GB · RTX 4090 (24 GB, with offload) · 2× A10G\n- **Inference speed (eval hardware):** avg **30.24 sec/case** on RTX 4090\n- **Multilingual:** English, French, German, Spanish, Italian, Portuguese, Russian, Chinese, Japanese\n\n> **Why 30.24 sec/case avg?** This is a 12B non-reasoning model at full F16 precision. The json family averaged **64.02 sec/case** due to multi-step token generation, with json\\_01 reaching **149.88 seconds** — a significant single-case outlier visible in both F16 and Q4\\_K\\_M runs on this model. The Q4\\_K\\_M variant averages **1.42 sec/case** (21.3× faster).\n\n---\n\n## PBH Applied Systems Evaluation — quant\\_eval v7.21\n\n> **Evaluation conducted by PBH Applied Systems, LLC using quant_eval v7.21**\n> Run ID: `20260211_014604` · Fixtures: `golden_oracle_fixtures_v7_21` (SHA256: `6d71a0b9147c...`) · Seed: 42\n> Hardware: NVIDIA RTX 4090 · Runner: `full_weight_transformers` (F16 only) · Total rows: 42\n\n**Note on aggregate scores:** The normalized aggregate dimensions (task completion, reasoning, coherence, instruction following) are computed from the combined F16 + Q4\\_K\\_M comparison run and are reported on the Q4\\_K\\_M card. This F16 card reports per-family pass rates from the `full_weight_transformers` runner.\n\n### Per-Family Pass Rates — F16 (`full_weight_transformers`)\n\n| Family | N | Pass Rate | Avg Secs | Min | Max | Notes |\n|---|---:|---:|---:|---:|---:|---|\n| json\\_multistep | 5 | 0.600 | 79.08 | 53.91 | 92.74 | ms\\_easy\\_02 + ms\\_hard\\_01 fail |\n| stateful\\_followup | 2 | **1.000** | 7.60 | 5.57 | 9.63 | Both turns exact match |\n| toolcall\\_only | 2 | **1.000**\\* | 9.54 | 9.31 | 9.76 | Gating passed; schema wrapper issue — see note |\n| mixed\\_brief\\_json | 2 | **1.000** | 10.59 | 10.39 | 10.79 | Answer line + JSON schema correct |\n| toolcall | 2 | **1.000** | 14.39 | 14.25 | 14.53 | Both cases bucket=11 — clean at F16 |\n| json | 4 | n/a | 64.02 | 33.87 | 149.88 | bucket\\_score avg = 10.000; json\\_01 outlier |\n| fuzz | 20 | n/a | 26.61 | 10.19 | 57.78 | bucket\\_score avg = 10.000 |\n| mcq | 5 | n/a | 0.46 | 0.45 | 0.48 | bucket\\_score avg = 0.600 — 2 failures |\n\n### json\\_multistep — Case-Level Breakdown\n\n| Case | Difficulty | Result | Secs | Failure Signals |\n|---|---|---|---:|---|\n| ms\\_easy\\_01 | Easy | ✅ PASS | 53.91 | — |\n| ms\\_easy\\_02 | Easy | ❌ FAIL | 76.46 | oracle\\_equiv\\_ok=0 |\n| ms\\_med\\_01 | Medium | ✅ PASS | 83.74 | — |\n| ms\\_med\\_02 | Medium | ✅ PASS | 88.53 | — |\n| ms\\_hard\\_01 | Hard | ❌ FAIL | 92.74 | checks\\_consistent\\_ok=0, oracle\\_equiv\\_ok=0 |\n\n**F16 passes both medium cases.** At Q4\\_K\\_M, ms\\_med\\_02 additionally fails on `checks_consistent_ok`, and ms\\_hard\\_01 degrades from a partial failure (cc+oracle) to a complete failure across all four gating signals. This is the direct capability argument for F16 in planning-adjacent workloads.\n\nThe json_multistep failures at F16 are scoped: ms\\_easy\\_02 fails only on oracle equivalence (the model plans incorrectly but is internally consistent), and ms\\_hard\\_01 fails on both consistency and oracle (a harder failure that tracks with the difficulty level). No schema or STOP semantics failures occur at F16.\n\n### toolcall — Fully Clean at F16\n\nBoth `tool_01` and `tool_02` pass at bucket=11 — the maximum score. Tool parse is valid, schema is valid, and the final answer matches expected output. This contrasts with the Q4\\_K\\_M variant where `tool_02` produces a `final_mismatch` (bucket=0) despite a valid tool dispatch.\n\n| Case | F16 bucket | F16 detail | Q4\\_K\\_M bucket | Q4\\_K\\_M detail |\n|---|---:|---|---:|---|\n| tool\\_01 | 11 | ok | 11 | ok |\n| tool\\_02 | **11** | **ok** | **0** | **final\\_mismatch** |\n\n**At F16, tool dispatch and post-tool answer accuracy are both reliable.** If your application depends on the model correctly processing tool outputs and reporting accurate results — not just calling the right tool — F16 is the safer choice.\n\n### ⚠️ toolcall\\_only — Schema Wrapper Non-Compliance (F16)\n\n`toolcall_only` passes gating at 1.000 (tool\\_name\\_ok=1, args\\_ok=1) but carries `schema_ok=0` on both cases with `detail=schema_error`. The model correctly identifies the tool and extracts valid arguments but wraps the output using `\"tool\"` as the outer key instead of the expected `\"tool_name\"`.\n\n| Signal | Rate | Tier |\n|---|---:|---|\n| tool\\_name\\_ok | 1.000 | Tier-1 (gating) ✅ |\n| args\\_ok | 1.000 | Tier-1 (gating) ✅ |\n| schema\\_ok | 0.000 | Non-gating (tracked) |\n\nThis is a schema discipline issue, not a capability failure. A one-line normalization step (`\"tool\"` → `\"tool_name\"`) in the response parser resolves it for strict schema enforcement environments. See the Q4\\_K\\_M companion card for the `normalize_tool_wrapper()` implementation pattern.\n\n### MCQ — A-Bias at F16\n\n`mcq_02` and `mcq_05` both fail with `wrong_choice got=A`. Both failures occur at 0.46 seconds — fast, confident, and wrong. This model defaults to option A when uncertain. The Q4\\_K\\_M variant adds `mcq_04` to the failure set (also `got=A`), but the A-bias is a model-level characteristic present at full precision, not a quantization artifact.\n\n| Case | Result | Detail |\n|---|---|---|\n| mcq\\_01 | ✅ PASS | ok |\n| mcq\\_02 | ❌ FAIL | wrong\\_choice got=A |\n| mcq\\_03 | ✅ PASS | ok |\n| mcq\\_04 | ✅ PASS | ok |\n| mcq\\_05 | ❌ FAIL | wrong\\_choice got=A |\n\n### Signal-Level Diagnostics (F16)\n\n#### json\\_multistep\n\n| Signal | Rate | Tier |\n|---|---:|---|\n| schema\\_ok | 1.000 | Tier-1 (gating) |\n| checks\\_consistent\\_ok | 0.800 | Tier-1 (gating) |\n| stop\\_semantics\\_ok | 1.000 | Tier-1 (gating) |\n| oracle\\_equiv\\_ok | 0.600 | Tier-1 (gating) |\n| final\\_consistent\\_ok | 0.000 | Tier-2 (tracked, non-gating) |\n| final\\_match\\_reported | 0.000 | Tier-2 (tracked, non-gating) |\n\n> `schema_ok=1.000` and `stop_semantics_ok=1.000` at F16 — both drop to 0.800 at Q4\\_K\\_M. These signal-level regressions are what drive the pass rate from 0.600 to 0.400 under quantization.\n\n#### stateful\\_followup\n\n| Signal | Rate |\n|---|---:|\n| turn1\\_parse\\_ok | 1.000 |\n| turn2\\_parse\\_ok | 1.000 |\n| turn1\\_exact\\_match | 1.000 |\n| turn2\\_exact\\_match | 1.000 |\n\n#### toolcall\\_only\n\n| Signal | Rate | Tier |\n|---|---:|---|\n| tool\\_name\\_ok | 1.000 | Tier-1 (gating) |\n| args\\_ok | 1.000 | Tier-1 (gating) |\n| schema\\_ok | 0.000 | Non-gating (tracked) |\n\n#### mixed\\_brief\\_json\n\n| Signal | Rate |\n|---|---:|\n| answer\\_line\\_ok | 1.000 |\n| json\\_parse\\_ok | 1.000 |\n| schema\\_ok | 1.000 |\n\n---\n\n## When to Deploy F16 vs. Q4\\_K\\_M\n\n| Criterion | F16 (this repo) | Q4\\_K\\_M |\n|---|---|---|\n| VRAM available | ~26 GB | ~10 GB |\n| Latency acceptable | ~30 sec/case avg | ~1.4 sec/case avg |\n| json\\_multistep pass rate | 0.600 | 0.400 |\n| ms\\_med\\_02 | ✅ PASS | ❌ FAIL |\n| ms\\_hard\\_01 | ❌ partial fail (cc+oracle) | ❌ complete fail (all 4) |\n| toolcall accuracy | ✅ Clean (both bucket=11) | ⚠️ tool\\_02 final\\_mismatch |\n| toolcall\\_only gating | ✅ 1.000 | ❌ 0.000 (args fail) |\n| Schema wrapper norm | Needed | Not applicable |\n| MCQ failures | 2/5 (mcq\\_02, mcq\\_05) | 3/5 (+mcq\\_04) |\n| 128K context at full precision | ✅ | Limited by VRAM |\n\n---\n\n## Recommended Use Cases — F16\n\n### ✅ Deploy with Confidence\n\n- **Tool-calling pipelines requiring answer accuracy** — Both `toolcall` cases pass at bucket=11 with no final\\_mismatch. F16 is the correct choice when tool execution results feed downstream computation that must be correct.\n- **Stateful multi-turn agents** — Perfect two-turn state retention (1.000) at 7.60 sec/case avg.\n- **Structured JSON outputs (single-step)** — `json` and `fuzz` both achieve bucket\\_score 10.000.\n- **Hybrid brief + JSON responses** — `mixed_brief_json` passes at 1.000.\n- **Medium-difficulty multi-step planning** — ms\\_med\\_01 and ms\\_med\\_02 both pass. F16 retains ms\\_med\\_02 that Q4\\_K\\_M loses.\n- **Long-document processing at full precision** — 128K context with F16 weights provides maximum fidelity for large-scale document Q&A, multi-document comparison, and long-form extraction.\n- **Multilingual structured tasks** — All 9 supported languages at maximum precision.\n- **Tool-only dispatch with schema normalization** — `toolcall_only` passes gating at 1.000 with a one-line wrapper key fix.\n\n### ⚠️ Use with Guardrails\n\n- **Hard multi-step planning** — ms\\_hard\\_01 fails at F16 too (cc+oracle), though less catastrophically than at Q4\\_K\\_M. Use with an external validator for hard planning tasks.\n- **MCQ with A-bias mitigation** — Two of five cases fail at F16 with wrong\\_choice got=A. Add chain-of-thought prompting or response validation for MCQ-style pipelines.\n\n### ❌ Not Recommended\n\n- **High-throughput pipelines** — At 30.24 sec/case average and up to 149.88 seconds on a single json case, F16 is not suitable for latency-sensitive or batch workloads.\n\n---\n\n## Hardware Requirements\n\n| Configuration | VRAM Required | Recommended GPU |\n|---|---|---|\n| F16 (this repo) · full GPU offload | ~26 GB | A100 40 GB · 2× A10G · RTX 4090 (partial) |\n| F16 · mixed CPU/GPU offload | 16–24 GB VRAM + 16 GB RAM | RTX 3090/4090 with `n_gpu_layers` tuning |\n| Q4\\_K\\_M (companion repo) · 8K context | ~10 GB | T4 16 GB · RTX 3080 |\n| Q4\\_K\\_M (companion repo) · 128K context | ~18 GB | A10G 24 GB · RTX 4090 |\n\n---\n\n## Usage\n\n### Installation\n\n```bash\npip install llama-cpp-python huggingface_hub\n```\n\nFor GPU acceleration (CUDA):\n\n```bash\nCMAKE_ARGS=\"-DGGML_CUDA=on\" pip install llama-cpp-python --force-reinstall --no-cache-dir\n```\n\n### Python — llama-cpp-python\n\n```python\nfrom huggingface_hub import hf_hub_download\nfrom llama_cpp import Llama\n\n# Note: 24.5 GB download — ensure sufficient disk space and ~26 GB VRAM\nmodel_path = hf_hub_download(\n repo_id=\"pbhappliedsystems/mistral-nemo-instruct-2407-gguf-F16\",\n filename=\"mistral-nemo-instruct-2407-gguf-F16.gguf\"\n)\n\nllm = Llama(\n model_path=model_path,\n n_ctx=32768, # Adjust to use case; supports up to 128K\n n_gpu_layers=-1, # -1 offloads all layers to GPU; reduce if VRAM < 26 GB\n verbose=False,\n)\n\nresponse = llm.create_chat_completion(\n messages=[\n {\n \"role\": \"system\",\n \"content\": \"You are a precise assistant. Follow instructions exactly and return structured outputs when requested.\"\n },\n {\n \"role\": \"user\",\n \"content\": \"Analyze the following document and return a JSON object with keys: summary, key_entities, sentiment, action_items.\"\n }\n ],\n temperature=0.3,\n max_tokens=1024,\n)\n\nprint(response[\"choices\"][0][\"message\"][\"content\"])\n```\n\nFor partial GPU offload when VRAM is between 16–24 GB:\n\n```python\nllm = Llama(\n model_path=model_path,\n n_ctx=16384,\n n_gpu_layers=20, # Tune based on available VRAM\n verbose=True, # Enable to monitor layer offload and memory usage\n)\n```\n\nFor tool-calling with schema normalization (addresses the `toolcall_only` wrapper issue):\n\n```python\nimport json, re\nfrom huggingface_hub import hf_hub_download\nfrom llama_cpp import Llama\n\nmodel_path = hf_hub_download(\n repo_id=\"pbhappliedsystems/mistral-nemo-instruct-2407-gguf-F16\",\n filename=\"mistral-nemo-instruct-2407-gguf-F16.gguf\"\n)\n\nllm = Llama(model_path=model_path, n_ctx=4096, n_gpu_layers=-1, verbose=False)\n\ndef normalize_tool_wrapper(raw: str) -> dict:\n \"\"\"\n Normalize F16 schema wrapper non-compliance.\n Maps non-standard 'tool' key -> 'tool_name' before validation.\n See quant_eval v7.21 toolcall_only: schema_ok=0 at F16 (gating passes).\n \"\"\"\n match = re.search(r'```(?:json)?\\s*([\\s\\S]*?)```', raw)\n payload = match.group(1).strip() if match else raw.strip()\n parsed = json.loads(payload)\n if \"tool\" in parsed and \"tool_name\" not in parsed:\n parsed[\"tool_name\"] = parsed.pop(\"tool\")\n assert \"tool_name\" in parsed and \"args\" in parsed\n return parsed\n\nresponse = llm.create_chat_completion(\n messages=[\n {\"role\": \"system\", \"content\": \"Respond only with a valid JSON tool call.\"},\n {\"role\": \"user\", \"content\": \"Add 5 and 10.\"}\n ],\n temperature=0.0,\n max_tokens=256,\n)\nresult = normalize_tool_wrapper(response[\"choices\"][0][\"message\"][\"content\"])\nprint(result)\n```\n\n### CLI — llama-cli\n\n```bash\nllama-cli \\\n --model mistral-nemo-instruct-2407-gguf-F16.gguf \\\n --chat-template mistral \\\n --system-prompt \"You are a precise assistant.\" \\\n --prompt \"Analyze the following and return a JSON object with keys: summary, risk_level, action_items.\" \\\n --n-predict 1024 \\\n --ctx-size 32768 \\\n --n-gpu-layers -1 \\\n --temp 0.3\n```\n\nFor server deployment:\n\n```bash\nllama-server \\\n --model mistral-nemo-instruct-2407-gguf-F16.gguf \\\n --chat-template mistral \\\n --ctx-size 32768 \\\n --n-gpu-layers -1 \\\n --port 8080 \\\n --host 0.0.0.0\n```\n\nQuery via the OpenAI-compatible API:\n\n```python\nfrom openai import OpenAI\n\nclient = OpenAI(base_url=\"http://localhost:8080/v1\", api_key=\"not-required\")\n\nresponse = client.chat.completions.create(\n model=\"mistral-nemo-instruct-2407-gguf-F16\",\n messages=[{\"role\": \"user\", \"content\": \"Your prompt here\"}],\n temperature=0.3,\n timeout=180, # Allow up to 3 minutes for long-context or complex json cases\n)\nprint(response.choices[0].message.content)\n```\n\n---\n\n## Artifact Provenance\n\n| Artifact | Format | Size | SHA256 |\n|---|---|---|---|\n| `mistral-nemo-instruct-2407-gguf-F16.gguf` | GGUF F16 | 24.5 GB | `cc7b8c5c3f129ad32aee562017b5e56f1284ee7d92445be292c644a42b3c9556` |\n| Q4\\_K\\_M *(companion repo)* | GGUF Q4\\_K\\_M | 7.48 GB | `5765024ff3361f6dc5b590b963b378bd2e87ac95eabe5823a08a3ad336b498c9` |\n\nThe F16 GGUF was converted from the `mistralai/Mistral-Nemo-Instruct-2407` HuggingFace snapshot using a custom-built llama.cpp conversion pipeline developed by PBH Applied Systems, without modification to model weights.\n\n**Two-pass evaluation architecture:** The F16 evaluation run (`20260211_014604`) operated in cache-generation mode (`skip_quant=true`), producing the `full_weight_cache.json` used as the reference baseline for the Q4\\_K\\_M comparison run (`20260211_022944`). This ensures that F16 and Q4\\_K\\_M results are measured against the identical fixture set under controlled, reproducible conditions.\n\n---\n\n## Evaluation Methodology\n\n**quant_eval v7.21** is a proprietary behavioral evaluation harness developed by PBH Applied Systems. The two-run architecture evaluates the full-precision (F16) model first, caches its results, then evaluates the quantized variant against the same fixture set.\n\n**Fixture set:** `golden_oracle_fixtures_v7_21` (SHA256: `6d71a0b9147c079371b02a94f3c149eb78a6adc03dc16ff6833b964fbf4174f0`)\n\n| Family | Description | Pass Signals |\n|---|---|---|\n| `fuzz` | Property-based regression; structured placement correctness | schema\\_ok, constraints\\_ok |\n| `json` | Single-step structured JSON with constraint rules | schema\\_ok, constraints\\_ok |\n| `json_multistep` | Multi-step planning with self-check and oracle verification | schema\\_ok, checks\\_consistent\\_ok, stop\\_semantics\\_ok, oracle\\_equiv\\_ok |\n| `mcq` | Multiple-choice extraction | choice\\_ok |\n| `stateful_followup` | Two-turn state tracking; turn-2 correct given turn-1 | turn1/2\\_parse\\_ok, turn1/2\\_exact\\_match |\n| `mixed_brief_json` | Hybrid: natural language answer + valid JSON block | answer\\_line\\_ok, json\\_parse\\_ok, schema\\_ok |\n| `toolcall` | Tool call embedded in response; parse + schema validation | stage1\\_tool\\_parse\\_ok, stage1\\_tool\\_schema\\_ok |\n| `toolcall_only` | Bare schema-only tool call; strict tool name + args check | tool\\_name\\_ok, args\\_ok |\n\n**Evaluation hardware:** NVIDIA RTX 4090 (24 GB VRAM)\n**F16 evaluation date:** February 11, 2026\n**quant_eval seed:** 42\n\n---\n\n## About PBH Applied Systems\n\n[**PBH Applied Systems, LLC**](https://pbhappliedsystems.com) is an Oklahoma City–based applied machine learning and AI systems company specializing in production-grade model evaluation, quantization pipelines, agentic AI infrastructure, and scalable AI-driven application development. The organization operates with a strong emphasis on engineering rigor, reproducibility, and real-world deployment constraints — particularly in environments where performance, cost efficiency, and reliability must be balanced against available hardware and budget.\n\n### Founder — Patrick Hill, M.S.\n\nPBH Applied Systems was founded by **Patrick Hill**, a Data Scientist and AI/ML Engineer with 10+ years of experience delivering advanced analytics, predictive modeling, and decision-support solutions across high-stakes operational environments. Patrick holds a **Master of Science in Software Engineering with concentrations in Artificial Intelligence and Machine Learning** (GPA: 4.0) and a B.S. in Business Finance.\n\n**Technical expertise spans:**\n\n- **Languages & Data:** Python, SQL, Linux, Pandas, NumPy, scikit-learn\n- **ML & Modeling:** Supervised and unsupervised learning, neural networks, NLP, transformers, regression, classification, forecasting, and feature engineering\n- **AI/ML Frameworks:** PyTorch, TensorFlow/Keras, HuggingFace Transformers, GGUF, llama.cpp, BitsAndBytes, PEFT, QLoRA\n- **Deployment & MLOps:** Flask APIs, Docker, CI/CD pipelines, REST endpoints, streaming inference, version control\n- **Data Platforms:** Jupyter, Databricks, Power BI, Matplotlib\n- **Quantization:** GGUF conversion, Q4\\_K\\_M / Q5\\_K\\_M / Q8\\_0 strategies, adapter-per-model evaluation architecture\n\n### Published Author\n\nPatrick is the author of **[Applied Machine Learning: Concepts, Tools, and Case Studies](https://a.co/d/05qat7Xz)** — a 1,200+ page practitioner-oriented textbook adopted as **required reading for CSC 373 – Machine Learning at the University of Advancing Technology**.\n\n### Core Service Areas\n\n**1. LLM Optimization & Deployment** — End-to-end GGUF conversion and quantization with custom llama.cpp pipelines and adapter-per-model architecture.\n\n**2. AI Evaluation Frameworks** — Proprietary behavioral evaluation via quant_eval: per-family pass rates, F16 vs. quantized delta analysis, failure cluster diagnostics, deployment recommendations.\n\n**3. Agentic AI Infrastructure** — LlamaIndex ReAct agents, Flask orchestration, serverless GPU inference, full pipeline from model selection to production serving.\n\n**4. Scalable AI Application Development** — Multimodal applications (quantized LLMs + Whisper + BLIP), Dockerized Flask APIs, advanced time-series forecasting with custom attention mechanisms, Bayesian hyperparameter optimization, and FinBERT sentiment fusion.\n\n**5. ML Pipeline Design & Analytics** — Feature engineering, forward-chaining cross-validation, KPI dashboards, analytical governance at scale.\n\n**6. Model & Agent Cataloging** — Structured catalog publishing with reproducible artifacts and clear performance tradeoff documentation.\n\n---\n\n## 📞 Work With PBH Applied Systems\n\nThis F16 card documents clean tool execution that degrades at Q4\\_K\\_M, medium planning cases that survive F16 but fail under quantization, and an MCQ A-bias that is a model characteristic — not a precision artifact. The [Q4\\_K\\_M companion card](https://huggingface.co/pbhappliedsystems/mistral-nemo-instruct-2407-gguf-Q4-K-M) maps exactly where each of these findings changes under quantization. **The decision between F16 and Q4\\_K\\_M should be made with this data, not guessed at.**\n\n👉 **[Book a Scoping Call](https://pbhappliedsystems.com)** — Discuss your model selection, quantization strategy, or deployment architecture directly with Patrick.\n\n👉 **[Request an Evaluation Report](https://pbhappliedsystems.com)** — A full quant_eval behavioral audit: per-family pass rates, F16 vs. quantized delta analysis, failure cluster diagnostics, and a deployment recommendation. Engagements from $2,500.\n\n### Connect\n\n| | |\n|---|---|\n| 🌐 **Website** | [pbhappliedsystems.com](https://pbhappliedsystems.com) |\n| 📧 **Email** | [patrick@pbhappliedsystems.com](mailto:patrick@pbhappliedsystems.com) |\n| 💼 **LinkedIn** | [PBH Applied Systems, LLC](https://www.linkedin.com/company/pbh-applied-systems-llc) |\n| ▶️ **YouTube** | [@pbhappliedsystems](https://www.youtube.com/@pbhappliedsystems) |\n| 📸 **Instagram** | [@pbhappliedsystems](https://www.instagram.com/pbhappliedsystems) |\n| 👍 **Facebook** | [pbhappliedsystems](https://www.facebook.com/pbhappliedsystems) |\n\n---\n\n## License\n\nThis GGUF repository inherits the license of the base model:\n**Apache 2.0** — [`mistralai/Mistral-Nemo-Instruct-2407`](https://huggingface.co/mistralai/Mistral-Nemo-Instruct-2407)\n\nThe quant_eval evaluation methodology, fixture set, and scoring framework are proprietary to PBH Applied Systems, LLC and are not included in this repository.\n\n---\n\n*GGUF conversion and behavioral evaluation performed by [PBH Applied Systems, LLC](https://pbhappliedsystems.com) · quant_eval v7.21 · F16 Run ID: `20260211_014604`*\n",
"related_quantizations": []
},
"tags": [
"nemo",
"gguf",
"f16",
"full-precision",
"mistral",
"instruct",
"llama-cpp",
"agentic",
"tool-calling",
"structured-output",
"multilingual",
"pbh-applied-systems",
"quant-eval",
"baseline",
"en",
"fr",
"de",
"es",
"it",
"pt",
"ru",
"zh",
"ja",
"base_model:mistralai/Mistral-Nemo-Instruct-2407",
"base_model:quantized:mistralai/Mistral-Nemo-Instruct-2407",
"license:apache-2.0",
"endpoints_compatible",
"region:us",
"conversational"
],
"likes": 0,
"downloads": 202,
"gated": false,
"private": false,
"last_modified": "2026-04-15T07:00:19.000Z",
"created_at": "2026-04-11T20:31:25.000Z",
"pipeline_tag": "",
"library_name": "nemo"
}
Source payload excerpt (from Hugging Face API)
{
"_id": "69daaf9dc75a8c6ca44de24a",
"id": "pbhappliedsystems/mistral-nemo-instruct-2407-gguf-F16",
"modelId": "pbhappliedsystems/mistral-nemo-instruct-2407-gguf-F16",
"sha": "739e08f0578b6a8bfcbd469b7e54c48b15dfc7cc",
"createdAt": "2026-04-11T20:31:25.000Z",
"lastModified": "2026-04-15T07:00:19.000Z",
"author": "pbhappliedsystems",
"downloads": 202,
"likes": 0,
"gated": false,
"private": false,
"pipeline_tag": "",
"library_name": "nemo",
"siblings_count": 7
}