GraySoft
Projects Models About FAQ Contact Download guIDE →
Model Intelligence Sheet

pbhappliedsystems/mistral-nemo-instruct-2407-gguf-q4-k-m overview

Quantized, converted, and evaluated by PBH Applied Systems, LLC — Applied AI/ML Consulting · LLM Optimization & Deployment · Quantized AI Infrastructure 🔬 This repository is part of a production-oriented evaluation series. Every model published under pbhappliedsystems has been independently evaluated using quant_eval v7.21 — a proprietary behavioral evaluation harness developed by PBH Applied Systems. Scores measure real agent-adjacent task performance across structured output, tool dispatch, multi-turn state retention, and multi-step planning families — not perplexity or benchmark leaderboard proxies. ---

nemoggufquantizedq4_k_mmistralinstructllama-cppagentictool-callingstructured-outputmultilingualpbh-applied-systemsquant-evalenfrdeesitptruzhjabase_model:mistralai/Mistral-Nemo-Instruct-2407base_model:quantized:mistralai/Mistral-Nemo-Instruct-2407license:apache-2.0endpoints_compatibleregion:usconversational
pbhappliedsystems/mistral-nemo-instruct-2407-gguf-q4-k-m visual
Downloads
1,045
Likes
0
Pipeline
Library
nemo
Visibility
Public
Access
Open

Repository Files & Downloads

1 files detected
Direct downloads for all repository files
FileTypeQuantizationSizeLink
mistral-nemo-instruct-2407-gguf-Q4-K-M.gguf GGUF 6.96 GB Download

Model Details Live

Model Slug
pbhappliedsystems/mistral-nemo-instruct-2407-gguf-q4-k-m
Author
pbhappliedsystems
Pipeline Task
Library
nemo
Created
2026-04-11
Last Modified
2026-04-15
Gated
No
Private
No
HF SHA
eeda50bbc8110885c5e08ccbe1e0e14f83604164
License
apache-2.0
Language
en, fr, de, es, it, pt, ru, zh, ja
Base Model
mistralai/Mistral-Nemo-Instruct-2407

Metadata Inspector

Normalized metadata (stored in metadata_json)
{
  "metadata": {},
  "card_data": {
    "language": [
      "en",
      "fr",
      "de",
      "es",
      "it",
      "pt",
      "ru",
      "zh",
      "ja"
    ],
    "license": "apache-2.0",
    "base_model": "mistralai/Mistral-Nemo-Instruct-2407",
    "tags": [
      "gguf",
      "quantized",
      "q4_k_m",
      "mistral",
      "nemo",
      "instruct",
      "llama-cpp",
      "agentic",
      "tool-calling",
      "structured-output",
      "multilingual",
      "pbh-applied-systems",
      "quant-eval"
    ],
    "frontmatter": {
      "language": [
        "en",
        "fr",
        "de",
        "es",
        "it",
        "pt",
        "ru",
        "zh",
        "ja"
      ],
      "license": "apache-2.0",
      "base_model": "mistralai/Mistral-Nemo-Instruct-2407",
      "tags": [
        "gguf",
        "quantized",
        "q4_k_m",
        "mistral",
        "nemo",
        "instruct",
        "llama-cpp",
        "agentic",
        "tool-calling",
        "structured-output",
        "multilingual",
        "pbh-applied-systems",
        "quant-eval"
      ]
    },
    "hero_image_url": "",
    "summary": "**Quantized, converted, and evaluated by PBH Applied Systems, LLC** — Applied AI/ML Consulting · LLM Optimization & Deployment · Quantized AI Infrastructure > 🔬 **This repository is part of a production-oriented evaluation series.** Every model published under pbhappliedsystems has been independently evaluated using **quant_eval v7.21** — a proprietary behavioral evaluation harness developed by PBH Applied Systems. Scores measure real agent-adjacent task performance across structured output, tool dispatch, multi-turn state retention, and multi-step planning families — not perplexity or benchmark leaderboard proxies. ---",
    "quick_links": [],
    "benchmark_table_html": "",
    "readme_markdown": "---\nlanguage:\n  - en\n  - fr\n  - de\n  - es\n  - it\n  - pt\n  - ru\n  - zh\n  - ja\nlicense: apache-2.0\nbase_model: mistralai/Mistral-Nemo-Instruct-2407\ntags:\n  - gguf\n  - quantized\n  - q4_k_m\n  - mistral\n  - nemo\n  - instruct\n  - llama-cpp\n  - agentic\n  - tool-calling\n  - structured-output\n  - multilingual\n  - pbh-applied-systems\n  - quant-eval\n---\n\n# Mistral-Nemo-Instruct-2407 · GGUF Q4\\_K\\_M\n\n**Quantized, converted, and evaluated by [PBH Applied Systems, LLC](https://pbhappliedsystems.com)**\n— Applied AI/ML Consulting · LLM Optimization & Deployment · Quantized AI Infrastructure\n\n> 🔬 **This repository is part of a production-oriented evaluation series.** Every model published under [`pbhappliedsystems`](https://huggingface.co/pbhappliedsystems) has been independently evaluated using **quant_eval v7.21** — a proprietary behavioral evaluation harness developed by PBH Applied Systems. Scores measure real agent-adjacent task performance across structured output, tool dispatch, multi-turn state retention, and multi-step planning families — not perplexity or benchmark leaderboard proxies.\n\n---\n\n## Model Description\n\nThis repository contains the **4-bit quantized (Q4\\_K\\_M)** GGUF of [`mistralai/Mistral-Nemo-Instruct-2407`](https://huggingface.co/mistralai/Mistral-Nemo-Instruct-2407), a 12-billion parameter instruction-tuned model developed by Mistral AI in collaboration with NVIDIA (July 2024 release). Mistral-Nemo features the Tekken tokenizer and a **128,000-token context window** — the largest context in the evaluated series, outside of the Qwen2.5-14B-Instruct-1M, which supports a 1 million token window, in the PBH Applied Systems evaluated series.\n\nThe Q4\\_K\\_M format applies 4-bit quantization with K-quant medium precision. As documented in the evaluation section below, Q4\\_K\\_M quantization produces measurable degradation on multi-step planning tasks — including a complete breakdown on the hardest planning case — while preserving strong performance on stateful, structured-output, and hybrid response tasks.\n\nThe full-precision F16 baseline is published separately at [`pbhappliedsystems/mistral-nemo-instruct-2407-gguf-F16`](https://huggingface.co/pbhappliedsystems/mistral-nemo-instruct-2407-gguf-F16).\n\n### Key Characteristics\n\n- **Parameters:** 12B\n- **Format:** GGUF Q4\\_K\\_M\n- **File size:** 7.48 GB\n- **SHA256:** `5765024ff3361f6dc5b590b963b378bd2e87ac95eabe5823a08a3ad336b498c9`\n- **Context window:** 128,000 tokens (Tekken tokenizer)\n- **Minimum VRAM (GPU inference):** ~10 GB (T4 class or better)\n- **Recommended GPU tier:** NVIDIA T4 (16 GB) · RTX 3080/4080 · A10G\n- **Inference speed (eval hardware):** avg **1.42 sec/case** on RTX 4090\n- **Multilingual:** English, French, German, Spanish, Italian, Portuguese, Russian, Chinese, Japanese\n\n---\n\n## PBH Applied Systems Evaluation — quant\\_eval v7.21\n\n> **Evaluation conducted by PBH Applied Systems, LLC using quant_eval v7.21**\n> Run ID: `20260211_022944` · Fixtures: `golden_oracle_fixtures_v7_21` (SHA256: `6d71a0b9147c...`) · Seed: 42\n> Hardware: NVIDIA RTX 4090 · Total rows evaluated: 84 (42 F16 · 42 Q4\\_K\\_M)\n\n### Aggregate Scores (Q4\\_K\\_M)\n\nScores are normalized to [0.0 – 1.0]. Higher is better.\n\n| Dimension | Score |\n|---|---:|\n| Task Completion | 0.6631 |\n| Reasoning | 0.7870 |\n| Coherence | 0.8836 |\n| Instruction Following | 0.9329 |\n| **Avg inference time** | **1.42 sec/case** |\n\n### Per-Family Pass Rates\n\n#### F16 Baseline (`full_weight_transformers`)\n\n| Family | N | Pass Rate | Avg Secs | Notes |\n|---|---:|---:|---:|---|\n| json\\_multistep | 5 | 0.600 | 79.08 | ms\\_easy\\_02 + ms\\_hard\\_01 fail |\n| stateful\\_followup | 2 | **1.000** | 7.60 | Both turns exact match |\n| toolcall\\_only | 2 | **1.000**\\* | 9.54 | Gating passed; schema wrapper issue — see note |\n| mixed\\_brief\\_json | 2 | **1.000** | 10.59 | Answer line + JSON schema correct |\n| toolcall | 2 | **1.000** | 14.39 | Tool parse + schema valid |\n| json | 4 | n/a | 64.02 | bucket\\_score avg = 10.000 (json\\_01 = 149.88s outlier) |\n| fuzz | 20 | n/a | 26.61 | bucket\\_score avg = 10.000 |\n| mcq | 5 | n/a | 0.46 | bucket\\_score avg = 0.600 — 2 failures |\n\n#### Q4\\_K\\_M (`quantized_llama_cpp`)\n\n| Family | N | Pass Rate | Δ vs F16 | Avg Secs | Notes |\n|---|---:|---:|---:|---:|---|\n| json\\_multistep | 5 | **0.400** | **−0.200** | 4.52 | ⚠️ Hard case complete breakdown |\n| stateful\\_followup | 2 | **1.000** | 0.000 | 0.40 | Perfect retention |\n| toolcall\\_only | 2 | 0.000 | −1.000 | 0.46 | tool\\_name\\_ok=1, args\\_ok=0 |\n| mixed\\_brief\\_json | 2 | **1.000** | 0.000 | 0.52 | No degradation |\n| toolcall | 2 | **1.000** | 0.000 | 0.62 | Pass rate holds; final\\_mismatch on tool\\_02 — see note |\n| json | 4 | n/a | — | 1.65 | bucket\\_score avg = 10.000 |\n| fuzz | 20 | n/a | — | 1.31 | bucket\\_score avg = 10.000 |\n| mcq | 5 | n/a | — | 0.03 | bucket\\_score avg = 0.400 — 3 failures |\n\n---\n\n## Key Findings\n\n### Finding 1: json\\_multistep — Complete Breakdown on Hard Case\n\nThe drop from 0.600 (F16) to 0.400 (Q4\\_K\\_M) is the largest json\\_multistep degradation in the PBH Applied Systems evaluated series to date. The case-level breakdown reveals why:\n\n| Case | Difficulty | F16 Result | Q4\\_K\\_M Result | Q4\\_K\\_M Signals |\n|---|---|---|---|---|\n| ms\\_easy\\_01 | Easy | ✅ PASS | ✅ PASS | All pass |\n| ms\\_easy\\_02 | Easy | ❌ FAIL | ❌ FAIL | cc=0, oe=0 |\n| ms\\_med\\_01 | Medium | ✅ PASS | ✅ PASS | All pass |\n| ms\\_med\\_02 | Medium | ✅ PASS | ❌ **FAIL** | cc=0 only |\n| ms\\_hard\\_01 | Hard | ❌ FAIL | ❌ **FAIL** | **ALL 4 signals fail** |\n\n**ms\\_hard\\_01 at Q4\\_K\\_M is a total failure:** `schema_ok=0`, `checks_consistent_ok=0`, `stop_semantics_ok=0`, `oracle_equiv_ok=0`. All four Tier-1 gating signals fail simultaneously. The model does not produce a parseable schema response, its intermediate checks are self-inconsistent, its STOP semantics are wrong, and the computed final state does not match the oracle. The F16 variant fails this case too — but only on consistency and oracle, not on schema or STOP semantics. Quantization turns a partial failure into a complete one.\n\n**ms\\_med\\_02** is a new failure at Q4\\_K\\_M that passes at F16: `checks_consistent_ok=0` with `oracle_equiv_ok=1` — the model arrives at the correct final state but its internal reasoning steps are self-inconsistent. This is a structural coherence regression under quantization.\n\n**Practical implication:** This model at Q4\\_K\\_M should not be used for multi-step planning tasks without an external validation layer. The hard case cannot be considered reliably solvable at this precision level.\n\n### Finding 2: toolcall — Final Mismatch on tool\\_02 (Q4\\_K\\_M)\n\n`toolcall` passes at 1.000 (both stage1 signals pass), but `tool_02` shows `detail=final_mismatch` with `bucket_score=0` at Q4\\_K\\_M. The tool call JSON is dispatched and validated correctly — the stage-1 parse and schema check both pass — but the model's final answer (the computed result returned after tool execution) does not match the expected output.\n\n| Case | F16 bucket | Q4\\_K\\_M bucket | Q4\\_K\\_M detail |\n|---|---:|---:|---|\n| tool\\_01 | 11 | 11 | ok |\n| tool\\_02 | 11 | **0** | **final\\_mismatch** |\n\nThis is not a gating failure — the pass rate remains 1.000 because Tier-1 only evaluates the dispatch quality, not the final answer. However, in a production pipeline where tool results feed downstream computation, a final\\_mismatch means the model called the tool correctly but gave a wrong answer when reporting the result. **For applications where post-tool reasoning accuracy matters, treat this as a deployment risk at Q4\\_K\\_M.**\n\nThe Q4\\_K\\_M `toolcall` bucket\\_score average of 5.5 (vs 11.0 at F16) directly reflects this: one perfect (11) and one complete failure (0) averaged together.\n\n### Finding 3: toolcall\\_only — Consistent args Failure with Stable Tool Name\n\n| Signal | F16 Rate | Q4\\_K\\_M Rate |\n|---|---:|---:|\n| tool\\_name\\_ok | 1.000 | 1.000 |\n| args\\_ok | 1.000 | 0.000 |\n| schema\\_ok | 0.000\\* | 0.000 |\n\nAt F16, `toolcall_only` passes gating (tool\\_name\\_ok=1, args\\_ok=1) but carries the same schema wrapper non-compliance observed across multiple models in this series (`schema_ok=0`, `detail=schema_error`). At Q4\\_K\\_M, `tool_name_ok` stays perfect at 1.000 — the model correctly identifies which tool to call — but `args_ok` drops to 0.000 on both cases. The quantized model knows the tool name but cannot construct a valid argument payload.\n\n> \\*F16 schema\\_ok=0 is a non-gating wrapper issue (uses `\"tool\"` instead of `\"tool_name\"` as outer key), not a capability failure. Both gating signals pass at F16.\n\n### Finding 4: MCQ \"got=A\" Bias\n\nBoth runners show a systematic bias toward selecting choice A on failures:\n\n| Case | F16 result | Q4\\_K\\_M result |\n|---|---|---|\n| mcq\\_01 | ✅ ok | ✅ ok |\n| mcq\\_02 | ❌ wrong\\_choice got=A | ❌ wrong\\_choice got=A |\n| mcq\\_03 | ✅ ok | ✅ ok |\n| mcq\\_04 | ✅ ok | ❌ wrong\\_choice got=A |\n| mcq\\_05 | ❌ wrong\\_choice got=A | ❌ wrong\\_choice got=A |\n\nEvery failure on both runners produces `got=A`. This is a model-level characteristic: when uncertain, this model defaults to option A. Q4\\_K\\_M extends this bias to mcq\\_04 (which F16 answers correctly), reducing the bucket\\_score from 0.600 to 0.400. For MCQ applications, be aware of this A-default tendency and consider instruction-tuning or chain-of-thought prompting to elicit more deliberate choice selection.\n\n---\n\n## Signal-Level Diagnostics (Q4\\_K\\_M)\n\n### json\\_multistep\n\n| Signal | F16 Rate | Q4\\_K\\_M Rate | Delta |\n|---|---:|---:|---:|\n| schema\\_ok | 1.000 | 0.800 | −0.200 |\n| checks\\_consistent\\_ok | 0.800 | **0.400** | **−0.400** |\n| stop\\_semantics\\_ok | 1.000 | 0.800 | −0.200 |\n| oracle\\_equiv\\_ok | 0.600 | 0.600 | 0.000 |\n| final\\_consistent\\_ok | 0.000 | 0.000 | 0.000 |\n| final\\_match\\_reported | 0.000 | 0.000 | 0.000 |\n\n`checks_consistent_ok` takes the largest hit (−0.400), dropping from 0.800 to 0.400. This signal measures whether the model's intermediate reasoning steps are internally self-consistent. The Q4\\_K\\_M degradation here is the root cause of the json\\_multistep pass rate drop: the model fails ms\\_med\\_02 on consistency alone, and ms\\_hard\\_01 on all signals.\n\n### stateful\\_followup\n\n| Signal | Rate |\n|---|---:|\n| turn1\\_parse\\_ok | 1.000 |\n| turn2\\_parse\\_ok | 1.000 |\n| turn1\\_exact\\_match | 1.000 |\n| turn2\\_exact\\_match | 1.000 |\n\n### toolcall\\_only (Q4\\_K\\_M)\n\n| Signal | Rate |\n|---|---:|\n| tool\\_name\\_ok | 1.000 |\n| args\\_ok | 0.000 |\n\n### mixed\\_brief\\_json\n\n| Signal | Rate |\n|---|---:|\n| answer\\_line\\_ok | 1.000 |\n| json\\_parse\\_ok | 1.000 |\n| schema\\_ok | 1.000 |\n\n---\n\n## Recommended Use Cases\n\n### ✅ Deploy with Confidence (Q4\\_K\\_M)\n\n- **Stateful multi-turn agents** — Perfect two-turn state retention (1.000). Reliable at 0.40 sec/case.\n- **Structured JSON outputs (single-step)** — `json` and `fuzz` both achieve bucket\\_score 10.000. Valid constraint-adherent outputs every case.\n- **Hybrid brief + JSON responses** — `mixed_brief_json` passes at 1.000. Fast at 0.52 sec/case.\n- **Long-context document processing** — 128K token context window is the key differentiator for this model. Suitable for full-document Q&A, multi-document comparison, and long-form extraction tasks.\n- **Multilingual structured tasks** — Trained on 9 languages. Reliable for non-English structured output pipelines.\n- **Tool-calling with response scaffolding** — `toolcall` pass rate holds at 1.000. Use stage-1 tool dispatch reliably; add final-answer validation for downstream computation (see tool\\_02 finding).\n\n### ⚠️ Use with Guardrails (Q4\\_K\\_M)\n\n- **Multi-step planning at easy-to-medium difficulty** — ms\\_easy\\_01 and ms\\_med\\_01 pass. ms\\_med\\_02 fails on internal consistency. Use with an external validation loop for any planning task beyond trivial difficulty.\n- **Post-tool answer validation** — `toolcall` dispatches correctly but tool\\_02 returns a wrong final answer. Validate model output after tool execution, not just the tool call itself.\n- **Bare tool-call dispatch** — `toolcall_only` fails on args (0.000). Add a schema enforcement layer or use scaffolded tool calling.\n\n### ❌ Not Recommended (Q4\\_K\\_M)\n\n- **Hard multi-step planning** — ms\\_hard\\_01 fails on all four gating signals simultaneously. Do not deploy for hard planning tasks without F16 or an external planner/verifier.\n- **MCQ without A-bias mitigation** — Three of five MCQ cases fail at Q4\\_K\\_M, all defaulting to A. Add chain-of-thought prompting or answer validation for MCQ-style pipelines.\n\n---\n\n## 128K Context Window — Deployment Considerations\n\nMistral-Nemo's 128K context window is a meaningful production advantage for document-intensive applications. At Q4\\_K\\_M (7.48 GB model weight), actual usable context depends on KV cache VRAM overhead:\n\n| Context Length | Approx. KV Cache | Total VRAM Needed | Fits on |\n|---|---|---|---|\n| 8K tokens | ~0.5 GB | ~10 GB | T4 16 GB |\n| 32K tokens | ~2 GB | ~12 GB | T4 16 GB · RTX 4080 |\n| 64K tokens | ~4 GB | ~14 GB | A10G 24 GB · RTX 4090 |\n| 128K tokens | ~8 GB | ~18 GB | A10G 24 GB · RTX 4090 |\n\nSet `n_ctx` in llama-cpp-python to the actual context length you need — do not default to 128K if your use case only needs 8K. Unnecessary context allocation wastes VRAM and slows inference.\n\n---\n\n## Hardware Requirements\n\n| Configuration | VRAM Required | Recommended GPU |\n|---|---|---|\n| Q4\\_K\\_M · 8K context (this repo) | ~10 GB | T4 16 GB · RTX 3080 |\n| Q4\\_K\\_M · 128K context | ~18 GB | A10G 24 GB · RTX 4090 |\n| F16 baseline (companion repo) | ~26 GB | A100 40 GB · RTX 4090 · 2× A10G |\n\n---\n\n## Usage\n\n### Installation\n\n```bash\npip install llama-cpp-python huggingface_hub\n```\n\nFor GPU acceleration (CUDA):\n\n```bash\nCMAKE_ARGS=\"-DGGML_CUDA=on\" pip install llama-cpp-python --force-reinstall --no-cache-dir\n```\n\n### Python — llama-cpp-python\n\n```python\nfrom huggingface_hub import hf_hub_download\nfrom llama_cpp import Llama\n\nmodel_path = hf_hub_download(\n    repo_id=\"pbhappliedsystems/mistral-nemo-instruct-2407-gguf-Q4-K-M\",\n    filename=\"mistral-nemo-instruct-2407-gguf-Q4-K-M.gguf\"\n)\n\nllm = Llama(\n    model_path=model_path,\n    n_ctx=32768,      # Adjust to your use case; model supports up to 128K\n    n_gpu_layers=-1,  # -1 offloads all layers to GPU\n    verbose=False,\n)\n\nresponse = llm.create_chat_completion(\n    messages=[\n        {\n            \"role\": \"system\",\n            \"content\": \"You are a precise assistant. Follow instructions exactly and return structured outputs when requested.\"\n        },\n        {\n            \"role\": \"user\",\n            \"content\": \"Analyze the following document and return a JSON object with keys: summary, key_entities, sentiment, action_items.\"\n        }\n    ],\n    temperature=0.3,\n    max_tokens=1024,\n)\n\nprint(response[\"choices\"][0][\"message\"][\"content\"])\n```\n\nFor long-document use (leveraging the 128K context window):\n\n```python\n# Load a large document and process it within a single context window\nwith open(\"large_document.txt\", \"r\") as f:\n    document = f.read()\n\nllm_long = Llama(\n    model_path=model_path,\n    n_ctx=65536,      # 64K context — adjust based on available VRAM\n    n_gpu_layers=-1,\n    verbose=False,\n)\n\nresponse = llm_long.create_chat_completion(\n    messages=[\n        {\"role\": \"system\", \"content\": \"You are a document analysis assistant.\"},\n        {\"role\": \"user\", \"content\": f\"Summarize the following document and extract all action items:\\n\\n{document}\"}\n    ],\n    temperature=0.3,\n    max_tokens=2048,\n)\nprint(response[\"choices\"][0][\"message\"][\"content\"])\n```\n\nFor tool-calling with post-tool answer validation (addresses tool\\_02 final\\_mismatch finding):\n\n```python\nimport json\nfrom huggingface_hub import hf_hub_download\nfrom llama_cpp import Llama\n\nmodel_path = hf_hub_download(\n    repo_id=\"pbhappliedsystems/mistral-nemo-instruct-2407-gguf-Q4-K-M\",\n    filename=\"mistral-nemo-instruct-2407-gguf-Q4-K-M.gguf\"\n)\n\nllm = Llama(model_path=model_path, n_ctx=4096, n_gpu_layers=-1, verbose=False)\n\ndef execute_tool(tool_name: str, args: dict) -> str:\n    \"\"\"Stub: replace with actual tool execution.\"\"\"\n    if tool_name == \"add\":\n        return str(args[\"a\"] + args[\"b\"])\n    raise ValueError(f\"Unknown tool: {tool_name}\")\n\ndef call_with_tool_and_validate(prompt: str) -> dict:\n    \"\"\"\n    Scaffolded tool dispatch with post-execution answer validation.\n    Addresses quant_eval v7.21 finding: tool_02 final_mismatch at Q4_K_M.\n    toolcall stage1 pass rate = 1.000; final answer accuracy is not guaranteed.\n    \"\"\"\n    response = llm.create_chat_completion(\n        messages=[\n            {\n                \"role\": \"system\",\n                \"content\": \"You are a tool-calling assistant. Emit a tool call JSON, then report the result.\"\n            },\n            {\"role\": \"user\", \"content\": prompt}\n        ],\n        temperature=0.0,\n        max_tokens=512,\n    )\n    raw = response[\"choices\"][0][\"message\"][\"content\"]\n\n    # Extract tool call\n    import re\n    match = re.search(r'\\{[^{}]*\"tool_name\"[^{}]*\\}', raw, re.DOTALL)\n    if not match:\n        raise ValueError(f\"No tool call found: {raw[:200]}\")\n    call = json.loads(match.group(0))\n\n    # Execute tool independently — do not trust model's reported result\n    actual_result = execute_tool(call[\"tool_name\"], call[\"args\"])\n    return {\"tool_call\": call, \"validated_result\": actual_result, \"model_raw\": raw}\n\nresult = call_with_tool_and_validate(\"What is 10 minus 4?\")\nprint(f\"Validated result: {result['validated_result']}\")\n```\n\n### CLI — llama-cli\n\n```bash\n# One-shot prompt\nllama-cli \\\n  --model mistral-nemo-instruct-2407-gguf-Q4-K-M.gguf \\\n  --chat-template mistral \\\n  --system-prompt \"You are a precise assistant.\" \\\n  --prompt \"Analyze the following and return a JSON object with keys: summary, risk_level, action_items.\" \\\n  --n-predict 1024 \\\n  --ctx-size 32768 \\\n  --n-gpu-layers -1 \\\n  --temp 0.3\n```\n\nFor server deployment (OpenAI-compatible endpoint):\n\n```bash\nllama-server \\\n  --model mistral-nemo-instruct-2407-gguf-Q4-K-M.gguf \\\n  --chat-template mistral \\\n  --ctx-size 32768 \\\n  --n-gpu-layers -1 \\\n  --port 8080 \\\n  --host 0.0.0.0\n```\n\nQuery via the OpenAI-compatible API:\n\n```python\nfrom openai import OpenAI\n\nclient = OpenAI(base_url=\"http://localhost:8080/v1\", api_key=\"not-required\")\n\nresponse = client.chat.completions.create(\n    model=\"mistral-nemo-instruct-2407-gguf-Q4-K-M\",\n    messages=[{\"role\": \"user\", \"content\": \"Your prompt here\"}],\n    temperature=0.3,\n)\nprint(response.choices[0].message.content)\n```\n\n---\n\n## Artifact Provenance\n\n| Artifact | Format | Size | SHA256 |\n|---|---|---|---|\n| `mistral-nemo-instruct-2407-gguf-Q4-K-M.gguf` | GGUF Q4\\_K\\_M | 7.48 GB | `5765024ff3361f6dc5b590b963b378bd2e87ac95eabe5823a08a3ad336b498c9` |\n| F16 *(companion repo)* | GGUF F16 | 24.5 GB | `cc7b8c5c3f129ad32aee562017b5e56f1284ee7d92445be292c644a42b3c9556` |\n\nBoth artifacts were produced from `mistralai/Mistral-Nemo-Instruct-2407` using a custom-built llama.cpp conversion and quantization pipeline developed by PBH Applied Systems.\n\n---\n\n## Evaluation Methodology\n\n**quant_eval v7.21** is a proprietary behavioral evaluation harness developed by PBH Applied Systems. It evaluates both the full-precision (F16) and quantized variants against an identical fixture set, enabling direct comparison of capability retention across quantization levels.\n\n**Fixture set:** `golden_oracle_fixtures_v7_21` (SHA256: `6d71a0b9147c079371b02a94f3c149eb78a6adc03dc16ff6833b964fbf4174f0`)\n\n| Family | Description | Pass Signals |\n|---|---|---|\n| `fuzz` | Property-based regression; structured placement correctness | schema\\_ok, constraints\\_ok |\n| `json` | Single-step structured JSON with constraint rules | schema\\_ok, constraints\\_ok |\n| `json_multistep` | Multi-step planning with self-check and oracle verification | schema\\_ok, checks\\_consistent\\_ok, stop\\_semantics\\_ok, oracle\\_equiv\\_ok |\n| `mcq` | Multiple-choice extraction | choice\\_ok |\n| `stateful_followup` | Two-turn state tracking; turn-2 correct given turn-1 | turn1/2\\_parse\\_ok, turn1/2\\_exact\\_match |\n| `mixed_brief_json` | Hybrid: natural language answer + valid JSON block | answer\\_line\\_ok, json\\_parse\\_ok, schema\\_ok |\n| `toolcall` | Tool call embedded in response; parse + schema validation | stage1\\_tool\\_parse\\_ok, stage1\\_tool\\_schema\\_ok |\n| `toolcall_only` | Bare schema-only tool call; strict tool name + args check | tool\\_name\\_ok, args\\_ok |\n\n**Evaluation hardware:** NVIDIA RTX 4090 (24 GB VRAM)\n**Evaluation date:** February 11, 2026\n**quant_eval seed:** 42\n\n---\n\n## About PBH Applied Systems\n\n[**PBH Applied Systems, LLC**](https://pbhappliedsystems.com) is an Oklahoma City–based applied machine learning and AI systems company specializing in production-grade model evaluation, quantization pipelines, agentic AI infrastructure, and scalable AI-driven application development. The organization operates with a strong emphasis on engineering rigor, reproducibility, and real-world deployment constraints — particularly in environments where performance, cost efficiency, and reliability must be balanced against available hardware and budget.\n\n### Founder — Patrick Hill, M.S.\n\nPBH Applied Systems was founded by **Patrick Hill**, a Data Scientist and AI/ML Engineer with 10+ years of experience delivering advanced analytics, predictive modeling, and decision-support solutions across high-stakes operational environments. Patrick holds a **Master of Science in Software Engineering with concentrations in Artificial Intelligence and Machine Learning** (GPA: 4.0) and a B.S. in Business Finance.\n\n**Technical expertise spans:**\n\n- **Languages & Data:** Python, SQL, Linux, Pandas, NumPy, scikit-learn\n- **ML & Modeling:** Supervised and unsupervised learning, neural networks, NLP, transformers, regression, classification, forecasting, and feature engineering\n- **AI/ML Frameworks:** PyTorch, TensorFlow/Keras, HuggingFace Transformers, GGUF, llama.cpp, BitsAndBytes, PEFT, QLoRA\n- **Deployment & MLOps:** Flask APIs, Docker, CI/CD pipelines, REST endpoints, streaming inference, version control\n- **Data Platforms:** Jupyter, Databricks, Power BI, Matplotlib\n- **Quantization:** GGUF conversion, Q4\\_K\\_M / Q5\\_K\\_M / Q8\\_0 strategies, adapter-per-model evaluation architecture\n\n### Published Author\n\nPatrick is the author of **[Applied Machine Learning: Concepts, Tools, and Case Studies](https://a.co/d/05qat7Xz)** — a 1,200+ page practitioner-oriented textbook adopted as **required reading for CSC 373 – Machine Learning at the University of Advancing Technology**.\n\n### Core Service Areas\n\n**1. LLM Optimization & Deployment** — End-to-end GGUF conversion and quantization with custom llama.cpp pipelines and adapter-per-model architecture.\n\n**2. AI Evaluation Frameworks** — Proprietary behavioral evaluation via quant_eval: per-family pass rates, F16 vs. quantized delta analysis, failure cluster diagnostics, deployment recommendations.\n\n**3. Agentic AI Infrastructure** — LlamaIndex ReAct agents, Flask orchestration, serverless GPU inference, full pipeline from model selection to production serving.\n\n**4. Scalable AI Application Development** — Multimodal applications (quantized LLMs + Whisper + BLIP), Dockerized Flask APIs, advanced time-series forecasting with custom attention mechanisms, Bayesian hyperparameter optimization, and FinBERT sentiment fusion.\n\n**5. ML Pipeline Design & Analytics** — Feature engineering, forward-chaining cross-validation, KPI dashboards, analytical governance at scale.\n\n**6. Model & Agent Cataloging** — Structured catalog publishing with reproducible artifacts and clear performance tradeoff documentation.\n\n---\n\n## 📞 Work With PBH Applied Systems\n\nThe complete breakdown of ms\\_hard\\_01 at Q4\\_K\\_M — all four gating signals failing simultaneously — and the tool\\_02 final\\_mismatch are findings that only appear when you run both the F16 and quantized variant against the same behavioral test suite. Neither shows up in standard benchmarks. Neither is visible from casual testing. Both have direct consequences for production deployment decisions.\n\n**A model that dispatches tools correctly but gives wrong answers, and that fails completely on hard planning cases, needs to be known before it goes to production — not after.**\n\n👉 **[Book a Scoping Call](https://pbhappliedsystems.com)** — Discuss your model selection, quantization strategy, or deployment architecture directly with Patrick.\n\n👉 **[Request an Evaluation Report](https://pbhappliedsystems.com)** — A full quant_eval behavioral audit: per-family pass rates, F16 vs. quantized delta analysis, failure cluster diagnostics, and a deployment recommendation. Engagements from $2,500.\n\n### Connect\n\n| | |\n|---|---|\n| 🌐 **Website** | [pbhappliedsystems.com](https://pbhappliedsystems.com) |\n| 📧 **Email** | [patrick@pbhappliedsystems.com](mailto:patrick@pbhappliedsystems.com) |\n| 💼 **LinkedIn** | [PBH Applied Systems, LLC](https://www.linkedin.com/company/pbh-applied-systems-llc) |\n| ▶️ **YouTube** | [@pbhappliedsystems](https://www.youtube.com/@pbhappliedsystems) |\n| 📸 **Instagram** | [@pbhappliedsystems](https://www.instagram.com/pbhappliedsystems) |\n| 👍 **Facebook** | [pbhappliedsystems](https://www.facebook.com/pbhappliedsystems) |\n\n---\n\n## License\n\nThis GGUF repository inherits the license of the base model:\n**Apache 2.0** — [`mistralai/Mistral-Nemo-Instruct-2407`](https://huggingface.co/mistralai/Mistral-Nemo-Instruct-2407)\n\nThe quant_eval evaluation methodology, fixture set, and scoring framework are proprietary to PBH Applied Systems, LLC and are not included in this repository.\n\n---\n\n*GGUF conversion, quantization, and behavioral evaluation performed by [PBH Applied Systems, LLC](https://pbhappliedsystems.com) · quant_eval v7.21 · Run ID: `20260211_022944`*\n",
    "related_quantizations": []
  },
  "tags": [
    "nemo",
    "gguf",
    "quantized",
    "q4_k_m",
    "mistral",
    "instruct",
    "llama-cpp",
    "agentic",
    "tool-calling",
    "structured-output",
    "multilingual",
    "pbh-applied-systems",
    "quant-eval",
    "en",
    "fr",
    "de",
    "es",
    "it",
    "pt",
    "ru",
    "zh",
    "ja",
    "base_model:mistralai/Mistral-Nemo-Instruct-2407",
    "base_model:quantized:mistralai/Mistral-Nemo-Instruct-2407",
    "license:apache-2.0",
    "endpoints_compatible",
    "region:us",
    "conversational"
  ],
  "likes": 0,
  "downloads": 1045,
  "gated": false,
  "private": false,
  "last_modified": "2026-04-15T07:00:50.000Z",
  "created_at": "2026-04-11T23:30:31.000Z",
  "pipeline_tag": "",
  "library_name": "nemo"
}
Source payload excerpt (from Hugging Face API)
{
  "_id": "69dad99797388ce31fa890c1",
  "id": "pbhappliedsystems/mistral-nemo-instruct-2407-gguf-Q4-K-M",
  "modelId": "pbhappliedsystems/mistral-nemo-instruct-2407-gguf-Q4-K-M",
  "sha": "eeda50bbc8110885c5e08ccbe1e0e14f83604164",
  "createdAt": "2026-04-11T23:30:31.000Z",
  "lastModified": "2026-04-15T07:00:50.000Z",
  "author": "pbhappliedsystems",
  "downloads": 1045,
  "likes": 0,
  "gated": false,
  "private": false,
  "pipeline_tag": "",
  "library_name": "nemo",
  "siblings_count": 7
}