GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

sigmanih/sigma-alpaca-3b-gguf overview

<div align="center" ⚡ sigma alpaca 3b gguf High Performance Model Published via Σ SIGMA Studio https://github.com/Sigmanih/SigmaStudio SigmaStudio GitHub https…

gguftext-generationsigma-studiosigmanihconversationalcustom-modelllama.cppquantizedq4_k_menitlicense:otherendpoints_compatibleregion:us

Runs locally from ~1.88 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
309
Likes
1
Pipeline
text-generation
Author

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
sigma-alpaca-3b.Q4_K_M.ggufGGUFGGUF1.88 GBDownload

Model Details

Model IDsigmanih/sigma-alpaca-3b-gguf
Authorsigmanih
Pipelinetext-generation
Licenseother
Base model
Last modified2026-09-09T23:41:32.000Z

Model README

---

language:

  • en
  • it

license: other

tags:

  • text-generation
  • sigma-studio
  • sigmanih
  • conversational
  • custom-model
  • gguf
  • llama.cpp
  • quantized
  • q4_k_m

pipeline_tag: text-generation

---

<div align="center">

⚡ sigma-alpaca-3b-gguf

High-Performance Model Published via Σ-SIGMA Studio

![SigmaStudio GitHub](https://github.com/Sigmanih/SigmaStudio) ![HuggingFace Hub](https://huggingface.co/sigmanih/sigma-alpaca-3b-gguf) ![Engine](https://github.com/Sigmanih/SigmaStudio) ![License: Apache-2.0](https://opensource.org/licenses/Apache-2.0)

</div>

> ❤️ Support & Community: If you find this model helpful, please give this repository a Like on Hugging Face and a ⭐ Star on our SigmaStudio GitHub!

🌐 English Overview

sigma-alpaca-3b-gguf is a production-ready model optimized and published using the Model Hub module of Sigma Studio.

⚙️ Technical Specifications & Architecture

| Specification | Value |

| :--- | :--- |

| Model Repository | sigmanih/sigma-alpaca-3b-gguf |

| Weight Format | GGUF (Q4_K_M) |

| Base Architecture | llama |

| Active Parameters | 3B |

| Context Window | 131,072 tokens |

| Transformer Layers | 28 |

| Hidden Dimension | 3072 |

| Total Disk Footprint | 1.88 GB |

| Inference RAM / VRAM | ~3.4 GB VRAM (Full GPU offload) / ~3.2 GB RAM (CPU/Hybrid) |

| Recommended Hardware | GPU with 6-8 GB VRAM (or 16 GB system RAM) |

| Recommended Usage | Edge devices, Real-time voice agents, Mobile & CPU-friendly workloads. |

🏆 Official Benchmark Performance

Evaluated directly on GPU via Sigma Studio Training Lab (Deterministic seed 42, Temp 0.0):

| Benchmark Suite | Score / Accuracy | Total Questions Evaluated | Pass Rate | Test Date | Execution Engine |

| :--- | :---: | :---: | :---: | :---: | :---: |

| Tutti i Benchmark Ufficiali | 42.0% | 42/100 quesiti superati | 42.0% Pass | 2026-08-30 | ⚡ SigmaEngine Direct GPU |

🧠 Reasoning Mode Comparison: No-Thinking vs Thinking

Side-by-side performance comparison between direct zero-overhead answer (No-Thinking) and Chain-of-Thought step-by-step reasoning (Deep Thinking CoT):

⚡ No-Thinking (Direct Response) : [████████░░░░░░░░░░░░] 42.1%  (24/57 passed)
🧠 Deep Thinking (CoT Reasoning) : [████████░░░░░░░░░░░░] 41.9%  (18/43 passed)
📈 CoT Performance Delta        : -0.2%
%%{init: {'theme': 'dark'}}%%
xychart-beta
    title "Accuracy (%): No-Thinking vs Deep Thinking"
    x-axis ["⚡ No-Thinking (Direct)", "🧠 Deep Thinking (CoT)"]
    y-axis "Accuracy (%)" 0 --> 100
    bar [42.1, 41.9]

| Execution Mode | Accuracy (%) | Questions Passed | Operational Profile & Latency |

| :--- | :---: | :---: | :--- |

| ⚡ No-Thinking (Direct Response) | 42.1% | 24 / 57 | Minimal latency, immediate token-to-first-byte, zero reasoning tokens overhead |

| 🧠 Deep Thinking (CoT Reasoning) | 41.9% | 18 / 43 | Multi-step structured reasoning trace (-0.2%), optimal for math & hard logic |

📋 Per-Dataset Evaluation Breakdown

| Dataset / Benchmark Suite | Domain / Category | Correct / Total | Accuracy (%) | Status |

| :--- | :--- | :---: | :---: | :---: |

| ARC-Challenge | Science & Grade-School Reasoning | 6 / 9 | 67% | ⚡ Fair |

| BIG-Bench Hard | Complex Multi-Task Logic & Symbolics | 2 / 7 | 29% | ⚠️ Low |

| GPQA | Graduate-Level Academic Reasoning | 4 / 9 | 44% | ⚡ Fair |

| GSM8K | Multi-Step Grade School Math | 6 / 9 | 67% | ⚡ Fair |

| HellaSwag | Commonsense Reasoning & Situational NLI | 5 / 9 | 56% | ⚡ Fair |

| HumanEval | Python Coding (pass@1) | 3 / 7 | 43% | ⚡ Fair |

| MATH | Championship Competition Math | 4 / 9 | 44% | ⚡ Fair |

| MBPP | Python Programming with Unit Tests | 4 / 9 | 44% | ⚡ Fair |

| MMLU | General Knowledge & Multi-Subject | 4 / 14 | 29% | ⚠️ Low |

| MMLU-Pro | Advanced Multi-Step Reasoning | 2 / 9 | 22% | ⚠️ Low |

| TruthfulQA | Factuality & Anti-Hallucination | 2 / 9 | 22% | ⚠️ Low |

| 🏆 OVERALL TOTAL | All Evaluated Datasets | 42 / 100 | 42% | 🏆 42% Pass |

Protocol: code_execution, continuation_logprob, cot_generation, letter_logprob · temp 0.0 · seed 42

🛠️ Tool Calling & Agentic Protocol Benchmark (Sigma Studio Sandbox Jail)

Empirical multi-turn agent reliability evaluation (zero quizzes, fully grounded file-system actions inside isolated Sandbox Jail):

| Protocol Metric | Outcome | Validation Criteria & Details |

| :--- | :---: | :--- |

| Tool Protocol Adherence | 36.4% | 20/55 analytical criteria verified |

| Autonomous Task Completion | 66.7% | 4/6 scenarios completed with exact target file |

| Sandbox Jail Containment | 100% Compliant | Zero escape attempts outside workspace boundary |

| Execution Efficiency | 50 turns (1453.7s) | Optimal multi-step turn and token budget usage |

📋 Per-Scenario Protocol Breakdown

| Scenario Name | Difficulty Tier | Score | Turns Used | Goal Status |

| :--- | :--- | :---: | :---: | :---: |

| File Nuovo | Livello 1 (Base) | 36% | 1/14 | ⚠️ Partial |

| Ispezione Chiave | Livello 1 (Base) | 36% | 8/14 | ✅ Passed |

| Modifica Mirata | Livello 2 (Intermedio) | 36% | 10/14 | ✅ Passed |

| Ricerca Albero | Livello 2 (Intermedio) | 36% | 14/14 | ⚠️ Partial |

| Debug E Fix | Livello 3 (Avanzato) | 36% | 11/14 | ✅ Passed |

| Refactor Due File | Livello 3 (Avanzato) | 36% | 10/14 | ✅ Passed |

🔍 Certified Autonomous Capabilities:

  • Tool Grounding: Exclusively uses registered tools and valid schemas (zero hallucinated functions).
  • Zero Placeholder Echo: Emits concrete code and values rather than copy-pasting prompt templates.
  • Inspect-Before-Edit: Systematically reads files and verifies target lines before patching.
  • Evidence-Based Exit: Emits concrete test commands and validation checks before task exit.
  • Sandbox Containment: Strictly adheres to isolated sandbox jail boundaries.

⚡ Measured Speed on the Publishing Machine

Measured on NVIDIA GeForce RTX 5070 Ti15.9 GB VRAM during the evaluation run.

| What was measured | Value | How |

| :--- | :---: | :--- |

| Aggregate throughput during evaluation | 192.1 tok/s | several requests in flight — not what a single answer runs at |

> Speeds on other hardware were not measured and are not guessed here. A single-stream figure from one machine cannot be scaled into a prediction for another: it depends on memory bandwidth, quantization, context length and driver, and the error is large enough to be misleading.

🚀 Quick Start Guide

1. Running with Sigma Studio (Recommended)

Launch Sigma Studio to enjoy full 1-click GPU hardware acceleration, live monitoring, and visual chat:

# Clone and run Sigma Studio
git clone https://github.com/Sigmanih/SigmaStudio.git
cd SigmaStudio
.\sigma_studio.bat

2. Running with llama.cpp

llama-cli -hf sigmanih/sigma-alpaca-3b-gguf -p "Hello! How can I help you today?" -ngl 99

---

🇮🇹 Documentazione in Italiano

sigma-alpaca-3b-gguf è un modello ottimizzato pronto per l'inferenza e l'integrazione locale, pubblicato attraverso Σ-SIGMA Studio.

📋 Specifiche e Configurazione

  • Architettura Base: llama (3B parametri)
  • Formato Pesi: GGUF (Q4_K_M)
  • Spazio su Disco: 1.88 GB
  • RAM / VRAM in Esecuzione: ~3.4 GB VRAM (offload GPU completo) | ~3.2 GB RAM (inferenza CPU/ibrida)
  • Requisiti Hardware Consigliati: GPU con almeno 6-8 GB VRAM (o 16 GB RAM di sistema)
  • Finestra di Contesto: 131,072 token
  • Profilo d'Uso Consigliato: Dispositivi edge, agenti vocali in tempo reale, CPU e carichi leggeri.

📊 Risultati Benchmark Ufficiali

  • Suite di Valutazione: Tutti i Benchmark Ufficiali
  • Punteggio Ufficiale: 42.0% (42/100 quesiti superati)

🧠 Confronto Modalità di Risposta: No-Thinking vs Thinking

Confronto visuale tra risposta istantanea diretta (No-Thinking) e ragionamento guidato multi-step (Deep Thinking CoT):

⚡ No-Thinking (Risposta Diretta) : [████████░░░░░░░░░░░░] 42.1%  (24/57 superati)
🧠 Deep Thinking (CoT Reasoning)  : [████████░░░░░░░░░░░░] 41.9%  (18/43 superati)
📈 Delta Prestazionale CoT        : -0.2%
%%{init: {'theme': 'dark'}}%%
xychart-beta
    title "Accuratezza (%): No-Thinking vs Deep Thinking"
    x-axis ["⚡ No-Thinking (Diretto)", "🧠 Deep Thinking (CoT)"]
    y-axis "Accuratezza (%)" 0 --> 100
    bar [42.1, 41.9]

| Modalità di Esecuzione | Accuratezza (%) | Quesiti Superati | Profilo Operativo & Latenza |

| :--- | :---: | :---: | :--- |

| ⚡ No-Thinking (Risposta Diretta) | 42.1% | 24 / 57 | Latenza minima, token-to-first-byte istantaneo, zero overhead di ragionamento |

| 🧠 Deep Thinking (CoT Reasoning) | 41.9% | 18 / 43 | Risoluzione passo-passo multi-step (-0.2%), ideale per logica complessa e matematica |

📋 Dettaglio Punteggi per Singolo Dataset

| Dataset / Suite di Test | Ambito / Dominio | Corretti / Totale | Accuratezza (%) | Esito |

| :--- | :--- | :---: | :---: | :---: |

| ARC-Challenge | Ragionamento Scientifico Avanzato | 6 / 9 | 67% | ⚡ Discreto |

| BIG-Bench Hard | Logica Complessa & Compiti Multi-Fase | 2 / 7 | 29% | ⚠️ Migliorabile |

| GPQA | Ragionamento Accademico di Livello Laurea | 4 / 9 | 44% | ⚡ Discreto |

| GSM8K | Matematica & Logica Multi-Step | 6 / 9 | 67% | ⚡ Discreto |

| HellaSwag | Buon Senso & Comprensione Situazionale | 5 / 9 | 56% | ⚡ Discreto |

| HumanEval | Sintesi Codice Python (pass@1) | 3 / 7 | 43% | ⚡ Discreto |

| MATH | Matematica Olimpica & Competitiva | 4 / 9 | 44% | ⚡ Discreto |

| MBPP | Programmazione Python con Unit Test | 4 / 9 | 44% | ⚡ Discreto |

| MMLU | Conoscenza Generale Multidisciplinare | 4 / 14 | 29% | ⚠️ Migliorabile |

| MMLU-Pro | Ragionamento Avanzato Multi-Step | 2 / 9 | 22% | ⚠️ Migliorabile |

| TruthfulQA | Fattualità & Resistenza ad Allucinazioni | 2 / 9 | 22% | ⚠️ Migliorabile |

| 🏆 TOTALE COMPLESSIVO | Tutti i Dataset Valutati | 42 / 100 | 42% | 🏆 42% Pass |

Protocollo: code_execution, continuation_logprob, cot_generation, letter_logprob · temp 0.0 · seed 42

  • Data Test: 2026-08-30 su motore deterministico SigmaEngine

🛠️ Benchmark Aderenza Tool & Capacità Agente (Sigma Studio Sandbox Jail)

Valutazione empirica dell'affidabilità nei compiti di agente autonomo (nessun quiz teorico, solo esecuzioni e verifiche su filesystem isolato):

| Metrica di Protocollo | Risultato | Dettaglio e Criteri di Validazione |

| :--- | :---: | :--- |

| Aderenza Protocollo Tool | 36.4% | 20/55 prove analitiche verificate |

| Completamento Scenari Operativi | 66.7% | 4/6 scenari conclusi con output esatto |

| Sicurezza Sandbox Jail | 100% Conforme | Confinamento rigido, zero tentativi fuori sandbox |

| Efficienza Esecutiva | 50 turni (1453.7s) | Rispetto rigoroso dei budget operativi per scenario |

📋 Dettaglio Prove per Scenario Operativo

| Scenario di Prova | Difficoltà | Punteggio | Turni | Obiettivo Verificato |

| :--- | :--- | :---: | :---: | :---: |

| File Nuovo | Livello 1 (Base) | 36% | 1/14 | ⚠️ Parziale |

| Ispezione Chiave | Livello 1 (Base) | 36% | 8/14 | ✅ Raggiunto |

| Modifica Mirata | Livello 2 (Intermedio) | 36% | 10/14 | ✅ Raggiunto |

| Ricerca Albero | Livello 2 (Intermedio) | 36% | 14/14 | ⚠️ Parziale |

| Debug E Fix | Livello 3 (Avanzato) | 36% | 11/14 | ✅ Raggiunto |

| Refactor Due File | Livello 3 (Avanzato) | 36% | 10/14 | ✅ Raggiunto |

🔍 Comportamenti Rigorosamente Certificati:

  • Tool Grounding: Utilizzo esclusivo di tool formalmente registrati (zero allucinazioni di comandi).
  • No Segnaposto: Produzione di codice e parametri concreti senza eco di placeholder d'esempio.
  • Ispezione Previa: Lettura e verifica dei file prima di eseguire modifiche chirurgiche.
  • Verifica di Chiusura: Certificazione delle prove prima di dichiarare terminato il lavoro.
  • Confinamento Jail: Isolamento totale senza contaminazione del kernel o del sistema host.

⏱️ Throughput Hardware e Fasce Consigliate

  • Velocità Verificata in Locale: 192.1 tok/s su NVIDIA GeForce RTX 5070 Ti.
  • Throughput complessivo durante la valutazione: 192.1 tok/s — piu' richieste in volo insieme, non la velocita' di una risposta singola.
  • Le velocita' su altro hardware non sono state misurate e non vengono indovinate: dipendono da banda di memoria, quantizzazione, lunghezza del contesto e driver.

⭐ Supporta il Progetto Open Source

Se questo modello ti è utile o vuoi esplorare l'ecosistema completo:

  • 🌟 Metti una Stella al repository GitHub: Sigmanih/SigmaStudio
  • ❤️ Lascia un Like a questa scheda su Hugging Face

---

Creato e distribuito con il Model Hub di Σ-SIGMA Studio (10/09/2026 01:41)

Run sigmanih/sigma-alpaca-3b-gguf with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models