sigmanih/ornith-ai-Ornith-1.0-35B-GGUF-Q4_K_M overview
<div align="center" ⚡ ornith ai Ornith 1.0 35B GGUF Q4 K M High Performance Model Published via Σ SIGMA Studio https://github.com/Sigmanih/SigmaStudio SigmaStu…
Runs locally from ~19.71 GB disk (24 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| ornith-ai--Ornith-1.0-35B.Q4_K_M.gguf | GGUF | GGUF | 19.71 GB | Download |
Model Details
| Model ID | sigmanih/ornith-ai-Ornith-1.0-35B-GGUF-Q4_K_M |
|---|---|
| Author | sigmanih |
| Pipeline | text-generation |
| License | other |
| Base model | ornith-ai/Ornith-1.0-35B |
| Last modified | 2026-09-09T23:40:49.000Z |
Model README
---
language:
- en
- it
license: other
base_model:
- ornith-ai/Ornith-1.0-35B
tags:
- text-generation
- sigma-studio
- sigmanih
- conversational
- custom-model
- gguf
- llama.cpp
- quantized
- q4_k_m
pipeline_tag: text-generation
---
<div align="center">
⚡ ornith-ai-Ornith-1.0-35B-GGUF-Q4_K_M
High-Performance Model Published via Σ-SIGMA Studio
   
</div>
> ❤️ Support & Community: If you find this model helpful, please give this repository a Like on Hugging Face and a ⭐ Star on our SigmaStudio GitHub!
🌐 English Overview
ornith-ai-Ornith-1.0-35B-GGUF-Q4_K_M is a production-ready model optimized and published using the Model Hub module of Sigma Studio.
⚙️ Technical Specifications & Architecture
| Specification | Value |
| :--- | :--- |
| Model Repository | sigmanih/ornith-ai-Ornith-1.0-35B-GGUF-Q4_K_M |
| Weight Format | GGUF (Q4_K_M) |
| Base Architecture | qwen35moe |
| Active Parameters | 35B |
| Context Window | 262,144 tokens |
| Transformer Layers | 40 |
| Hidden Dimension | 2048 |
| Total Disk Footprint | 19.71 GB |
| Inference RAM / VRAM | ~20.8 GB VRAM (Full GPU offload) / ~20.3 GB RAM (CPU/Hybrid) |
| Recommended Hardware | GPU with 24 GB VRAM (e.g. RTX 3090 / 4090 or 32-64 GB RAM) |
| Recommended Usage | Flagship frontier intelligence, Deep research & multi-step mathematics. |
🏆 Official Benchmark Performance
Evaluated directly on GPU via Sigma Studio Training Lab (Deterministic seed 42, Temp 0.0):
| Benchmark Suite | Score / Accuracy | Total Questions Evaluated | Pass Rate | Test Date | Execution Engine |
| :--- | :---: | :---: | :---: | :---: | :---: |
| Tutti i Benchmark Ufficiali | 72.0% | 72/100 quesiti superati | 72.0% Pass | 2026-09-08 | ⚡ SigmaEngine Direct GPU |
🧠 Reasoning Mode Comparison: No-Thinking vs Thinking
Side-by-side performance comparison between direct zero-overhead answer (No-Thinking) and Chain-of-Thought step-by-step reasoning (Deep Thinking CoT):
⚡ No-Thinking (Direct Response) : [████████████████░░░░] 80.7% (46/57 passed)
🧠 Deep Thinking (CoT Reasoning) : [████████████░░░░░░░░] 60.5% (26/43 passed)
📈 CoT Performance Delta : -20.2%
%%{init: {'theme': 'dark'}}%%
xychart-beta
title "Accuracy (%): No-Thinking vs Deep Thinking"
x-axis ["⚡ No-Thinking (Direct)", "🧠 Deep Thinking (CoT)"]
y-axis "Accuracy (%)" 0 --> 100
bar [80.7, 60.5]
| Execution Mode | Accuracy (%) | Questions Passed | Operational Profile & Latency |
| :--- | :---: | :---: | :--- |
| ⚡ No-Thinking (Direct Response) | 80.7% | 46 / 57 | Minimal latency, immediate token-to-first-byte, zero reasoning tokens overhead |
| 🧠 Deep Thinking (CoT Reasoning) | 60.5% | 26 / 43 | Multi-step structured reasoning trace (-20.2%), optimal for math & hard logic |
📋 Per-Dataset Evaluation Breakdown
| Dataset / Benchmark Suite | Domain / Category | Correct / Total | Accuracy (%) | Status |
| :--- | :--- | :---: | :---: | :---: |
| ARC-Challenge | Science & Grade-School Reasoning | 9 / 9 | 100% | ✅ Passed |
| BIG-Bench Hard | Complex Multi-Task Logic & Symbolics | 5 / 7 | 71% | ✅ Passed |
| GPQA | Graduate-Level Academic Reasoning | 2 / 9 | 22% | ⚠️ Low |
| GSM8K | Multi-Step Grade School Math | 8 / 9 | 89% | ✅ Passed |
| HellaSwag | Commonsense Reasoning & Situational NLI | 6 / 9 | 67% | ⚡ Fair |
| HumanEval | Python Coding (pass@1) | 7 / 7 | 100% | ✅ Passed |
| MATH | Championship Competition Math | 5 / 9 | 56% | ⚡ Fair |
| MBPP | Python Programming with Unit Tests | 7 / 9 | 78% | ✅ Passed |
| MMLU | General Knowledge & Multi-Subject | 9 / 14 | 64% | ⚡ Fair |
| MMLU-Pro | Advanced Multi-Step Reasoning | 6 / 9 | 67% | ⚡ Fair |
| TruthfulQA | Factuality & Anti-Hallucination | 8 / 9 | 89% | ✅ Passed |
| 🏆 OVERALL TOTAL | All Evaluated Datasets | 72 / 100 | 72% | 🏆 72% Pass |
Protocol: code_execution, continuation_logprob, cot_generation, letter_logprob · temp 0.0 · seed 42
🛠️ Tool Calling & Agentic Protocol Benchmark (Sigma Studio Sandbox Jail)
Empirical multi-turn agent reliability evaluation (zero quizzes, fully grounded file-system actions inside isolated Sandbox Jail):
| Protocol Metric | Outcome | Validation Criteria & Details |
| :--- | :---: | :--- |
| Tool Protocol Adherence | 100.0% | 10/10 analytical criteria verified |
| Autonomous Task Completion | 0.0% | 0/1 scenarios completed with exact target file |
| Sandbox Jail Containment | 100% Compliant | Zero escape attempts outside workspace boundary |
| Execution Efficiency | 5 turns (12.0s) | Optimal multi-step turn and token budget usage |
🔍 Certified Autonomous Capabilities:
- ✅ Tool Grounding: Exclusively uses registered tools and valid schemas (zero hallucinated functions).
- ✅ Zero Placeholder Echo: Emits concrete code and values rather than copy-pasting prompt templates.
- ✅ Inspect-Before-Edit: Systematically reads files and verifies target lines before patching.
- ✅ Evidence-Based Exit: Emits concrete test commands and validation checks before task exit.
- ✅ Sandbox Containment: Strictly adheres to isolated sandbox jail boundaries.
⚡ Measured Speed on the Publishing Machine
Measured on NVIDIA GeForce RTX 5070 Ti • 15.9 GB VRAM during the evaluation run.
| What was measured | Value | How |
| :--- | :---: | :--- |
| Aggregate throughput during evaluation | 35.4 tok/s | several requests in flight — not what a single answer runs at |
> Speeds on other hardware were not measured and are not guessed here. A single-stream figure from one machine cannot be scaled into a prediction for another: it depends on memory bandwidth, quantization, context length and driver, and the error is large enough to be misleading.
🚀 Quick Start Guide
1. Running with Sigma Studio (Recommended)
Launch Sigma Studio to enjoy full 1-click GPU hardware acceleration, live monitoring, and visual chat:
# Clone and run Sigma Studio
git clone https://github.com/Sigmanih/SigmaStudio.git
cd SigmaStudio
.\sigma_studio.bat
2. Running with llama.cpp
llama-cli -hf sigmanih/ornith-ai-Ornith-1.0-35B-GGUF-Q4_K_M -p "Hello! How can I help you today?" -ngl 99
---
🇮🇹 Documentazione in Italiano
ornith-ai-Ornith-1.0-35B-GGUF-Q4_K_M è un modello ottimizzato pronto per l'inferenza e l'integrazione locale, pubblicato attraverso Σ-SIGMA Studio.
📋 Specifiche e Configurazione
- Architettura Base:
qwen35moe(35B parametri) - Formato Pesi:
GGUF(Q4_K_M) - Spazio su Disco:
19.71 GB - RAM / VRAM in Esecuzione:
~20.8 GB VRAM(offload GPU completo) |~20.3 GB RAM(inferenza CPU/ibrida) - Requisiti Hardware Consigliati: GPU con 24 GB VRAM (es. RTX 3090 / 4090 o 32-64 GB RAM)
- Finestra di Contesto:
262,144 token - Profilo d'Uso Consigliato: Intelligenza di frontiera, ricerca approfondita e matematica multi-step.
📊 Risultati Benchmark Ufficiali
- Suite di Valutazione:
Tutti i Benchmark Ufficiali - Punteggio Ufficiale:
72.0%(72/100 quesiti superati)
🧠 Confronto Modalità di Risposta: No-Thinking vs Thinking
Confronto visuale tra risposta istantanea diretta (No-Thinking) e ragionamento guidato multi-step (Deep Thinking CoT):
⚡ No-Thinking (Risposta Diretta) : [████████████████░░░░] 80.7% (46/57 superati)
🧠 Deep Thinking (CoT Reasoning) : [████████████░░░░░░░░] 60.5% (26/43 superati)
📈 Delta Prestazionale CoT : -20.2%
%%{init: {'theme': 'dark'}}%%
xychart-beta
title "Accuratezza (%): No-Thinking vs Deep Thinking"
x-axis ["⚡ No-Thinking (Diretto)", "🧠 Deep Thinking (CoT)"]
y-axis "Accuratezza (%)" 0 --> 100
bar [80.7, 60.5]
| Modalità di Esecuzione | Accuratezza (%) | Quesiti Superati | Profilo Operativo & Latenza |
| :--- | :---: | :---: | :--- |
| ⚡ No-Thinking (Risposta Diretta) | 80.7% | 46 / 57 | Latenza minima, token-to-first-byte istantaneo, zero overhead di ragionamento |
| 🧠 Deep Thinking (CoT Reasoning) | 60.5% | 26 / 43 | Risoluzione passo-passo multi-step (-20.2%), ideale per logica complessa e matematica |
📋 Dettaglio Punteggi per Singolo Dataset
| Dataset / Suite di Test | Ambito / Dominio | Corretti / Totale | Accuratezza (%) | Esito |
| :--- | :--- | :---: | :---: | :---: |
| ARC-Challenge | Ragionamento Scientifico Avanzato | 9 / 9 | 100% | ✅ Superato |
| BIG-Bench Hard | Logica Complessa & Compiti Multi-Fase | 5 / 7 | 71% | ✅ Superato |
| GPQA | Ragionamento Accademico di Livello Laurea | 2 / 9 | 22% | ⚠️ Migliorabile |
| GSM8K | Matematica & Logica Multi-Step | 8 / 9 | 89% | ✅ Superato |
| HellaSwag | Buon Senso & Comprensione Situazionale | 6 / 9 | 67% | ⚡ Discreto |
| HumanEval | Sintesi Codice Python (pass@1) | 7 / 7 | 100% | ✅ Superato |
| MATH | Matematica Olimpica & Competitiva | 5 / 9 | 56% | ⚡ Discreto |
| MBPP | Programmazione Python con Unit Test | 7 / 9 | 78% | ✅ Superato |
| MMLU | Conoscenza Generale Multidisciplinare | 9 / 14 | 64% | ⚡ Discreto |
| MMLU-Pro | Ragionamento Avanzato Multi-Step | 6 / 9 | 67% | ⚡ Discreto |
| TruthfulQA | Fattualità & Resistenza ad Allucinazioni | 8 / 9 | 89% | ✅ Superato |
| 🏆 TOTALE COMPLESSIVO | Tutti i Dataset Valutati | 72 / 100 | 72% | 🏆 72% Pass |
Protocollo: code_execution, continuation_logprob, cot_generation, letter_logprob · temp 0.0 · seed 42
- Data Test:
2026-09-08su motore deterministico SigmaEngine
🛠️ Benchmark Aderenza Tool & Capacità Agente (Sigma Studio Sandbox Jail)
Valutazione empirica dell'affidabilità nei compiti di agente autonomo (nessun quiz teorico, solo esecuzioni e verifiche su filesystem isolato):
| Metrica di Protocollo | Risultato | Dettaglio e Criteri di Validazione |
| :--- | :---: | :--- |
| Aderenza Protocollo Tool | 100.0% | 10/10 prove analitiche verificate |
| Completamento Scenari Operativi | 0.0% | 0/1 scenari conclusi con output esatto |
| Sicurezza Sandbox Jail | 100% Conforme | Confinamento rigido, zero tentativi fuori sandbox |
| Efficienza Esecutiva | 5 turni (12.0s) | Rispetto rigoroso dei budget operativi per scenario |
🔍 Comportamenti Rigorosamente Certificati:
- ✅ Tool Grounding: Utilizzo esclusivo di tool formalmente registrati (zero allucinazioni di comandi).
- ✅ No Segnaposto: Produzione di codice e parametri concreti senza eco di placeholder d'esempio.
- ✅ Ispezione Previa: Lettura e verifica dei file prima di eseguire modifiche chirurgiche.
- ✅ Verifica di Chiusura: Certificazione delle prove prima di dichiarare terminato il lavoro.
- ✅ Confinamento Jail: Isolamento totale senza contaminazione del kernel o del sistema host.
⏱️ Throughput Hardware e Fasce Consigliate
- Velocità Verificata in Locale:
35.4 tok/ssuNVIDIA GeForce RTX 5070 Ti. - Throughput complessivo durante la valutazione:
35.4 tok/s— piu' richieste in volo insieme, non la velocita' di una risposta singola. - Le velocita' su altro hardware non sono state misurate e non vengono indovinate: dipendono da banda di memoria, quantizzazione, lunghezza del contesto e driver.
⭐ Supporta il Progetto Open Source
Se questo modello ti è utile o vuoi esplorare l'ecosistema completo:
- 🌟 Metti una Stella al repository GitHub: Sigmanih/SigmaStudio
- ❤️ Lascia un Like a questa scheda su Hugging Face
---
Creato e distribuito con il Model Hub di Σ-SIGMA Studio (10/09/2026 01:40)
Run sigmanih/ornith-ai-Ornith-1.0-35B-GGUF-Q4_K_M with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models