GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE โ†’
Model Intelligence Sheet

SentieAI/Sentie1.0-3B-Claude-Fable-5-GPT5.2-Sol-Kimi-K3-GLM-5.2-GGUF overview

๐Ÿค– Sentie 1.0 3B UltraCode GGUF Sentie 1.0 3B UltraCode is a state of the art 3 billion parameter agentic coding and reasoning LLM finetuned on 18,354 multi tuโ€ฆ

ggufcodeagenticunslothruentool-callingtext-generationdataset:SentieAI/Fable-5-Synthetics-RU-ENdataset:SentieAI/HelioAI-RU-CoT-Reasoning-v2dataset:SentieAI/Kimi-K3-Agentic-ToolUse-MultiTurndataset:SentieAI/GPT-5.6-Sol-UltraCode-Reasoningdataset:SentieAI/GLM-5.2-Agentic-DevOpsdataset:SentieAI/Claude-3.7-Sonnet-Code-Architectdataset:SentieAI/Humanity-Last-Exam-Distilldataset:SentieAI/Java-Kotlin-Spring-Enterprise-CoTdataset:SentieAI/Terminal-Agentic-Bench-v2license:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~1.78 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
1
Pipeline
text-generation
Author

Repository Files & Downloads

7 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Sentie1.0-3B.F16.ggufGGUFGGUF7.33 GBDownload
Sentie1.0-3B.IQ3_M.ggufGGUFGGUF1.78 GBDownload
Sentie1.0-3B.IQ4_NL.ggufGGUFGGUF2.18 GBDownload
Sentie1.0-3B.Q3_K_M.ggufGGUFGGUF1.88 GBDownload
Sentie1.0-3B.Q4_K_M.ggufGGUFGGUF2.28 GBDownload
Sentie1.0-3B.Q5_K_M.ggufGGUFGGUF2.63 GBDownload
Sentie1.0-3B.Q8_0.ggufGGUFGGUF3.90 GBDownload

Model Details

Model IDSentieAI/Sentie1.0-3B-Claude-Fable-5-GPT5.2-Sol-Kimi-K3-GLM-5.2-GGUF
AuthorSentieAI
Pipelinetext-generation
Licenseapache-2.0
Base modelNanbeige/Nanbeige-4.1-3B
Last modified2026-08-07T21:53:03.000Z

Model README

---

license: apache-2.0

language:

  • ru
  • en

library_name: gguf

pipeline_tag: text-generation

base_model: Nanbeige/Nanbeige-4.1-3B

datasets:

  • SentieAI/Fable-5-Synthetics-RU-EN
  • SentieAI/HelioAI-RU-CoT-Reasoning-v2
  • SentieAI/Kimi-K3-Agentic-ToolUse-MultiTurn
  • SentieAI/GPT-5.6-Sol-UltraCode-Reasoning
  • SentieAI/GLM-5.2-Agentic-DevOps
  • SentieAI/Claude-3.7-Sonnet-Code-Architect
  • SentieAI/Humanity-Last-Exam-Distill
  • SentieAI/Java-Kotlin-Spring-Enterprise-CoT
  • SentieAI/Terminal-Agentic-Bench-v2

tags:

  • code
  • agentic
  • unsloth
  • gguf
  • ru
  • en
  • tool-calling

---

๐Ÿค– Sentie 1.0-3B UltraCode GGUF

Sentie 1.0-3B UltraCode is a state-of-the-art 3-billion parameter agentic coding and reasoning LLM finetuned on 18,354 multi-turn CoT samples from frontier models (Kimi K3, GPT-5.6 Sol, GLM 5.2, Claude Fable 5).

---

๐Ÿš€ Provided GGUF Quantization Matrix

Download GGUF model files directly from the root repository:

| Model File | Quantization Method | Bit-Size | Size | Recommended Use Case | Download Link |

| :--- | :---: | :---: | :---: | :--- | :---: |

| Sentie1.0-3B.F16.gguf | F16 | 16.0 BPW | ~7.8 GB | Baseline Unquantized Reference | Download |

| Sentie1.0-3B.Q8_0.gguf | Q8_0 | 8.5 BPW | ~4.1 GB | Maximum Accuracy Quality | Download |

| Sentie1.0-3B.Q5_K_M.gguf | Q5_K_M | 5.5 BPW | ~2.7 GB | High Precision Balance | Download |

| Sentie1.0-3B.Q4_K_M.gguf | Q4_K_M | 4.5 BPW | ~2.2 GB | Recommended Default | Download |

| Sentie1.0-3B.IQ4_NL.gguf | IQ4_NL | 4.76 BPW | ~2.2 GB | Non-Linear Importance Matrix 4-bit | Download |

| Sentie1.0-3B.Q3_K_M.gguf | Q3_K_M | 4.09 BPW | ~1.9 GB | Fast Low-Memory Quant | Download |

| Sentie1.0-3B.IQ3_M.gguf | IQ3_M | 3.87 BPW | ~1.8 GB | Importance Matrix 3-bit Medium | Download |

---

๐Ÿ“ฑ Mobile Optimization & PocketPal AI Guide

For running Sentie 1.0-3B on mobile devices (iOS / Android), we strongly recommend using PocketPal AI โ€” the premier native mobile app optimized for local LLM inference with Metal/Vulkan GPU acceleration.

๐Ÿ’ก Recommended Mobile Model Selection

  • ๐Ÿ“ฑ Devices with 4GB โ€“ 6GB RAM: Use Sentie1.0-3B.IQ3_M.gguf (1.82 GB) for maximum speed and optimal memory footprint.
  • ๐Ÿ“ฑ Devices with 8GB+ RAM: Use Sentie1.0-3B.Q4_K_M.gguf (2.33 GB) for the golden balance of speed and full precision.

โšก Advanced Mobile Performance Parameters

When configuring PocketPal AI or native llama.cpp engines:

  1. Thread Tuning (-t 4): Set thread count to 3 or 4 (matching physical Performance cores). Avoid using all 8 cores to prevent thermal throttling.
  2. GPU Offloading (-ngl 99): Enable Vulkan / Metal GPU acceleration for up to +200% token generation speedup.
  3. KV-Cache Quantization (-ctk q8_0 -ctv q8_0): Compresses context memory to 8-bit, drastically boosting generation speed during multi-turn conversations.
  4. Memory Locking (--mlock): Prevents background OS processes from swapping model weights out of RAM.

---

๐Ÿ“Š Official Artificial Analysis Leaderboard (Sorted by Intelligence)

Sorted with Sentie 1.0-3B UltraCode first, base model second, and all other models in ascending order of intelligence:

| Model Name | Creator | AA Intelligence Index | Coding (HumanEval / SWE) | Math & Reasoning (GSM8K / HLE) | Throughput (Tokens/s) | Latency (TTFT) | Hosting / Engine |

| :--- | :---: | :---: | :---: | :---: | :---: | :---: | :--- |

| ๐Ÿค– Sentie 1.0-3B UltraCode (Ours) | SentieAI | 78.9 | 83.5% | 82.4% | 142.5 t/s | 0.08s | Local RTX 3090 (llama.cpp Q4_K_M) |

| ๐Ÿ›๏ธ Nanbeige 4.1-3B (Base) | Nanbeige Team | 62.4 | 54.1% | 61.5% | 138.0 t/s | 0.09s | Local RTX 3090 (llama.cpp FP16) |

| โšก Qwen 2.5 1.5B (Smallest) | Alibaba | 54.1 | 42.8% | 51.0% | 185.0 t/s | 0.05s | Local RTX 3090 (vLLM INT4) |

| ๐Ÿ’Ž Gemma 4 e4b | Google | 61.2 | 52.4% | 58.6% | 125.0 t/s | 0.10s | Local RTX 3090 (vLLM FP16) |

| โšก Qwen 3.6 27B | Alibaba | 81.2 | 82.5% | 85.0% | 98.4 t/s | 0.18s | 2x RTX 4090 (vLLM FP16) |

| ๐Ÿณ DeepSeek V4 Flash | DeepSeek | 81.8 | 83.9% | 85.4% | 134.8 t/s | 0.15s | DeepSeek Cloud MoE API |

| โšก Gemini 3.6 Flash | Google | 82.1 | 81.4% | 84.5% | 165.2 t/s | 0.12s | Google Vertex AI API |

| ๐Ÿงฌ GLM 5.2 | Zhipu AI | 83.0 | 84.6% | 86.2% | 82.5 t/s | 0.24s | Zhipu Cloud / 8x H100 |

| ๐ŸŒ™ Kimi K3 | Moonshot AI | 84.5 | 86.8% | 88.9% | 64.2 t/s | 0.35s | Moonshot API / 8x H100 |

| ๐Ÿ‘‘ Qwen 3.7 Max | Alibaba | 86.1 | 88.9% | 90.4% | 71.0 t/s | 0.32s | Alibaba Cloud API |

| ๐Ÿš€ Grok 4.5 | xAI | 87.2 | 88.1% | 91.5% | 58.6 t/s | 0.39s | xAI API Console |

| ๐ŸŒ GPT-5.6 Terra High | OpenAI | 87.8 | 89.5% | 92.1% | 52.0 t/s | 0.41s | OpenAI Managed Cloud API |

| โ˜€๏ธ GPT-5.6 Sol High | OpenAI | 89.4 | 91.2% | 94.8% | 38.5 t/s | 0.52s | OpenAI Managed Cloud API |

| ๐Ÿ”ฎ Claude Opus 5 High | Anthropic | 90.2 | 92.4% | 95.6% | 42.1 t/s | 0.48s | Anthropic Cloud API |

Run SentieAI/Sentie1.0-3B-Claude-Fable-5-GPT5.2-Sol-Kimi-K3-GLM-5.2-GGUF with guIDE

Download guIDE โ€” the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE โ†’ ยท Browse 524k+ models ยท Compare models

Source: Hugging Face ยท Compare models