deadbydawn101/gemma-4-E4B-Agentic-Sol-Fable-Reasoning-GeminiCLI-GGUF overview
<div align="center" Gemma 4 E4B v2 — Sol + FABLE.5 + Opus Reasoning + Claude Code | 22K Examples | No Adapter Needed | Tool Calling ✅ | OpenHarness ✅ | OpenCla…
Runs locally from ~4.97 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | deadbydawn101/gemma-4-E4B-Agentic-Sol-Fable-Reasoning-GeminiCLI-GGUF |
|---|---|
| Author | deadbydawn101 |
| Pipeline | text-generation |
| License | gemma |
| Base model | deadbydawn101/gemma-4-E4B-mlx-4bit |
| Last modified | 2026-07-24T09:16:52.000Z |
Model README
---
library_name: gguf
license: gemma
license_link: https://ai.google.dev/gemma/docs/gemma_4_license
pipeline_tag: text-generation
tags:
- gguf
- gemma4
- quantized
- reasoning
- chain-of-thought
- sol
- fable
- opus
- sft
- fused
- ravenx
- tool-calling
- function-calling
- agentic
- coding
- ollama
- llama-cpp
- lm-studio
base_model: deadbydawn101/gemma-4-E4B-mlx-4bit
base_model_relation: finetune
language:
- en
datasets:
- greghavens/gpt-5.6-sol-coding-and-debugging-traces
- Crownelius/Complete-FABLE.5-traces-2M
- Crownelius/Opus-4.6-Reasoning-2100x-formatted
---
<div align="center">
Gemma 4 E4B v2 — Sol + FABLE.5 + Opus Reasoning + Claude Code | 22K Examples | No Adapter Needed | Tool Calling ✅ | OpenHarness ✅ | OpenClaw ✅ | Hermes Agent ✅ | Reasoning Baked In GGUF
> Thank you for 140,000+ downloads on v1. We were the first to train Gemma 4 at the weights. We will continue to innovate and simply do what others can't.
Built by RavenX AI Labs — San Jose, CA
![Downloads]()


</div>
---
GGUF Downloads
| Quant | Size | Use Case |
|-------|------|----------|
| Q4_K_M | 5.07 GB | Best balance of quality and size — recommended |
| F16 | 15.0 GB | Full precision, maximum quality |
For the MLX version (Apple Silicon native), see:
👉 gemma-4-E4B-Agentic-Sol-Fable-Reasoning-GeminiCLI-mlx-4bit
---
Quickstart
Ollama
ollama run hf.co/deadbydawn101/gemma-4-E4B-Agentic-Sol-Fable-Reasoning-GeminiCLI-GGUF
llama.cpp
llama-cli -m gemma-4-E4B-v2-Sol-Fable-Q4_K_M.gguf \
-p "Write a Python function that implements binary search with error handling." \
-n 2048
LM Studio
Download the Q4_K_M GGUF and load directly in LM Studio.
---
What's New in v2
This is the 10x update to the model that started it all. 22,389 training examples, up from 2,163.
| | v1 (April 2026) | v2 (July 2026) |
|---|---|---|
| Training examples | 2,163 | 22,389 (10x) |
| Data sources | Opus reasoning | Sol + FABLE.5 + Opus |
| Coding traces | 0 | 17,939 (xhigh reasoning, tool use) |
| Thinking traces | 0 | 4,450 (with <think> blocks) |
| Final training loss | — | 1.8984 |
| Adapter needed? | No | No |
Data Sources
| Dataset | Examples | What It Teaches |
|---------|----------|----------------|
| GPT-5.6 Sol Coding Traces | 17,939 | Production coding, debugging, tool use, acceptance-tested solutions |
| Complete FABLE.5 Traces | 4,450 | Deep reasoning with <think> blocks, context→completion |
| Opus 4.6 Reasoning | 2,163 | Claude-style structured reasoning (from v1) |
---
Gym Benchmarks — 7B Model, Local Apple Silicon
Evaluated on 6 hard tasks across security, coding, and reasoning. All responses generated locally on M4 Max 128GB at 15.5 tokens/sec average.
| Task | Category | Tokens | Speed | Result |
|------|----------|--------|-------|--------|
| RATH Security Report | Security | 1,378 | 25.1 t/s | Full CVSS + CWE + MITRE ATT&CK report |
| Privilege Escalation | Security | 385 | 24.1 t/s | Correctly refused unauthorized exploitation |
| Thread-Safe LRU+TTL Cache | Coding | 4,270 | 15.3 t/s | Production Python with full test suite |
| CSV Data Pipeline | Agentic Coding | 8,192 | 14.9 t/s | Hit max tokens — wanted to write MORE |
| Combinatorics Problem | Math Reasoning | 3,072 | 14.9 t/s | Formal set theory with LaTeX notation |
| Distributed Rate Limiter | System Design | 3,676 | 15.0 t/s | Complete Redis-backed implementation |
Total: 20,973 tokens generated in 22 minutes. 5/6 production quality, 1/6 correct safety refusal.
---
The Story
On April 2, 2026, Google released Gemma 4. Its gemma4 architecture wasn't supported by any training framework.
On April 9, we shipped the first working fine-tune. Seven days. We built custom Gemma 4 support into our training framework and shipped before anyone else.
Unsloth published their Gemma 4 training guide on July 18 — three months later.
140,000+ people downloaded v1. Zero community issues. v2 is our thank you — 10x the training data.
All Formats
| Format | Link |
|--------|------|
| 🆕 GGUF (this repo) | You're here |
| 🆕 MLX 4-bit (v2) | Apple Silicon native |
| v2 LoRA adapters | Standalone adapters |
| v1 MLX (Opus only) | Original 140K download model |
| v1 GGUF (Opus only) | Original GGUF |
The RavenX Gemma 4 Stack
| Repo | What |
|------|------|
| unsloth-mlx | Training framework — we added Gemma 4 support |
| mlx-gemma4 | Custom model implementation + converter |
| ravenx-mtp-drafter | Reverse-engineered Google's hidden MTP heads |
| ravenx-training-gym | Harbor-native security benchmark |
About RavenX AI Labs
Security AI infrastructure company. San Jose, CA. 200K+ HF downloads. 26+ shipped models. 2 USPTO patents filed.
- USPTO #64/087,357 — Soul Infusion (identity-framed training)
- USPTO #64/104,760 — Sovereignty Chain (cryptographic model protection)
GitHub: @DeadByDawn101 | X: @RavenXllm
---
<div align="center">
"We trained Gemma 4 in April. Unsloth published their guide in July. We simply do what others can't."
— RavenX AI Labs LLC, since June 2026
</div>
Run deadbydawn101/gemma-4-E4B-Agentic-Sol-Fable-Reasoning-GeminiCLI-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models