jashepp/Qwable-v2-35B-A3B-MXFP4_MOE_Hybrid-Imatrix-GGUF overview
π Qwable v2 35B A3B Custom Mixed Precision GGUFs with Imatrix Qwen + Fable, second iteration Β· An open weights agentic coding model.\ 35B Mixture of Experts 3β¦
Runs locally from ~183.3 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | jashepp/Qwable-v2-35B-A3B-MXFP4_MOE_Hybrid-Imatrix-GGUF |
|---|---|
| Author | jashepp |
| Pipeline | text-generation |
| License | agpl-3.0 |
| Base model | lordx64/Qwable-v2 |
| Last modified | 2026-07-03T12:52:04.000Z |
Model README
---
license: agpl-3.0
language:
- en
library_name: transformers
base_model_relation: quantized
pipeline_tag: text-generation
base_model:
- lordx64/Qwable-v2
tags:
- qwen
- qwen3
- qwen3.6
- moe
- distillation
- chain-of-thought
- agentic
- claude-fable-5
- claude-opus-4.7
- tool-use
- chained-distill
- imatrix
- GGUF
- mxfp4
- quantized
---
π Qwable-v2-35B-A3B - Custom Mixed Precision GGUFs with Imatrix
> Qwen + Fable, second iteration Β· An open-weights agentic coding model.\
> 35B Mixture-of-Experts (3B active), built by layering Claude Fable-5 agentic tool-use behavior on top of a Claude Opus 4.7 reasoning distill of Qwen3.6-35B-A3B. Trained with 4Γ the LoRA capacity and 2Γ the SFT data of Qwable-v1.



This repository contains custom, highly optimized, multi-tier mixed precision GGUF weights for lordx64/Qwable-v2.
> [!TIP]
> βΉοΈ For advanced agentic and programming tasks, I highly recommend upgrading to Ornith-1.0-35B for significantly better performance.
Qwable-v2 is a 35B Mixture-of-Experts (3B active) hybrid architecture alternating between standard Attention and Mamba State-Space (SSM) blocks.\
Distilled heavily from Claude 4.7 Opus reasoning and Claude Fable-5 agentic traces, on top of Qwen3.6-35B-A3B.
These quants were generated using manual layer targeting to maximize quality while shrinking the massive VRAM footprint of the Mixture of Experts layers.
π Importance Matrix (Imatrix)
The following datasets were used for the imatrix:
- Custom Target Matrix (500 chunks), up to a max of
10MBof each:
- eaddario/imatrix-calibration - tools_huge, code_huge, math_medium
- Glint-Research/Fable-5-traces
- lordx64/fable-sft-combined-v2
- osunlp/QUEST-SFT-Data-Open-ended
π GGUF Files
In order of quality:
| Filename | Size | Quants |
| :--- | :--- | :--- |
| Qwable-v2-35B-A3B-MXFP4_MOE_Q8_0_F16-Imatrix.gguf | 21.3 GB | MXFP4_MOE + Q8_0 + F16 |
| Qwable-v2-35B-A3B-MXFP4_MOE_Q8_0-Imatrix.gguf | 20.3 GB | MXFP4_MOE + Q8_0 |
---
π Precision Matrix & Flavor Variations
Standard global quantization presets (like stock MXFP4_MOE) compress the backbone layers uniformly, which degrades the delicate reasoning capabilities of advanced agent models.\
This repository provides two distinct manual configuration layouts to balance precision and memory constraints:
1. The Tri-Quant Hybrid Flavor (MXFP4 + Q8_0 + F16)
Qwable-v2-35B-A3B-MXFP4_MOE_Q8_0_F16-Imatrix.gguf - Designed for maximum quality preservation, this layout implements a strict 3-Tier Precision Matrix:
- Tier 1 (Core & Mamba Gating - F16 Precision):
- token_embd.weight, output.weight - Protects the critical input/output vocabulary mappings. Adds ~1GB to the file size but dramatically prevents text degradation.
- ssm_alpha, ssm_beta - Protects the integrity of the Mamba state-space calculations across long-range context tokens.
- Tier 2 (Backbone & Shared - Q8_0 Precision):
ssm_out,*._shexp- Keeps the attention mechanics, and all trailing shared experts at high quality, to protect the logical research loops. - Tier 3 (Routed Experts - MXFP4 Precision):
ffn_down_exps,ffn_gate_exps,ffn_up_exps- Shrink the massive background expert parameters directly toMXFP4.
2. The Dual-Quant Hybrid Flavor (MXFP4 + Q8_0)
Qwable-v2-35B-A3B-MXFP4_MOE_Q8_0-Imatrix.gguf - Designed for a slightly leaner memory profile, this layout utilizes 2-Tier Precision:
- Tier 1 (Backbone - Q8_0 Precision): All attention blocks, Mamba structures, vocabulary embeddings, and internal routers use the universal
Q8_0format. - Tier 2 (Experts - MXFP4 Precision): The heavy sparse expert blocks are target-quantized directly to
MXFP4.
---
π Exact Conversion Details
Because Qwable-v2 utilizes trailing shared-expert vectors at the boundary edge of its alternating architecture (Layer 40 boundary), standard conversion requires boundary bypass instructions using .*_shexp\.weight=Q8_0. These files were converted via llama-quantize utilizing the following manual recipe parameters:
Convert SafeTensors to GGUF:
# Requires python3.12, with `pip install --upgrade transformers`
python convert_hf_to_gguf.py "Qwable-v2/" --outtype f16 --outfile "Qwable-v2_F16.gguf"
Generate Tri-Quant MXFP4_MOE + Q8_0 + F16:
llama-quantize \
--tensor-type ".*_shexp\.weight=Q8_0" \
--tensor-type "token_embd\.weight=F16" \
--tensor-type "^output\.weight=F16" \
--tensor-type "blk\..*\.(ssm_alpha|ssm_beta)\.weight=F16" \
--tensor-type "blk\..*\.(ffn_down_exps|ffn_gate_exps|ffn_up_exps)\.weight=MXFP4" \
--imatrix "imatrix.gguf" \
"Qwable-v2_F16.gguf" \
"Qwable-v2-35B-A3B-MXFP4_MOE_Q8_0_F16-Imatrix.gguf" \
Q8_0
Generate Dual-Quant MXFP4_MOE + Q8_0:
llama-quantize \
--tensor-type ".*_shexp\.weight=Q8_0" \
--tensor-type "blk\..*\.(ffn_down_exps|ffn_gate_exps|ffn_up_exps)\.weight=MXFP4" \
--imatrix "imatrix.gguf" \
"Qwable-v2_F16.gguf" \
"Qwable-v2-35B-A3B-MXFP4_MOE_Q8_0-Imatrix.gguf" \
Q8_0
---
π Local Deployment & llama-server Configuration (config.ini)
To maintain the rock-solid reasoning loop depth of the Fable+Opus distillation and prevent agents from falling into repetitive tool-calling deadlocks, use the following server parameter recommendations.
# --- Samplers (Dynamic & Expressive) ---
temperature = 0.65
top-k = 40
top-p = 0.90
min-p = 0.08
# --- Penalties (Prevent Syntax & Reasoner Corruption) ---
repeat-penalty = 1.00
presence-penalty = 0.00
# --- DRY Sampler (Protects Indentation & Structural Boilerplate) ---
dry-multiplier = 0.8
dry-base = 1.75
dry-allowed-length = 8
dry-penalty-last-n = 1024
dry-sequence-breaker = ["\n", ":", " ", "\t", "\"", ","]
# --- Enforced Execution Graph ---
samplers = temp;top_k;top_p;min_p;dry
Keep reasoning off for this model.
reasoning = off
reasoning-budget = 0
reasoning-format = none
This works well with 256k context window.
---
βΉοΈ Misc Details
I'm doing this as a side hobby, with my AMD 5900X, 64GB DDR4, RTX 3060 12GB & RTX 5060 Ti 16GB.
In addition to the above configuration, I also use:
slots = 1
parallel = 1
no-warmup = true
flash-attn = on
mlock = false
no-mmap = false
no-context-shift = true
batch-size = 2048
ubatch-size = 256
fit = on
fit-target = 768
main-gpu = 0
split-mode = layer
n-gpu-layers = 999
n-cpu-moe = 0
tensor-split = 16,12
override-tensor = (token_embd)=CUDA0,(vision|vpm|nextn)=CPU
cache-type-k = q8_0
cache-type-v = q8_0
jinja = true
chat-template = jinja
chat-template-file = chat_template.jinja
For further quality and better ssm behaviour, this configuration can help:
context-shift = false
cache-type-k = f16
cache-type-v = f16
---
π€ Support the Journey
As a passionate developer, I'm always programming, automating, or experimenting with new ideas.\
I love building open-source tools, trying out new web tech, and creating things that don't yet exist, including local AI & quantizing models.
I love sharing these creations to give back to the community.\
If my projects have saved you time or helped you out, consider supporting my work below!
π Support me on Ko-fi
---
β¨ Acknowledgments
- lordx64 for the exceptional
Qwable-v2base model.
π License
See lordx64/Qwable-v2#license--terms.
π Citation
@misc{lordx64_qwable_v2_2026,
title = {Qwable-v2: Agentic coding distillation from Claude Fable-5 onto Qwen3.6-35B-A3B with LoRA r=64 + combined corpus},
author = {lordx64},
year = {2026},
howpublished = {\url{https://huggingface.co/lordx64/Qwable-v2}},
}Run jashepp/Qwable-v2-35B-A3B-MXFP4_MOE_Hybrid-Imatrix-GGUF with guIDE
Download guIDE β the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face Β· Compare models