Koshkasa/Gryphe_WorldSim-Opus-3.6-35B-A3B-iQ_compact_APEX-GGUF overview
What's that? A modification of the default APEX i Compact Qwen 3.6 35B A3B recipe utilizing IQ3 S instead of Q3 K, IQ4 NL instead of Q4 K, and Q8 0 instead of …
Runs locally from ~183.3 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
Model Details
| Model ID | Koshkasa/Gryphe_WorldSim-Opus-3.6-35B-A3B-iQ_compact_APEX-GGUF |
|---|---|
| Author | Koshkasa |
| Pipeline | text-generation |
| License | apache-2.0 |
| Base model | Gryphe/WorldSim-Opus-3.6-35B-A3B |
| Last modified | 2026-08-14T13:44:09.000Z |
Model README
---
license: apache-2.0
base_model:
- Gryphe/WorldSim-Opus-3.6-35B-A3B
library_name: llama.cpp
pipeline_tag: text-generation
tags:
- gguf
- quantized
- llama.cpp
- roleplay
- mixed precision
- APEX
quantized_by: Koshkasa
base_model_relation: quantized
---
What's that?
A modification of the default APEX i-Compact Qwen 3.6 35B A3B recipe utilizing IQ3_S instead of Q3_K, IQ4_NL instead of Q4_K, and Q8_0 instead of Q6_K in shared exps, applied to Gryphe/WorldSim-Opus-3.6-35B-A3B. Though slower on older hardware, they should preserve much more brain due to better outlier handling.
Deliberately stepping away from mixed math/article/story/rp soup datasets, imatrix dataset is a random conversation 250000-token prune of Squish42/bluemoon-fandom-1-1-rp-cleaned. For RP-specific finetunes not aimed at being a "general assistant", the weights engaged in creative writing must be preserved first.
At least that's how the theory goes!
Disclosure
My only contribution is compute. This is not my merge. Have fun.
Model card incomplete. Tests and comparisons may be uploaded at a later date.
Run Koshkasa/Gryphe_WorldSim-Opus-3.6-35B-A3B-iQ_compact_APEX-GGUF with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models