GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Koshkasa/Gryphe_WorldSim-Opus-3.6-35B-A3B-IQ4_KT-GGUF overview

REQUIRES ik llama.cpp or its derivatives INCOMPATIBLE with mainline llama.cpp and its derivatives as of August 16th, 2026 What's that? A ik llama.cpp compatibl…

ik_llama.cppggufquantizedllama.cpproleplaymixed precisiontrellisdataset:Squish42/bluemoon-fandom-1-1-rp-cleanedbase_model:Gryphe/WorldSim-Opus-3.6-35B-A3Bbase_model:quantized:Gryphe/WorldSim-Opus-3.6-35B-A3Blicense:apache-2.0endpoints_compatibleregion:usimatrixconversational

Runs locally from ~16.91 GB disk (24 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
Author

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
Gryphe_WorldSim-Opus-3.6-35B-A3B-IQ4_KT.ggufGGUFIQ4_KT16.91 GBDownload

Model Details

Model IDKoshkasa/Gryphe_WorldSim-Opus-3.6-35B-A3B-IQ4_KT-GGUF
AuthorKoshkasa
Pipeline
Licenseapache-2.0
Base modelGryphe/WorldSim-Opus-3.6-35B-A3B
Last modified2026-08-16T06:53:39.000Z

Model README

---

license: apache-2.0

base_model:

  • Gryphe/WorldSim-Opus-3.6-35B-A3B

library_name: ik_llama.cpp

tags:

  • gguf
  • quantized
  • llama.cpp
  • roleplay
  • mixed precision
  • trellis

quantized_by: Koshkasa

base_model_relation: quantized

datasets:

  • Squish42/bluemoon-fandom-1-1-rp-cleaned

---

REQUIRES ik_llama.cpp or its derivatives

INCOMPATIBLE with mainline llama.cpp and its derivatives (as of August 16th, 2026)

What's that?

A ik_llama.cpp compatible quantization of Gryphe/WorldSim-Opus-3.6-35B-A3B, utilizing:

  • Q8_0 for SSM tensors
  • IQ6_K for embeddings, output, shared experts, attention tensors in global self-attn layers
  • IQ4_KT for sparse experts and attention in local attn layers.

The intention behind such a recipe is squeezing the most sparse experts brainpower in the least amount of space, with absolutely no compromise on long-context attention.

Deliberately stepping away from mixed math/article/story/rp soup datasets, imatrix dataset is a random conversation ~250000-token prune of Squish42/bluemoon-fandom-1-1-rp-cleaned with some turn-based conversation data added from bartowski's calibration data. The imatrix is provided.

For RP-specific finetunes not aimed at being a "general assistant", the weights engaged in creative writing must be preserved first.

At least that's how the theory goes!

Disclosure

WYSIWYG. My only contribution is compute. This is not my merge. Assume WTFPL licensing where no other license implicitly applies. Have fun.

Model card incomplete. Tests and comparisons may be uploaded at a later date.

Run Koshkasa/Gryphe_WorldSim-Opus-3.6-35B-A3B-IQ4_KT-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models