GraySoft
Projects Models About FAQ Contact Download guIDE →

jackbinary/qwen-3.5-10b-frankenmerge-opus-4.6-distill-gguf Q4_K_M GGUF - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.

Model Intelligence Sheet

jackbinary/qwen-3.5-10b-frankenmerge-opus-4.6-distill-gguf overview

| Category | Base (Qwen3.5-9B-Base-Q80) | Qwen-3.5-10B-Frankenmerge-Opus-4.6-Distill | Δ | |---|---|---|---| | Factual Knowledge | 85.0% B | 85.0% B | = | | Reasoning | 88.0% B | 60.0% C | ↓ −28.0% | | Coding | 56.0% D | 80.0% B | ↑ +24.0% | | Instruction Following | 100.0% A | 30.0% F | ↓ −70.0% | | Language | 100.0% A | 70.0% C | ↓ −30.0% | | Safety Calibration | 66.7% C | 66.7% C | = | | Overall | 82.4% B | 65.6% C | ↓ −16.8% | Method: Layer surgery on Qwen3.5-9B-Base-Q80 followed by fine-tuning. Benchmarks run at temperature=0, seed=42 Coding capability improved significantly (+24%) at the cost of instruction-following and language tasks --- This model was GGUF format using Unsloth. Example usage:

ggufqwen3_5_textllama.cppunslothqwen3_5frankenmergelayer-surgeryreasoningchain-of-thoughtabliteratedendataset:Jackrong/Qwen3.5-reasoning-700xdataset:nohurry/Opus-4.6-Reasoning-3000x-filteredbase_model:JackBinary/Qwen-3.5-10B-Frankenmerge-Opus-4.6-Distillbase_model:quantized:JackBinary/Qwen-3.5-10B-Frankenmerge-Opus-4.6-Distilllicense:apache-2.0endpoints_compatibleregion:usconversational
jackbinary/qwen-3.5-10b-frankenmerge-opus-4.6-distill-gguf visual
Downloads
576
Likes
0
Pipeline
Library
Visibility
Public
Access
Open

Repository Files & Downloads

3 files detected
Direct downloads for all repository files
FileTypeQuantizationSizeLink
Qwen-3.5-10B-Frankenmerge-Opus-4.6-Distill.Q4_K_M.gguf GGUF Q4_K_M 5.73 GB Download
Qwen-3.5-10B-Frankenmerge-Opus-4.6-Distill.Q6_K.gguf GGUF Q6_K 7.52 GB Download
Qwen-3.5-10B-Frankenmerge-Opus-4.6-Distill.Q8_0.gguf GGUF 9.73 GB Download

Model Details Live

Model Slug
jackbinary/qwen-3.5-10b-frankenmerge-opus-4.6-distill-gguf
Author
JackBinary
Pipeline Task
Library
Created
2026-03-17
Last Modified
2026-03-19
Gated
No
Private
No
HF SHA
392b0c566ffbde26eac8463b36c12e98f7e56f95
License
apache-2.0
Language
en
Base Model
JackBinary/Qwen-3.5-10B-Frankenmerge-Opus-4.6-Distill

Metadata Inspector

Normalized metadata (stored in metadata_json)
{
  "metadata": {},
  "card_data": {
    "license": "apache-2.0",
    "base_model": "JackBinary/Qwen-3.5-10B-Frankenmerge-Opus-4.6-Distill",
    "tags": [
      "gguf",
      "llama.cpp",
      "unsloth",
      "qwen3_5",
      "frankenmerge",
      "layer-surgery",
      "reasoning",
      "chain-of-thought",
      "unsloth",
      "abliterated"
    ],
    "language": [
      "en"
    ],
    "datasets": [
      "Jackrong/Qwen3.5-reasoning-700x",
      "nohurry/Opus-4.6-Reasoning-3000x-filtered"
    ],
    "frontmatter": {
      "license": "apache-2.0",
      "base_model": "JackBinary/Qwen-3.5-10B-Frankenmerge-Opus-4.6-Distill",
      "tags": [
        "gguf",
        "llama.cpp",
        "unsloth",
        "qwen3_5",
        "frankenmerge",
        "layer-surgery",
        "reasoning",
        "chain-of-thought",
        "unsloth",
        "abliterated"
      ],
      "language": [
        "en"
      ],
      "datasets": [
        "Jackrong/Qwen3.5-reasoning-700x",
        "nohurry/Opus-4.6-Reasoning-3000x-filtered"
      ]
    },
    "hero_image_url": "https://raw.githubusercontent.com/unslothai/unsloth/main/images/unsloth%20made%20with%20love.png",
    "summary": "| Category | Base (Qwen3.5-9B-Base-Q8_0) | Qwen-3.5-10B-Frankenmerge-Opus-4.6-Distill | Δ | |---|---|---|---| | Factual Knowledge | 85.0% B | 85.0% B | = | | Reasoning | 88.0% B | 60.0% C | ↓ −28.0% | | Coding | 56.0% D | 80.0% B | ↑ +24.0% | | Instruction Following | 100.0% A | 30.0% F | ↓ −70.0% | | Language | 100.0% A | 70.0% C | ↓ −30.0% | | Safety Calibration | 66.7% C | 66.7% C | = | | **Overall** | **82.4% B** | **65.6% C** | **↓ −16.8%** | > **Method:** Layer surgery on Qwen3.5-9B-Base-Q8_0 followed by fine-tuning. > Benchmarks run at temperature=0, seed=42 > Coding capability improved significantly (+24%) at the cost of instruction-following and language tasks --- This model was GGUF format using Unsloth. **Example usage**:",
    "quick_links": [],
    "benchmark_table_html": "",
    "readme_markdown": "---\nlicense: apache-2.0\nbase_model: JackBinary/Qwen-3.5-10B-Frankenmerge-Opus-4.6-Distill\ntags:\n  - gguf\n  - llama.cpp\n  - unsloth\n  - qwen3_5\n  - frankenmerge\n  - layer-surgery\n  - reasoning\n  - chain-of-thought\n  - unsloth\n  - abliterated\nlanguage:\n  - en\ndatasets:\n  - Jackrong/Qwen3.5-reasoning-700x\n  - nohurry/Opus-4.6-Reasoning-3000x-filtered\n---\n\n# Qwen3.5-10B-Frankenmerge-Opus-4.6-Distill\n\n| Category | Base (Qwen3.5-9B-Base-Q8_0) | Qwen-3.5-10B-Frankenmerge-Opus-4.6-Distill | Δ |\n|---|---|---|---|\n| Factual Knowledge | 85.0% B | 85.0% B | = |\n| Reasoning | 88.0% B | 60.0% C | ↓ −28.0% |\n| Coding | 56.0% D | 80.0% B | ↑ +24.0% |\n| Instruction Following | 100.0% A | 30.0% F | ↓ −70.0% |\n| Language | 100.0% A | 70.0% C | ↓ −30.0% |\n| Safety Calibration | 66.7% C | 66.7% C | = |\n| **Overall** | **82.4% B** | **65.6% C** | **↓ −16.8%** |\n\n> **Method:** Layer surgery on Qwen3.5-9B-Base-Q8_0 followed by fine-tuning.  \n> Benchmarks run at `temperature=0, seed=42`  \n> Coding capability improved significantly (+24%) at the cost of instruction-following and language tasks\n\n---\n\nThis model was GGUF format using [Unsloth](https://github.com/unslothai/unsloth).\n\n**Example usage**:\n- For text only LLMs:    `llama-cli -hf JackBinary/Qwen-3.5-10B-Frankenmerge-Opus-4.6-Distill-GGUF --jinja`\n- For multimodal models: `llama-mtmd-cli -hf JackBinary/Qwen-3.5-10B-Frankenmerge-Opus-4.6-Distill-GGUF --jinja`\n\n## Available Model files:\n- `Qwen-3.5-10B-Frankenmerge-Opus-4.6-Distill.Q6_K.gguf`\n- `Qwen-3.5-10B-Frankenmerge-Opus-4.6-Distill.Q8_0.gguf`\n- `Qwen-3.5-10B-Frankenmerge-Opus-4.6-Distill.Q4_K_M.gguf`\nThis was trained 2x faster with [Unsloth](https://github.com/unslothai/unsloth)\n[<img src=\"https://raw.githubusercontent.com/unslothai/unsloth/main/images/unsloth%20made%20with%20love.png\" width=\"200\"/>](https://github.com/unslothai/unsloth)\n\nA DIY frankenmerge of Qwen3.5-9B with duplicated reasoning layers, then fine-tuned on high-quality reasoning data. 36 layers instead of 32. ~10B parameters. Text-only, thinking mode supported.\n\n## What this is\n\nI took [llmfan46/Qwen3.5-9B-ultra-heretic](https://huggingface.co/llmfan46/Qwen3.5-9B-ultra-heretic) (an abliterated Qwen3.5-9B), duplicated layers 24-27 to give it an extra reasoning block, then trained it sequentially on two datasets to make the new layers earn their keep.\n\nThe original 9B has 32 layers arranged as 8 blocks of `DeltaNet × 3 + Attention × 1`. After surgery, it has 36 layers: 9 complete blocks. The duplicated block starts as an exact copy but diverges during training, giving the model more depth for complex reasoning without changing anything about the input/output behavior.\n\nAfter the merge, two rounds of SFT with high-rank LoRA (r=128, alpha=256):\n\n1. **Stage 1:** [Jackrong/Qwen3.5-reasoning-700x](https://huggingface.co/datasets/Jackrong/Qwen3.5-reasoning-700x) (633 examples) at LR 2e-4. Reasoning distillation from Qwen3.5-27B. Gets the frankenmerge coherent and stabilizes the duplicated layers.\n2. **Stage 2:** [nohurry/Opus-4.6-Reasoning-3000x-filtered](https://huggingface.co/datasets/nohurry/Opus-4.6-Reasoning-3000x-filtered) (~3000 examples) at LR 5e-5. Claude Opus 4.6 reasoning traces. Strengthens the model's actual problem-solving ability.\n\n## Why frankenmerge + train?\n\n[David Noel Ng's RYS work](https://dnhkng.github.io/posts/rys) showed you can top the Open LLM Leaderboard by duplicating middle \"reasoning\" layers of a model without changing a single weight. The idea: early layers handle input encoding, late layers handle output decoding, and the middle layers do the actual thinking. Give the model more layers to think with, it thinks better.\n\n[RockTalk/Qwen3.5-9B-Franken-L24-27](https://huggingface.co/RockTalk/Qwen3.5-9B-Franken-L24-27) applied this to Qwen3.5-9B and showed improvements without any post-training. [A reddit post on layer surgery](https://www.reddit.com/r/LocalLLaMA/comments/1rvxmnh/i_spent_a_weekend_doing_layer_surgery_on_6/) explored similar ideas.\n\nThen I saw [Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled](https://huggingface.co/Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled), which showed that distilling structured reasoning from Claude Opus into Qwen3.5 massively reduces the overthinking/looping problem and makes the model more coherent and autonomous.\n\nSo the logic was: frankenmerge for extra capacity, then train the new capacity on high-quality reasoning data. Layer surgery gives you the architecture; SFT teaches the duplicated layers what to do with themselves.\n\n## The surgery, specifically\n\nQwen3.5-9B's 32 layers follow a repeating pattern:\n\n```\nBlock 0: layers  0- 3  (DeltaNet, DeltaNet, DeltaNet, Attention)\nBlock 1: layers  4- 7  (DeltaNet, DeltaNet, DeltaNet, Attention)\n...\nBlock 6: layers 24-27  (DeltaNet, DeltaNet, DeltaNet, Attention)  ← duplicated\nBlock 7: layers 28-31  (DeltaNet, DeltaNet, DeltaNet, Attention)\n```\n\nAfter surgery:\n\n```\nBlocks 0-6: layers  0-27  (original, unchanged)\nBlock 6':  layers 28-31  (deep copy of layers 24-27)\nBlock 7:   layers 32-35  (original layers 28-31, shifted)\n```\n\nThe copy is done with `copy.deepcopy` in PyTorch from clean bf16 weights. No quantization artifacts, no weight key remapping hacks.\n\n## Training details\n\n| | Stage 1 | Stage 2 |\n|---|---|---|\n| Dataset | Qwen3.5-reasoning-700x | Opus-4.6-Reasoning-3000x-filtered |\n| Examples | 633 | 2326 |\n| Learning rate | 2e-4 | 5e-5 |\n| Schedule | Cosine | Cosine |\n| Epochs | 1 | 1 |\n| Effective batch | 8 | 8 |\n| LoRA rank | 128 | 128 |\n| LoRA alpha | 256 | 256 |\n| RSLoRA | Yes | Yes |\n| Precision | bf16 | bf16 |\n\nTrained on a single G4 using Unsloth. Response-only masking (instruction tokens masked with -100). Sequential training: Stage 1 completes fully before Stage 2 begins. The LoRA adapters accumulate both stages.\n\n## Usage\n\n```python\nfrom transformers import AutoModelForCausalLM, AutoTokenizer\n\nmodel = AutoModelForCausalLM.from_pretrained(\n    \"YOUR_USERNAME/Qwen3.5-9B-Franken-L24-27-Reasoning\",\n    torch_dtype=\"auto\",\n    device_map=\"auto\",\n    trust_remote_code=True,\n)\ntokenizer = AutoTokenizer.from_pretrained(\n    \"YOUR_USERNAME/Qwen3.5-9B-Franken-L24-27-Reasoning\",\n    trust_remote_code=True,\n)\n\nmessages = [{\"role\": \"user\", \"content\": \"Prove that the square root of 2 is irrational.\"}]\ntext = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)\ninputs = tokenizer([text], return_tensors=\"pt\").to(model.device)\noutputs = model.generate(**inputs, max_new_tokens=2048, temperature=0.7, top_p=0.8, top_k=20)\nprint(tokenizer.decode(outputs[0], skip_special_tokens=True))\n```\n\n## Acknowledgments\n\nThis model wouldn't exist without the work of:\n\n- **[David Noel Ng (dnhkng)](https://dnhkng.github.io/posts/rys)** for the RYS research proving layer duplication works, and for writing such a clear explanation of the \"LLM neuroanatomy\" concept\n- **[RockTalk](https://huggingface.co/RockTalk/Qwen3.5-9B-Franken-L24-27)** for demonstrating the frankenmerge on Qwen3.5-9B specifically (even though the weights turned out to be 4-bit under the hood, the idea was sound)\n- **[Jackrong](https://huggingface.co/Jackrong/Qwen3.5-27B-Claude-4.6-Opus-Reasoning-Distilled)** for both the Opus-distilled model showing how well reasoning distillation works on Qwen3.5, and for the [Qwen3.5-reasoning-700x](https://huggingface.co/datasets/Jackrong/Qwen3.5-reasoning-700x) dataset\n- **[nohurry](https://huggingface.co/datasets/nohurry/Opus-4.6-Reasoning-3000x-filtered)** for the filtered Opus 4.6 reasoning dataset\n- **[llmfan46](https://huggingface.co/llmfan46/Qwen3.5-9B-ultra-heretic)** for the ultra-heretic abliteration, which gave me a clean, uncensored base to build on\n- **[r/LocalLLaMA](https://www.reddit.com/r/LocalLLaMA/comments/1rvxmnh/i_spent_a_weekend_doing_layer_surgery_on_6/)** for the collective insanity that makes all of this happen\n- The **Qwen team** at Alibaba for the base Qwen3.5 architecture\n- **[Unsloth](https://unsloth.ai/)** for making training on a single GPU actually feasible\n\n## License\n\nApache 2.0, same as the base Qwen3.5 model.\n",
    "related_quantizations": []
  },
  "tags": [
    "gguf",
    "qwen3_5_text",
    "llama.cpp",
    "unsloth",
    "qwen3_5",
    "frankenmerge",
    "layer-surgery",
    "reasoning",
    "chain-of-thought",
    "abliterated",
    "en",
    "dataset:Jackrong/Qwen3.5-reasoning-700x",
    "dataset:nohurry/Opus-4.6-Reasoning-3000x-filtered",
    "base_model:JackBinary/Qwen-3.5-10B-Frankenmerge-Opus-4.6-Distill",
    "base_model:quantized:JackBinary/Qwen-3.5-10B-Frankenmerge-Opus-4.6-Distill",
    "license:apache-2.0",
    "endpoints_compatible",
    "region:us",
    "conversational"
  ],
  "likes": 0,
  "downloads": 576,
  "gated": false,
  "private": false,
  "last_modified": "2026-03-19T18:49:55.000Z",
  "created_at": "2026-03-17T19:58:31.000Z",
  "pipeline_tag": "",
  "library_name": ""
}
Source payload excerpt (from Hugging Face API)
{
  "_id": "69b9b26770d1dd20c6c9a2d7",
  "id": "JackBinary/Qwen-3.5-10B-Frankenmerge-Opus-4.6-Distill-GGUF",
  "modelId": "JackBinary/Qwen-3.5-10B-Frankenmerge-Opus-4.6-Distill-GGUF",
  "sha": "392b0c566ffbde26eac8463b36c12e98f7e56f95",
  "createdAt": "2026-03-17T19:58:31.000Z",
  "lastModified": "2026-03-19T18:49:55.000Z",
  "author": "JackBinary",
  "downloads": 576,
  "likes": 0,
  "gated": false,
  "private": false,
  "pipeline_tag": "",
  "library_name": "",
  "siblings_count": 6
}