GraySoft
Projects Models About FAQ Contact Download guIDE →

avalon2244/qwen3.5-4b-claude-opus-4.6-distilled-gguf Q4_K_M GGUF - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.

Model Intelligence Sheet

avalon2244/qwen3.5-4b-claude-opus-4.6-distilled-gguf overview

This terribly named model was a quick finetune of Qwen3.5-4B on the nohurry/Opus-4.6-Reasoning-3000x-filtered dataset. It tends to have cleaner reasoning traces than the original Qwen3.5-4B, and is around as accurate. I haven't tested it, though. This model was finetuned and converted to GGUF format using Unsloth. It's a bit inconsistent with reasoning. It's far less likely to enter endless loops, and is uses far fewer tokens than the original model. But it's still a 4B model that's been finetuned on one dataset, so it's not fantastic. Example usage:

ggufqwen3_5llama.cppunslothvision-language-modeldataset:nohurry/Opus-4.6-Reasoning-3000x-filteredbase_model:Qwen/Qwen3.5-4Bbase_model:quantized:Qwen/Qwen3.5-4Bendpoints_compatibleregion:usconversational
avalon2244/qwen3.5-4b-claude-opus-4.6-distilled-gguf visual
Downloads
1,560
Likes
9
Pipeline
Library
Visibility
Public
Access
Open

Repository Files & Downloads

4 files detected
Direct downloads for all repository files
FileTypeQuantizationSizeLink
Qwen3.5-4B.BF16-mmproj.gguf GGUF BF16 644.27 MB Download
Qwen3.5-4B.Q4_K_M.gguf GGUF Q4_K_M 2.52 GB Download
Qwen3.5-4B.Q5_K_M.gguf GGUF Q5_K_M 2.90 GB Download
Qwen3.5-4B.Q8_0.gguf GGUF 4.17 GB Download

Model Details Live

Model Slug
avalon2244/qwen3.5-4b-claude-opus-4.6-distilled-gguf
Author
avalon2244
Pipeline Task
Library
Created
2026-03-04
Last Modified
2026-03-06
Gated
No
Private
No
HF SHA
31aa833697c069c7d7e937cb7c574c74c3d6b065
License
Unknown
Language
Unknown
Base Model
unsloth/Qwen3.5-4B, Qwen/Qwen3.5-4B

Metadata Inspector

Normalized metadata (stored in metadata_json)
{
  "metadata": {},
  "card_data": {
    "tags": [
      "gguf",
      "llama.cpp",
      "unsloth",
      "vision-language-model"
    ],
    "datasets": [
      "nohurry/Opus-4.6-Reasoning-3000x-filtered"
    ],
    "base_model": [
      "unsloth/Qwen3.5-4B",
      "Qwen/Qwen3.5-4B"
    ],
    "frontmatter": {
      "tags": [
        "gguf",
        "llama.cpp",
        "unsloth",
        "vision-language-model"
      ],
      "datasets": [
        "nohurry/Opus-4.6-Reasoning-3000x-filtered"
      ],
      "base_model": [
        "unsloth/Qwen3.5-4B",
        "Qwen/Qwen3.5-4B"
      ]
    },
    "hero_image_url": "https://raw.githubusercontent.com/unslothai/unsloth/main/images/unsloth%20made%20with%20love.png",
    "summary": "This terribly named model was a quick finetune of Qwen3.5-4B on the nohurry/Opus-4.6-Reasoning-3000x-filtered dataset. It tends to have cleaner reasoning traces than the original Qwen3.5-4B, and is around as accurate. I haven't tested it, though. This model was finetuned and converted to GGUF format using Unsloth. It's a bit inconsistent with reasoning. It's far less likely to enter endless loops, and is uses far fewer tokens than the original model. But it's still a 4B model that's been finetuned on one dataset, so it's not fantastic. **Example usage**:",
    "quick_links": [],
    "benchmark_table_html": "",
    "readme_markdown": "---\ntags:\n- gguf\n- llama.cpp\n- unsloth\n- vision-language-model\ndatasets:\n- nohurry/Opus-4.6-Reasoning-3000x-filtered\nbase_model:\n- unsloth/Qwen3.5-4B\n- Qwen/Qwen3.5-4B\n---\n\n# Qwen3.5-4B-Claude-Opus-4.6-Distilled-GGUF\n\nThis terribly named model was a quick finetune of Qwen3.5-4B on the nohurry/Opus-4.6-Reasoning-3000x-filtered dataset. It tends to have cleaner reasoning traces than the original Qwen3.5-4B, and is around as accurate. I haven't tested it, though.\nThis model was finetuned and converted to GGUF format using [Unsloth](https://github.com/unslothai/unsloth).\n\nIt's a bit inconsistent with reasoning. It's far less likely to enter endless loops, and is uses far fewer tokens than the original model. But it's still a 4B model that's been finetuned on one dataset, so it's not fantastic.\n\n**Example usage**:\n- For text only LLMs:    `llama-cli -hf avalon2244/Qwen3.5-4B-Claude-Opus-4.6-Distilled-GGUF --jinja`\n- For multimodal models: `llama-mtmd-cli -hf avalon2244/Qwen3.5-4B-Claude-Opus-4.6-Distilled-GGUF --jinja`\n\n## Available Model files:\n- `Qwen3.5-4B.Q5_K_M.gguf`\n- `Qwen3.5-4B.Q8_0.gguf`\n- `Qwen3.5-4B.Q4_K_M.gguf`\n- `Qwen3.5-4B.BF16-mmproj.gguf`\nThis was trained 2x faster with [Unsloth](https://github.com/unslothai/unsloth)\n[<img src=\"https://raw.githubusercontent.com/unslothai/unsloth/main/images/unsloth%20made%20with%20love.png\" width=\"200\"/>](https://github.com/unslothai/unsloth)",
    "related_quantizations": []
  },
  "tags": [
    "gguf",
    "qwen3_5",
    "llama.cpp",
    "unsloth",
    "vision-language-model",
    "dataset:nohurry/Opus-4.6-Reasoning-3000x-filtered",
    "base_model:Qwen/Qwen3.5-4B",
    "base_model:quantized:Qwen/Qwen3.5-4B",
    "endpoints_compatible",
    "region:us",
    "conversational"
  ],
  "likes": 9,
  "downloads": 1560,
  "gated": false,
  "private": false,
  "last_modified": "2026-03-06T09:52:16.000Z",
  "created_at": "2026-03-04T06:09:30.000Z",
  "pipeline_tag": "",
  "library_name": ""
}
Source payload excerpt (from Hugging Face API)
{
  "_id": "69a7cc9ad48a908cb7eb2a16",
  "id": "avalon2244/Qwen3.5-4B-Claude-Opus-4.6-Distilled-GGUF",
  "modelId": "avalon2244/Qwen3.5-4B-Claude-Opus-4.6-Distilled-GGUF",
  "sha": "31aa833697c069c7d7e937cb7c574c74c3d6b065",
  "createdAt": "2026-03-04T06:09:30.000Z",
  "lastModified": "2026-03-06T09:52:16.000Z",
  "author": "avalon2244",
  "downloads": 1560,
  "likes": 9,
  "gated": false,
  "private": false,
  "pipeline_tag": "",
  "library_name": "",
  "siblings_count": 7
}