GraySoft
Projects Models About FAQ Contact Download guIDE →

jackrong/gpt-distill-qwen3-8b-thinking-gguf Q3_K_L GGUF - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.

Model Intelligence Sheet

jackrong/gpt-distill-qwen3-8b-thinking-gguf overview

!Base Model !License !Language !Unsloth !Context !Parameters !Reasoning !Feature !Distillation

ggufunslothsftchain-of-thoughtreasoningqwengenerated_from_trainerdataset:Jackrong/Natural-Reasoning-gpt-oss-120B-S1dataset:Jackrong/Chinese-Qwen3-235B-Thinking-2507-Distill-100kdataset:Jackrong/ShareGPT-gpt-oss-120B-reasoningdataset:Jackrong/gpt-oss-120b-Reasoning-Instructionbase_model:Qwen/Qwen3-8Bbase_model:quantized:Qwen/Qwen3-8Blicense:apache-2.0endpoints_compatibleregion:usconversational
jackrong/gpt-distill-qwen3-8b-thinking-gguf visual
Downloads
239
Likes
0
Pipeline
Library
Visibility
Public
Access
Open

Repository Files & Downloads

12 files detected
Direct downloads for all repository files
FileTypeQuantizationSizeLink
GPT-Distill-Qwen3-8B-Thinking-IQ4_XS.gguf GGUF IQ4_XS 4.28 GB Download
GPT-Distill-Qwen3-8B-Thinking-Q2_K.gguf GGUF Q2_K 3.06 GB Download
GPT-Distill-Qwen3-8B-Thinking-Q3_K_L.gguf GGUF Q3_K_L 4.13 GB Download
GPT-Distill-Qwen3-8B-Thinking-Q3_K_M.gguf GGUF Q3_K_M 3.84 GB Download
GPT-Distill-Qwen3-8B-Thinking-Q3_K_S.gguf GGUF Q3_K_S 3.51 GB Download
GPT-Distill-Qwen3-8B-Thinking-Q4_K_M.gguf GGUF Q4_K_M 4.68 GB Download
GPT-Distill-Qwen3-8B-Thinking-Q4_K_S.gguf GGUF Q4_K_S 4.47 GB Download
GPT-Distill-Qwen3-8B-Thinking-Q5_K_M.gguf GGUF Q5_K_M 5.45 GB Download
GPT-Distill-Qwen3-8B-Thinking-Q5_K_S.gguf GGUF Q5_K_S 5.33 GB Download
GPT-Distill-Qwen3-8B-Thinking-Q6_K.gguf GGUF Q6_K 6.26 GB Download
GPT-Distill-Qwen3-8B-Thinking-Q8_0.gguf GGUF 8.11 GB Download
GPT-Distill-Qwen3-8B-Thinking-f16.gguf GGUF F16 15.26 GB Download

Model Details Live

Model Slug
jackrong/gpt-distill-qwen3-8b-thinking-gguf
Author
Jackrong
Pipeline Task
Library
Created
2025-11-30
Last Modified
2025-11-30
Gated
No
Private
No
HF SHA
7d240a33f7ccb427b017b7618fcdbccb6bf0e208
License
apache-2.0
Language
Unknown
Base Model
Qwen/Qwen3-8B

Metadata Inspector

Normalized metadata (stored in metadata_json)
{
  "metadata": {},
  "card_data": {
    "license": "apache-2.0",
    "base_model": "Qwen/Qwen3-8B",
    "tags": [
      "unsloth",
      "sft",
      "chain-of-thought",
      "reasoning",
      "qwen",
      "generated_from_trainer"
    ],
    "datasets": [
      "Jackrong/Natural-Reasoning-gpt-oss-120B-S1",
      "Jackrong/Chinese-Qwen3-235B-Thinking-2507-Distill-100k",
      "Jackrong/ShareGPT-gpt-oss-120B-reasoning",
      "Jackrong/gpt-oss-120b-Reasoning-Instruction"
    ],
    "model-index": [
      {
        "name": "GPT-Distill-Qwen3-8B-Thinking",
        "results": []
      }
    ],
    "frontmatter": {
      "license": "apache-2.0",
      "base_model": "Qwen/Qwen3-8B",
      "tags": [
        "unsloth",
        "sft",
        "chain-of-thought",
        "reasoning",
        "qwen",
        "generated_from_trainer"
      ],
      "datasets": [
        "Jackrong/Natural-Reasoning-gpt-oss-120B-S1",
        "Jackrong/Chinese-Qwen3-235B-Thinking-2507-Distill-100k",
        "Jackrong/ShareGPT-gpt-oss-120B-reasoning",
        "Jackrong/gpt-oss-120b-Reasoning-Instruction",
        "name: GPT-Distill-Qwen3-8B-Thinking"
      ]
    },
    "hero_image_url": "https://cdn-uploads.huggingface.co/production/uploads/66309bd090589b7c65950665/5izo9jTYcx-QuVZEKwFgE.png",
    "summary": "!Base Model !License !Language !Unsloth !Context !Parameters !Reasoning !Feature !Distillation",
    "quick_links": [],
    "benchmark_table_html": "",
    "readme_markdown": "---\nlicense: apache-2.0\nbase_model: Qwen/Qwen3-8B\ntags:\n- unsloth\n- sft\n- chain-of-thought\n- reasoning\n- qwen\n- generated_from_trainer\ndatasets:\n- Jackrong/Natural-Reasoning-gpt-oss-120B-S1\n- Jackrong/Chinese-Qwen3-235B-Thinking-2507-Distill-100k\n- Jackrong/ShareGPT-gpt-oss-120B-reasoning\n- Jackrong/gpt-oss-120b-Reasoning-Instruction\nmodel-index:\n- name: GPT-Distill-Qwen3-8B-Thinking\n  results: []\n---\n\n# GPT-Distill-Qwen3-8B-Thinking\n![Base Model](https://img.shields.io/badge/Base_Model-Qwen3--8B-0088CC?style=flat)\n![License](https://img.shields.io/badge/License-Apache_2.0-green?style=flat)\n![Language](https://img.shields.io/badge/Language-English_%7C_Chinese-blue?style=flat)\n![Unsloth](https://img.shields.io/badge/Unsloth-Fine--Tuning-blue?style=flat&logo=unsloth)\n![Context](https://img.shields.io/badge/Context-16K_Tokens-success?style=flat)\n![Parameters](https://img.shields.io/badge/Parameters-8B-lightgrey?style=flat)\n![Reasoning](https://img.shields.io/badge/Task-Reasoning_%26_CoT-8A2BE2?style=flat)\n![Feature](https://img.shields.io/badge/Feature-Thinking_Process_%3Cthink%3E-FF4500?style=flat)\n![Distillation](https://img.shields.io/badge/Distilled_From-120B%2B_Teacher_Models-orange?style=flat)\n\n<div align=\"center\">\n  <img src=\"https://cdn-uploads.huggingface.co/production/uploads/66309bd090589b7c65950665/5izo9jTYcx-QuVZEKwFgE.png\" width=\"500\"/>\n  <img src=\"https://cdn-uploads.huggingface.co/production/uploads/66309bd090589b7c65950665/n53llfb_o1mDsY2k007lk.jpeg\" width=\"800\"/>\n</div>\n\n\n## 1. Model Overview\n* **Model Name:** `Jackrong/GPT-Distill-Qwen3-8B-Thinking`\n* **Model Type:** Instruction-tuned & Reasoning-enhanced LLM\n* **Base Model:** `Qwen/Qwen3-8B`\n* **Parameters:** ~8B\n* **Context Length:** Up to **16K tokens** (`max_seq_length = 16384`)\n* **Supported Languages:** Chinese, English, Mixed inputs/outputs\n* **Training Method:**\n    * Supervised Fine-Tuning (SFT) using **Unsloth** (LoRA adapter merged to 16bit)\n    * Knowledge Distillation from large-scale reasoning models (120B/235B class)\n    * **Thinking/CoT Integration:** Explicitly trained to generate internal reasoning chains wrapped in `<think>...</think>` tags.\n\n### Description\nThis model is a specialized fine-tune of **Qwen3-8B**, designed to excel at complex reasoning and instruction following. It utilizes a **16k context window** and has been distilled from high-intelligence teacher models (GPT-OSS-120B and Qwen3-235B).\n\nA key feature is its **\"Thinking\" capability**: the model is trained to output a Chain-of-Thought (CoT) process before providing the final answer, significantly improving performance on math, logic, and scientific tasks.\n\n---\n\n## 2. Intended Use Cases\n\n### ✅ Recommended:\n* **Complex Reasoning:** Math problems, logical puzzles, and scientific derivations using the `<think>` mechanism.\n* **Long-Context Tasks:** Processing documents or conversations up to 16k tokens.\n* **Instruction Following:** High adherence to complex user constraints.\n* **Chinese/English NLP:** Fluent generation in both languages, including cultural nuances.\n* **Knowledge Distillation:** Acting as a lightweight student model that mimics the reasoning patterns of 100B+ parameter models.\n\n### ⚠️ Not Suitable For:\n* **High-risk decision-making:** Medical diagnosis, legal advice without professional oversight.\n* **Real-time Factual Updates:** The model's knowledge is static and based on the training data cutoff.\n\n> **Note:** To trigger the reasoning capabilities effectively, the model may spontaneously use `<think>` tags, or you can prompt it to \"think step-by-step\".\n\n---\n\n## 3. Training Data & Distillation Process\n\nThe model was trained on a curated mix of **~88,000 high-quality examples**, filtered for length and quality.\n\n### Key Datasets\n\n#### (1) Reasoning & Thinking (CoT)\n* **Sources:** `Jackrong/Natural-Reasoning-gpt-oss-120B-S1` & `Jackrong/Chinese-Qwen3-235B-Thinking-2507-Distill-100k`\n* **Format:** `Input` -> `<think> Reasoning Chain </think>` -> `Answer`\n* **Purpose:** Distills the \"thinking process\" of massive models (120B/235B) into the 8B architecture, teaching the model *how* to solve problems, not just the answer.\n\n#### (2) ShareGPT & Conversation\n* **Source:** `Jackrong/ShareGPT-gpt-oss-120B-reasoning`\n* **Purpose:** Ensures natural multi-turn conversational flow and user intent understanding.\n\n#### (3) Instruction Following\n* **Source:** `Jackrong/gpt-oss-120b-Reasoning-Instruction`\n* **Purpose:** Enhances the ability to follow specific formatting and constraint instructions.\n\n### Training Configuration\n* **Framework:** Unsloth + TRL (SFTTrainer)\n* **Hardware:** NVIDIA H100 80GB\n* **Batch Size:** Global batch size of 32 (per_device=4, grad_accum=8)\n* **Optimization:** AdamW 8-bit, Learning Rate `2e-5`\n* **LoRA Config:** Rank `r=32`, Alpha `32`, target modules = all linear layers\n* **Data Strategy:** Trained on responses only (`train_on_responses_only`) to strictly model assistant behavior.\n\n---\n\n## 4. Key Features Summary\n\n| Feature | Description |\n| :--- | :--- |\n| **Thinking Process** | Embeds CoT reasoning in `<think>` tags for explainable outputs. |\n| **Distilled Intelligence** | Inherits reasoning patterns from 120B+ parameter teacher models. |\n| **Efficient 8B Size** | High performance with low VRAM usage (optimized via Unsloth). |\n| **Long Context** | 16,384 token context window for extensive document processing. |\n\n---\n\n## 5. Acknowledgements\n\nWe thank:\n* **The Unsloth Team** for their efficient fine-tuning library that made this training possible.\n* **Qwen Team** for the powerful Qwen3-8B base model.\n* **Jackrong** for the curation of the distillation datasets (ShareGPT-OSS, Natural-Reasoning, etc.).\n\n*This project is an open research effort aimed at democratizing high-level reasoning capabilities in smaller, accessible language models.*",
    "related_quantizations": []
  },
  "tags": [
    "gguf",
    "unsloth",
    "sft",
    "chain-of-thought",
    "reasoning",
    "qwen",
    "generated_from_trainer",
    "dataset:Jackrong/Natural-Reasoning-gpt-oss-120B-S1",
    "dataset:Jackrong/Chinese-Qwen3-235B-Thinking-2507-Distill-100k",
    "dataset:Jackrong/ShareGPT-gpt-oss-120B-reasoning",
    "dataset:Jackrong/gpt-oss-120b-Reasoning-Instruction",
    "base_model:Qwen/Qwen3-8B",
    "base_model:quantized:Qwen/Qwen3-8B",
    "license:apache-2.0",
    "endpoints_compatible",
    "region:us",
    "conversational"
  ],
  "likes": 0,
  "downloads": 239,
  "gated": false,
  "private": false,
  "last_modified": "2025-11-30T15:14:51.000Z",
  "created_at": "2025-11-30T08:41:39.000Z",
  "pipeline_tag": "",
  "library_name": ""
}
Source payload excerpt (from Hugging Face API)
{
  "_id": "692c034395512abda396b62b",
  "id": "Jackrong/GPT-Distill-Qwen3-8B-Thinking-GGUF",
  "modelId": "Jackrong/GPT-Distill-Qwen3-8B-Thinking-GGUF",
  "sha": "7d240a33f7ccb427b017b7618fcdbccb6bf0e208",
  "createdAt": "2025-11-30T08:41:39.000Z",
  "lastModified": "2025-11-30T15:14:51.000Z",
  "author": "Jackrong",
  "downloads": 239,
  "likes": 0,
  "gated": false,
  "private": false,
  "pipeline_tag": "",
  "library_name": "",
  "siblings_count": 14
}