GraySoft
Projects Models About FAQ Contact Download guIDE →

jackrong/gpt-5-distill-qwen3-4b-instruct-gguf Q4_K_S GGUF - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.

Model Intelligence Sheet

jackrong/gpt-5-distill-qwen3-4b-instruct-gguf overview

!Base Model !Distillation !Language !Context !Format !License Model Type: Instruction-tuned conversational LLM Supports LoRA adapters and full-finetuned models for inference This model is trained on ShareGPT-Qwen3 instruction datasets and distilled toward the conversational style and quality of GPT-5. It aims to achieve high-quality, natural-sounding dialogues with low computational overhead—perfect for lightweight applications without sacrificing responsiveness. ---

ggufqwen3llama.cppenzhdataset:Jackrong/ShareGPT-Qwen3-235B-A22B-Instuct-2507base_model:Qwen/Qwen3-4B-Instruct-2507base_model:quantized:Qwen/Qwen3-4B-Instruct-2507license:apache-2.0endpoints_compatibleregion:usconversational
jackrong/gpt-5-distill-qwen3-4b-instruct-gguf visual
Downloads
3,200
Likes
17
Pipeline
Library
Visibility
Public
Access
Open

Repository Files & Downloads

12 files detected
Direct downloads for all repository files
FileTypeQuantizationSizeLink
GPT-5-Distill-Qwen3-4B-Instruct-IQ4_XS.gguf GGUF IQ4_XS 2.13 GB Download
GPT-5-Distill-Qwen3-4B-Instruct-Q2_K.gguf GGUF Q2_K 1.55 GB Download
GPT-5-Distill-Qwen3-4B-Instruct-Q3_K_L.gguf GGUF Q3_K_L 2.09 GB Download
GPT-5-Distill-Qwen3-4B-Instruct-Q3_K_M.gguf GGUF Q3_K_M 1.93 GB Download
GPT-5-Distill-Qwen3-4B-Instruct-Q3_K_S.gguf GGUF Q3_K_S 1.76 GB Download
GPT-5-Distill-Qwen3-4B-Instruct-Q4_K_S.gguf GGUF Q4_K_S 2.22 GB Download
GPT-5-Distill-Qwen3-4B-Instruct-Q5_K_M.gguf GGUF Q5_K_M 2.69 GB Download
GPT-5-Distill-Qwen3-4B-Instruct-Q5_K_S.gguf GGUF Q5_K_S 2.63 GB Download
GPT-5-Distill-Qwen3-4B-Instruct-Q6_K.gguf GGUF Q6_K 3.08 GB Download
GPT-5-Distill-Qwen3-4B-Instruct-f16.gguf GGUF F16 7.50 GB Download
qwen3-4b-instruct-2507.Q4_K_M.gguf GGUF Q4_K_M 2.33 GB Download
qwen3-4b-instruct-2507.Q8_0.gguf GGUF 3.99 GB Download

Model Details Live

Model Slug
jackrong/gpt-5-distill-qwen3-4b-instruct-gguf
Author
Jackrong
Pipeline Task
Library
Created
2025-11-21
Last Modified
2025-11-30
Gated
No
Private
No
HF SHA
6f08fdd42bc4d884cb031cab5d35939daf4841e5
License
apache-2.0
Language
en, zh
Base Model
Qwen/Qwen3-4B-Instruct-2507

Metadata Inspector

Normalized metadata (stored in metadata_json)
{
  "metadata": {},
  "card_data": {
    "tags": [
      "gguf",
      "llama.cpp"
    ],
    "license": "apache-2.0",
    "datasets": [
      "Jackrong/ShareGPT-Qwen3-235B-A22B-Instuct-2507"
    ],
    "language": [
      "en",
      "zh"
    ],
    "base_model": [
      "Qwen/Qwen3-4B-Instruct-2507"
    ],
    "frontmatter": {
      "tags": [
        "gguf",
        "llama.cpp"
      ],
      "license": "apache-2.0",
      "datasets": [
        "Jackrong/ShareGPT-Qwen3-235B-A22B-Instuct-2507"
      ],
      "language": [
        "en",
        "zh"
      ],
      "base_model": [
        "Qwen/Qwen3-4B-Instruct-2507"
      ]
    },
    "hero_image_url": "https://cdn-uploads.huggingface.co/production/uploads/66309bd090589b7c65950665/sk5gVFD15S0UNMek3gU0o.png",
    "summary": "!Base Model !Distillation !Language !Context !Format !License    **Model Type**: Instruction-tuned conversational LLM Supports LoRA adapters and full-finetuned models for inference This model is trained on ShareGPT-Qwen3 instruction datasets and distilled toward the conversational style and quality of GPT-5. It aims to achieve high-quality, natural-sounding dialogues with low computational overhead—perfect for lightweight applications without sacrificing responsiveness. ---",
    "quick_links": [],
    "benchmark_table_html": "",
    "readme_markdown": "---\ntags:\n- gguf\n- llama.cpp\nlicense: apache-2.0\ndatasets:\n- Jackrong/ShareGPT-Qwen3-235B-A22B-Instuct-2507\nlanguage:\n- en\n- zh\nbase_model:\n- Qwen/Qwen3-4B-Instruct-2507\n---\n\n\n# GPT-5-Distill-Qwen3-4B-Instruct-2507\n\n![Base Model](https://img.shields.io/badge/Base_Model-Qwen3--4B--Instruct-0088CC?style=flat)\n![Distillation](https://img.shields.io/badge/Distillation-GPT--5_Responses-8A2BE2?style=flat)\n![Language](https://img.shields.io/badge/Language-English_%7C_Chinese-blue?style=flat)\n![Context](https://img.shields.io/badge/Context-32K_Tokens-success?style=flat)\n![Format](https://img.shields.io/badge/Format-GGUF_%7C_llama.cpp-yellow?style=flat)\n![License](https://img.shields.io/badge/License-Apache_2.0-green?style=flat)\n\n<img src=\"https://cdn-uploads.huggingface.co/production/uploads/66309bd090589b7c65950665/sk5gVFD15S0UNMek3gU0o.png\" width=\"800\"/>\n\n<img src=\"https://cdn-uploads.huggingface.co/production/uploads/66309bd090589b7c65950665/vGzi5hSHJJ72ysJuM5EAv.png\" width=\"800\"/>\n\n<img src=\"https://cdn-uploads.huggingface.co/production/uploads/66309bd090589b7c65950665/j39PSDVoQmK4EI9pLANpa.png\" width=\"800\"/>\n\n**Model Type**: Instruction-tuned conversational LLM  \n  Supports LoRA adapters and full-finetuned models for inference\n- **Base Model**: `Qwen/Qwen3-4B-Instruct-2507`\n- **Parameters**:  4B \n- **Training Method**:\n  - Supervised Fine-Tuning (SFT) on ShareGPT data\n  - Knowledge distillation from LMSYS GPT-5 responses\n- **Supported Languages**: Chinese, English, mixed inputs/outputs\n- **Max Context Length**: Up to **32K tokens** (`max_seq_length = 32768`)\n\nThis model is trained on ShareGPT-Qwen3 instruction datasets and distilled toward the conversational style and quality of GPT-5. It aims to achieve high-quality, natural-sounding dialogues with low computational overhead—perfect for lightweight applications without sacrificing responsiveness.\n\n---\n\n## 2. Intended Use Cases\n\n### ✅ Recommended:\n\n- Casual chat in Chinese/English\n- General knowledge explanations & reasoning guidance\n- Code suggestions and simple debugging tips\n- Writing assistance: editing, summarizing, rewriting\n- Role-playing conversations (with well-designed prompts)\n\n### ⚠️ Not Suitable For:\n\n- High-risk decision-making:\n  - Medical diagnosis, mental health support\n  - Legal advice, financial investment recommendations\n- Real-time factual tasks (e.g., news, stock updates)\n- Authoritative judgment on sensitive topics\n\n> **Note**: Outputs are for reference only and not intended as the sole basis for critical decisions.\n\n---\n\n## 3. Training Data & Distillation Process\n\n### Key Datasets:\n\n#### (1) ds1: ShareGPT-Qwen3 Instruction Dataset  \n- Source: `Jackrong/ShareGPT-Qwen3-235B-A22B-Instuct-2507`  \n- Purpose:\n  - Provides diverse instruction-response pairs\n  - Supports multi-turn dialogues and context awareness\n- Processing:\n  - Cleaned for quality and relevance\n  - Standardized into `instruction`, `input`, `output` format\n\n#### (2) ds2: LMSYS GPT-5 Teacher Response Data  \n- Source: `ytz20/LMSYS-Chat-GPT-5-Chat-Response`  \n- Filtering:\n  - Only kept samples with `flaw == \"normal\"`\n  - Removed hallucinations and inconsistent responses\n- Purpose:\n  - Distillation target for conversational quality\n  - Enhances clarity, coherence, and fluency\n\n### Training Flow:\n\n1. Prepare unified Chat-formatted dataset\n2. Fine-tune base Qwen3-4B-Instruct-2507 via SFT\n3. Conduct knowledge distillation using GPT-5's normal responses as teacher outputs\n4. Balance style imitation with semantic fidelity to ensure robustness\n\n> ⚖️ **Note**: This work is based on publicly available, non-sensitive datasets and uses them responsibly under fair use principles.\n\n---\n\n## 4. Key Features Summary\n\n| Feature | Description |\n|--------|-------------|\n| **Lightweight** | ~4B parameter model – fast inference, low resource usage |\n| **Distillation-Style Responses** | Mimics GPT-5’s conversational fluency and helpfulness |\n| **Highly Conversational** | Excellent for chatbot-style interactions with rich dialogue flow |\n| **Multilingual Ready** | Seamless support for Chinese and English |\n\n---\n\n## 5. Acknowledgements\n\nWe thank:\n- LMSYS team for sharing GPT-5 response data\n- Jackrong for the ShareGPT-Qwen3 dataset\n- Qwen team for releasing `Qwen3-4B-Instruct`\n\nThis project is an open research effort aimed at making high-quality conversational AI accessible with smaller models.\n\n---",
    "related_quantizations": []
  },
  "tags": [
    "gguf",
    "qwen3",
    "llama.cpp",
    "en",
    "zh",
    "dataset:Jackrong/ShareGPT-Qwen3-235B-A22B-Instuct-2507",
    "base_model:Qwen/Qwen3-4B-Instruct-2507",
    "base_model:quantized:Qwen/Qwen3-4B-Instruct-2507",
    "license:apache-2.0",
    "endpoints_compatible",
    "region:us",
    "conversational"
  ],
  "likes": 17,
  "downloads": 3200,
  "gated": false,
  "private": false,
  "last_modified": "2025-11-30T15:16:56.000Z",
  "created_at": "2025-11-21T10:26:36.000Z",
  "pipeline_tag": "",
  "library_name": ""
}
Source payload excerpt (from Hugging Face API)
{
  "_id": "69203e5c91bc5b1c155a44d1",
  "id": "Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF",
  "modelId": "Jackrong/GPT-5-Distill-Qwen3-4B-Instruct-GGUF",
  "sha": "6f08fdd42bc4d884cb031cab5d35939daf4841e5",
  "createdAt": "2025-11-21T10:26:36.000Z",
  "lastModified": "2025-11-30T15:16:56.000Z",
  "author": "Jackrong",
  "downloads": 3200,
  "likes": 17,
  "gated": false,
  "private": false,
  "pipeline_tag": "",
  "library_name": "",
  "siblings_count": 16
}