GraySoft
Projects Models About FAQ Contact Download guIDE →

cstr/qwen3-asr-1.7b-gguf q8_0 GGUF - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.

Model Intelligence Sheet

cstr/qwen3-asr-1.7b-gguf overview

GGUF conversions of Qwen/Qwen3-ASR-1.7B, the larger of the two Qwen3-ASR speech-LLMs (Whisper-style audio encoder + 1.7B Qwen3 LLM head). Runs through CrispASR's --backend qwen3 path with no architecture changes — the same C++ qwen3 backend that already supported the 0.6B variant loads the 1.7B directly because the converter and the runtime both read sizes from GGUF metadata rather than hardcoding them.

ggmlggufasrspeech-to-textaudiotranscriptionspeech-llmqwen3automatic-speech-recognitionzhenyueardefresptiditkoruthvijatrhimsnlsvda
cstr/qwen3-asr-1.7b-gguf visual
Downloads
387
Likes
1
Pipeline
automatic-speech-recognition
Library
ggml
Visibility
Public
Access
Open

Repository Files & Downloads

3 files detected
Direct downloads for all repository files
FileTypeQuantizationSizeLink
qwen3-asr-1.7b-f16.gguf GGUF F16 4.38 GB Download
qwen3-asr-1.7b-q4_k.gguf GGUF Q4_K 1.24 GB Download
qwen3-asr-1.7b-q8_0.gguf GGUF 2.33 GB Download

Model Details Live

Model Slug
cstr/qwen3-asr-1.7b-gguf
Author
cstr
Pipeline Task
automatic-speech-recognition
Library
ggml
Created
2026-04-11
Last Modified
2026-04-11
Gated
No
Private
No
HF SHA
4593c9d871d43870c73b3500723399409f31a3de
License
apache-2.0
Language
zh, en, yue, ar, de, fr, es, pt, id, it, ko, ru, th, vi, ja, tr, hi, ms, nl, sv, da, fi, pl, cs, fil, fa, el, ro, hu, mk
Base Model
Qwen/Qwen3-ASR-1.7B

Metadata Inspector

Normalized metadata (stored in metadata_json)
{
  "metadata": {},
  "card_data": {
    "license": "apache-2.0",
    "base_model": "Qwen/Qwen3-ASR-1.7B",
    "language": [
      "zh",
      "en",
      "yue",
      "ar",
      "de",
      "fr",
      "es",
      "pt",
      "id",
      "it",
      "ko",
      "ru",
      "th",
      "vi",
      "ja",
      "tr",
      "hi",
      "ms",
      "nl",
      "sv",
      "da",
      "fi",
      "pl",
      "cs",
      "fil",
      "fa",
      "el",
      "ro",
      "hu",
      "mk"
    ],
    "tags": [
      "asr",
      "speech-to-text",
      "audio",
      "transcription",
      "gguf",
      "ggml",
      "speech-llm",
      "qwen3"
    ],
    "pipeline_tag": "automatic-speech-recognition",
    "library_name": "ggml",
    "frontmatter": {
      "license": "apache-2.0",
      "base_model": "Qwen/Qwen3-ASR-1.7B",
      "language": [
        "zh",
        "en",
        "yue",
        "ar",
        "de",
        "fr",
        "es",
        "pt",
        "id",
        "it",
        "ko",
        "ru",
        "th",
        "vi",
        "ja",
        "tr",
        "hi",
        "ms",
        "nl",
        "sv",
        "da",
        "fi",
        "pl",
        "cs",
        "fil",
        "fa",
        "el",
        "ro",
        "hu",
        "mk"
      ],
      "tags": [
        "asr",
        "speech-to-text",
        "audio",
        "transcription",
        "gguf",
        "ggml",
        "speech-llm",
        "qwen3"
      ],
      "pipeline_tag": "automatic-speech-recognition",
      "library_name": "ggml"
    },
    "hero_image_url": "",
    "summary": "GGUF conversions of Qwen/Qwen3-ASR-1.7B, the larger of the two Qwen3-ASR speech-LLMs (Whisper-style audio encoder + 1.7B Qwen3 LLM head). Runs through CrispASR's --backend qwen3 path with no architecture changes — the same C++ qwen3 backend that already supported the 0.6B variant loads the 1.7B directly because the converter and the runtime both read sizes from GGUF metadata rather than hardcoding them.",
    "quick_links": [],
    "benchmark_table_html": "",
    "readme_markdown": "---\nlicense: apache-2.0\nbase_model: Qwen/Qwen3-ASR-1.7B\nlanguage:\n  - zh\n  - en\n  - yue\n  - ar\n  - de\n  - fr\n  - es\n  - pt\n  - id\n  - it\n  - ko\n  - ru\n  - th\n  - vi\n  - ja\n  - tr\n  - hi\n  - ms\n  - nl\n  - sv\n  - da\n  - fi\n  - pl\n  - cs\n  - fil\n  - fa\n  - el\n  - ro\n  - hu\n  - mk\ntags:\n  - asr\n  - speech-to-text\n  - audio\n  - transcription\n  - gguf\n  - ggml\n  - speech-llm\n  - qwen3\npipeline_tag: automatic-speech-recognition\nlibrary_name: ggml\n---\n\n# Qwen3-ASR-1.7B — GGUF (CrispASR)\n\nGGUF conversions of [`Qwen/Qwen3-ASR-1.7B`](https://huggingface.co/Qwen/Qwen3-ASR-1.7B), the larger of the two Qwen3-ASR speech-LLMs (Whisper-style audio encoder + 1.7B Qwen3 LLM head). Runs through [CrispASR](https://github.com/CrispStrobe/CrispASR)'s `--backend qwen3` path with no architecture changes — the same C++ qwen3 backend that already supported the 0.6B variant loads the 1.7B directly because the converter and the runtime both read sizes from GGUF metadata rather than hardcoding them.\n\n## What's in the box\n\n| File | Size | Quantization | Notes |\n|---|---|---|---|\n| `qwen3-asr-1.7b-f16.gguf` | 4.71 GB | F16 | Reference precision; matches PyTorch bfloat16 within float-noise tolerance |\n| `qwen3-asr-1.7b-q8_0.gguf` | 2.51 GB | Q8_0 | Effectively lossless; same transcript on jfk.wav as F16 |\n| `qwen3-asr-1.7b-q4_k.gguf` | 1.33 GB | Q4_K | 3.5× compressed; minor punctuation differences on tested clips |\n\nAll three contain:\n\n* The full audio encoder (24 layers, d_model 1024, 16 heads, 4096 ff)\n* The Qwen3 1.7B LLM (28 layers, d_model 2048, 16 heads / 8 KV heads, 6144 ff, 152K vocab, RoPE θ=1e6)\n* The full GPT-2-style BPE vocab + merges, the audio mel filterbank, and the Hann window — everything the C++ side needs to run inference end-to-end without re-fetching the original safetensors.\n\n## Use with CrispASR (no Python at runtime)\n\n```bash\n# Build crispasr (one-time, no Python deps for runtime — only the converter\n# was Python).\ngit clone https://github.com/CrispStrobe/CrispASR\ncd CrispASR\ncmake -B build -DCMAKE_BUILD_TYPE=Release\ncmake --build build -j$(nproc) --target whisper-cli\n\n# Auto-download the recommended quant on first use:\n./build/bin/crispasr --backend qwen3 -m auto -f my_audio.wav\n\n# Or point at a local file:\n./build/bin/crispasr --backend qwen3 \\\n    -m qwen3-asr-1.7b-q4_k.gguf \\\n    -f my_audio.wav\n\n# Auto language detection (whisper-tiny LID pre-step):\n./build/bin/crispasr --backend qwen3 \\\n    -m qwen3-asr-1.7b-q4_k.gguf \\\n    -f my_audio.wav -l auto\n\n# Speech translation to German via the system-prompt instruction path:\n./build/bin/crispasr --backend qwen3 \\\n    -m qwen3-asr-1.7b-q4_k.gguf \\\n    -f my_audio.wav --translate -tl de\n\n# Word-level timestamps via the canary-ctc-aligner second pass:\n./build/bin/crispasr --backend qwen3 \\\n    -m qwen3-asr-1.7b-q4_k.gguf \\\n    -f my_audio.wav -am canary-ctc-aligner-q5_0.gguf -osrt -ml 1\n```\n\nPure C++ runtime: no Python, no PyTorch, no ONNX Runtime. The only build dep beyond a C++17 compiler and CMake is `libcurl` or `wget` for the auto-download path.\n\n## Languages\n\nSame as upstream Qwen3-ASR: 30 languages plus 22 Chinese dialects, including English, Chinese, Cantonese, German, French, Spanish, Japanese, Korean, Russian, Arabic, Hindi, and most major European languages. See the [base model card](https://huggingface.co/Qwen/Qwen3-ASR-1.7B) for the full matrix and accuracy figures.\n\n## How it was made\n\n```bash\n# 1. Download the base model from HF\nhf download Qwen/Qwen3-ASR-1.7B --local-dir ./Qwen3-ASR-1.7B\n\n# 2. Convert to F16 GGUF\npython models/convert-qwen3-asr-to-gguf.py \\\n    --input ./Qwen3-ASR-1.7B \\\n    --output qwen3-asr-1.7b-f16.gguf\n\n# 3. Quantize\n./build/bin/crispasr-quantize qwen3-asr-1.7b-f16.gguf qwen3-asr-1.7b-q8_0.gguf q8_0\n./build/bin/crispasr-quantize qwen3-asr-1.7b-f16.gguf qwen3-asr-1.7b-q4_k.gguf q4_k\n```\n\nThe converter is shared with the 0.6B variant — it reads all sizes (audio encoder layers, LLM hidden size, KV head count, vocab size, head dim) from `config.json` rather than hardcoding them, so the same script handles both checkpoints. Total: 708 tensors (348 F16, 360 F32) per file.\n\n## Bit-identical regression check\n\nBoth quants converge on samples/jfk.wav under CrispASR's qwen3 backend:\n\n```\nqwen3-asr-1.7b-f16  → \"And so, my fellow Americans, ask not what your country can do for you. Ask what you can do for your country.\"\nqwen3-asr-1.7b-q8_0 → \"And so, my fellow Americans, ask not what your country can do for you. Ask what you can do for your country.\"\nqwen3-asr-1.7b-q4_k → \"And so, my fellow Americans, ask not what your country can do for you; ask what you can do for your country.\"\n```\n\n## License\n\nApache-2.0, same as the upstream Qwen3-ASR-1.7B model.\n\n## Citation\n\n```bibtex\n@misc{qwen3asr,\n    title  = {Qwen3-ASR},\n    author = {Qwen Team},\n    year   = {2026},\n    url    = {https://huggingface.co/Qwen/Qwen3-ASR-1.7B}\n}\n```\n",
    "related_quantizations": []
  },
  "tags": [
    "ggml",
    "gguf",
    "asr",
    "speech-to-text",
    "audio",
    "transcription",
    "speech-llm",
    "qwen3",
    "automatic-speech-recognition",
    "zh",
    "en",
    "yue",
    "ar",
    "de",
    "fr",
    "es",
    "pt",
    "id",
    "it",
    "ko",
    "ru",
    "th",
    "vi",
    "ja",
    "tr",
    "hi",
    "ms",
    "nl",
    "sv",
    "da",
    "fi",
    "pl",
    "cs",
    "fil",
    "fa",
    "el",
    "ro",
    "hu",
    "mk",
    "base_model:Qwen/Qwen3-ASR-1.7B",
    "base_model:quantized:Qwen/Qwen3-ASR-1.7B",
    "license:apache-2.0",
    "region:us"
  ],
  "likes": 1,
  "downloads": 387,
  "gated": false,
  "private": false,
  "last_modified": "2026-04-11T17:41:11.000Z",
  "created_at": "2026-04-11T17:37:52.000Z",
  "pipeline_tag": "automatic-speech-recognition",
  "library_name": "ggml"
}
Source payload excerpt (from Hugging Face API)
{
  "_id": "69da86f0e7d734b4eec72601",
  "id": "cstr/qwen3-asr-1.7b-GGUF",
  "modelId": "cstr/qwen3-asr-1.7b-GGUF",
  "sha": "4593c9d871d43870c73b3500723399409f31a3de",
  "createdAt": "2026-04-11T17:37:52.000Z",
  "lastModified": "2026-04-11T17:41:11.000Z",
  "author": "cstr",
  "downloads": 387,
  "likes": 1,
  "gated": false,
  "private": false,
  "pipeline_tag": "automatic-speech-recognition",
  "library_name": "ggml",
  "siblings_count": 5
}