GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

cstr/cohere-transcribe-03-2026-GGUF overview

cohere transcribe 03 2026 — GGUF GGUF weights for CohereLabs/cohere transcribe 03 2026 https://huggingface.co/CohereLabs/cohere transcribe 03 2026 — Cohere's o…

ggufaudiospeech-recognitiontranscriptionconformercrispasrautomatic-speech-recognitionardeelenesfritjakonlplptvizhbase_model:CohereLabs/cohere-transcribe-03-2026base_model:quantized:CohereLabs/cohere-transcribe-03-2026license:apache-2.0

Runs locally from ~1.41 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
2,671
Likes
9
Pipeline
automatic-speech-recognition
Author

Repository Files & Downloads

6 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
cohere-transcribe-q4_k.ggufGGUFQ4_K1.41 GBDownload
cohere-transcribe-q5_0.ggufGGUFQ5_01.62 GBDownload
cohere-transcribe-q5_1.ggufGGUFQ5_11.73 GBDownload
cohere-transcribe-q6_k.ggufGGUFQ6_K1.85 GBDownload
cohere-transcribe-q8_0.ggufGGUFQ8_02.26 GBDownload
cohere-transcribe.ggufGGUFGGUF3.85 GBDownload

Model Details

Model IDcstr/cohere-transcribe-03-2026-GGUF
Authorcstr
Pipelineautomatic-speech-recognition
Licenseapache-2.0
Base modelCohereLabs/cohere-transcribe-03-2026
Last modified2026-08-07T04:53:10.000Z

Model README

---

license: apache-2.0

language:

  • ar
  • de
  • el
  • en
  • es
  • fr
  • it
  • ja
  • ko
  • nl
  • pl
  • pt
  • vi
  • zh

pipeline_tag: automatic-speech-recognition

tags:

  • audio
  • speech-recognition
  • transcription
  • gguf
  • conformer
  • crispasr

base_model: CohereLabs/cohere-transcribe-03-2026

---

cohere-transcribe-03-2026 — GGUF

GGUF weights for CohereLabs/cohere-transcribe-03-2026 — Cohere's open-source 2B-parameter ASR model, #1 on the Open ASR Leaderboard (avg WER 5.42, as of March 2026).

This conversion enables high-performance CPU inference via CrispASR — a crispasr-style C++ runtime for the Cohere Conformer-encoder / Transformer-decoder architecture.

> License: Apache 2.0 (inherited from source model). See original model card for full terms.

---

Files

| File | Size | Type | RTFx (8 threads) |

|------|------|------|------------------|

| cohere-transcribe.gguf | 3.85 GB | F16 | 0.80x |

| cohere-transcribe-q8_0.gguf | 2.05 GB | Q8_0 | 1.03x |

| cohere-transcribe-q6_k.gguf | 1.62 GB | Q6_K | 1.05x |

| cohere-transcribe-q5_1.gguf | 1.45 GB | Q5_1 | 1.06x |

| cohere-transcribe-q5_0.gguf | 1.38 GB | Q5_0 | 1.07x |

| cohere-transcribe-q4_k.gguf | 1.21 GB | Q4_K | 1.08x |

RTFx measured on jfk.wav (11s) using 8 CPU threads. Higher is faster. 1.0x means real-time.

---

Quick Start

1. Build CrispASR

git clone --recursive https://github.com/CrispStrobe/CrispASR
cd CrispASR && mkdir build && cd build
cmake .. -DCMAKE_BUILD_TYPE=Release
cmake --build . -j $(nproc) --target crispasr-cli

2. Download a GGUF

huggingface-cli download cstr/cohere-transcribe-03-2026-GGUF \
    cohere-transcribe-q4_k.gguf \
    --local-dir .

3. Transcribe

./bin/crispasr --backend cohere \
    -m cohere-transcribe-q4_k.gguf \
    -f audio.wav \
    -l en \
    -t 8

---

Implementation Notes (Critical for Correctness)

The language whitelist is metadata, not vocab

config.json lists 14 supported languages (`en fr de es it pt nl pl el ar ja

zh vi ko), but the tokenizer carries all 183 ISO-639-1 <|xx|>` tokens. So

<|ru|> decodes without error, and the model answers a wrong language

fluently rather than failing — there is no runtime signal at all. The

whitelist therefore has to travel with the weights: these GGUFs carry it as

cohere_transcribe.supported_languages, and CrispASR substitutes loudly when

-l names something outside it. For a GGUF converted before that key existed,

declare it with CRISPASR_COHERE_LANGS=en,fr,de,….

This matters most because -l auto runs an external detector that knows 99

languages against a model that accepts 14.

Mel normalization

Per-feature normalization uses biased standard deviation std = sqrt(mean(diff²) + ε), matching the ONNX reference. Using the Bessel-corrected (unbiased) formula produces a sqrt(T) ≈ 20× larger denominator for T ≈ 417 frames and completely corrupts the encoder output.

Conformer Attention Scaling

The self-attention mechanism in the Conformer encoder must be scaled by 1/sqrt(head_dim) before the softmax. Omitting this results in saturated attention scores and repetitive "garbage" output (e.g., "what what what...").

Encoder preprocessing

  1. Pre-emphasis: y[n] = x[n] - 0.97·x[n-1]
  2. Center-pad: n_fft/2 = 256 samples on each side
  3. STFT: Hann window (length 400, zero-padded to 512), hop 160, rfft → power spectrum
  4. Mel Filterbank: 128 bins → log → per-feature norm (biased std)

Conv subsampling

5 convolutions with 3 stride-2 steps reducing T_mel → T_enc ≈ T_mel/8:

conv0(ReLU) → conv2(DW) → conv3(PW,ReLU) → conv5(DW) → conv6(PW,ReLU) → linear(d=1280)

Cross-Attention Pre-computation

For high performance, cross-attention Key and Value tensors are pre-computed once per utterance from the encoder output. In this implementation, these projections are performed as part of the encoder's GGML compute graph to leverage backend acceleration.

Decoder activation

Transformer decoder FFN uses ReLU (not SiLU/Swish).

---

Architecture

| Component | Details |

|-----------|---------|

| Encoder | 48-layer Conformer, d=1280, heads=8, head_dim=160, ffn=5120, conv_kernel=9 |

| Decoder | 8-layer causal Transformer, d=1024, heads=8, head_dim=128, ffn=4096, max_ctx=1024 |

| Vocab | 16,384 SentencePiece tokens |

| Audio | 16 kHz mono, 128 mel bins, n_fft=512, hop=160, win=400 |

| Parameters | ~2B |

---

Related

Sister GGUF releases in the same family

Use case → which runtime?

| Need | Right tool |

| --- | --- |

| Lowest English WER (Open ASR Leaderboard #1) | --backend cohere ← this repo |

| Multilingual ASR + free word timestamps | parakeet-main (cstr/parakeet-tdt-0.6b-v3-GGUF) |

| Multilingual ASR + speech translation + explicit language control | canary-main (cstr/canary-1b-v2-GGUF) |

| Multilingual subword forced alignment of any transcript | nfa-align (cstr/canary-ctc-aligner-GGUF) |

| English-only character-level forced alignment (~30 ms MAE) | cohere-align (uses wav2vec2-large-xlsr-53-english) |

Provenance and EU AI Act Art. 53 note

  • Upstream model: CohereLabs/cohere-transcribe-03-2026 — published by CohereLabs.
  • Upstream licence: apache-2.0. This repository redistributes under the same terms; it grants no rights the upstream licence does not.
  • What was done here: format conversion and/or quantisation only (GGUF). No training, no fine-tuning, no merging, no distillation, no change to architecture, vocabulary or capability. Only the numeric representation of the upstream weights differs.
  • Training data: documented — where it is documented at all — by the upstream provider; see the upstream model card. No training data was used, added or selected by this repository.
  • Provider status: under Regulation (EU) 2024/1689 the upstream authors remain the provider of this model. Converting the serialisation format does not make this repository the provider of a new general-purpose AI model, and no such claim is made. Questions about training content, copyright policy or model capability belong upstream.

Run cstr/cohere-transcribe-03-2026-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models