GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

cstr/wav2vec2-large-xlsr-53-japanese-GGUF overview

Wav2Vec2 Large Xlsr 53 Japanese GGUF GGUF conversions and quantisations of jonatasgrosman/wav2vec2 large xlsr 53 japanese https://huggingface.co/jonatasgrosman…

ggmlggufaudiospeech-recognitiontranscriptionwav2vec2crispasrjapaneseautomatic-speech-recognitionjabase_model:jonatasgrosman/wav2vec2-large-xlsr-53-japanesebase_model:quantized:jonatasgrosman/wav2vec2-large-xlsr-53-japaneselicense:apache-2.0region:us

Runs locally from ~220.8 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
512
Likes
0
Pipeline
automatic-speech-recognition
Author

Repository Files & Downloads

3 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
wav2vec2-large-xlsr-53-japanese-f16.ggufGGUFF16635.5 MBDownload
wav2vec2-large-xlsr-53-japanese-q4_k.ggufGGUFQ4_K220.8 MBDownload
wav2vec2-large-xlsr-53-japanese-q8_0.ggufGGUFQ8_0365.0 MBDownload

Model Details

Model IDcstr/wav2vec2-large-xlsr-53-japanese-GGUF
Authorcstr
Pipelineautomatic-speech-recognition
Licenseapache-2.0
Base modeljonatasgrosman/wav2vec2-large-xlsr-53-japanese
Last modified2026-08-02T15:44:13.000Z

Model README

---

license: apache-2.0

language:

  • ja

pipeline_tag: automatic-speech-recognition

tags:

  • audio
  • speech-recognition
  • transcription
  • gguf
  • wav2vec2
  • crispasr
  • japanese

library_name: ggml

base_model: jonatasgrosman/wav2vec2-large-xlsr-53-japanese

---

Wav2Vec2 Large Xlsr 53 Japanese -- GGUF

GGUF conversions and quantisations of jonatasgrosman/wav2vec2-large-xlsr-53-japanese for use with CrispStrobe/CrispASR.

Available variants

| File | Quant | Size | Notes |

|---|---|---|---|

| wav2vec2-large-xlsr-53-japanese-f16.gguf | F16 | ~635 MB | Reference conversion |

| wav2vec2-large-xlsr-53-japanese-q4_k.gguf | Q4_K | ~221 MB | Default auto-download target |

| wav2vec2-large-xlsr-53-japanese-q8_0.gguf | Q8_0 | ~365 MB | Higher precision |

Model details

  • Architecture: Wav2Vec2ForCTC
  • Hidden size: 1024
  • Attention heads: 16
  • Transformer layers: 24
  • CTC vocabulary: 2341 tokens
  • Language: Japanese
  • Base model: jonatasgrosman/wav2vec2-large-xlsr-53-japanese

Usage with CrispASR

crispasr --backend wav2vec2 -m wav2vec2-large-xlsr-53-japanese-q4_k.gguf -f audio.wav -l ja

# Auto-download via registry alias once published:
crispasr --backend wav2vec2 -m auto --auto-download -l ja -f audio.wav

Provenance

Converted from the upstream Hugging Face checkpoint with

models/convert-wav2vec2-to-gguf.py, then quantized with

build-ninja-compile/bin/crispasr-quantize.

License

Apache-2.0 — same as the upstream model card.

Provenance and EU AI Act Art. 53 note

  • Upstream model: jonatasgrosman/wav2vec2-large-xlsr-53-japanese — published by jonatasgrosman.
  • Upstream licence: apache-2.0. This repository redistributes under the same terms; it grants no rights the upstream licence does not.
  • What was done here: format conversion and/or quantisation only (GGUF/GGML). No training, no fine-tuning, no merging, no distillation, no change to architecture, vocabulary or capability. Only the numeric representation of the upstream weights differs.
  • Training data: documented — where it is documented at all — by the upstream provider; see the upstream model card. No training data was used, added or selected by this repository.
  • Provider status: under Regulation (EU) 2024/1689 the upstream authors remain the provider of this model. Converting the serialisation format does not make this repository the provider of a new general-purpose AI model, and no such claim is made. Questions about training content, copyright policy or model capability belong upstream.

Run cstr/wav2vec2-large-xlsr-53-japanese-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models