GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

cstr/crepe-GGUF overview

CREPE — GGUF GGUF conversions of CREPE , a convolutional pitch F0 estimator, for use with CrispASR https://github.com/CrispStrobe/CrispASR 's ggml runtime. CRE…

crispasrggufpitch-estimationf0crepeaudiomusic-information-retrievalarxiv:1802.06182license:mitregion:us

Runs locally from ~0.3 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
1,119
Likes
0
Pipeline
Author

Repository Files & Downloads

6 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
crepe-full-f16.ggufGGUFF1642.4 MBDownload
crepe-full-q4_k.ggufGGUFQ4_K12.0 MBDownload
crepe-full-q8_0.ggufGGUFQ8_022.6 MBDownload
crepe-tiny-f16.ggufGGUFF160.9 MBDownload
crepe-tiny-q4_k.ggufGGUFQ4_K0.3 MBDownload
crepe-tiny-q8_0.ggufGGUFQ8_00.5 MBDownload

Model Details

Model IDcstr/crepe-GGUF
Authorcstr
Pipeline
Licensemit
Base model
Last modified2026-08-02T15:21:07.000Z

Model README

---

license: mit

library_name: crispasr

tags:

- pitch-estimation

- f0

- crepe

- gguf

- audio

- music-information-retrieval

---

CREPE — GGUF

GGUF conversions of CREPE, a convolutional pitch (F0) estimator, for use

with CrispASR's ggml runtime.

CREPE runs directly on the raw waveform — no STFT, no CQT — and emits a

360-bin pitch activation per frame.

Files

| file | capacity | quant | size |

|---|---|---|---|

| crepe-tiny-f16.gguf | tiny | f16 | 0.93 MB |

| crepe-tiny-q8_0.gguf | tiny | q8_0 | 0.50 MB |

| crepe-tiny-q4_k.gguf | tiny | q4_k | 0.27 MB |

| crepe-full-f16.gguf | full | f16 | 42.4 MB |

| crepe-full-q8_0.gguf | full | q8_0 | 22.6 MB |

| crepe-full-q4_k.gguf | full | q4_k | 12.0 MB |

tiny is the recommended default; full is ~38× more compute per frame for a

modest accuracy gain and is the right choice for offline work.

Input / output contract

  • Input: 16 kHz mono audio. The model consumes 1024-sample frames, each

normalized per-frame (subtract mean, divide by max(std, 1e-10)). Reference

hop is 10 ms.

  • Output: 360 activations per frame, sigmoid-valued. The bins are spaced

20 cents apart:

```

cents = 20 * bin + 1997.3794084376191

Hz = 10 2 * (cents / 1200)

```

Bin 0 ≈ 32.7 Hz, bin 359 ≈ 1975.5 Hz. Decode with the original CREPE

weighted-local-average around the argmax; the activation peak value doubles as

a voicing confidence.

Quantization

Only conv*.weight and classifier.weight are quantized. The per-channel

affine parameters — conv.bias, conv_BN.scale, conv*_BN.offset,

classifier.bias — are kept at F32 deliberately: in CREPE the ReLU comes

before the BatchNorm, so the BN cannot be folded into the conv and ships as a

standalone per-channel affine. Rounding those would apply a multiplicative

error to an entire channel.

Note on q4_k: the conv2conv6 kernels are 64 taps wide, and 64 is not a

multiple of Q4_K's 256-element super-block, so those five tensors fall back to

Q4_0 (32-element blocks). conv1 (512 taps) and classifier are true Q4_K.

There is no size penalty — Q4_0 and Q4_K are both 4.5 bits per weight.

Measured fidelity

Two independent measurements. Prefer f16 or q8_0.

Per-frame, against the model's own f16 (crispasr-diff crepe on 1101 frames

of real speech) — cos_min and the fraction of frames whose argmax pitch bin

is unchanged:

| | f16 | q8_0 | q4_k |

|---|---|---|---|

| tiny | 0.999999 · 100% | 0.999807 · 98.5% | 0.961643 · 85.2% |

| full | 1.000000 · 100% | 0.999937 · 99.5% | 0.992563 · 91.4% |

q4_k does not meet a 0.999 cosine bar at either capacity. For tiny-q4_k,

roughly 1 frame in 7 lands on a different pitch bin than f16. Ship q4_k only

if size genuinely dominates and you post-filter by voicing confidence; it is not

a drop-in for f16/q8_0. q8_0 is effectively lossless and is the right choice

whenever f16's size is inconvenient.

The f16 files themselves score cos = 1.0 against torchcrepe (max abs error

~2e-5 tiny / ~4e-6 full, i.e. f16 weight rounding).

Accuracy on real music

Evaluated on 10 monophonic instrumental recordings (violin arco + pizzicato,

piano, glockenspiel, carillon, cello, flute, three folk melodies, brass). With no

hand-labelled F0, the proxies are tiny-vs-full octave disagreement and the

in-tessitura rate over frames with voiced_prob >= 0.5:

| | tiny | full |

|---|---|---|

| in-tessitura | 89.6% | 89.0% |

| octave disagreement tiny-vs-full | 2.3% | — |

tiny is not meaningfully worse than full on monophonic music, despite

being ~38x cheaper — so tiny is the recommended default. Known domain limits,

shared by both capacities: plucked/percussive attacks with fast decay (violin

pizzicato scored ~50%, most frames having no sustained pitch) and **inharmonic

sources such as bells**, where the model correctly abstains — a carillon clip

marked only 39/1501 frames voiced at tiny — rather than inventing pitch.

Caveat: the tessitura bounds are hand-chosen, so the absolute percentages are

soft; the tiny-vs-full comparison is the robust part, both being scored

identically. A labelled MIR dataset is still needed for an absolute note-F.

Performance

Measured on an Apple M1 (quiet box), 10 s of audio at the reference 10 ms hop:

| model | Metal | CPU |

|---|---|---|

| tiny | RTF 0.28 | RTF ~2.4 |

| full | RTF 2.0 | RTF ~40 |

CREPE is genuinely expensive per frame (≈7.3 GFLOP per second of audio for

tiny, ≈282 GFLOP/s for full). Neither capacity is real-time on CPU — the

GPU path is not optional here.

Provenance and license

MIT, at every step of the chain:

  • Original model: Jong Wook Kim, Justin Salamon, Peter Li, Juan Pablo Bello,

"CREPE: A Convolutional Representation for Pitch Estimation", ICASSP 2018.

Released under the MIT license.

(paper ·

code)

by Max Morrison (MIT), which is itself a port of the original CREPE Keras

weights.

  • This conversion: models/convert-crepe-to-gguf.py in CrispASR (MIT).

If you use CREPE, please cite the original paper:

@inproceedings{kim2018crepe,
  title     = {{CREPE}: A Convolutional Representation for Pitch Estimation},
  author    = {Kim, Jong Wook and Salamon, Justin and Li, Peter and Bello, Juan Pablo},
  booktitle = {IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
  year      = {2018}
}

Provenance and EU AI Act Art. 53 note

  • Upstream model: CREPE (marl/crepe) — paper named, upstream repo not linked.
  • Upstream licence: mit. This repository redistributes under the same terms; it grants no rights the upstream licence does not.
  • What was done here: format conversion and/or quantisation only (GGUF). No training, no fine-tuning, no merging, no distillation, no change to architecture, vocabulary or capability. Only the numeric representation of the upstream weights differs.
  • Training data: documented — where it is documented at all — by the upstream provider; see the upstream model card. No training data was used, added or selected by this repository.
  • Provider status: under Regulation (EU) 2024/1689 the upstream authors remain the provider of this model. Converting the serialisation format does not make this repository the provider of a new general-purpose AI model, and no such claim is made. Questions about training content, copyright policy or model capability belong upstream.

Run cstr/crepe-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models