GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

bidubr/repente-v0.7-GGUF overview

Repente v0.7 A language model that writes Pure Data patches and SuperCollider code, and runs locally. Qwen2.5 Coder 7B Instruct fine tuned with QLoRA, exported…

ggufpure-datasupercollidercomputer-musiccode-generationqlorallama-cpptext-generationenbase_model:Qwen/Qwen2.5-Coder-7B-Instructbase_model:finetune:Qwen/Qwen2.5-Coder-7B-Instructlicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~4.36 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
text-generation
Author

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
repente-v0.7-Q4_K_M.ggufGGUFQ4_K_M4.36 GBDownload

Model Details

Model IDbidubr/repente-v0.7-GGUF
Authorbidubr
Pipelinetext-generation
Licenseapache-2.0
Base modelQwen/Qwen2.5-Coder-7B-Instruct
Last modified2026-08-19T23:18:25.000Z

Model README

---

license: apache-2.0

base_model: Qwen/Qwen2.5-Coder-7B-Instruct

base_model_relation: finetune

pipeline_tag: text-generation

language:

  • en

tags:

  • gguf
  • pure-data
  • supercollider
  • computer-music
  • code-generation
  • qlora
  • llama-cpp

---

Repente v0.7

A language model that writes Pure Data patches and SuperCollider code, and runs

locally. Qwen2.5-Coder-7B-Instruct fine-tuned with QLoRA, exported to GGUF at

Q4_K_M.

This is the current release. A companion repository,

bidubr/repente-v0.5-GGUF, holds the

earlier v0.5, kept available because the experiments in the paper were run on it.

What it does

Give it a description of a sound and it writes the code that produces it. Ask it about

an existing patch and it explains the signal flow.

make a sine wave at 440 Hz connected to output
#N canvas 0 0 450 300 12;
#X obj 100 100 osc~ 440;
#X obj 100 150 dac~;
#X connect 0 0 1 0;
#X connect 0 0 1 1;

That is a Pure Data file. Save it, open it, and it makes a sound.

How good it is, measured

Five Pure Data generation prompts, each sampled 30 times at temperature 0.7, scored by

whether the output carries a canvas header, an output object, and connections.

| Model | Expected score out of 5 | 95% interval |

|---|---|---|

| Qwen2.5-Coder-7B, unmodified | 0.03 | [0.00, 0.10] |

| Repente v0.5 | 2.67 | [2.37, 2.97] |

| Repente v0.7 | 3.83 | [3.53, 4.13] |

The base model emits a valid canvas header in 2.7% of attempts, yet names an output

object in 68.7% of them, more often than v0.5 does. Its deficit is serialization, not

intent, and that is the gap fine-tuning closes.

Version 0.7 is the strongest of the nine checkpoints trained, and is the only one the

battery separates from v0.5 by disjoint intervals. Full protocol in Appendix C of the

paper.

Running it

With Ollama, straight from this repository:

ollama run hf.co/bidubr/repente-v0.7-GGUF:Q4_K_M

With llama.cpp:

llama-server -m repente-v0.7-Q4_K_M.gguf -ngl 99 -c 4096

4.5 GB on disk. Fits in 8 GB of VRAM with room for a 4096-token context.

Prompting

The system prompt used throughout the reported experiments is short:

You are Repente, a musical programming expert.

Format validity improves substantially when three worked examples and a short

chain-of-thought instruction are supplied together. Measured on frozen weights, the

combination raises format validity from 75% to 100% across four prompt difficulty

levels and eliminates output truncation, at no compute cost. Few-shot exemplars

supplied alone introduce a context-leak failure mode in which the model continues the

turn structure of the examples until the token limit; the reasoning instruction

suppresses it. Details in Section 8 of the paper.

Known limitations

ELSE objects are not generated. Across 1,350 measured generations, an object from

the ELSE library appears once. Five cycles of corpus weighting, including one that

weighted ELSE explicitly, did not change this. Treat library coverage as a retrieval

problem at inference time, not something the weights will supply.

Analysis responses are short. The unmodified base model produces analyses averaging

776 tokens; this model produces 151. The compression is a self-distillation artifact of

using each version's own output to train the next, and it is documented rather than

fixed.

Patch validity is not patch quality. The measurements above score a generation as

passing when it carries a canvas header, an output object, and connections. That is

necessary for a patch to make sound and not sufficient for it to make the right sound.

Related

  • Paper: [arXiv:[ARXIV_ID]](https://arxiv.org/abs/[ARXIV_ID])
  • Code, data and measurement protocol: https://repente.net
  • pd-repente: a PlugData fork with a prompt bar in the patching window,

https://github.com/dobidu/plugdata

Citation

@misc{repente2026,
  author       = {Batista, Carlos Eduardo Coelho Freire},
  title        = {Repente: A Specialized LLM for Musical Programming Languages
                  with SonicUnit Knowledge Architecture},
  year         = {2026},
  eprint       = {[ARXIV_ID]},
  archivePrefix= {arXiv},
  primaryClass = {cs.SD}
}

Licensed Apache 2.0, inherited from the base model.

Run bidubr/repente-v0.7-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models