SkibidiBreaddd/qwen2.5-coder-1.5b-t2sql-gguf overview
SQLCoder — Text2SQL on a small LLM Fine tune a small open LLM Qwen2.5 Coder 1.5B , ≤ 3B to turn natural language questions into executable SQLite queries — the…
Runs locally from ~940.4 MB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).
Repository Files & Downloads
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| Qwen2.5-Coder-1.5B-Instruct.Q4_K_M.gguf | GGUF | GGUF | 940.4 MB | Download |
Model Details
| Model ID | SkibidiBreaddd/qwen2.5-coder-1.5b-t2sql-gguf |
|---|---|
| Author | SkibidiBreaddd |
| Pipeline | — |
| License | — |
| Base model | Qwen/Qwen2.5-Coder-1.5B-Instruct |
| Last modified | 2026-07-26T08:53:49.000Z |
Model README
---
tags:
- gguf
- llama.cpp
- unsloth
datasets:
- SkibidiBreaddd/synsql-t2sql-subset
base_model:
- Qwen/Qwen2.5-Coder-1.5B-Instruct
---
SQLCoder — Text2SQL on a small LLM
Fine-tune a small open LLM (Qwen2.5-Coder-1.5B, ≤ 3B) to turn natural-language questions into
executable SQLite queries — the engine for a fintech chatbot that lets non-technical teams pull
data without writing SQL.
The whole project is built around free resources: training on a Google Colab T4, inference on
an ordinary laptop CPU via a quantized GGUF served through Ollama — no GPU required to run
it. A LoRA adapter (~1% of parameters trained) sits on top of the frozen base model and is exported
to a ~1 GB GGUF for local use.
- What it does: given a database schema (DDL) + a question, it returns one runnable SQLite query.
- What it's for: a data-access chatbot — plain English in, executable SQL out.
Result of the fine-tune
Evaluated on 200 held-out test examples, greedy decoding, identical prompts for both models.
The numbers below are the final max_new_tokens=512 run (Result/*_preds_512.json).
| Model | Valid rate | Executable rate |
|---|---|---|
| Baseline (no fine-tune) | 99.0% | 42.5% |
| Fine-tuned (LoRA, 512-tok) | 99.5% | 75.5% |
| Gold queries (ceiling) | 100.0% | 99.5% |
Fine-tuning raised the executable-query rate from 42.5% → 75.5% (+33 pp absolute, **+78%
relative**), recovering roughly half the gap to the gold ceiling.
Using the model (Ollama)
The published artifacts on the Hugging Face Hub:
| Artifact | Repo | Use |
|---|---|---|
| Quantized GGUF | SkibidiBreaddd/qwen2.5-coder-1.5b-t2sql-gguf | local CPU inference |
| LoRA adapter | SkibidiBreaddd/qwen2.5-coder-1.5b-t2sql-lora | GPU / further training |
Fastest path — pull straight from the Hub
Ollama downloads the GGUF for you, no manual steps:
ollama run hf.co/SkibidiBreaddd/qwen2.5-coder-1.5b-t2sql-gguf
Recommended — build with the baked-in system prompt
This applies the Text2SQL system prompt and temperature 0 from report/Modelfile,
so you get deterministic, prompt-correct output:
# 1. download just the GGUF (~1 GB)
huggingface-cli download SkibidiBreaddd/qwen2.5-coder-1.5b-t2sql-gguf \
--include "*.gguf" --local-dir ./gguf
# 2. build a local Ollama model from the Modelfile
ollama create t2sql -f report/Modelfile
# 3. run it
ollama run t2sql
You then also get an OpenAI-compatible HTTP API on localhost:11434. Prompt it with the schema DDL
followed by the question (same order used in training).
Compute requirements
To run it (inference) — no GPU needed:
| Resource | Minimum | Comfortable |
|---|---|---|
| RAM | 4 GB free | 8 GB |
| Disk | 2 GB | 5 GB |
| CPU | any x86-64 with AVX2, 2 cores | 4–8 cores (Apple Silicon works natively) |
| GPU | none | optional (llama.cpp offloads layers if present) |
Run SkibidiBreaddd/qwen2.5-coder-1.5b-t2sql-gguf with guIDE
Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.
Source: Hugging Face · Compare models