GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

Bhuvandesai/phi3-text-to-sql-gguf overview

Phi 3 mini Text to SQL — GGUF quantized for CPU Quantized GGUF builds of the fine tuned Phi 3 mini Text to SQL https://huggingface.co/Bhuvandesai/phi3 text to …

gguftext-to-sqlsqlllama-cppquantizedphi-3text-generationenbase_model:Bhuvandesai/phi3-text-to-sql-adapterbase_model:quantized:Bhuvandesai/phi3-text-to-sql-adapterlicense:mitendpoints_compatibleregion:usconversational

Runs locally from ~2.23 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
127
Likes
0
Pipeline
text-generation

Repository Files & Downloads

2 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
phi3-text-to-sql-Q4_K_M.ggufGGUFQ4_K_M2.23 GBDownload
phi3-text-to-sql-Q5_K_M.ggufGGUFQ5_K_M2.57 GBDownload

Model Details

Model IDBhuvandesai/phi3-text-to-sql-gguf
AuthorBhuvandesai
Pipelinetext-generation
Licensemit
Base modelBhuvandesai/phi3-text-to-sql-adapter
Last modified2026-06-22T08:54:37.000Z

Model README

---

license: mit

base_model: Bhuvandesai/phi3-text-to-sql-adapter

pipeline_tag: text-generation

language:

- en

tags:

- text-to-sql

- sql

- gguf

- llama-cpp

- quantized

- phi-3

---

Phi-3-mini Text-to-SQL — GGUF (quantized for CPU)

Quantized GGUF builds of the fine-tuned Phi-3-mini Text-to-SQL model (LoRA already merged into the base weights), for fast CPU inference with llama.cpp.

| File | Size | Effective bits/weight | vs f16 |

|---|---:|---:|---:|

| phi3-text-to-sql-Q4_K_M.gguf ⭐ recommended | 2.40 GB | 5.01 | −68.6% (3.2× smaller) |

| phi3-text-to-sql-Q5_K_M.gguf | 2.76 GB | 5.76 | −64.0% (2.8× smaller) |

> Note: "Q4" K-quants average ~5 effective bits/weight (embeddings and some tensors stay higher-precision), so the file is larger than a literal 4-bit×params calculation.

Which one?

Use Q4_K_M. On this task it matched Q5_K_M on quality while being smaller and faster.

Benchmarks (measured)

CPU = Intel i7-13650HX, 14 threads, llama-bench, build 9637:

| Model | Prompt processing (pp256) | Token generation (tg64) |

|---|---:|---:|

| Q4_K_M | 91.4 tok/s | 20.1 tok/s |

| Q5_K_M | 59.6 tok/s | 18.5 tok/s |

Task quality (12 held-out questions, execution-match against a live SQLite DB):

| Model | Execution-match | Valid SQL |

|---|---:|---:|

| Q4_K_M | 75.0% | 100% |

| Q5_K_M | 75.0% | 100% |

4-bit quantization cost no measurable task accuracy vs 5-bit here.

Run it

# CLI
llama-cli -m phi3-text-to-sql-Q4_K_M.gguf -p "<|user|>\n<schema + question><|end|>\n<|assistant|>\n" -n 150 --temp 0

# Server (OpenAI-compatible)
llama-server -m phi3-text-to-sql-Q4_K_M.gguf -c 2048 -t 14 --port 8080

The model expects Phi-3 chat formatting; include the database schema in the user turn (see the adapter card for the exact prompt). It outputs raw SQLite.

License: MIT.

Run Bhuvandesai/phi3-text-to-sql-gguf with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models