GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

ajvikram/toolcall-2b-gguf overview

Toolcall 2B — GGUF Quantized builds of ajvikram/toolcall 2b https://huggingface.co/ajvikram/toolcall 2b , a 2B function calling model fine tuned from Qwen3.5 2…

gguffunction-callingtool-useagentsllama.cppenbase_model:ajvikram/toolcall-2bbase_model:quantized:ajvikram/toolcall-2blicense:apache-2.0endpoints_compatibleregion:usconversational

Runs locally from ~1.22 GB disk (4 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
0
Likes
0
Pipeline
Author

Repository Files & Downloads

4 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
toolcall-2b-Q4_K_M.ggufGGUFQ4_K_M1.22 GBDownload
toolcall-2b-Q5_K_M.ggufGGUFQ5_K_M1.35 GBDownload
toolcall-2b-Q8_0.ggufGGUFQ8_01.93 GBDownload
toolcall-2b-f16.ggufGGUFF163.63 GBDownload

Model Details

Model IDajvikram/toolcall-2b-gguf
Authorajvikram
Pipeline
Licenseapache-2.0
Base modelajvikram/toolcall-2b
Last modified2026-09-09T00:54:41.000Z

Model README

---

license: apache-2.0

base_model: ajvikram/toolcall-2b

tags:

- gguf

- function-calling

- tool-use

- agents

- llama.cpp

language:

- en

---

Toolcall-2B — GGUF

Quantized builds of ajvikram/toolcall-2b,

a 2B function-calling model fine-tuned from Qwen3.5-2B for local agent tool routing.

Full results, training details and limitations are on the parent model's card.

| File | Size | Use |

|---|---|---|

| toolcall-2b-Q4_K_M.gguf | 1.22 GB | Default. Smallest sensible quality loss, runs on a laptop CPU. |

| toolcall-2b-Q5_K_M.gguf | 1.35 GB | A little closer to full precision for modest extra memory. |

| toolcall-2b-Q8_0.gguf | 1.93 GB | Near-lossless; use when you have the memory. |

| toolcall-2b-f16.gguf | 3.63 GB | Unquantized source for making your own quants. |

Measured on the benchmark harness (safetensors, bf16): 36.35 overall on BFCL v4

against 33.85 for the Qwen3.5-2B base, with every group ahead of the base. The

quantized builds are not separately scored.

Run it

llama-server -m toolcall-2b-Q4_K_M.gguf --jinja -c 8192
ollama run hf.co/ajvikram/toolcall-2b-gguf:Q4_K_M

The model uses Qwen3.5's native XML tool-call format, so any client that already

parses Qwen3.5 tool calls works unchanged:

<tool_call>
<function=get_weather>
<parameter=city>
Berlin
</parameter>
</function>
</tool_call>

Verified with llama-cli on CPU: the Q4_K_M build loads, generates at roughly 33

tokens per second on an ARM CPU, and returns the call above for a get_weather

tool given "What is the weather in Berlin?".

Thinking is off by default, matching how the model was trained and evaluated.

Notes

  • Built with llama.cpp (September 2026), which added Qwen3.5 conversion support;

older builds cannot convert this architecture.

  • These are text-only builds. The base architecture is vision-capable, but this

model was trained and evaluated purely on text tool calling.

Run ajvikram/toolcall-2b-gguf with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models