GraySoft
Projects Models Compare Cloud benchmarks FAQ Download guIDE →
Model Intelligence Sheet

cstr/aya-expanse-8b-Q4_K_M-GGUF overview

cstr/aya expanse 8b Q4 K M GGUF This model was converted to GGUF format from CohereForAI/aya expanse 8b https://huggingface.co/CohereForAI/aya expanse 8b using…

transformersggufllama-cppgguf-my-repoenfrdeesitptjakozharelfaplidcshehinlroru

Runs locally from ~4.71 GB disk (8 GB VRAM class GPUs with llama.cpp / guIDE).

Downloads
217
Likes
0
Pipeline
Author

Repository Files & Downloads

1 GGUF files detected
Direct downloads for local inference
FileTypeQuantizationSizeLink
aya-expanse-8b-q4_k_m.ggufGGUFQ4_K_M4.71 GBDownload

Model Details

Model IDcstr/aya-expanse-8b-Q4_K_M-GGUF
Authorcstr
Pipeline
Licensecc-by-nc-4.0
Base modelCohereForAI/aya-expanse-8b
Last modified2026-08-02T15:18:51.000Z

Model README

---

inference: false

library_name: transformers

language:

  • en
  • fr
  • de
  • es
  • it
  • pt
  • ja
  • ko
  • zh
  • ar
  • el
  • fa
  • pl
  • id
  • cs
  • he
  • hi
  • nl
  • ro
  • ru
  • tr
  • uk
  • vi

license: cc-by-nc-4.0

extra_gated_prompt: By submitting this form, you agree to the License Agreement and

acknowledge that the information you provide will be collected, used, and shared

in accordance with Cohere’s Privacy Policy. You’ll

receive email updates about C4AI and Cohere research, events, products and services.

You can unsubscribe at any time.

extra_gated_fields:

Name: text

Affiliation: text

Country: country

I agree to use this model for non-commercial use ONLY: checkbox

base_model: CohereForAI/aya-expanse-8b

tags:

  • llama-cpp
  • gguf-my-repo

---

cstr/aya-expanse-8b-Q4_K_M-GGUF

This model was converted to GGUF format from CohereForAI/aya-expanse-8b using llama.cpp via the ggml.ai's GGUF-my-repo space.

Refer to the original model card for more details on the model.

Use with llama.cpp

Install llama.cpp through brew (works on Mac and Linux)

brew install llama.cpp

Invoke the llama.cpp server or the CLI.

CLI:

llama-cli --hf-repo cstr/aya-expanse-8b-Q4_K_M-GGUF --hf-file aya-expanse-8b-q4_k_m.gguf -p "The meaning to life and the universe is"

Server:

llama-server --hf-repo cstr/aya-expanse-8b-Q4_K_M-GGUF --hf-file aya-expanse-8b-q4_k_m.gguf -c 2048

Note: You can also use this checkpoint directly through the usage steps listed in the Llama.cpp repo as well.

Step 1: Clone llama.cpp from GitHub.

git clone https://github.com/ggerganov/llama.cpp

Step 2: Move into the llama.cpp folder and build it with LLAMA_CURL=1 flag along with other hardware-specific flags (for ex: LLAMA_CUDA=1 for Nvidia GPUs on Linux).

cd llama.cpp && LLAMA_CURL=1 make

Step 3: Run inference through the main binary.

./llama-cli --hf-repo cstr/aya-expanse-8b-Q4_K_M-GGUF --hf-file aya-expanse-8b-q4_k_m.gguf -p "The meaning to life and the universe is"

or

./llama-server --hf-repo cstr/aya-expanse-8b-Q4_K_M-GGUF --hf-file aya-expanse-8b-q4_k_m.gguf -c 2048

ollama

use a modelfile like:

FROM aya-expanse-8b-q4_k_m.gguf
TEMPLATE "{{ if .System }}<|START_OF_TURN_TOKEN|><|SYSTEM_TOKEN|>{{ .System }}<|END_OF_TURN_TOKEN|>
{{ end }}{{ if .Prompt }}<|START_OF_TURN_TOKEN|><|USER_TOKEN|>{{ .Prompt }}<|END_OF_TURN_TOKEN|>
{{ end }}<|START_OF_TURN_TOKEN|><|CHATBOT_TOKEN|>"

PARAMETER num_ctx 8192
PARAMETER stop "<|END_OF_TURN_TOKEN|>"
PARAMETER stop "<|START_OF_TURN_TOKEN|>"
PARAMETER stop "|END_OF_TURN_TOKEN"
PARAMETER stop "|START_OF_TURN_TOKEN"

SYSTEM "You are Aya, a brilliant, sophisticated, multilingual AI-assistant trained to assist human users by providing thorough responses. You are able to interact and respond to questions in 23 languages and you are powered by a multilingual model built by Cohere For AI."

Provenance and EU AI Act Art. 53 note

  • Upstream model: CohereForAI/aya-expanse-8b — published by CohereForAI.
  • Upstream licence: cc-by-nc-4.0. This repository redistributes under the same terms; it grants no rights the upstream licence does not.
  • What was done here: format conversion and/or quantisation only (GGUF, INT4 precision). No training, no fine-tuning, no merging, no distillation, no change to architecture, vocabulary or capability. Only the numeric representation of the upstream weights differs.
  • Training data: documented — where it is documented at all — by the upstream provider; see the upstream model card. No training data was used, added or selected by this repository. No training-content summary was found on the upstream model card at the time of writing; that documentation gap is upstream's and is not filled here.
  • Provider status: under Regulation (EU) 2024/1689 the upstream authors remain the provider of this model. Converting the serialisation format does not make this repository the provider of a new general-purpose AI model, and no such claim is made. Questions about training content, copyright policy or model capability belong upstream.

Run cstr/aya-expanse-8b-Q4_K_M-GGUF with guIDE

Download guIDE — the AI-native code editor with local LLM inference and 69 built-in tools.

Download guIDE → · Browse 524k+ models · Compare models

Source: Hugging Face · Compare models