Model Intelligence Sheet

richarderkhov/openchat_-_openchat-3.5-0106-gemma-gguf overview

Quantization made by Richard Erkhov. Github Discord Request more models openchat-3.5-0106-gemma - GGUF | Name | Quant method | Size | | ---- | ---- | ---- | | openchat-3.5-0106-gemma.Q2K.gguf | Q2K | 3.24GB | | openchat-3.5-0106-gemma.IQ3XS.gguf | IQ3XS | 3.54GB | | openchat-3.5-0106-gemma.IQ3S.gguf | IQ3S | 3.71GB | | openchat-3.5-0106-gemma.Q3KS.gguf | Q3KS | 3.71GB | | openchat-3.5-0106-gemma.IQ3M.gguf | IQ3M | 3.82GB | | openchat-3.5-0106-gemma.Q3K.gguf | Q3K | 4.07GB | | openchat-3.5-0106-gemma.Q3KM.gguf | Q3KM | 4.07GB | | openchat-3.5-0106-gemma.Q3KL.gguf | Q3KL | 4.39GB | | openchat-3.5-0106-gemma.IQ4XS.gguf | IQ4XS | 4.48GB | | openchat-3.5-0106-gemma.Q40.gguf | Q40 | 4.67GB | | openchat-3.5-0106-gemma.IQ4NL.gguf | IQ4NL | 4.69GB | | openchat-3.5-0106-gemma.Q4KS.gguf | Q4KS | 4.7GB | | openchat-3.5-0106-gemma.Q4K.gguf | Q4K | 4.96GB | | openchat-3.5-0106-gemma.Q4KM.gguf | Q4KM | 4.96GB | | openchat-3.5-0106-gemma.Q41.gguf | Q41 | 5.12GB | | openchat-3.5-0106-gemma.Q50.gguf | Q50 | 5.57GB | | openchat-3.5-0106-gemma.Q5KS.gguf | Q5KS | 5.57GB | | openchat-3.5-0106-gemma.Q5K.gguf | Q5K | 5.72GB | | openchat-3.5-0106-gemma.Q5KM.gguf | Q5KM | 5.72GB | | openchat-3.5-0106-gemma.Q51.gguf | Q51 | 6.02GB | | openchat-3.5-0106-gemma.Q6K.gguf | Q6K | 6.53GB | | openchat-3.5-0106-gemma.Q80.gguf | Q80 | 8.45GB | Original model description: --- license: other licensename: gemma-terms-of-use licenselink: https://ai.google.dev/gemma/terms ---

ggufarxiv:2309.11235endpoints_compatibleregion:usconversational

richarderkhov/openchat_-_openchat-3.5-0106-gemma-gguf visual

Downloads

197

Likes

Pipeline

—

Library

—

Visibility

Public

Access

Open

Repository Files & Downloads

22 files detected

Direct downloads for all repository files

File	Type	Quantization	Size	Link
openchat-3.5-0106-gemma.IQ3_M.gguf	GGUF	IQ3_M	3.82 GB	Download
openchat-3.5-0106-gemma.IQ3_S.gguf	GGUF	IQ3_S	3.71 GB	Download
openchat-3.5-0106-gemma.IQ3_XS.gguf	GGUF	IQ3_XS	3.54 GB	Download
openchat-3.5-0106-gemma.IQ4_NL.gguf	GGUF	IQ4_NL	4.69 GB	Download
openchat-3.5-0106-gemma.IQ4_XS.gguf	GGUF	IQ4_XS	4.48 GB	Download
openchat-3.5-0106-gemma.Q2_K.gguf	GGUF	Q2_K	3.24 GB	Download
openchat-3.5-0106-gemma.Q3_K.gguf	GGUF	Q3_K	4.07 GB	Download
openchat-3.5-0106-gemma.Q3_K_L.gguf	GGUF	Q3_K_L	4.39 GB	Download
openchat-3.5-0106-gemma.Q3_K_M.gguf	GGUF	Q3_K_M	4.07 GB	Download
openchat-3.5-0106-gemma.Q3_K_S.gguf	GGUF	Q3_K_S	3.71 GB	Download
openchat-3.5-0106-gemma.Q4_0.gguf	GGUF	—	4.67 GB	Download
openchat-3.5-0106-gemma.Q4_1.gguf	GGUF	—	5.12 GB	Download
openchat-3.5-0106-gemma.Q4_K.gguf	GGUF	Q4_K	4.96 GB	Download
openchat-3.5-0106-gemma.Q4_K_M.gguf	GGUF	Q4_K_M	4.96 GB	Download
openchat-3.5-0106-gemma.Q4_K_S.gguf	GGUF	Q4_K_S	4.70 GB	Download
openchat-3.5-0106-gemma.Q5_0.gguf	GGUF	—	5.57 GB	Download
openchat-3.5-0106-gemma.Q5_1.gguf	GGUF	—	6.02 GB	Download
openchat-3.5-0106-gemma.Q5_K.gguf	GGUF	Q5_K	5.72 GB	Download
openchat-3.5-0106-gemma.Q5_K_M.gguf	GGUF	Q5_K_M	5.72 GB	Download
openchat-3.5-0106-gemma.Q5_K_S.gguf	GGUF	Q5_K_S	5.57 GB	Download
openchat-3.5-0106-gemma.Q6_K.gguf	GGUF	Q6_K	6.53 GB	Download
openchat-3.5-0106-gemma.Q8_0.gguf	GGUF	—	8.45 GB	Download

Model Details Live

Model Slug

richarderkhov/openchat_-_openchat-3.5-0106-gemma-gguf

Author

RichardErkhov

Pipeline Task

—

Library

—

Created

2024-10-07

Last Modified

2024-10-07

Gated

Private

HF SHA

671c7416232ac16d21403ccb686e2fb377308ee9

License

Unknown

Language

Unknown

Base Model

Unknown

Metadata Inspector

Normalized metadata (stored in metadata_json)

{
  "metadata": {},
  "card_data": {
    "frontmatter": {},
    "hero_image_url": "https://cdn-uploads.huggingface.co/production/uploads/63972847b3e2256c9ce1307b/Ez9cDw8xstbTKlFtBgbVs.png",
    "summary": "Quantization made by Richard Erkhov. Github Discord Request more models openchat-3.5-0106-gemma - GGUF | Name | Quant method | Size | | ---- | ---- | ---- | | openchat-3.5-0106-gemma.Q2_K.gguf | Q2_K | 3.24GB | | openchat-3.5-0106-gemma.IQ3_XS.gguf | IQ3_XS | 3.54GB | | openchat-3.5-0106-gemma.IQ3_S.gguf | IQ3_S | 3.71GB | | openchat-3.5-0106-gemma.Q3_K_S.gguf | Q3_K_S | 3.71GB | | openchat-3.5-0106-gemma.IQ3_M.gguf | IQ3_M | 3.82GB | | openchat-3.5-0106-gemma.Q3_K.gguf | Q3_K | 4.07GB | | openchat-3.5-0106-gemma.Q3_K_M.gguf | Q3_K_M | 4.07GB | | openchat-3.5-0106-gemma.Q3_K_L.gguf | Q3_K_L | 4.39GB | | openchat-3.5-0106-gemma.IQ4_XS.gguf | IQ4_XS | 4.48GB | | openchat-3.5-0106-gemma.Q4_0.gguf | Q4_0 | 4.67GB | | openchat-3.5-0106-gemma.IQ4_NL.gguf | IQ4_NL | 4.69GB | | openchat-3.5-0106-gemma.Q4_K_S.gguf | Q4_K_S | 4.7GB | | openchat-3.5-0106-gemma.Q4_K.gguf | Q4_K | 4.96GB | | openchat-3.5-0106-gemma.Q4_K_M.gguf | Q4_K_M | 4.96GB | | openchat-3.5-0106-gemma.Q4_1.gguf | Q4_1 | 5.12GB | | openchat-3.5-0106-gemma.Q5_0.gguf | Q5_0 | 5.57GB | | openchat-3.5-0106-gemma.Q5_K_S.gguf | Q5_K_S | 5.57GB | | openchat-3.5-0106-gemma.Q5_K.gguf | Q5_K | 5.72GB | | openchat-3.5-0106-gemma.Q5_K_M.gguf | Q5_K_M | 5.72GB | | openchat-3.5-0106-gemma.Q5_1.gguf | Q5_1 | 6.02GB | | openchat-3.5-0106-gemma.Q6_K.gguf | Q6_K | 6.53GB | | openchat-3.5-0106-gemma.Q8_0.gguf | Q8_0 | 8.45GB | Original model description: --- license: other license_name: gemma-terms-of-use license_link: https://ai.google.dev/gemma/terms ---",
    "quick_links": [],
    "benchmark_table_html": "",
    "readme_markdown": "Quantization made by Richard Erkhov.\n\n[Github](https://github.com/RichardErkhov)\n\n[Discord](https://discord.gg/pvy7H8DZMG)\n\n[Request more models](https://github.com/RichardErkhov/quant_request)\n\n\nopenchat-3.5-0106-gemma - GGUF\n- Model creator: https://huggingface.co/openchat/\n- Original model: https://huggingface.co/openchat/openchat-3.5-0106-gemma/\n\n\n| Name | Quant method | Size |\n| ---- | ---- | ---- |\n| [openchat-3.5-0106-gemma.Q2_K.gguf](https://huggingface.co/RichardErkhov/openchat_-_openchat-3.5-0106-gemma-gguf/blob/main/openchat-3.5-0106-gemma.Q2_K.gguf) | Q2_K | 3.24GB |\n| [openchat-3.5-0106-gemma.IQ3_XS.gguf](https://huggingface.co/RichardErkhov/openchat_-_openchat-3.5-0106-gemma-gguf/blob/main/openchat-3.5-0106-gemma.IQ3_XS.gguf) | IQ3_XS | 3.54GB |\n| [openchat-3.5-0106-gemma.IQ3_S.gguf](https://huggingface.co/RichardErkhov/openchat_-_openchat-3.5-0106-gemma-gguf/blob/main/openchat-3.5-0106-gemma.IQ3_S.gguf) | IQ3_S | 3.71GB |\n| [openchat-3.5-0106-gemma.Q3_K_S.gguf](https://huggingface.co/RichardErkhov/openchat_-_openchat-3.5-0106-gemma-gguf/blob/main/openchat-3.5-0106-gemma.Q3_K_S.gguf) | Q3_K_S | 3.71GB |\n| [openchat-3.5-0106-gemma.IQ3_M.gguf](https://huggingface.co/RichardErkhov/openchat_-_openchat-3.5-0106-gemma-gguf/blob/main/openchat-3.5-0106-gemma.IQ3_M.gguf) | IQ3_M | 3.82GB |\n| [openchat-3.5-0106-gemma.Q3_K.gguf](https://huggingface.co/RichardErkhov/openchat_-_openchat-3.5-0106-gemma-gguf/blob/main/openchat-3.5-0106-gemma.Q3_K.gguf) | Q3_K | 4.07GB |\n| [openchat-3.5-0106-gemma.Q3_K_M.gguf](https://huggingface.co/RichardErkhov/openchat_-_openchat-3.5-0106-gemma-gguf/blob/main/openchat-3.5-0106-gemma.Q3_K_M.gguf) | Q3_K_M | 4.07GB |\n| [openchat-3.5-0106-gemma.Q3_K_L.gguf](https://huggingface.co/RichardErkhov/openchat_-_openchat-3.5-0106-gemma-gguf/blob/main/openchat-3.5-0106-gemma.Q3_K_L.gguf) | Q3_K_L | 4.39GB |\n| [openchat-3.5-0106-gemma.IQ4_XS.gguf](https://huggingface.co/RichardErkhov/openchat_-_openchat-3.5-0106-gemma-gguf/blob/main/openchat-3.5-0106-gemma.IQ4_XS.gguf) | IQ4_XS | 4.48GB |\n| [openchat-3.5-0106-gemma.Q4_0.gguf](https://huggingface.co/RichardErkhov/openchat_-_openchat-3.5-0106-gemma-gguf/blob/main/openchat-3.5-0106-gemma.Q4_0.gguf) | Q4_0 | 4.67GB |\n| [openchat-3.5-0106-gemma.IQ4_NL.gguf](https://huggingface.co/RichardErkhov/openchat_-_openchat-3.5-0106-gemma-gguf/blob/main/openchat-3.5-0106-gemma.IQ4_NL.gguf) | IQ4_NL | 4.69GB |\n| [openchat-3.5-0106-gemma.Q4_K_S.gguf](https://huggingface.co/RichardErkhov/openchat_-_openchat-3.5-0106-gemma-gguf/blob/main/openchat-3.5-0106-gemma.Q4_K_S.gguf) | Q4_K_S | 4.7GB |\n| [openchat-3.5-0106-gemma.Q4_K.gguf](https://huggingface.co/RichardErkhov/openchat_-_openchat-3.5-0106-gemma-gguf/blob/main/openchat-3.5-0106-gemma.Q4_K.gguf) | Q4_K | 4.96GB |\n| [openchat-3.5-0106-gemma.Q4_K_M.gguf](https://huggingface.co/RichardErkhov/openchat_-_openchat-3.5-0106-gemma-gguf/blob/main/openchat-3.5-0106-gemma.Q4_K_M.gguf) | Q4_K_M | 4.96GB |\n| [openchat-3.5-0106-gemma.Q4_1.gguf](https://huggingface.co/RichardErkhov/openchat_-_openchat-3.5-0106-gemma-gguf/blob/main/openchat-3.5-0106-gemma.Q4_1.gguf) | Q4_1 | 5.12GB |\n| [openchat-3.5-0106-gemma.Q5_0.gguf](https://huggingface.co/RichardErkhov/openchat_-_openchat-3.5-0106-gemma-gguf/blob/main/openchat-3.5-0106-gemma.Q5_0.gguf) | Q5_0 | 5.57GB |\n| [openchat-3.5-0106-gemma.Q5_K_S.gguf](https://huggingface.co/RichardErkhov/openchat_-_openchat-3.5-0106-gemma-gguf/blob/main/openchat-3.5-0106-gemma.Q5_K_S.gguf) | Q5_K_S | 5.57GB |\n| [openchat-3.5-0106-gemma.Q5_K.gguf](https://huggingface.co/RichardErkhov/openchat_-_openchat-3.5-0106-gemma-gguf/blob/main/openchat-3.5-0106-gemma.Q5_K.gguf) | Q5_K | 5.72GB |\n| [openchat-3.5-0106-gemma.Q5_K_M.gguf](https://huggingface.co/RichardErkhov/openchat_-_openchat-3.5-0106-gemma-gguf/blob/main/openchat-3.5-0106-gemma.Q5_K_M.gguf) | Q5_K_M | 5.72GB |\n| [openchat-3.5-0106-gemma.Q5_1.gguf](https://huggingface.co/RichardErkhov/openchat_-_openchat-3.5-0106-gemma-gguf/blob/main/openchat-3.5-0106-gemma.Q5_1.gguf) | Q5_1 | 6.02GB |\n| [openchat-3.5-0106-gemma.Q6_K.gguf](https://huggingface.co/RichardErkhov/openchat_-_openchat-3.5-0106-gemma-gguf/blob/main/openchat-3.5-0106-gemma.Q6_K.gguf) | Q6_K | 6.53GB |\n| [openchat-3.5-0106-gemma.Q8_0.gguf](https://huggingface.co/RichardErkhov/openchat_-_openchat-3.5-0106-gemma-gguf/blob/main/openchat-3.5-0106-gemma.Q8_0.gguf) | Q8_0 | 8.45GB |\n\n\n\n\nOriginal model description:\n---\nlicense: other\nlicense_name: gemma-terms-of-use\nlicense_link: https://ai.google.dev/gemma/terms\n---\n\n<div align=\"center\">\n  <a>\n    <img src=\"https://cdn-uploads.huggingface.co/production/uploads/63972847b3e2256c9ce1307b/Ez9cDw8xstbTKlFtBgbVs.png\" >\n  </a>\n</div>\n\n\n## The highest performing Gemma model in the world. Trained with OpenChat's C-RLFT on openchat-3.5-0106 data. Achieving similar performance to Mistral-based openchat, and much better than Gemma-7b and Gemma-7b-it.\n\nPlease refer to [openchat-3.5-0106](https://huggingface.co/openchat/openchat-3.5-0106) for details.\n\n> P.S.: 6T pre-training tokens + 0.003 init std dev + C-RLFT is the secret sauce?\n>\n> P.P.S.: @Google team, we know your model is great, but please use an OSI-approved license like Mistral (or even Phi and Orca).\n\n## Benchmarks\n\n| Model                       | # Params | Average  | MT-Bench | HumanEval | BBH MC   | AGIEval  | TruthfulQA | MMLU     | GSM8K    | BBH CoT  |\n|-----------------------------|----------|----------|----------|-----------|----------|----------|------------|----------|----------|----------|\n| **OpenChat-3.5-0106 Gemma** | **7B**   | 64.4     | 7.83     | 67.7      | **52.7** | **50.2** | 55.4       | 65.7     | **81.5** | 63.7     |\n| OpenChat-3.5-0106 Mistral   | **7B**   | **64.5** | 7.8      | **71.3**  | 51.5     | 49.1     | **61.0**   | 65.8     | 77.4     | 62.2     |\n| ChatGPT (March)             | ???B     | 61.5     | **7.94** | 48.1      | 47.6     | 47.1     | 57.7       | **67.3** | 74.9     | **70.1** |\n|                             |          |          |          |           |          |          |            |          |          |          |\n| Gemma-7B                    | 7B       | -        | -        | 32.3      | -        | 41.7     | -          | 64.3     | 46.4     | -        |\n| Gemma-7B-it *               | 7B       | 25.4     | -        | 28.0      | 38.4     | 32.5     | 34.1       | 26.5     | 10.8     | 7.6      |\n| OpenHermes 2.5              | 7B       | 59.3     | 7.54     | 48.2      | 49.4     | 46.5     | 57.5       | 63.8     | 73.5     | 59.9     |\n\n*: `Gemma-7b-it` failed to understand and follow most few-shot templates.\n\n## Usage\n\nTo use this model, we highly recommend installing the OpenChat package by following the [installation guide](https://github.com/imoneoi/openchat#installation) in our repository and using the OpenChat OpenAI-compatible API server by running the serving command from the table below. The server is optimized for high-throughput deployment using [vLLM](https://github.com/vllm-project/vllm) and can run on a consumer GPU with 24GB RAM. To enable tensor parallelism, append `--tensor-parallel-size N` to the serving command.\n\nOnce started, the server listens at `localhost:18888` for requests and is compatible with the [OpenAI ChatCompletion API specifications](https://platform.openai.com/docs/api-reference/chat). Please refer to the example request below for reference. Additionally, you can use the [OpenChat Web UI](https://github.com/imoneoi/openchat#web-ui) for a user-friendly experience.\n\nIf you want to deploy the server as an online service, you can use `--api-keys sk-KEY1 sk-KEY2 ...` to specify allowed API keys and `--disable-log-requests --disable-log-stats --log-file openchat.log` for logging only to a file. For security purposes, we recommend using an [HTTPS gateway](https://fastapi.tiangolo.com/es/deployment/concepts/#security-https) in front of the server.\n\n| Model                   | Size | Context | Weights                                                                | Serving                                                                                                                |\n|-------------------------|------|---------|------------------------------------------------------------------------|------------------------------------------------------------------------------------------------------------------------|\n| OpenChat-3.5-0106-Gemma | 7B   | 8192    | [Huggingface](https://huggingface.co/openchat/openchat-3.5-0106-gemma) | `python -m ochat.serving.openai_api_server --model openchat/openchat-3.5-0106-gemma --engine-use-ray --worker-use-ray` |\n\n<details>\n  <summary>Example request (click to expand)</summary>\n\n```bash\ncurl http://localhost:18888/v1/chat/completions \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\n    \"model\": \"openchat_3.5_gemma_new\",\n    \"messages\": [{\"role\": \"user\", \"content\": \"You are a large language model named OpenChat. Write a poem to describe yourself\"}]\n  }'\n```\n\n\n</details>\n\n## Conversation template\n\n⚠️ **Notice:** This is different from the Mistral version. End-of-turn token is `<end_of_turn>` now (Mistral version is `<|end_of_turn|>`). Remember to set `<end_of_turn>` as end of generation token.\n\n```\nGPT4 Correct User: Hello<end_of_turn>GPT4 Correct Assistant: Hi<end_of_turn>GPT4 Correct User: How are you today?<end_of_turn>GPT4 Correct Assistant:\n```\n\nWith system message (**NOT** recommended, may degrade performance)\n\n```\nYou are a helpful assistant.<end_of_turn>GPT4 Correct User: Hello<end_of_turn>GPT4 Correct Assistant: Hi<end_of_turn>GPT4 Correct User: How are you today?<end_of_turn>GPT4 Correct Assistant:\n```\n\n## Hallucination of Non-existent Information\nOpenChat may sometimes generate information that does not exist or is not accurate, also known as \"hallucination\". Users should be aware of this possibility and verify any critical information obtained from the model.\n\n## Safety\nOpenChat may sometimes generate harmful, hate speech, biased responses, or answer unsafe questions. It's crucial to apply additional AI safety measures in use cases that require safe and moderated responses.\n\n<div align=\"center\">\n<h2> License </h2>\n</div>\n\nOur OpenChat 3.5 code and models are distributed under the Apache License 2.0.\n\n## Citation \n\n```\n@article{wang2023openchat,\n  title={OpenChat: Advancing Open-source Language Models with Mixed-Quality Data},\n  author={Wang, Guan and Cheng, Sijie and Zhan, Xianyuan and Li, Xiangang and Song, Sen and Liu, Yang},\n  journal={arXiv preprint arXiv:2309.11235},\n  year={2023}\n}\n```\n\n<div align=\"center\">\n<h2> 💌 Contact </h2>\n</div>\n\n**Project Lead:**\n- Guan Wang [imonenext at gmail dot com]\n- [Alpay Ariyak](https://github.com/alpayariyak) [aariyak at wpi dot edu]\n\n\n",
    "related_quantizations": []
  },
  "tags": [
    "gguf",
    "arxiv:2309.11235",
    "endpoints_compatible",
    "region:us",
    "conversational"
  ],
  "likes": 0,
  "downloads": 197,
  "gated": false,
  "private": false,
  "last_modified": "2024-10-07T23:01:35.000Z",
  "created_at": "2024-10-07T20:03:04.000Z",
  "pipeline_tag": "",
  "library_name": ""
}

Source payload excerpt (from Hugging Face API)

{
  "_id": "67043e7825e14f0786e1ba1c",
  "id": "RichardErkhov/openchat_-_openchat-3.5-0106-gemma-gguf",
  "modelId": "RichardErkhov/openchat_-_openchat-3.5-0106-gemma-gguf",
  "sha": "671c7416232ac16d21403ccb686e2fb377308ee9",
  "createdAt": "2024-10-07T20:03:04.000Z",
  "lastModified": "2024-10-07T23:01:35.000Z",
  "author": "RichardErkhov",
  "downloads": 197,
  "likes": 0,
  "gated": false,
  "private": false,
  "pipeline_tag": "",
  "library_name": "",
  "siblings_count": 24
}