GraySoft
Projects Models About FAQ Contact Download guIDE โ†’

liquidai/lfm2-colbert-350m-gguf Q5_K_M GGUF - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.

Model Intelligence Sheet

liquidai/lfm2-colbert-350m-gguf overview

LFM2-ColBERT-350M is a late interaction retriever with excellent multilingual performance. It allows you to store documents in one language (for example, a product description in English) and retrieve them in many languages with high accuracy. Find more information about LFM2-ColBERT-350M in our blog post. ๐Ÿš€ Try our demo: https://huggingface.co/spaces/LiquidAI/LFM2-ColBERT

sentence-transformersggufliquidlfm2edgeColBERTPyLatesentence-similarityfeature-extractionllama.cppenarzhfrdejakoesbase_model:LiquidAI/LFM2-ColBERT-350Mbase_model:quantized:LiquidAI/LFM2-ColBERT-350Mlicense:otherendpoints_compatibleregion:usconversational
liquidai/lfm2-colbert-350m-gguf visual
Downloads
583
Likes
14
Pipeline
sentence-similarity
Library
sentence-transformers
Visibility
Public
Access
Open

Repository Files & Downloads

7 files detected
Direct downloads for all repository files
FileTypeQuantizationSizeLink
LFM2-ColBERT-350M-BF16.gguf GGUF BF16 676.53 MB Download
LFM2-ColBERT-350M-F16.gguf GGUF F16 676.53 MB Download
LFM2-ColBERT-350M-Q4_0.gguf GGUF โ€” 208.29 MB Download
LFM2-ColBERT-350M-Q4_K_M.gguf GGUF Q4_K_M 217.83 MB Download
LFM2-ColBERT-350M-Q5_K_M.gguf GGUF Q5_K_M 247.47 MB Download
LFM2-ColBERT-350M-Q6_K.gguf GGUF Q6_K 278.96 MB Download
LFM2-ColBERT-350M-Q8_0.gguf GGUF โ€” 360.58 MB Download

Model Details Live

Model Slug
liquidai/lfm2-colbert-350m-gguf
Author
LiquidAI
Pipeline Task
sentence-similarity
Library
sentence-transformers
Created
2026-01-05
Last Modified
2026-01-05
Gated
No
Private
No
HF SHA
a799c05750fa5786d0d5f47b1b5df05fe5a4a6b2
License
other
Language
en, ar, zh, fr, de, ja, ko, es
Base Model
LiquidAI/LFM2-ColBERT-350M

Metadata Inspector

Normalized metadata (stored in metadata_json)
{
  "metadata": {},
  "card_data": {
    "language": [
      "en",
      "ar",
      "zh",
      "fr",
      "de",
      "ja",
      "ko",
      "es"
    ],
    "tags": [
      "liquid",
      "lfm2",
      "edge",
      "ColBERT",
      "PyLate",
      "sentence-transformers",
      "sentence-similarity",
      "feature-extraction",
      "llama.cpp",
      "gguf"
    ],
    "pipeline_tag": "sentence-similarity",
    "license": "other",
    "license_name": "lfm1.0",
    "license_link": "LICENSE",
    "base_model": [
      "LiquidAI/LFM2-ColBERT-350M"
    ],
    "frontmatter": {
      "language": [
        "en",
        "ar",
        "zh",
        "fr",
        "de",
        "ja",
        "ko",
        "es"
      ],
      "tags": [
        "liquid",
        "lfm2",
        "edge",
        "ColBERT",
        "PyLate",
        "sentence-transformers",
        "sentence-similarity",
        "feature-extraction",
        "llama.cpp",
        "gguf"
      ],
      "pipeline_tag": "sentence-similarity",
      "license": "other",
      "license_name": "lfm1.0",
      "license_link": "LICENSE",
      "base_model": [
        "LiquidAI/LFM2-ColBERT-350M"
      ]
    },
    "hero_image_url": "https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/2b08LKpev0DNEk6DlnWkY.png",
    "summary": "LFM2-ColBERT-350M is a late interaction retriever with excellent multilingual performance. It allows you to store documents in one language (for example, a product description in English) and retrieve them in many languages with high accuracy. Find more information about LFM2-ColBERT-350M in our blog post. > [!NOTE] > ๐Ÿš€ Try our demo: https://huggingface.co/spaces/LiquidAI/LFM2-ColBERT",
    "quick_links": [],
    "benchmark_table_html": "",
    "readme_markdown": "---\nlanguage:\n- en\n- ar\n- zh\n- fr\n- de\n- ja\n- ko\n- es\ntags:\n- liquid\n- lfm2\n- edge\n- ColBERT\n- PyLate\n- sentence-transformers\n- sentence-similarity\n- feature-extraction\n- llama.cpp\n- gguf\npipeline_tag: sentence-similarity\nlicense: other\nlicense_name: lfm1.0\nlicense_link: LICENSE\nbase_model:\n- LiquidAI/LFM2-ColBERT-350M\n---\n\n<center>\n<div style=\"text-align: center;\">\n  <img \n    src=\"https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/2b08LKpev0DNEk6DlnWkY.png\" \n    alt=\"Liquid AI\"\n    style=\"width: 100%; max-width: 100%; height: auto; display: inline-block; margin-bottom: 0.5em; margin-top: 0.5em;\"\n  />\n</div>\n<div style=\"display: flex; justify-content: center; gap: 0.5em;\">\n  <a href=\"https://playground.liquid.ai/chat\">\n<a href=\"https://playground.liquid.ai/\"><strong>Try LFM</strong></a> โ€ข <a href=\"https://docs.liquid.ai/lfm\"><strong>Documentation</strong></a> โ€ข <a href=\"https://leap.liquid.ai/\"><strong>LEAP</strong></a></a>\n</div>\n</center>\n\n# LFM2-ColBERT-350M\n\nLFM2-ColBERT-350M is a late interaction retriever with excellent multilingual performance. It allows you to store documents in one language (for example, a product description in English) and retrieve them in many languages with high accuracy.\n\n- LFM2-ColBERT-350M offers **best-in-class accuracy** across different languages.\n- Inference speed is **on par with models 2.3 times smaller**, thanks to the efficient LFM2 backbone.\n- You can use it as a **drop-in replacement** in your current RAG pipelines to improve performance.\n\nFind more information about LFM2-ColBERT-350M in our [blog post](http://www.liquid.ai/blog/lfm2-colbert-350m-one-model-to-embed-them-all).\n\n\n> [!NOTE]\n> ๐Ÿš€ Try our demo: https://huggingface.co/spaces/LiquidAI/LFM2-ColBERT\n\n## ๐Ÿƒ How to run\n\nExample usage with [llama.cpp](https://github.com/ggml-org/llama.cpp):\n\nStart llama-server\n```bash\nllama-server -hf LiquidAI/LFM2-ColBERT-350M-GGUF --embeddings\n```\n\nMake requests to embed queries and documents, and compute similarity scores\n\n```bash\nโฏ uv run colbert-rerank.py\n\nScore: 29.69 | Q: What is panda? | D: hi\nScore: 29.83 | Q: What is panda? | D: it is a bear\nScore: 30.47 | Q: What is panda? | D: The giant panda (Ailuropoda melanoleuca), sometimes called a panda bear or simply panda, is a bear species endemic to China.\n```\n\n```python\n# /// script\n# requires-python = \">=3.10\"\n# dependencies = [\n#     \"transformers\",\n#     \"huggingface-hub\",\n#     \"numpy\",\n#     \"requests\",\n#     \"torch\",\n# ]\n# ///\n\n# colbert-rerank.py\nfrom transformers import AutoTokenizer\nfrom huggingface_hub import hf_hub_download\nimport numpy as np, requests, torch, torch.nn.functional as F, json\n\n\nmodel_id = \"LiquidAI/LFM2-ColBert-350M\"\ntokenizer = AutoTokenizer.from_pretrained(model_id)\nconfig = json.load(open(hf_hub_download(model_id, \"config_sentence_transformers.json\")))\nskiplist = set(\n    t\n    for w in config[\"skiplist_words\"]\n    for t in tokenizer.encode(w, add_special_tokens=False)\n)\n\n\ndef maxsim(q, d):\n    return (q @ d.T).max(dim=1).values.sum().item()\n\n\ndef preprocess(text, is_query):\n    prefix = config[\"query_prefix\"] if is_query else config[\"document_prefix\"]\n    toks = tokenizer.encode(prefix + text)\n    max_len = config[\"query_length\"] if is_query else config[\"document_length\"]\n    if is_query:\n        toks += [tokenizer.pad_token_id] * (max_len - len(toks))\n    else:\n        toks = toks[:max_len]\n    mask = None if is_query else [t not in skiplist for t in toks]\n    return toks, mask\n\n\ndef embed(content, mask=None):\n    emb = np.array(\n        requests.post(\n            \"http://localhost:8080/embedding\",\n            json={\"content\": content},\n        ).json()[0][\"embedding\"]\n    )\n    if mask:\n        emb = emb[mask]\n    emb = torch.from_numpy(emb)\n    emb = F.normalize(emb, p=2, dim=-1)  # L2 normalize each token embedding\n    return emb.unsqueeze(0)\n\n\ndocs = [\n    \"hi\",\n    \"it is a bear\",\n    \"The giant panda (Ailuropoda melanoleuca), sometimes called a panda bear or simply panda, is a bear species endemic to China.\",\n]\nquery = \"What is panda?\"\n\nq = embed(*preprocess(query, True))\nd = [embed(*preprocess(doc, False)) for doc in docs]\ns = [(query, doc, maxsim(q.squeeze(), di.squeeze())) for doc, di in zip(docs, d)]\nfor q_text, d_text, score in s:\n    print(f\"Score: {score:.2f} | Q: {q_text} | D: {d_text}\")\n\n```\n\nFind more details in the original model card: https://huggingface.co/LiquidAI/LFM2-ColBERT-350M\n",
    "related_quantizations": []
  },
  "tags": [
    "sentence-transformers",
    "gguf",
    "liquid",
    "lfm2",
    "edge",
    "ColBERT",
    "PyLate",
    "sentence-similarity",
    "feature-extraction",
    "llama.cpp",
    "en",
    "ar",
    "zh",
    "fr",
    "de",
    "ja",
    "ko",
    "es",
    "base_model:LiquidAI/LFM2-ColBERT-350M",
    "base_model:quantized:LiquidAI/LFM2-ColBERT-350M",
    "license:other",
    "endpoints_compatible",
    "region:us",
    "conversational"
  ],
  "likes": 14,
  "downloads": 583,
  "gated": false,
  "private": false,
  "last_modified": "2026-01-05T22:56:07.000Z",
  "created_at": "2026-01-05T22:35:53.000Z",
  "pipeline_tag": "sentence-similarity",
  "library_name": "sentence-transformers"
}
Source payload excerpt (from Hugging Face API)
{
  "_id": "695c3cc9f7fa80c597bcaabf",
  "id": "LiquidAI/LFM2-ColBERT-350M-GGUF",
  "modelId": "LiquidAI/LFM2-ColBERT-350M-GGUF",
  "sha": "a799c05750fa5786d0d5f47b1b5df05fe5a4a6b2",
  "createdAt": "2026-01-05T22:35:53.000Z",
  "lastModified": "2026-01-05T22:56:07.000Z",
  "author": "LiquidAI",
  "downloads": 583,
  "likes": 14,
  "gated": false,
  "private": false,
  "pipeline_tag": "sentence-similarity",
  "library_name": "sentence-transformers",
  "siblings_count": 11
}