liquidai/lfm2-colbert-350m-gguf Q8_0 GGUF - Free GGUF Download is indexed on GraySoft with repository links, GGUF quant files, and Hugging Face metadata. This page helps you pick a local model for guIDE or other runtimes. See related models in the same shard below.
Model Intelligence Sheet
liquidai/lfm2-colbert-350m-gguf overview
LFM2-ColBERT-350M is a late interaction retriever with excellent multilingual performance. It allows you to store documents in one language (for example, a product description in English) and retrieve them in many languages with high accuracy. Find more information about LFM2-ColBERT-350M in our blog post. ๐ Try our demo: https://huggingface.co/spaces/LiquidAI/LFM2-ColBERT
Downloads
583
Likes
14
Pipeline
sentence-similarity
Library
sentence-transformers
Visibility
Public
Access
Open
Repository Files & Downloads
7 files detected
Direct downloads for all repository files
| File | Type | Quantization | Size | Link |
|---|---|---|---|---|
| LFM2-ColBERT-350M-BF16.gguf | GGUF | BF16 | 676.53 MB | Download |
| LFM2-ColBERT-350M-F16.gguf | GGUF | F16 | 676.53 MB | Download |
| LFM2-ColBERT-350M-Q4_0.gguf | GGUF | โ | 208.29 MB | Download |
| LFM2-ColBERT-350M-Q4_K_M.gguf | GGUF | Q4_K_M | 217.83 MB | Download |
| LFM2-ColBERT-350M-Q5_K_M.gguf | GGUF | Q5_K_M | 247.47 MB | Download |
| LFM2-ColBERT-350M-Q6_K.gguf | GGUF | Q6_K | 278.96 MB | Download |
| LFM2-ColBERT-350M-Q8_0.gguf | GGUF | โ | 360.58 MB | Download |
Model Details Live
Metadata Inspector
Normalized metadata (stored in metadata_json)
{
"metadata": {},
"card_data": {
"language": [
"en",
"ar",
"zh",
"fr",
"de",
"ja",
"ko",
"es"
],
"tags": [
"liquid",
"lfm2",
"edge",
"ColBERT",
"PyLate",
"sentence-transformers",
"sentence-similarity",
"feature-extraction",
"llama.cpp",
"gguf"
],
"pipeline_tag": "sentence-similarity",
"license": "other",
"license_name": "lfm1.0",
"license_link": "LICENSE",
"base_model": [
"LiquidAI/LFM2-ColBERT-350M"
],
"frontmatter": {
"language": [
"en",
"ar",
"zh",
"fr",
"de",
"ja",
"ko",
"es"
],
"tags": [
"liquid",
"lfm2",
"edge",
"ColBERT",
"PyLate",
"sentence-transformers",
"sentence-similarity",
"feature-extraction",
"llama.cpp",
"gguf"
],
"pipeline_tag": "sentence-similarity",
"license": "other",
"license_name": "lfm1.0",
"license_link": "LICENSE",
"base_model": [
"LiquidAI/LFM2-ColBERT-350M"
]
},
"hero_image_url": "https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/2b08LKpev0DNEk6DlnWkY.png",
"summary": "LFM2-ColBERT-350M is a late interaction retriever with excellent multilingual performance. It allows you to store documents in one language (for example, a product description in English) and retrieve them in many languages with high accuracy. Find more information about LFM2-ColBERT-350M in our blog post. > [!NOTE] > ๐ Try our demo: https://huggingface.co/spaces/LiquidAI/LFM2-ColBERT",
"quick_links": [],
"benchmark_table_html": "",
"readme_markdown": "---\nlanguage:\n- en\n- ar\n- zh\n- fr\n- de\n- ja\n- ko\n- es\ntags:\n- liquid\n- lfm2\n- edge\n- ColBERT\n- PyLate\n- sentence-transformers\n- sentence-similarity\n- feature-extraction\n- llama.cpp\n- gguf\npipeline_tag: sentence-similarity\nlicense: other\nlicense_name: lfm1.0\nlicense_link: LICENSE\nbase_model:\n- LiquidAI/LFM2-ColBERT-350M\n---\n\n<center>\n<div style=\"text-align: center;\">\n <img \n src=\"https://cdn-uploads.huggingface.co/production/uploads/61b8e2ba285851687028d395/2b08LKpev0DNEk6DlnWkY.png\" \n alt=\"Liquid AI\"\n style=\"width: 100%; max-width: 100%; height: auto; display: inline-block; margin-bottom: 0.5em; margin-top: 0.5em;\"\n />\n</div>\n<div style=\"display: flex; justify-content: center; gap: 0.5em;\">\n <a href=\"https://playground.liquid.ai/chat\">\n<a href=\"https://playground.liquid.ai/\"><strong>Try LFM</strong></a> โข <a href=\"https://docs.liquid.ai/lfm\"><strong>Documentation</strong></a> โข <a href=\"https://leap.liquid.ai/\"><strong>LEAP</strong></a></a>\n</div>\n</center>\n\n# LFM2-ColBERT-350M\n\nLFM2-ColBERT-350M is a late interaction retriever with excellent multilingual performance. It allows you to store documents in one language (for example, a product description in English) and retrieve them in many languages with high accuracy.\n\n- LFM2-ColBERT-350M offers **best-in-class accuracy** across different languages.\n- Inference speed is **on par with models 2.3 times smaller**, thanks to the efficient LFM2 backbone.\n- You can use it as a **drop-in replacement** in your current RAG pipelines to improve performance.\n\nFind more information about LFM2-ColBERT-350M in our [blog post](http://www.liquid.ai/blog/lfm2-colbert-350m-one-model-to-embed-them-all).\n\n\n> [!NOTE]\n> ๐ Try our demo: https://huggingface.co/spaces/LiquidAI/LFM2-ColBERT\n\n## ๐ How to run\n\nExample usage with [llama.cpp](https://github.com/ggml-org/llama.cpp):\n\nStart llama-server\n```bash\nllama-server -hf LiquidAI/LFM2-ColBERT-350M-GGUF --embeddings\n```\n\nMake requests to embed queries and documents, and compute similarity scores\n\n```bash\nโฏ uv run colbert-rerank.py\n\nScore: 29.69 | Q: What is panda? | D: hi\nScore: 29.83 | Q: What is panda? | D: it is a bear\nScore: 30.47 | Q: What is panda? | D: The giant panda (Ailuropoda melanoleuca), sometimes called a panda bear or simply panda, is a bear species endemic to China.\n```\n\n```python\n# /// script\n# requires-python = \">=3.10\"\n# dependencies = [\n# \"transformers\",\n# \"huggingface-hub\",\n# \"numpy\",\n# \"requests\",\n# \"torch\",\n# ]\n# ///\n\n# colbert-rerank.py\nfrom transformers import AutoTokenizer\nfrom huggingface_hub import hf_hub_download\nimport numpy as np, requests, torch, torch.nn.functional as F, json\n\n\nmodel_id = \"LiquidAI/LFM2-ColBert-350M\"\ntokenizer = AutoTokenizer.from_pretrained(model_id)\nconfig = json.load(open(hf_hub_download(model_id, \"config_sentence_transformers.json\")))\nskiplist = set(\n t\n for w in config[\"skiplist_words\"]\n for t in tokenizer.encode(w, add_special_tokens=False)\n)\n\n\ndef maxsim(q, d):\n return (q @ d.T).max(dim=1).values.sum().item()\n\n\ndef preprocess(text, is_query):\n prefix = config[\"query_prefix\"] if is_query else config[\"document_prefix\"]\n toks = tokenizer.encode(prefix + text)\n max_len = config[\"query_length\"] if is_query else config[\"document_length\"]\n if is_query:\n toks += [tokenizer.pad_token_id] * (max_len - len(toks))\n else:\n toks = toks[:max_len]\n mask = None if is_query else [t not in skiplist for t in toks]\n return toks, mask\n\n\ndef embed(content, mask=None):\n emb = np.array(\n requests.post(\n \"http://localhost:8080/embedding\",\n json={\"content\": content},\n ).json()[0][\"embedding\"]\n )\n if mask:\n emb = emb[mask]\n emb = torch.from_numpy(emb)\n emb = F.normalize(emb, p=2, dim=-1) # L2 normalize each token embedding\n return emb.unsqueeze(0)\n\n\ndocs = [\n \"hi\",\n \"it is a bear\",\n \"The giant panda (Ailuropoda melanoleuca), sometimes called a panda bear or simply panda, is a bear species endemic to China.\",\n]\nquery = \"What is panda?\"\n\nq = embed(*preprocess(query, True))\nd = [embed(*preprocess(doc, False)) for doc in docs]\ns = [(query, doc, maxsim(q.squeeze(), di.squeeze())) for doc, di in zip(docs, d)]\nfor q_text, d_text, score in s:\n print(f\"Score: {score:.2f} | Q: {q_text} | D: {d_text}\")\n\n```\n\nFind more details in the original model card: https://huggingface.co/LiquidAI/LFM2-ColBERT-350M\n",
"related_quantizations": []
},
"tags": [
"sentence-transformers",
"gguf",
"liquid",
"lfm2",
"edge",
"ColBERT",
"PyLate",
"sentence-similarity",
"feature-extraction",
"llama.cpp",
"en",
"ar",
"zh",
"fr",
"de",
"ja",
"ko",
"es",
"base_model:LiquidAI/LFM2-ColBERT-350M",
"base_model:quantized:LiquidAI/LFM2-ColBERT-350M",
"license:other",
"endpoints_compatible",
"region:us",
"conversational"
],
"likes": 14,
"downloads": 583,
"gated": false,
"private": false,
"last_modified": "2026-01-05T22:56:07.000Z",
"created_at": "2026-01-05T22:35:53.000Z",
"pipeline_tag": "sentence-similarity",
"library_name": "sentence-transformers"
}
Source payload excerpt (from Hugging Face API)
{
"_id": "695c3cc9f7fa80c597bcaabf",
"id": "LiquidAI/LFM2-ColBERT-350M-GGUF",
"modelId": "LiquidAI/LFM2-ColBERT-350M-GGUF",
"sha": "a799c05750fa5786d0d5f47b1b5df05fe5a4a6b2",
"createdAt": "2026-01-05T22:35:53.000Z",
"lastModified": "2026-01-05T22:56:07.000Z",
"author": "LiquidAI",
"downloads": 583,
"likes": 14,
"gated": false,
"private": false,
"pipeline_tag": "sentence-similarity",
"library_name": "sentence-transformers",
"siblings_count": 11
}