Skip to main content

Embedding model choice

Which embedding model to index a project with. Three models are worth considering; the provider decides which of them you can run.

ModelRoleProviderSpeed vs jinaDimsContext per slot
CodeRankEmbed (Q8_0)Recommended for llama-serverllama-server0.93×7682048
jina v2 code (f16)Ollama default, the baselineOllama, llama-server1.00×7688192
Muninn-small (f16)Fast optionllama-server2.37×3848192

Every number on this page comes from Embedding model comparison: the hardware, corpora, query construction, the full tables and the caveats are there. Speed is texts per second on one GPU host (three llama-server instances on an RX 7800M plus one on an Arc 140T iGPU); quality is known-item retrieval MRR over 200 natural-language queries per corpus. Differences below about 0.02 MRR are noise.

Candidates​

CodeRankEmbed​

nomic-ai/CodeRankEmbed, 137M parameters. The recommended model when you run llama-server.

When to pick it. Any project you index through llama-server, and Ruby projects in particular. It was the best of the fast models on all three corpora we measured.

Trained languages (model card): Python, Java, JavaScript, Go, PHP, Ruby (the CoRNStack dataset). TypeScript is not on the list.

Measured here (dense MRR, Q8_0; jina v2 code in brackets):

CorpusCodeRankEmbedjina v2 code
TypeScript, TeaRAGs src/0.9180.850
Ruby, mastodon app/0.8900.832
Ruby, RSpec sample of a 3.5M LoC production monolith0.8700.683
Ruby, method source → its specs (recall@10)86.582.4

TypeScript works although it is not in the training list. The query prefix from the model card changed MRR by at most ±0.01 on every corpus; TeaRAGs sends no prefix.

Speed. 359 texts/s against 385 for jina (0.93×). f16 runs at 356 texts/s with the same quality, so Q8_0 is the smaller file at no cost.

Dimensions768
Context per slot2048 tokens (the model's training context)
LicenseMIT
GGUF146 MB (Q8_0), 274 MB (f16)

How to run it. CodeRankEmbed is not in the Ollama library. Take the community GGUF from handwoven8588/CodeRankEmbed-GGUF, or convert nomic-ai/CodeRankEmbed yourself with llama.cpp's convert_hf_to_gguf.py. Print the launch lines with tea-rags llama-server command --model <path to the GGUF on the host>, then replace the printed -c 32768 -b 8192 -ub 8192 with the values for its 2048-token context — four slots of 2048:

llama-server -m CodeRankEmbed-Q8_0.gguf --embedding -ngl 999 -fa on \
-np 4 -c 8192 -b 2048 -ub 2048 --host 0.0.0.0 --port 8081
export EMBEDDING_PROVIDER=llama-server
export EMBEDDING_MODEL=nomic-ai/CodeRankEmbed
export EMBEDDING_BASE_URL=http://192.168.1.71:8081,8082,8083,8084

TeaRAGs knows the 768-dimension width from its model table and derives the chunk size from the 2048-token per-slot window, so chunks are smaller than with jina. Keep CodeRankEmbed in the GGUF file name: the served-model check retires an endpoint whose file does not look like EMBEDDING_MODEL.

jina v2 code​

jinaai/jina-embeddings-v2-base-code, 161M parameters, served as unclemusclez/jina-embeddings-v2-base-code:latest. The default for Ollama and llama-server, and the baseline every other model here is compared against.

When to pick it. You run Ollama: of the three, it is the only one in the Ollama library. Also when the project's languages are outside both other models' training lists.

Trained languages (model card): 30 programming languages.

Measured here (dense MRR, f16): 0.850 on TypeScript, 0.832 on mastodon, 0.683 on the RSpec corpus, 82.4 recall@10 on method → specs. It is the weakest of the three on every corpus except mastodon, where Muninn-small is lower.

Speed. 385 texts/s on the GPU host — the reference for the 1.00× column. Q8_0 (370) and Q4_K_M (324) were slower than f16, with the same quality.

Dimensions768
Context per slot8192 tokens
LicenseApache-2.0
GGUF323 MB (f16)

How to run it. Ollama pulls it on first use. For llama-server, tea-rags llama-server command prints the download and the launch lines with -np 4 -c 32768 -b 8192 -ub 8192 as they are; see llama-server.

Muninn-small​

brokkai/Muninn-small, 47M parameters. The fast option: 2.37× jina's throughput. TeaRAGs indexes its own repository with it.

When to pick it. TypeScript- or Python-heavy projects without Ruby, a small GPU host, or when full reindex time matters more than the last points of ranking quality. A full --force reindex of the TeaRAGs repository (about 42,000 chunks) took 115 s with Muninn-small against 230 s with jina; embedding alone 101 s against 215 s.

Trained languages (model card): C, C++, C#, Go, Java, JavaScript, TypeScript, PHP, Python, Rust, Scala. No Ruby.

Measured here (dense MRR, f16):

CorpusMuninn-smalljina v2 code
TypeScript, TeaRAGs src/0.8890.850
Ruby, mastodon app/0.7860.832
Ruby, RSpec sample of a 3.5M LoC production monolith0.7650.683
Ruby, method source → its specs (recall@10)82.182.4

Ruby is mixed — better than jina on one corpus, worse on another — so do not use it for Ruby. The model card's query and document prefixes moved MRR by +0.04, −0.03 and −0.04 on the three corpora; TeaRAGs sends no prefixes.

Speed. 914 texts/s against 385 for jina. Q8_0 was slower (844 texts/s); use f16.

Dimensions384
Context per slot8192 tokens
LicenseApache-2.0
GGUF97 MB (f16)

How to run it. There is no public GGUF. Convert it with llama.cpp:

git clone https://huggingface.co/brokkai/Muninn-small
python llama.cpp/convert_hf_to_gguf.py Muninn-small \
--outtype f16 --outfile Muninn-small-f16.gguf

With an older transformers the conversion fails on the tokenizer: the model's tokenizer_config.json declares "tokenizer_class": "TokenizersBackend". Change it to "PreTrainedTokenizerFast" and run the conversion again.

The printed launch lines fit as they are — four slots of 8192:

llama-server -m Muninn-small-f16.gguf --embedding -ngl 999 -fa on \
-np 4 -c 32768 -b 8192 -ub 8192 --host 0.0.0.0 --port 8081
export EMBEDDING_PROVIDER=llama-server
export EMBEDDING_MODEL=brokkai/Muninn-small

Decision table​

Project languagesOllama onlyllama-server, GPU hostllama-server, small host or reindex time first
Ruby, or Ruby mixed with othersjina v2 codeCodeRankEmbedCodeRankEmbed
TypeScript, JavaScript, Pythonjina v2 codeCodeRankEmbedMuninn-small
Go, Java, PHPjina v2 codeCodeRankEmbed (not measured)Muninn-small (not measured)
C, C++, C#, Rust, Scalajina v2 codeMuninn-small (not measured)Muninn-small (not measured)
Other languagesjina v2 codejina v2 code (not measured)jina v2 code (not measured)

"Not measured" rows follow the training lists on the model cards; we measured only TypeScript and Ruby.

Switching models​

  • A model switch needs a full --force reindex of the project. The vectors change, the width may change (768 → 384 for Muninn-small), and the per-slot context changes the chunk size. The collection's embedding model guard refuses to mix vectors of two models.
  • One model per llama-server process group. A llama-server process serves the model it loaded. To serve two models on one host, run a separate group of instances on its own port range for each; see several models on one host.
  • The fallback serves the same model. EMBEDDING_FALLBACK_URL must point at a llama-server with the same GGUF as the peers. The served-model check compares every endpoint's GGUF file name with EMBEDDING_MODEL and takes an endpoint on another model out of the run; when none serving the configured model is left, indexing fails with INFRA_LLAMA_SERVER_MODEL_MISMATCH instead of writing vectors of the wrong model.
  • Projects on different models share one MCP server. Each project is searched with the model and endpoints it was indexed with, as recorded in the registry; the server's EMBEDDING_MODEL and EMBEDDING_BASE_URL are only defaults for unregistered collections. TeaRAGs indexes itself with Muninn-small while the server default is CodeRankEmbed. See which model a project uses.

Models we measured and do not recommend​

Full numbers are in the comparison.

  • mxbai-embed-large — 2.4× slower than jina and a 512-token context; strong on natural language, behind CodeRankEmbed on code.
  • jina-code-embeddings-0.5b — 8.5× slower than jina, and its CC-BY-NC license rules out commercial use.
  • BGE-Code-v1, Qodo-Embed-1-1.5B — 1.5B parameters, 22–23 texts/s (0.06× jina); without an instruction prefix BGE-Code-v1 ranks like CodeRankEmbed.
  • Nomic Embed Code — 7B parameters, 3.3 texts/s (0.009× jina); three instances do not fit on a 12 GB GPU.
  • CodeSage-small-v2 — cannot run on llama-server: custom architecture with no GGUF converter.

Why the large models buy little for an agent: Why strong embedding models pay off little for an LLM agent.