Skip to main content

Embedding Model Comparison

Fifteen GGUF configurations of eight embedding models, served by llama-server on one GPU host, measured for throughput and for retrieval quality on three code corpora (2026-10-03). The practical outcome — which model to index with — is on Embedding model choice. How llama-server itself compares with Ollama on the same host is on llama-server vs Ollama: measured configurations.

Method​

Hardware​

  • GPU host: Windows 11 mini-PC (Intel Core Ultra 9 285H) with an AMD Radeon RX 7800M eGPU (12 GB) and an Intel Arc 140T iGPU. llama.cpp build 11222, Vulkan.
  • Layout "3 RX + 1 Arc": three llama-server instances on the RX, one on the Arc, each -np 4 -fa on. The per-slot context is the model's training context: 8192 for jina v2 code, 2048 for CodeRankEmbed, 8192 for Muninn-small.

Speed​

Synthetic load from a LAN client: batches of 64 production chunks (about 670 characters on average), four requests in flight per server, about 45 s per configuration. The figure is aggregate texts per second. Every model was measured in the same session.

Quality​

Known-item retrieval. For each corpus an LLM wrote 200 natural-language queries, each from one sampled chunk, describing its behaviour without copying identifiers. The target of a query is that exact chunk among all chunks of the corpus.

  • Metrics: MRR and Recall@1, @5, @10.
  • dense is cosine similarity only; hybrid is dense fused with BM25 by reciprocal rank fusion (k = 60).
  • Set B (monolith sample only): 145 queries, each the source of a production Ruby method; the targets are the spec chunks that mention it. Metric: recall@10.

Corpora​

CorpusLanguageChunksFilesQuery
Sample of a 3.5M LoC production monolithRuby, RSpec tests165260 of its spec filesParaphrased test scenario (set A); method source (set B)
mastodonRuby, production app/2543414 (every third file)Behaviour description
TeaRAGsTypeScript, production src/, tests excluded3130353 (every fourth file)Behaviour description

Noise floor and caveats​

  • The corpora are samples of 1.6k–3.1k chunks; the full TeaRAGs index holds about 42k chunks and the monolith's about 180k. A full index has far more near-miss chunks, so absolute R@k there is lower than in these tables. The comparison between models holds: every model searched the same chunks with the same queries.
  • With 200 queries, one query is 0.5 percentage points. Differences below about 3 pp of R@1 or 0.02 MRR are noise.
  • For a few very long chunks the LLM read only the first 1.2–1.8k characters when writing their queries (5 on TeaRAGs, some on mastodon).
  • mxbai-embed-large inputs were truncated to 510 tokens: 78 of the 1024 texts in the speed set, 68 of the 1652 RSpec chunks.

Speed​

Layout 3 RX + 1 Arc unless noted.

ModelParamsGGUFtexts/svs jina
jina-embeddings-v2-base-code f16 (current default)161M323 MB3851.00×
jina v2 code Q8_0173 MB3700.96×
jina v2 code Q4_K_M109 MB3240.84×
CodeRankEmbed f16137M274 MB3560.92×
CodeRankEmbed Q8_0146 MB3590.93×
Muninn-small f1647M97 MB9142.37×
Muninn-small Q8_052 MB8442.19×
mxbai-embed-large f16 (510-token inputs)335M670 MB1610.42×
mxbai-embed-large Q8_0357 MB1640.43×
jina-code-embeddings-0.5b f16 / Q8_0494M994 / 531 MB45 / 450.12×
BGE-Code-v1 Q8_0 / Q4_01.5B1646 / 935 MB22 / 210.06×
Qodo-Embed-1-1.5B Q8_01.5B1646 MB230.06×
Nomic Embed Code Q4_K_M (1 RX + 1 Arc; three instances do not fit)7B4377 MB3.30.009×

A real index​

A full --force reindex of the TeaRAGs repository from the GPU host over the LAN. The two runs are a few hours apart, so the repository grew in between:

ModelFiles / chunksTotalEmbeddingEnrichmentchunks/s (embedding)
jina v2 code3741 / 41,626230 s215 s14 s193
Muninn-small f163765 / 41,990115 s101 s12 s417

The synthetic 2.37× shows up as 2.2× on embedding throughput in a real run. Both runs predate work stealing: the client still split each batch across servers by measured speed, which leaves fast servers idle at the end of a batch.

Quality​

Dense retrieval. MRR is 0–1; R@k is in percent.

TeaRAGs (TypeScript, production src/)​

ModelMRRR@1R@5R@10
jina v2 code f160.85076.097.099.5
CodeRankEmbed Q8_00.91885.599.5100
CodeRankEmbed f160.91384.599.5100
Muninn-small f160.88982.597.099.0
Muninn-small f16 + card prefixes0.85276.595.096.5
BGE-Code-v1 Q8_00.91786.098.099.0
BGE-Code-v1 Q8_0 + instruction0.95091.599.599.5

mastodon (Ruby, production app/)​

ModelMRRR@1R@5R@10
jina v2 code f160.83276.590.593.5
CodeRankEmbed f16 / Q8_00.891 / 0.89082.597.099.0 / 98.5
Muninn-small f160.78669.088.594.5
Muninn-small + card prefixes0.75364.589.091.5
BGE-Code-v1 Q8_00.89683.598.098.0
BGE-Code-v1 Q8_0 + instruction0.94390.599.0100

Sample of a 3.5M LoC production monolith (Ruby, RSpec tests)​

Set A: natural-language scenario → the test. Set B: method source → its tests (recall@10).

ModelA MRRA R@1A R@10B recall@10
jina v2 code f160.68355.595.582.4
jina Q8_0 / Q4_K_M0.682 / 0.69255.0 / 56.095.5 / 95.082.3 / 83.5
CodeRankEmbed f16 / Q8_00.865 / 0.87079.0 / 79.598.586.2 / 86.5
Muninn-small f160.76564.098.082.1
Muninn-small + card prefixes0.80571.097.579.6
mxbai-embed-large f16 / Q8_00.768 / 0.77764.5 / 66.097.080.1 / 80.3
jina-code-embeddings-0.5b f16 + prefixes0.78368.098.088.1
BGE-Code-v1 Q8_0 + instruction0.90085.099.086.7
Qodo-Embed-1-1.5B Q8_0 + instruction0.88682.010085.9
Nomic Embed Code 7B Q4_K_M + prefix0.93389.099.588.1

Among the models within 0.92–2.37× of jina's speed, CodeRankEmbed is first on all three corpora: +0.068 MRR over jina on TypeScript, +0.058 on mastodon, +0.187 on the RSpec corpus.

Findings​

Query prefixes​

Several model cards ask for a query prefix. CodeRankEmbed's prefix changed MRR by at most ±0.01 on every corpus. Muninn-small's card prefixes moved it by +0.04 on the RSpec corpus, −0.03 on mastodon and −0.04 on TypeScript — no consistent gain, so they are not used. TeaRAGs sends no prefixes.

Quantization​

Quantization never sped a model up on this GPU: jina Q8_0 and Q4_K_M ran at 0.96× and 0.84× of f16, Muninn-small Q8_0 at 0.92× of f16, CodeRankEmbed Q8_0 level with f16. Under Vulkan, dequantizing the weights costs more than the weight reads it saves. Quality did not move either (jina f16 / Q8_0 / Q4_K_M: 0.683 / 0.682 / 0.692 MRR on the RSpec corpus). Quantize only to save disk or VRAM.

Decoder embedders inside llama-server​

The decoder-based embedders (Qwen2 architecture) run about 4× below their raw GPU speed inside llama-server. llama-bench pp512 of jina-code-embeddings-0.5b reaches 22.5k tokens/s on the RX; the server delivers about 5k tokens/s. The gap does not depend on the backend (Vulkan or ROCm), the context, the slot count or a unified KV cache.

Hybrid search on identifier-free queries​

Hybrid (dense fused with BM25) scored far below dense for every model on these queries — jina on mastodon 0.594 against 0.832, CodeRankEmbed on TypeScript 0.657 against 0.918. This is a property of the benchmark: the queries deliberately contain no identifiers, so BM25 adds mostly noise. BM25 helped only set B, where the query is code. Do not read these numbers as a verdict on hybrid search with real agent queries, which usually do carry identifiers.

Why strong embedding models pay off little for an LLM agent​

The largest models — BGE-Code-v1, Qodo-Embed-1-1.5B, Nomic Embed Code 7B — are the most accurate in the tables above. For an LLM agent that calls the search, the evidence says they buy little:

  • R@10 converges. On TeaRAGs every model reaches 99–100% R@10; on mastodon 93.5–100%. The difference between models is almost entirely in R@1 — the order inside the top 10. An agent reads the whole top-10 page and picks the hit itself.
  • The big models' largest gain is the instruction prefix. BGE-Code-v1 goes from 0.917 to 0.950 MRR on TypeScript and from 0.896 to 0.943 on mastodon with it. Without it, the 1.5B BGE-Code-v1 ranks like the 137M CodeRankEmbed: 0.917 against 0.918 on TypeScript, 0.896 against 0.891 on mastodon. An instruction is a reformulation of the query — what an agent already does when it rewrites a query, retries with other words, or navigates from a hit with find_symbol, callers or similar-code search.
  • The cost is 15–110× in throughput. 22, 23 and 3.3 texts/s against 356–385. At these synthetic rates, embedding a 180k-chunk project takes about 2.3 h with BGE-Code-v1 and about 15 h with Nomic Embed Code 7B, against about 8 minutes with CodeRankEmbed. Their vectors are also 2–4.7× wider (1536–3584 dimensions against 768), which costs index size and search latency.

What we did not measure. The agent loop itself: how often an agent rewrites a query, and whether it finishes a task more often with one model than with another. The argument above rests on R@10 parity and on the instruction effect, not on an agent benchmark.

Models not covered​

CodeSage-small-v2 could not be measured: it uses a custom architecture with no GGUF converter, so llama-server cannot run it.