Contents
Map

12 · RAG

Retrieval Evaluation

View as:

Code Lab 01 - Retrieval Evaluation

Measure retrieval instead of guessing: build BM25, dense, hybrid, reranked, learned-sparse and quantized retrievers over the same public benchmark, compute nDCG@10, Recall@k and MRR yourself, and see which "best practices" actually help on this data.

← Back to Overview: RAG · Concepts: Embeddings & Vector Search · Retrieval & Reranking · RAG Evaluation

Learning objectives 2 hours
By the end of this page you will be able to:
  • Implement nDCG@k, Recall@k and MRR from their definitions and run them on labelled queries
  • Build BM25, dense, hybrid (RRF) and reranked retrieval over one corpus and compare them fairly
  • Measure the quality cost of int8 and binary embedding quantization with rescoring
  • Explain from your own results why hybrid retrieval helps and why an out-of-domain reranker may not
Prerequisites
  • RAG Evaluation - retrieval metrics
  • Python 3.10+; about 2 GB free RAM (4 GB with --splade)

What's In This Lab

PropertyDetail
DatasetBEIR SciFact: 5,183 scientific abstracts, 300 test claims with relevance labels (downloaded from Hugging Face)
SystemsBM25 (bm25s), dense (all-MiniLM-L6-v2), hybrid RRF, hybrid + cross-encoder rerank, optional SPLADE, optional int8 / binary quantization
MetricsnDCG@10, Recall@10, Recall@100, MRR@10 - implemented in the script
VerifiedFull run on an Apple-silicon laptop (CPU/MPS): ~2.5 min for the default systems, ~8 min with --splade, using sentence-transformers 6.1 and bm25s 0.3
Files01-Retrieval-Evaluation/{retrieval_lab.py, requirements.txt}
flowchart LR
    C["📚 SciFact corpus<br/>+ 300 labelled claims"] --> B["🔤 BM25"]
    C --> D["🧬 Dense"]
    C --> S["🧮 SPLADE (optional)"]
    B --> H["🔀 RRF"]
    D --> H
    H --> R["🎯 Cross-encoder<br/>rerank top 50"]
    D --> Q["🗜️ int8 / binary<br/>+ rescore (optional)"]
    B & D & S & H & R & Q --> M["📏 nDCG@10, R@10,<br/>R@100, MRR@10"]

    style H fill:#d8dfe8,stroke:#b0bac8
    style R fill:#e8e0d4,stroke:#c8b89a
    style M fill:#dde4dc,stroke:#b0c4b0

Run It

cd 12-RAG/CodeLabs/01-Retrieval-Evaluation
pip install -r requirements.txt

python retrieval_lab.py --limit 50              # quick run on 50 queries
python retrieval_lab.py                         # BM25, dense, hybrid, rerank
python retrieval_lab.py --quantization --splade # everything
python retrieval_lab.py --embed-model BAAI/bge-small-en-v1.5

A reference run (300 queries):

system                                  nDCG@10    R@10   R@100  MRR@10
BM25                                      0.662   0.774   0.876   0.631
Dense                                     0.648   0.788   0.925   0.607
Hybrid (RRF)                              0.690   0.819   0.955   0.655
Hybrid + rerank top-50                    0.686   0.809   0.955   0.655
SPLADE (learned sparse)                   0.710   0.820   0.942   0.679
dense int8 (4x smaller)                   0.622   0.782   0.922   0.576
dense binary + rescore (32x smaller)      0.612   0.754   0.900   0.572

The BM25 and SPLADE numbers are close to published BEIR results for SciFact (about 0.665 and 0.70), which is a useful sanity check that the metrics are implemented correctly.

Walkthrough - What to Look At

  1. BM25 beats the small dense model at the top of the ranking (nDCG@10 0.662 vs 0.648). Scientific claims are full of precise terms; a general-purpose 22M-parameter embedding model trained on web data doesn't know them. Dense wins on Recall@100 (0.925 vs 0.876) - it finds more of the relevant documents, just not always at the top.
  2. Hybrid beats both on every metric. The two retrievers make different mistakes, and RRF rewards documents both agree on. Recall@100 of 0.955 means a reranker has almost everything it needs in the candidate set.
  3. The reranker doesn't help here. ms-marco-MiniLM-L6-v2 was trained on web search queries; scientific claim verification is a different task and domain. A reranker is a model with a training distribution - evaluate it rather than assuming it helps.
  4. SPLADE is the best single retriever - learned term weights plus expansion - but encoding the corpus took ~4 minutes on a laptop versus ~25 s for the small dense model, and needs more memory (the vocabulary-sized output per token).
  5. Quantization costs a few points with this model: int8 keeps 96% and binary-plus-rescore 94% of dense nDCG@10, for 4× and 32× less memory. Models trained for quantization lose less.
  6. Read ndcg_at_k, recall_at_k and mrr_at_k - each is a few lines, and writing them yourself removes the mystery from leaderboard numbers.

Check Yourself

Check yourself
0 / 3 answered
  1. Hybrid has Recall@100 = 0.955 but nDCG@10 = 0.690. What does that tell you about where to invest next?
  2. Why does the reranker in this lab slightly lower nDCG@10?
  3. Why is matching published BEIR numbers for BM25 a useful check?

Exercises

Exercise - A better embedding model

Re-run with --embed-model BAAI/bge-small-en-v1.5 and with a larger model of your choice (check its model card for query prefixes - some expect "Represent this sentence for searching relevant passages: " or "query: "). Add the prefix handling to the script where needed. Does dense now beat BM25? Does hybrid still help?

Hint

SentenceTransformer.encode accepts prompt= for a query prefix; some models define named prompts in their config.

Solution

In our run, bge-small-en-v1.5 (without a prefix) reached nDCG@10 0.721 - well above BM25's 0.662 - but equal-weight RRF hybrid scored 0.715, slightly below dense alone, while Recall@100 still rose (0.955 → 0.965). When one retriever is much stronger, equal-weight fusion can dilute its top ranks; weight the fusion toward the stronger retriever, or use hybrid as the candidate generator for a good reranker. Check the model card for query prefixes - they typically add a little more.

Exercise - Swap the reranker

Replace the reranker with a larger, more general one (e.g. BAAI/bge-reranker-v2-m3; a GPU helps) and vary --rerank-top-n between 20, 50 and 100. Plot nDCG@10 against rerank time.

Solution

A stronger general reranker should now improve on hybrid, with diminishing returns as top-n grows (Recall@100 caps what reranking can reach). The plot is the latency-quality curve you'd use to choose top-n in production.

Exercise - Your own corpus

Replace load_scifact() with your own documents and 50 labelled queries (query id → relevant doc ids). Run all systems and write a one-paragraph recommendation for your domain.

Solution

The recommendation should cite your numbers: which retriever wins at nDCG@10, whether hybrid helps, whether a reranker helps in your domain, and whether quantization is acceptable - rather than defaulting to the configuration that won on SciFact.

References

Last reviewed: 2026-09

⚡AI-assisted content - always verify, always explore multiple perspectives·