What is a reranker in a RAG system?
A reranker takes the candidates returned by search and scores them again for relevance to the question, so the answering model receives a better ordered shortlist. It can improve which retrieved passages reach the prompt, but it cannot select a relevant document that the first search never found.
Level 4: How AI systems are built · 4.1 Retrieval and RAG
RAG & RetrievalModel Architecture
Search broadly, then inspect the shortlist
Retrieval has to find useful candidates in a collection. Reranking works on the candidates already found. In a common implementation, the first stage uses embeddings to search efficiently, while a cross-encoder scores each question and candidate passage together. The Sentence Transformers tutorial demonstrates that two-stage design.
In episode 49, Fergal Reid describes retrieval and reranking as separate parts of the system that supplies context to an answering model. That distinction is useful even if you use off-the-shelf components. You can investigate whether the evidence was absent from search results or present but ranked too low.
Watch an exception move up the list
Illustrative example: an employee asks whether a damaged laptop can be replaced before its normal replacement date. Search returns a general equipment page, a purchasing guide and a damage-exception procedure. All discuss laptops, but only one directly addresses the early replacement question.
A reranker could move the exception procedure above the general pages. The answering model then receives that procedure with the relevant ordinary replacement rule. The exact improvement must be measured on your documents; adding the component does not guarantee this ordering.
Keep the original candidate list and the reranked list for the test. If the exception procedure was never in the candidate list, the reranker cannot rescue this request. Inspect the search query, filters, source coverage and chunking instead. If the procedure was present and still ranked low, you have a more focused ranking problem to investigate.
Measure what reaches the answering model
Build a small set of questions with reviewed supporting passages. Include exact terminology, ordinary paraphrases and questions whose answer depends on an exception. Decide how many passages your application will pass to the model, then compare that shortlist with and without reranking.
Track whether the necessary evidence survives the selection. Also inspect the final answers: promoting a relevant passage is useful only if the answer handles it correctly. A single end-to-end score would hide which step changed, so keep retrieval and generation checks separate.
Measure the additional response time and compute cost. Try different candidate counts within your application’s budget. Inspect long passages as well as short ones, including whether the chosen implementation truncates away the relevant section. These are experiment settings, not universal numbers that every RAG system should copy.
Where it falls short
Relevance is not the same as authority or permission. A passage can match a question closely while being outdated, unapproved or unavailable to that user. Apply those rules deliberately rather than expecting a relevance score to settle them.
Reranking can also favor a passage that answers part of a question while missing a required qualification elsewhere. Check the assembled context, including how related passages fit together. Start with RAG evaluation to establish where failures occur, and add a reranker when your evidence points to selection or ordering as a worthwhile place to improve.
Hear it from the guest
“Then we have a custom retrieval model.”
“And we have a custom re-ranker model, it's completely custom re-ranker,”
Quotes lightly edited to remove filler words.
Go deeper
- Retrieve & Re-Rank Walks through first-stage retrieval followed by a cross-encoder that scores query-document pairs.
From the conversation
This explainer is drawn from these episodes — each carries its full transcript.