Writing

When BM25 beat the embeddings

In short: on legal text, a keyword ranking function from the 1990s wins a whole class of queries that dense embeddings lose: citations, section numbers, defined terms. I run both and fuse the…

published
read time
5 min
words
905
lang
en
filed under
Engineering

In short: on legal text, a keyword ranking function from the 1990s wins a whole class of queries that dense embeddings lose: citations, section numbers, defined terms. I run both and fuse the rankings, and I decide by query type, not by an average.

The query that started it

The first version of search I built over legal text was embeddings only. Dense embeddings over chunked legislation and decisions, a vector store, the top results into the model. Plain-language questions work well in a setup like that. Now type a neutral citation, the year, court code and number a lawyer uses to name a case. A common failure looks like a handful of decisions about similar things, and not the case you named.

That's not a bug in the embedding model. It's what embeddings are for. They map text to meaning, and a citation string has almost no meaning. "Section 7" and "section 15" sit very close together in that space, because they look alike and are about the same kind of thing. To a lawyer they're different rules.

BM25 doesn't care about meaning. It scores a passage by how many of the query's exact terms it contains, weighted by how rare each term is across the corpus and adjusted for passage length. A rare token like a case number or a defined term is exactly what it's good at.

Which queries go which way

Reading the misses from a retrieval eval, the split is easy to see once you sort queries by kind. This is the general pattern, not a measured benchmark from my system.

Query typeShape of the queryUsually winsWhy
Citation lookupA neutral citation or a statute and section numberBM25Rare exact tokens, no meaning to embed
Defined termA term the act defines, in quotesBM25The exact string is the signal
Party nameA company or person named in a caseBM25Names are rare tokens
Plain-language question"Can my employer cut my hours without notice?"EmbeddingsNone of the statute's words appear in it
Concept search"duty to accommodate, undue hardship"Close, hybrid bestShared terms and shared meaning

A user doesn't know which kind of query they're typing, and a lawyer types all five before lunch. So picking one retriever means losing a row of that table every time.

bm25 dense hybrid bm25 dense hybrid citation queries plain questions 0 recall
The shape, not my numbers: each retriever wins one group, and fusion holds up in both.

Fusing two rankings

BM25 scores and cosine similarities live on different scales, so adding them is meaningless unless you normalise, and normalising is fiddly. Reciprocal rank fusion skips the problem. It throws the scores away and only uses each passage's position in each list.

def rrf(rankings, k=60, limit=10):
    """rankings: lists of passage IDs, best first, one list per retriever."""
    scores = {}
    for ranking in rankings:
        for rank, pid in enumerate(ranking, start=1):
            scores[pid] = scores.get(pid, 0.0) + 1.0 / (k + rank)
    return sorted(scores, key=scores.get, reverse=True)[:limit]

dense = vector_search(query, limit=50)
sparse = bm25_search(query, limit=50)
top = rrf([dense, sparse])

A passage that ranks well in either list gets a good fused score, and one that ranks well in both gets the best. The constant k damps the advantage of being first; 60 is the usual default and I haven't found a reason to move it. Take more candidates from each retriever than you need at the end, so a passage at rank 30 in one list still has a chance.

Backfilling the keyword side

The awkward part was not the fusion. It was that I already had an index full of embedded chunks and no keyword index beside it. Adding BM25 meant a backfill: walk every existing chunk, tokenise it the same way queries will be tokenised, and write it to the keyword index under the same passage ID as the vector.

Three things I'd tell myself before starting:

  • Same IDs on both sides. Fusion joins on passage ID. If one side uses chunk IDs and the other uses document IDs, nothing lines up and the fused list quietly degrades to whichever side has more hits.
  • Tokenise for law, not for English. A default tokenizer that splits "s. 15(1)" into loose pieces, or drops numbers as noise, throws away exactly the tokens BM25 was brought in for. Keep section numbers and citation parts as tokens.
  • Make it resumable. A backfill over a large corpus will stop halfway at least once. Write in batches, record the last ID done, and let a rerun skip what's already there.

Decide by query type, not by the average

The easy mistake is to run the eval, see that embeddings have higher overall recall than BM25, and stop there. An average like that is dominated by plain-language questions, because that's most of what people write when they build a question set. Citation queries are a minority of the rows, so embeddings can lose most of them and still win the average. A lawyer who types a citation and doesn't get the case won't type a second one.

So here is what I'd do in your place. Tag every question in your retrieval eval with its kind. Run BM25 alone, embeddings alone, and the fused list, and print recall at your k for each kind separately. If one retriever wins some rows and loses others, you don't need to pick. Fuse them, and keep the per-kind numbers in front of you every time you change something.

related

Keep reading