Qdrant BM25 average document length #8908
|
Hi, I’ve been reading through the documentation on setting up a collection for BM25 text search: https://qdrant.tech/documentation/search/text-search/ 1) Why does average document length need to be set for each individual point during upsert? 2) How should updates be handled if the average document length changes significantly over time? Thank you for your time and clarification. |
Replies: 3 comments 9 replies
The main reason is, that once document is converted into sparse vector representation, document length doesn't have much sense anymore. Document length is not equal to number of elements in sparse vector, as there are models which generate more than one element per token (see SPLADE and MiniCOIL). Document length only exist and text domain, and qdrant doesn't operate with texts directly.
The assumption is, that document length depends on the chunker settings, and even if you change documents, each chunk would stay about the same size. If this is not true for your case, it is unfortunately only possible to change by re-uploading dataset. |
|
Adding the operational layer on top of @generall's explanation — what to actually compute and how to keep it healthy in production. 1. Pre-compute # 1-pass over your corpus; avg_len in the same units the indexer counts
# (tokens AFTER your chunker / tokenizer step, not raw chars)
from statistics import mean
def chunk_token_lengths(corpus, chunker):
for doc in corpus:
for chunk in chunker(doc):
yield len(chunk.tokens) # whatever your tokenizer/sparse-encoder uses
AVG_LEN = round(mean(chunk_token_lengths(corpus, my_chunker)), 2)
# pin AVG_LEN alongside chunker_version, model_version in your configStoring it next to 2. Send Many users put a per-point computation here, which fights the design. The expected pattern is: points.append(PointStruct(
id=chunk_id,
vector={"text-sparse": sparse_vec},
payload={"text": chunk_text, "doc_id": doc_id, "avg_len": AVG_LEN}, # same constant for every point in this corpus
))If you have heterogeneous corpora (e.g., short titles + long bodies in the same collection), give each a separate 3. Drift detection — alert before re-uploading. The maintainer's point about "chunks stay about the same size" holds only as long as your chunker and content distribution don't shift. Watch for drift cheaply: # Run weekly against a fresh sample of newly upserted chunks
sample = client.scroll("my_collection", limit=10_000, with_payload=["text"]).points
current_avg = mean(len(my_chunker_tokenize(p.payload["text"])) for p in sample)
drift_pct = abs(current_avg - AVG_LEN) / AVG_LEN
# alert when drift > 0.2 — empirically the point where BM25 ranking starts to noticeably miscalibrateWhat "wrong" 4. Re-upload pattern that avoids downtime. If drift triggers a re-upload, blue-green at the collection level beats in-place updates: Doing it in-place via per-point payload updates also works ( 5. If your corpus is genuinely heterogeneous (titles + bodies + descriptions in one collection), split by source. In a single collection where chunk length actually varies by an order of magnitude per source, BM25's single- Recipe: pin |
Adding the operational layer on top of @generall's explanation — what to actually compute and how to keep it healthy in production.
1. Pre-compute
avg_lenonce from your corpus, store it next to your chunker config.