Skip to content
Discussion options

You must be logged in to vote

Adding the operational layer on top of @generall's explanation — what to actually compute and how to keep it healthy in production.

1. Pre-compute avg_len once from your corpus, store it next to your chunker config.

# 1-pass over your corpus; avg_len in the same units the indexer counts
# (tokens AFTER your chunker / tokenizer step, not raw chars)
from statistics import mean

def chunk_token_lengths(corpus, chunker):
    for doc in corpus:
        for chunk in chunker(doc):
            yield len(chunk.tokens)   # whatever your tokenizer/sparse-encoder uses

AVG_LEN = round(mean(chunk_token_lengths(corpus, my_chunker)), 2)
# pin AVG_LEN alongside chunker_version, model_version in your config

Replies: 3 comments 9 replies

Comment options

You must be logged in to vote
3 replies
@Topo1995
Comment options

@generall
Comment options

@Topo1995
Comment options

Comment options

You must be logged in to vote
2 replies
@Topo1995
Comment options

@ibondarenko1
Comment options

Answer selected by Topo1995
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Category
Q&A
Labels
None yet
3 participants