System prompt selection by qdrant matching in Synaplan #8949
orgaralf
started this conversation in
Show and tell
Replies: 1 comment
|
We've seen similar improvements in our RAG systems by fine-tuning the confidence threshold for vector matching. 45% confidence seems reasonable, depending on the embedding model and dataset. In our production experience with Qdrant, we've used Hugging Face's sentence-transformers to generate embeddings, and adjusting the confidence threshold helped balance precision and recall. For instance, we use from qdrant_client import QdrantClient
client = QdrantClient(url="http://localhost:6333")
search_result = client.search(
collection_name="system_prompts",
query_vector=query_embedding,
limit=1,
score_threshold=0.4
)For larger prompts, you might consider chunking the input or using a more advanced embedding model to maintain precision. We use a combination of both to handle diverse user requests. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Hi folks,
we went a bit crazy with our "prompt routing" and it works quite well: https://github.com/metadist/synaplan/releases/tag/v3.0.0
After integrating qdrant, we have kicked out gpt-oss-20b and 120b as a routing LLM and let the Vector DB decide, what the correct system prompt and model for the user request is.
That work really good. The trick is to accept a confidence level of 45% or so...
If you prompt "Create an image of a cat", this is routed to the configured image model correctly. If you put in larger prompts, the precision goes down, but that was expected.
Good work, coders - thanks for a pretty fast VDB!
All reactions