Replies: 2 comments
|
The "many collections is an antipattern" note in the docs targets per-tenant / per-user collections at scale (hundreds–thousands of them), not two logically distinct corpora. Splitting into a Two reasonable options: A — single collection + payload filter. Keep everything in one collection, tag points with B — two collections. Justified when the two sets have different lifecycles, and yours do: system data is shipped/replaced wholesale on app updates, user data must survive untouched. Keeping them separate lets you snapshot and replace the system corpus independently. For shipping new system-data versions: snapshots are per-collection, so option B fits your delivery plan directly. Combine it with collection aliases for zero-downtime swaps — load the new snapshot into So given your update-via-snapshot requirement, two collections + aliases is the cleaner choice here; reach for the single-collection payload-filter approach only if you'd rather avoid managing two collections and the lifecycles converge. |
|
I would decide this from lifecycle boundaries first, then from Qdrant collection count. For your case, two collections is a reasonable design because the two corpora have different ownership and update semantics:
The anti-pattern is usually one collection per user/project/document at high scale. Two stable collections for two lifecycles is not the same problem. A practical setup:
If you keep everything in one collection, use payload fields such as So my default recommendation for this exact scenario would be two collections: |
Uh oh!
There was an error while loading. Please reload this page.
Hey everyone!
I am currently planning on creating a RAG application using Langchain4J and I would like to get some input. The system will have two kinds of data:
Data provided by the system. This can be a a manual or other documents. It is data that will be cleaned up, split into chunks and ingested into qdrant. It might be a good idea to prepare all this data during development (or through some pipeline).
Data created by the user of the application. Users are able to create articles (e.g FAQ-texts) which we want to also add into our RAG data. If a user deletes such an article, we can delete it from qdrant, modify it etc.
Now when there is an update to the software, there might also be a new version of the manual and related data, which we will have to update then. (Maybe by delivering a new qdrant-snapshot file to replace the old collection or having some scripts to replace the data in the collection).
I thought about creating two collections: one for the users data and one for the systems data. The FAQ says that having multiple collections might be considered an antipattern: https://qdrant.tech/documentation/faq/qdrant-fundamentals/#how-many-collections-can-i-create
So I would like to get some input if that is a terrible idea or a valid plan for my use case. Are there problems with delivering qdrant-snapshots for data updates?
All reactions