What is the best text embedding model for ecommerce product search (short, noisy user queries)? #7679
Replies: 2 comments
|
We have positive experience using snowflake-arctic-embed-m-v2.0 for e-commerce, but Marqo-Ecommerce-L would be a better match for your requirements. |
|
I would avoid choosing a single model before testing the retrieval design, because these queries contain two different signals:
A dense-only system is likely to blur the second category. My starting baseline would be:
For an English, GPU-backed benchmark, I would test at least:
Marqo publishes ecommerce evaluation sets ( Build a small judged set from real queries/clicks and include hard negatives from the same category with one wrong specification (2 hp vs 1 hp, M12 vs M10, 1000 L vs 500 L). Compare Recall@20/100 and nDCG or MRR, then measure end-to-end p95 latency and GPU throughput at your 100 QPS target. Fine-tune only after inspecting which failure class remains; hard-negative training on specification mismatches is usually more valuable than generic extra pairs. For typo handling, normalize/expand the query before retrieval (units, common abbreviations, spelling), but search with both the original and normalized form so a correction cannot erase an exact product code. Qdrant references for the retrieval shape:
Model cards/evaluation data: |
Uh oh!
There was an error while loading. Please reload this page.
I am integrating a vector-based semantic search system into a large ecommerce platform's product search, and I want to select the right text embedding model.
Use Case
User queries are often:
Each product has:
I need embeddings that capture semantic meaning across these fields and match them with noisy, spec-heavy queries.
Constraints / Setup
Questions
Example queries
Any guidance or practical experience with embedding models for ecommerce search would be appreciated.
All reactions