qdrant-search-quality-diagnosis
Diagnoses Qdrant search quality issues. Use when someone reports 'results are bad', 'wrong results', 'not relevant results', 'missing matches', 'recall is low', 'approximate search worse than exact', 'which embedding model', 'should I fine-tune my embedding model', 'quality dropped after quantizatio
- 0
- Installs
- —
- Rating
- —
- Success rate
- 1
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 dc23c5326145d60f… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
How to Diagnose Bad Search Quality
Before diagnosing or evaluating search quality, establish two different ground truths: a labeled query set measures whether the results are the right ones; exact KNN measures whether ANN finds what brute force would (no labels needed). The first diagnoses relevance, the second diagnoses the index.
Need a Labeled Baseline to Score Quality or Validate a Gain
Use when: user has no golden set, asks "how do I know if my search is good?", or needs to gate releases on a retrieval metric. Every fix in the sections below should be validated this way.
- Pick the metric by usage:
Recall@kfor RAG,MRR/Hits@1for single-answer,NDCG@kfor re-ranking Choosing the metric - Build a labeled query set (human, log-based, or LLM-synthetic) and score retrieval with
ranxMeasuring Retrieval Relevance - When tuning and evaluating search quality, set a minimum target gain, then make sure your labeled query set is large enough to measure it reliably. Quantify the uncertainty around a measured gain with a confidence interval (e.g. bootstrap the per-query gains, 95%): if it includes zero, the improvement is inconclusive and may be due to query-to-query variation. Adding labeled queries narrows the interval; how many you need depends on the size of the gain you are targeting and on query-to-query variation, not on collection size Before tuning a collection
- If you use the labeled queries to tune or select a configuration, evaluate the final choice on held-out queries. A gain that doesn't survive the held-out evaluation isn't reliable evidence of an improvement.
- For full RAG pipelines, also score generation with Ragas and use the retrieval-vs-generation 2x2 to isolate regressions Pipeline Output Quality
- Gate CI on a per-metric threshold to catch regressions from embedding-model swaps, prompt changes, or index config changes
Don't Know What's Wrong Yet
Use when: results are irrelevant or missing expected matches and you need to isolate the cause.
- For a no-code quick check, use the Web UI's ANN Recall tab to compare approximate vs exact
recall@kWeb UI ANN Recall - For the same comparison in code (CI gating, regression tests), run each query twice: once approximate, once with
exact=true. Computerecall@kfrom the overlap ANN recall in CI - Target >95%
recall@kin production. Exact search bad = model, data, or search pipeline problem. Exact good, approximate bad = tune HNSW. - Match the embedding model to your data: the context window should fit your chunk length with room for special tokens, the model must cover every language in your corpus, and its training domain should resemble yours (e.g. a code-trained model for code) How to choose an embedding model
- Make sure documents are properly chunked; splitting chunks mid-sentence alone can drop quality by 30-40%
- Check if quantization degrades quality (compare with and without)
- Check if filters are too restrictive (then you might need to use ACORN)
- If duplicate results from chunked documents, use Grouping API to deduplicate Grouping
Payload filtering and sparse vector search are different things. Metadata (dates, categories, tags) goes in payload for filtering. Text content goes in sparse vectors for search.
Approximate Search Worse Than Exact
Use when: exact search returns good results but HNSW approximation misses them.
hnsw_efcontrols ANN search breadth; increase it while recall is still climbing, stop at the lowest value that hits your recall target inside your latency budget Search params Raisehnsw_efonly when recall is still climbing- Increase
ef_construct(200+ for high quality) HNSW config - Increase
m(16 default, 32 for high recall) HNSW config - Enable oversampling + rescore with quantization Search with quantization
- ACORN for filtered queries (v1.16+) ACORN
Binary quantization requires rescore. Without it, quality loss is severe. Use oversampling to recover recall: the docs report 0.98 recall with 2x oversampling on 4096-dimensional and 4x on 1536-dimensional embeddings, so start around 2-4x and tune on your data. Always test quantization impact on your data before production. Quantization
Wrong Embedding Model
Use when: exact search also returns bad results.
-
Check Qdrant team recommendations on how to choose an embedding model.
-
Test top 3 MTEB models on 100-1000 sample queries Hosted Qdrant inference. Score them against a labeled set to compare apples to apples Measuring Retrieval Relevance.
-
If your data is strongly hierarchical (taxonomies, product catalogs, part-whole relationships), consider hyperbolic (Poincaré) embeddings. They capture tree structure in far fewer dimensions than flat ones. In Qdrant, use Euclidean HNSW to pull a candidate set from the original Poincaré coordinates, then a Formula Query to rescore with the real hyperbolic distance. How to serve hyperbolic embeddings with Qdrant.
-
Consider fine-tuning an embedding model for your specific use case only after trying better-suited models and retrieval/pipeline tuning and confirming that the embedding model remains the bottleneck. Fine-tuning is most useful when general-purpose embeddings fail to capture important domain- or task-specific distinctions and you have good labeled query-document pairs. Fine-tuning requires re-embedding and re-indexing the collection Model migration.
Unoptimized Search Pipeline
Use when: exact search also returns bad results and model choice is confirmed by user.
- Optimize search according to the advanced search-strategies skill
- Check candidate depth: if your retrieval pipeline includes a first-stage retriever that feeds a reranker or fusion stage, test whether increasing the prefetch limit improves your quality metric. A downstream ranker cannot recover relevant documents that never enter its candidate set. For hybrid search, start around
limit=100-200and test larger values against your labeled queries Candidate depth
What NOT to Do
- Tune Qdrant before verifying the model is right for the task (most quality issues are model issues)
- Use binary quantization without rescore (severe quality loss)
- Set
hnsw_eflower than results requested (guaranteed bad recall) - Skip payload indexes on filtered fields then blame quality (HNSW can't traverse filtered-out nodes, and filterable HNSW is built only if payload indexes were set up prior)
- Deploy without baseline recall or other search relevance metrics (no way to measure regressions)
- Compare two configs on a query set too small to resolve the difference between them (the result is noise, not evidence)
- Confuse payload filtering with sparse vector search (different things, different config)
Files
1- SKILL.md
e72f37c45e9.1 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from qdrant/skills8
Qdrant provides client SDKs for various programming languages, allowing easy integration with Qdrant deployments.
Guides Qdrant deployment selection. Use when someone asks 'how to deploy Qdrant', 'Docker vs Cloud', 'local mode', 'embedded Qdrant', 'Qdrant EDGE', 'which deployment option', 'self-hosted vs cloud', or 'need lowest latency deployment'. Also use when choosing between deployment types for a new proje
Guides building on Qdrant Edge, the embedded in-process shard. Use when someone asks 'how to sync Edge with the server', 'keep a local shard in sync with Qdrant Cloud', 'BM25 or keyword search on Edge', 'hybrid search on Edge', 'embeddings on device', 'Edge snapshots', 'apply a partial snapshot', 'w
Diagnoses and guides Qdrant horizontal scaling decisions. Use when someone asks 'vertical or horizontal?', 'how many nodes?', 'how many shards?', 'how to add nodes', 'resharding', 'data doesn't fit', or 'need more capacity'. Also use when data growth outpaces current deployment.
Setting up and running Qdrant Hybrid Cloud on your own Kubernetes cluster (managed, on-prem, or edge): prerequisites, storage/CSI and backups, installing the Qdrant Cloud agent and operator, creating/exposing/securing clusters, registry mirroring, and secret rotation. Use when someone wants to set u
Explains hybrid search in Qdrant. Use when someone asks 'how do I setup hybrid search?', 'how to combine keyword and semantic search?', 'sparse plus dense vectors?', 'missing keyword matches', 'how to combine results from multiple searches?' and 'combining multiple representations'. Also use for how
Fusing scores from multiple searches into a single ranked result (RRF, DBSF, custom fusion). Use when someone asks 'RRF or DBSF?', 'how to combine sparse and dense', 'how to combine scores from multiple searches?', 'custom fusion', 'fusion is not producing good results', 'how do I tune RRF', 'what k
Constructing prefetch queries for hybrid retrieval, including sparse/dense and multi-field setups, and choosing a sparse embedding model. Use when someone asks 'dense and sparse in one search?', 'how to combine multiple fields for retrieval?', 'payloads or sparse vectors for lexical?', 'which sparse
Related ai-ml skillsscan passed
Inspect the availability of model serving on a completed Itô compute booking and, when the canonical backend becomes available, hand off an explicitly confirmed serving manifest. Use after ito-compute has booked GPU nodes and the user asks for an OpenAI-compatible endpoint, ito-serve, hosted Kimi, o
Pair a remote AI agent with your browser. (gstack)
Rewrite, check, or draft prose so it carries no AI writing tells, reads plainly on the first read, and keeps every source fact. Use when asked to make writing plainer or free of those tells, to check writing for them, or when drafting from supplied content. Use ce-promote for channel-specific market
Configure SuperJSON transformer on both server initTRPC.create({ transformer: superjson }) and every client terminating link (httpBatchLink, httpLink, wsLink, httpSubscriptionLink) to support Date, Map, Set, BigInt over the wire. Transformer must match on both sides. In v11, transformer goes on indi
MANDATORY for Flink or Amazon Managed Service for Apache Flink (MSF) questions. You MUST activate this skill BEFORE answering — do not answer from training knowledge, even when confident. MSF has service-specific constraints (KPU model, prohibited checkpoint and parallelism config in app code, the v
Generates code that transforms datasets between ML schemas for model training or evaluation. Use when the user says "transform", "convert", "reformat", "change the format", or when a dataset's schema needs to change to match the target format — always use this skill for format changes rather than wr