qdrant-search-speed-optimization
Diagnoses and fixes slow Qdrant search. Use when someone reports 'search is slow', 'high latency', 'queries take too long', 'low QPS', 'throughput too low', 'filtered search is slow', or 'search was fast but now it's slow'. Also use when search performance degrades after config changes or data growt
- 0
- Installs
- —
- Rating
- —
- Success rate
- 1
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 d5a60989eb5d927f… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Diagnose a problem
There the multiple possible reasons for search performance degradation. The most common ones are:
- Memory pressure: if the working set exceeds available RAM
- Complex requests (e.g. high
hnsw_ef, complex filters without payload index) - Competing background processes (e.g. optimizer still running after bulk upload)
- Problem with the cluster (e.g. network issues, hardware degradation)
Single Query Too Slow (Latency)
Use when: individual queries take too long regardless of load.
Diagnostic steps:
- Check if second run of the same request is significantly faster (indicates memory pressure)
- Try the same query with
with_payload: falseandwith_vectors: falseto see if payload retrieval is the bottleneck - If request uses filters, try to remove them one by one to identify if a specific filter condition is the bottleneck
Common fixes:
- Tune HNSW parameters: Fine-tuning search
- Enable in-memory quantization: Scalar quantization
- Reduce Vector Dimensionality with Matryoshka Models: Matryoshka Models
- Use oversampling + rescore for high-dimensional vectors Search with quantization
- Enable io_uring for disk-heavy workloads on Linux io_uring
Can't Handle Enough QPS (Throughput)
Use when: system can't serve enough queries per second under load.
- Reduce segment count (
default_segment_numberto 2) Maximizing throughput - Use batch search API instead of single queries Batch search
- Enable quantization to reduce CPU cost Scalar quantization
- Add replicas to distribute read load Replication
Filtered Search Is Slow
Use when: filtered search is significantly slower than unfiltered. Most common SA complaint after memory.
- Create payload index on the filtered field Payload index
- Use
is_tenant=truefor primary filtering condition: Tenant index - Try ACORN algorithm for complex filters: ACORN
- Avoid using
nestedfiltering conditions as a primary filter. It might force qdrant to read raw payload values instead of using index. - If payload index was added after HNSW build, trigger re-index to create filterable subgraph links
Optimize search performance with parallel updates
Diagnostic steps
- Try to run the same query with
indexed_only=trueparameter, if the query is significantly faster, it means that the optimizer is still running and has not yet indexed all segments. - If CPU or IO usage is high even with no queries, it also indicates that the optimizer is still running.
Recommended configuration changes
- reduce
optimizer_cpu_budgetto reserve more CPU for queries - Use
prevent_unoptimized=trueto prevent creating segments with a large amount of unindexed data for searches. Instead, once a segment reaches the so called indexing_threshold, all additional points will be added in ‘deferred state’.
Learn more here
What NOT to Do
- Set
always_ram=falseon quantization (disk thrashing on every search) - Put HNSW on disk for latency-sensitive production (only for cold storage)
- Increase segment count for throughput (opposite: fewer = better)
- Create payload indexes on every field (wastes memory)
- Blame Qdrant before checking optimizer status
Files
1- SKILL.md
ede068e0534.6 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from github/awesome-copilot8
Use this skill when the user explicitly asks to map, document, or onboard into an existing codebase. Trigger for prompts like "map this codebase", "document this architecture", "onboard me to this repo", or "create codebase docs". Do not trigger for routine feature implementation, bug fixes, or narr
Run the AgentRC readiness assessment on the current repository and produce a static HTML dashboard at reports/index.html. Wraps `npx github:microsoft/agentrc readiness` and hands off rendering to the @ai-readiness-reporter custom agent. Supports policies (--policy) for org-specific scoring. Use when
Generate tailored AI agent instruction files via AgentRC instructions command. Produces .github/copilot-instructions.md (default, recommended for Copilot in VS Code) plus optional per-area .instructions.md files with applyTo globs for monorepos. Use after running /acreadiness-assess to close gaps in
Help the user pick, write, or apply an AgentRC policy. Policies customise readiness scoring by disabling irrelevant checks, overriding impact/level, setting pass-rate thresholds, or chaining org baselines with team overrides. Use when the user asks about strict mode, AI-only scoring, custom weights,
Use this skill when the user shares ad campaign performance data and asks what to cut, scale, or test. Trigger for prompts like "analyze my ad campaigns", "where am I wasting ad spend", "reallocate my ad budget", "which ads are actually working", or "ROAS analysis". Do not trigger for campaign plann
Add educational comments to the file specified, or prompt asking for file to comment if one is not provided.
Write, debug, and optimize Adobe Illustrator automation scripts using ExtendScript (JavaScript/JSX). Use when creating or modifying scripts that manipulate documents, layers, paths, text frames, colors, symbols, artboards, or any Illustrator DOM objects. Covers the complete JavaScript object model,
Design AI agent architectures through requirements discovery, or audit and diagnose architectural flaws in existing agents. Architecture only; excludes implementation and general code review.
Related tooling skillsscan passed
GAN-inspired Generator-Evaluator agent harness for building high-quality applications autonomously. Based on Anthropic's March 2026 harness design paper. Use when a feature should be built autonomously through generator and evaluator iteration until it clears a quality bar.
Web performance regression detection. (gstack)
This skill should be used when the user asks to "demonstrate skills", "show skill format", "create a skill template", or discusses skill development patterns. Provides a reference template for creating Claude Code plugin skills.
Helps you build and check a color system for your project. It generates palettes, names semantic tokens, converts between formats and measures contrast.
Creates a new Angular app using the Angular CLI. This skill should be used whenever a user wants to create a new Angular application and contains important guidelines for how to effectively create a modern Angular application.
Audit, diagnose, or optimize website loading and interaction performance, Core Web Vitals, and Lighthouse performance scores.