eval
Evaluate a plugin or skill for quality
- 0
- Installs
- —
- Rating
- —
- Success rate
- 1
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 185ac3dc05adb55f… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
eval.md
Run the PluginEval quality evaluation on a plugin or skill directory.
Usage
/eval — evaluate at standard depth (static + LLM judge) /eval --depth quick — static analysis only (instant)
Process
Step 1: Run Static Analysis (Layer 1)
cd "${CLAUDE_PLUGIN_ROOT}"
uv run plugin-eval score {argument} --depth quick --output json
Parse the JSON output to get composite.score, composite.dimensions, and layers[0].anti_patterns.
Step 2: LLM Judge (Layer 2) — if NOT --depth quick
Dispatch the eval-judge agent with the skill path:
Evaluate the skill at: {resolved_path} Read the SKILL.md file and any references/ files, then score it on all 4 dimensions. Return your scores as JSON.
The judge returns scores for: triggering_accuracy, orchestration_fitness, output_quality, scope_calibration.
Step 3: Compute Final Score
If quick depth: Report the Layer 1 results directly from the CLI output.
If standard depth: Blend Layer 1 and Layer 2 scores.
For each dimension, use these blend weights (Static:Judge):
- triggering_accuracy: 0.375:0.625
- orchestration_fitness: 0.125:0.875
- output_quality: 0.0:1.0 (judge only)
- scope_calibration: 0.353:0.647
- progressive_disclosure: 1.0:0.0 (static only)
- token_efficiency: 0.8:0.2
- robustness: 0.0:1.0 (judge only)
- structural_completeness: 0.9:0.1
- code_template_quality: 0.3:0.7
- ecosystem_coherence: 0.85:0.15
Dimension weights: triggering(0.25), orchestration(0.20), output(0.15), scope(0.12), disclosure(0.10), efficiency(0.06), robustness(0.05), structural(0.03), code_quality(0.02), coherence(0.02)
Final = sum(weight * blended_score) * 100 * anti_pattern_penalty
Step 4: Present Results
## Overall Score: {score}/100 {badge}
## Layer Breakdown
| Layer | Score |
|-------|-------|
## Dimension Scores
| Dimension | Weight | Score | Grade |
|-----------|--------|-------|-------|
## Anti-Patterns Detected
## Recommendations
Badge thresholds: Platinum(90+), Gold(80+), Silver(70+), Bronze(60+)
Files
1- eval.md
c5058d82d72.1 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from wshobson/agents8
Audit web accessibility for WCAG compliance with automated axe-core tests, keyboard and screen reader checks, and remediation guidance
Audit UI code for WCAG compliance
Build AI assistant application with NLU, dialog management, and integrations
Run an AI-assisted code review that combines static analysis tools with AI review of security, performance, and architecture
Build realistic API mock servers with request stubbing, dynamic data, test scenarios, and contract testing
Open a review-action approval window by creating the ./.review-approved flag file. Takes an optional reason string that is recorded in the flag file and an unsigned approval log.
Verify every receipt in ./receipts/receipts.jsonl against the signer's public key. Detects tampered or malformed receipts across the audit trail.
Set up PreToolUse hook to block --no-verify and other git bypass flags in Claude Code projects
Related knowledge skillsscan passed
Onboard a Code-with-Claude Makers Cardputer — fetch the build-with-claude repo, flash firmware, and install the Claude Buddy apps.
Explain Stripe error codes and provide solutions with code examples
Define and enforce this project's quality bar — interview, sane defaults, CONSTRAINTS.md
Detects timing side-channels in cryptographic code
Ask a question about the repository using wiki context and source file references
Run Sanity TypeGen and troubleshoot type generation issues.