model-evaluation
Generates python code that evaluates SageMaker models. Supports two evaluation types: LLM-as-Judge and Custom Scorer. Use when the user says "evaluate my model", "run a benchmark", "test model performance", "how did my model perform", "compare models", or other similar requests.
- 0
- Installs
- —
- Rating
- —
- Success rate
- 15
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 cfcc286726e9d9a7… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Model Evaluation
Generate code that evaluates a SageMaker model.
Prerequisites
- The SDK environment has been verified (SDK version, region, execution role). If not done, activate the
sdk-getting-startedskill first.
Principles
- One thing at a time. Each response advances exactly one decision. Never combine multiple questions in a single turn.
- Confirm before proceeding. Wait for the user to agree before moving to the next step.
- Don't read files until you need them. Only read reference files when you've reached the step that requires them.
- Don't ask what you already know. If the answer is in conversation history, workflow_state.json, plan.md, or any file you've already read — use it. Confirm if unsure, but don't re-ask.
- No narration. Share outcomes and ask questions. Keep responses short.
- No repetition. If you said something before a tool call, don't repeat it after.
Scope
This skill supports the evaluation feature for SageMaker Serverless Model Customization. It can evaluate any base or fine-tuned model supported by SageMaker serverless model customization — both OSS models (Llama, Mistral, Qwen, etc.) and Nova models.
Tell the user when the skill is activated:
"I can help evaluate any base or fine-tuned model supported by SageMaker serverless model customization."
If the user requests help evaluating a model that isn't supported by SageMaker serverless model customization, explain that it is not supported by this skill.
Evaluation Types
There are two evaluation types:
- LLM-as-Judge — an LLM grades your model's responses. (OSS models only — not supported for Nova.)
- Custom Scorer — programmatic evaluation via Lambda function (includes built-in math and code scorers). Works with both OSS and Nova models.
Workflow
Step 1: Determine evaluation type
Do you already know which evaluation type to use?
Check conversation history, plan.md, workflow_state.json, or anything else you've already read.
If yes: confirm with the user.
"It sounds like you want to run [evaluation type]. Is that right?"
⏸ Wait for confirmation. If confirmed → go to Step 2.
If no: ask.
"What kind of evaluation would you like to run? I support:
- LLM-as-Judge — an LLM grades your model's responses
- Custom Scorer — programmatic scoring (math, code, or your own logic)
Pick one, or say 'help me decide' if you're not sure."
⏸ Wait for user.
- If user picks one → go to Step 2.
- If user indicates uncertainty, by saying something like "help me decide," "whatever you think," "I'm not sure" → read
references/evaluation-type-guide.mdand follow its instructions. It will guide the user to a choice and then return here. You MUST NEVER make a recommendation to the user on eval type without readingreferences/evaluation-type-guide.md.
Step 2: Validate and hand off to evaluation workflow
Before reading the reference file, validate that the chosen evaluation type is compatible with the user's situation. You may already know these answers from conversation context — don't ask if you don't need to.
LLM-as-Judge validation
- What model type are we evaluating? LLM-as-Judge is not supported for Nova models. To determine model type (if you don't already know it):
- If you have the training job name or ARN, use the AWS MCP tool
list-tagson the training job ARN and look for thesagemaker-studio:jumpstart-model-idtag. Contains "nova" → Nova. Anything else → OSS. - If you have a Model Package ARN, use the AWS MCP tool
describe-model-packageand check the model description or source tags. - If neither is available, ask the user.
- If you have the training job name or ARN, use the AWS MCP tool
- Does the user have an evaluation dataset? LLM-as-Judge requires one.
Custom Scorer validation
- Does the user have an evaluation dataset? Custom Scorer requires one. (Works with both OSS and Nova models, though for Nova only custom lambdas are supported.)
If validation fails, tell the user which requirement(s) aren't met and offer alternatives:
"[Evaluation type] won't work because [reason]."
If the failure reason was lack of an eval dataset, there's nothing we can do. Inform the user:
"Unfortunately all of the supported eval types require an eval dataset. I can't help you with model evaluation."
If the failure reason is something else, offer to help them pick a different evaluation type.
⏸ Wait for user.
If they say they do want help choosing a different eval type → read references/evaluation-type-guide.md.
If validation passes, read the corresponding reference file:
| User chose | Read |
|---|---|
| LLM-as-Judge | references/llmaaj-evaluation.md |
| Custom Scorer | references/custom-scorer-evaluation.md |
Follow the reference file's instructions from the beginning.
Files
15- SKILL.md
66e66b73d15.2 KB - code_templates/custom_scorer_evaluator.py
7991adc6d62.5 KB - code_templates/llmaaj_evaluator.py
5c24e900b62.4 KB - references/code_output_guide.md
c87b2b43803.3 KB - references/create-reward-function.md
b4839ba9e43.0 KB - references/custom-lambda-scorer.md
00231f83bf4.4 KB - references/custom-scorer-evaluation.md
ff131cde2e10.2 KB - references/evaluation-type-guide.md
4d129fb21e8.0 KB - references/llmaaj-builtin-evaluation.md
26d658729a5.0 KB - references/llmaaj-custom-evaluation.md
01833bb5fc2.3 KB - references/llmaaj-evaluation.md
0ae94f659e13.7 KB - references/supported-judge-models.md
8959efb5542.5 KB - scripts/nova_reward_function_source_template.py
82399ab55913.0 KB - scripts/reward_function_source_template.py
536d0bbf4e9.3 KB - scripts/validate_custom_metrics.py
9002cfbbcb3.6 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from awslabs/agent-plugins8
Integrates Amazon Location Service APIs for AWS applications. Use this skill when users want to add maps (interactive MapLibre or static images); geocode addresses to coordinates or reverse geocode coordinates to addresses; calculate routes, travel times, or service areas; find places and businesses
Build and deploy full-stack web and mobile apps with AWS Amplify Gen2
Build, manage, and operate APIs with Amazon API Gateway (REST, HTTP, and WebSocket). Triggers on phrases like: API Gateway, REST API, HTTP API, WebSocket API, custom domain, Lambda authorizer, usage plan, throttling, CORS, VPC link, private API. Also covers troubleshooting API Gateway errors (4xx, 5
Generate validated AWS architecture diagrams as draw.io XML using official AWS4 icon libraries. Use this skill whenever the user wants to create, generate, or design AWS architecture diagrams, cloud infrastructure diagrams, or system design visuals. Also triggers for requests to visualize existing i
Design, build, deploy, test, and debug serverless applications with AWS Lambda. Triggers on phrases like: Lambda function, event source, serverless application, API Gateway, EventBridge, Step Functions, serverless API, event-driven architecture, Lambda trigger. For deploying non-serverless apps to A
Build resilient, long-running, multi-step applications with AWS Lambda durable functions with automatic state persistence, retry logic, and orchestration for long-running executions. Covers the critical replay model, step operations, wait/callback patterns, error handling with saga pattern, testing
Evaluate, configure, and migrate workloads to AWS Lambda Managed Instances (LMI). Triggers on: Lambda Managed Instances, LMI, capacity provider, multi-concurrency Lambda, dedicated instance Lambda, EC2-backed Lambda, cold start elimination, Graviton Lambda, instance type for Lambda, scheduled scalin
Build, run, debug, and operate applications on AWS Lambda MicroVMs — Firecracker-isolated, snapshot-resumable serverless compute environments that run inside a container with up to 8-hour lifetimes. Triggers on: Lambda MicroVMs, Firecracker isolation, snapshot-resumable compute, suspend/resume, sand
Related ai-ml skillsscan passed
Prevent AI style drift on legacy projects by scanning the codebase for implicit conventions, resolving conflicts with the operator one at a time, and writing an enforceable .ai-style-rules.md (Golden Files, naming rules, DONTs) plus an optional CLAUDE.md hook. Use when onboarding an AI agent onto a
Pair a remote AI agent with your browser. (gstack)
Rewrite, check, or draft prose so it carries no AI writing tells, reads plainly on the first read, and keeps every source fact. Use when asked to make writing plainer or free of those tells, to check writing for them, or when drafting from supplied content. Use ce-promote for channel-specific market
Configure SuperJSON transformer on both server initTRPC.create({ transformer: superjson }) and every client terminating link (httpBatchLink, httpLink, wsLink, httpSubscriptionLink) to support Date, Map, Set, BigInt over the wire. Transformer must match on both sides. In v11, transformer goes on indi
Store and query vector embeddings using Amazon S3 Vectors, a cost-effective long-term vector storage service with its own API namespace (s3vectors). Triggers on: create S3 vector bucket, vector index, store embeddings, semantic search, RAG vector storage, similarity search, vector database, migrate
Builds voice and chat AI agents with LiveKit Agents and LiveKit Cloud. Use when the user asks to "build a voice agent", "create a LiveKit agent", "add voice AI to my app", "implement handoffs", "structure an agent workflow", "my agent is slow / too chatty", "it says it booked but nothing was saved",