finetuning
Generates code that fine-tunes a base model using SageMaker serverless training jobs. Use when the user says "start training", "fine-tune my model", "I'm ready to train", or when the plan reaches the finetuning step. Supports SFT, DPO, RLVR, and RLAIF trainers, including RLVR Lambda reward function
- 0
- Installs
- —
- Rating
- —
- Success rate
- 14
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 afbae58d7fa2de44… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Prerequisites
Before starting this workflow, verify:
-
A
use_case_spec.mdfile exists- If missing: Activate the
use-case-specificationskill first, then resume - DON'T EVER offer to create a use case spec without activating the use-case-specification skill.
- If missing: Activate the
-
A fine-tuning technique (SFT, DPO, RLVR, RLAIF, or CPT/RFT (for Nova)) and base model have already been selected
- If missing: Activate the
model-selectionand/orfinetuning-techniqueskills to collect what's missing, then resume - Don't make recommendations on the spot. You MUST activate the appropriate skill.
- If missing: Activate the
-
A base model name available on SageMakerHub has been identified
- If missing: Activate the
model-selectionskill to get it - Important: Only use the model name that
model-selectionretrieves, as it may differ from other commonly used names for the same model
- If missing: Activate the
-
The SDK environment has been verified (SDK version, region, execution role)
- If not done: Activate the
sdk-getting-startedskill first, then resume
- If not done: Activate the
-
A training dataset uploaded to a bucket in the environment's default region.
- If not met: Help the user upload the dataset to the correct S3
Critical Rules
Code Generation Rules
- ✅ Use EXACTLY the imports shown in each code template
- ❌ Do NOT add additional imports even if they seem helpful
- ❌ Do NOT create variables before they're needed in that section
- 📋 Copy the code structure precisely - no improvisation
- 🎯 Follow the minimal code principle strictly
- ✅ When writing code, make sure the indentation and f strings are correct
User Communication Rules
- ❌ NEVER offer to move on to a downstream skill while training is in progress (logically impossible)
- ❌ NEVER set ACCEPT_EULA to True without explicit user confirmation in the conversation
- ✅ Always mention both the number AND title of sections you reference
- ✅ If user asks how to run (notebook): If
run_cellis available, offer to run it. Otherwise, tell them to run cells one by one (mention ipykernel requirement). - ✅ If user asks how to run (script): Tell them to run with
python3 <script>.py
Workflow
1. Code Generation Setup
1.1 Directory Setup
- Identify project directory from conversation context
- If unclear (multiple relevant directories exist) → Ask user which folder to use
- If no project directory exists → activate the directory-management skill to set one up
⏸ Wait for user.
1.2 Select Code Template
Read references/code_output_guide.md for output format rules, then read the code template matching the finetuning strategy:
- SFT →
code_templates/sft.py - DPO →
code_templates/dpo.py - RLVR →
code_templates/rlvr.py - RLAIF with built-in rewards →
code_templates/rlaif_builtin.py - RLAIF with custom prompt →
code_templates/rlaif_custom_prompt.py
The template is a Python file where each # Cell N: Label comment marks the start of a new section. Split on these markers — everything between one marker and the next becomes one unit of output.
1.3 Generate Code
- Write the code from the template following the rules in
code_output_guide.md - Use same order, dependencies, and imports as the template
- DO NOT improvise or add extra code
- If the model is NOT a Meta/Llama model (model ID does NOT start with
meta-):- Omit the
ACCEPT_EULA = Falseline from the config cell - Omit the
accept_eula=ACCEPT_EULA,line from the trainer call
- Omit the
- If the model is from the Nova family, omit any code containing
max_epochsorlr_warmup_steps_ratiofrom the Configure Trainer section and the Hyperparameter Overrides section
1.4 Auto-Generate Configuration Values
In the 'Setup & Credentials' cell, populate:
-
BASE_MODEL
- Use the exact SageMakerHub model name from context
-
MODEL_PACKAGE_GROUP_NAME
- Generate from use case (read
use_case_spec.mdif needed) - Format rules:
- Lowercase, alphanumeric with hyphens only
- 1-63 characters
- Pattern:
[a-zA-Z0-9](-*[a-zA-Z0-9]){0,62} - Example: "Customer Support Chatbot" →
customer-support-chatbot-v1
- Generate from use case (read
-
Save notebook
2. RLVR Reward Function (for RLVR only, skip this section if technique is SFT or DPO)
2.1 Check Reward Function Status
- Ask if user has a reward function already, or would like help creating one.
- If user says they have one → Ask for the SageMaker Hub Evaluator ARN. Only proceed to Section 2.3 once the user provides a valid Evaluator ARN. If they don't have it registered as a SageMaker Hub Evaluator, continue to 2.2.
- If user says they do not have one → Continue to 2.2
2.2 Generate Reward Function From Template
- Follow workflow in
references/rlvr_reward_function.mdsection "Helping Users Create Custom Reward Functions"
2.3 Set CUSTOM_REWARD_FUNCTION value
- Set the value for
CUSTOM_REWARD_FUNCTIONin the Notebook with the ARN of the reward function (either given directly by the user, or from the function generation code asevaluator.arn).
3. RLAIF (for RLAIF only, skip this section if technique is not RLAIF)
Read references/rlaif_guide.md and follow its instructions.
4. EULA review and acceptance
- Look up the official license link for the selected base model from references/eula_links.md
- Display the license to the user following the phrasing in references/eula_links.md. For OSS models: "This model is licensed under {License}. Please review the license terms here: {URL}." For Nova models: "This model is subject to the AWS Service Terms: {URL}."
- Check if the selected base model is a Meta/Llama model (model ID starts with
meta-)- If Meta/Llama: Tell the user they must read and agree to the EULA before using this model. Ask: "Do you accept the license terms? (yes/no)". If the user confirms, set
ACCEPT_EULA = Trueand uncommentaccept_eula=ACCEPT_EULAin the generated notebook. If the user declines, leaveACCEPT_EULA = Falseand warn that training will fail without acceptance. - If non-Meta: Inform the user of the license for their awareness. No code-level action needed — the
ACCEPT_EULAvariable andaccept_eulaparameter should already be omitted from the notebook (see Step 1.3).
- If Meta/Llama: Tell the user they must read and agree to the EULA before using this model. Ask: "Do you accept the license terms? (yes/no)". If the user confirms, set
5. Post-Generation
After generating the code, offer to run it. Training can take hours depending on your dataset and model.
Notebook mode: If run_cell is available, offer to run the cells. Otherwise tell the user to run cells themselves.
Script mode: Present the user with options:
"Would you like me to:
- Leave it to you — run with
python scripts/[script_name]- Run it and wait until it's done
- Start it but don't wait — we can check status later"
- Option 1: Done. Wait for user to come back.
- Option 2: Execute the script as-is.
trainer.train(wait=True)blocks until complete. Report final status. - Option 3: Change
wait=Truetowait=Falsein the script, execute, report the training job name.
Checking status:
describe-training-job --training-job-name NAME→TrainingJobStatus,FailureReason,SecondaryStatusTransitions- For model package ARN after completion:
list-model-packages --model-package-group-name GROUP_NAME --sort-by CreationTime --sort-order Descending --max-results 1
Showing results after completion:
- Use
scripts/mlflow_reference.pyas the pattern to query MLflow metrics - Present loss by epoch as a text table (total_loss, val_eval_total_loss for SFT; rewards/margins for DPO; critic/rewards/mean for RLVR)
CRITICAL:
- DON'T suggest moving to next steps before training completes
- DON'T elaborate on the next steps unless the user specifically asks you about them.
6. Continuous Customization
If the user wants to finetune a model they had already customized, follow the instructions in references/continuous_customization.md
References
rlvr_reward_function.md- Lambda reward function creation guide (RLVR only)templates/rlvr_reward_function_source_template.py- Lambda reward function source template for open-weights models (RLVR only)templates/nova_rlvr_reward_function_source_template.py- Lambda reward function source template for Nova 2.0 Lite (RLVR only)code_templates/sft.py- Complete notebook template for Supervised Fine-Tuning (OSS path)code_templates/dpo.py- Complete notebook template for Direct Preference Optimization (OSS path)code_templates/rlvr.py- Complete notebook template for Reinforcement Learning from Verifiable Rewards (OSS path)references/continuous_customization.md- Instructions on fine-tuning an already fine-tuned model.rlaif_guide.md- instructions on RLAIF finetuning optionsrlaif_builtin.py- Code template for RLAIF with built-in judge promptrlaif_custom_prompt.py- Code template for RLAIF with custom judge prompt
Files
14- SKILL.md
02cacd845c9.1 KB - code_templates/dpo.py
41f60720035.7 KB - code_templates/rlaif_builtin.py
b315cd12cf5.5 KB - code_templates/rlaif_custom_prompt.py
0225b518066.0 KB - code_templates/rlvr.py
c10ebde4316.2 KB - code_templates/sft.py
45a483cc2f5.5 KB - references/code_output_guide.md
efbff6e9003.2 KB - references/continuous_customization.md
3757a770e18.5 KB - references/eula_links.md
a26bccb7947.8 KB - references/rlaif_guide.md
95b12bfaef3.8 KB - references/rlvr_reward_function.md
7831d68cd310.0 KB - scripts/mlflow_reference.py
db144579cf840 B - templates/nova_rlvr_reward_function_source_template.py
dad5e1603e12.7 KB - templates/rlvr_reward_function_source_template.py
f81126b7529.2 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from awslabs/agent-plugins8
Integrates Amazon Location Service APIs for AWS applications. Use this skill when users want to add maps (interactive MapLibre or static images); geocode addresses to coordinates or reverse geocode coordinates to addresses; calculate routes, travel times, or service areas; find places and businesses
Build and deploy full-stack web and mobile apps with AWS Amplify Gen2
Build, manage, and operate APIs with Amazon API Gateway (REST, HTTP, and WebSocket). Triggers on phrases like: API Gateway, REST API, HTTP API, WebSocket API, custom domain, Lambda authorizer, usage plan, throttling, CORS, VPC link, private API. Also covers troubleshooting API Gateway errors (4xx, 5
Generate validated AWS architecture diagrams as draw.io XML using official AWS4 icon libraries. Use this skill whenever the user wants to create, generate, or design AWS architecture diagrams, cloud infrastructure diagrams, or system design visuals. Also triggers for requests to visualize existing i
Design, build, deploy, test, and debug serverless applications with AWS Lambda. Triggers on phrases like: Lambda function, event source, serverless application, API Gateway, EventBridge, Step Functions, serverless API, event-driven architecture, Lambda trigger. For deploying non-serverless apps to A
Build resilient, long-running, multi-step applications with AWS Lambda durable functions with automatic state persistence, retry logic, and orchestration for long-running executions. Covers the critical replay model, step operations, wait/callback patterns, error handling with saga pattern, testing
Evaluate, configure, and migrate workloads to AWS Lambda Managed Instances (LMI). Triggers on: Lambda Managed Instances, LMI, capacity provider, multi-concurrency Lambda, dedicated instance Lambda, EC2-backed Lambda, cold start elimination, Graviton Lambda, instance type for Lambda, scheduled scalin
Build, run, debug, and operate applications on AWS Lambda MicroVMs — Firecracker-isolated, snapshot-resumable serverless compute environments that run inside a container with up to 8-hour lifetimes. Triggers on: Lambda MicroVMs, Firecracker isolation, snapshot-resumable compute, suspend/resume, sand
Related ai-ml skillsscan passed
Prevent AI style drift on legacy projects by scanning the codebase for implicit conventions, resolving conflicts with the operator one at a time, and writing an enforceable .ai-style-rules.md (Golden Files, naming rules, DONTs) plus an optional CLAUDE.md hook. Use when onboarding an AI agent onto a
Pair a remote AI agent with your browser. (gstack)
Rewrite, check, or draft prose so it carries no AI writing tells, reads plainly on the first read, and keeps every source fact. Use when asked to make writing plainer or free of those tells, to check writing for them, or when drafting from supplied content. Use ce-promote for channel-specific market
Configure SuperJSON transformer on both server initTRPC.create({ transformer: superjson }) and every client terminating link (httpBatchLink, httpLink, wsLink, httpSubscriptionLink) to support Date, Map, Set, BigInt over the wire. Transformer must match on both sides. In v11, transformer goes on indi
Store and query vector embeddings using Amazon S3 Vectors, a cost-effective long-term vector storage service with its own API namespace (s3vectors). Triggers on: create S3 vector bucket, vector index, store embeddings, semantic search, RAG vector storage, similarity search, vector database, migrate
Builds voice and chat AI agents with LiveKit Agents and LiveKit Cloud. Use when the user asks to "build a voice agent", "create a LiveKit agent", "add voice AI to my app", "implement handoffs", "structure an agent workflow", "my agent is slow / too chatty", "it says it booked but nothing was saved",