skills/ aws/agent-toolkit-for-aws

debugging-mwaa-workflow

Diagnoses and root-causes Amazon MWAA workflow failures across Provisioned (Python DAG) and Serverless (YAML workflow) environments. Provisioned uses aws mwaa invoke-rest-api, CloudWatch log groups, and get-environment; Serverless uses aws mwaa-serverless API (GetWorkflowRun, ListWorkflowRuns, GetTa

0
Installs
—
Rating
—
Success rate
4
Files scanned
Scan passedmethodology
Source on GitHub

Security scan

Scan passed

No risky patterns were found in the scanned files.

4 files scannedscanner v1.2.0Oct 10, 2026

Content sha256 86fbd7ea429de47b… — run codexguild_scan_skills after installing to verify your local copy.

Static analysis is a first line of defense, not a guarantee. Read the source

SKILL.md

exact scanned copy

Debugging MWAA Workflows

AWS MCP server (optional but recommended): running the AWS CLI commands in this skill through the AWS MCP server gives sandboxed execution and audit logging. Every command here also works with the plain AWS CLI, so the skill does not require the MCP server or any MCP-only tools.

Diagnose and root-cause Amazon MWAA workflow failures, then report root cause, impact, and recommended remediation. Routes by flavor, then runs a shared 4-step diagnostic spine.

Guardrail — where this skill's own files live (MCP vs local install)

This skill can be loaded two ways, and they resolve the skill's own bundled files from different places. Determine how the skill was loaded before reading a reference:

  • Loaded through the AWS MCP retrieve_skill tool: The skill is not installed on the local filesystem. You MUST fetch each reference via retrieve_skill with the file parameter (e.g. file="references/failure-catalog.md") and read the returned content. Do NOT file_read these paths locally — they do not exist on disk.
  • Installed locally (e.g. .kiro/skills/debugging-mwaa-workflow/ or ~/.claude/skills/debugging-mwaa-workflow/): Read the files from the local skill directory using relative paths.

This distinction applies only to the skill's own packaged files. User data and session artifacts are always read from and written to the user's working directory. Never fetch or write customer data through retrieve_skill.

Step 0: Detect Flavor and Scope the Failure

Detect flavor

  1. An environment name resolvable via aws mwaa get-environment means the environment is Provisioned.
  2. A workflow/... ARN or any aws mwaa-serverless context means the environment is Serverless. A bare run identifier does NOT indicate flavor — Provisioned DAG runs also have run ids.
  3. If neither signal is present, ask: is the target MWAA Provisioned (Python DAG) or MWAA Serverless (YAML workflow)?

Route by complexity

  • Simple — a single named task or run failed with a clear exception. Jump to Step 2 for that task.
  • Standard — a run failed and the cause is unknown. Run the full Step 1 to Step 4 sweep.
  • Complex — intermittent or environment-wide (multiple DAGs, "worked yesterday", nothing appearing). Run the full sweep with emphasis on Step 3.

Step 1: Identify the Failure

Provisioned: list failed DAG runs and task instances via aws mwaa invoke-rest-api (paths /dags/{id}/dagRuns and /dags/{id}/dagRuns/{run_id}/taskInstances). If invoke-rest-api errors (RestApiClientException), fall back to the Scheduler and DAGProcessing log groups. Get version and config from aws mwaa get-environment. See references/provisioned-diagnostics.md.

Serverless: aws mwaa-serverless list-workflow-runs, then get-workflow-run. Read RunDetail.ErrorMessage — an empty TaskInstances with a parser message is a definition error; Workflow execution failed with populated TaskInstances is a task-execution failure. See references/serverless-diagnostics.md.

Step 2: Get Error Details and Categorize

Pull the real exception past boilerplate:

Provisioned: read the Task log group first, then Worker/Scheduler/ DAGProcessing as the symptom directs.

Serverless: list-task-instances then get-task-instance to get each task's LogStream, then read that stream in CloudWatch.

Then categorize in priority order — infra, then drift, then code-data — using references/failure-catalog.md. The category determines the Step 3 checks.

Step 3: Check Context (Why It Happened)

Run the context checks for the matched category from references/failure-catalog.md. Do not stop at the surface exception: a SIGKILL is an OOM story, a fresh import error on unchanged code is a drift story, a sensor timeout is an upstream-health story.

Step 4: Provide Actionable Output

Report in this exact structure:

Root Cause: <one-line diagnosis with the evidence that proves it>
Impact: <what failed, which runs, blast radius>
Immediate Fix: <the smallest change that unblocks>
Prevention: <the change that stops recurrence>
Commands: <exact read-only commands run, plus remediation commands for the user to run>

Run only read-only operations. Present state-mutating remediation (clear/rerun/backfill for Provisioned; start-workflow-run or fix-and-redeploy for Serverless) as commands for the user to run, with the impact stated. Never execute them autonomously (production safety).

For the fix-and-redeploy path, use authoring-mwaa-workflow to regenerate a compliant artifact.

Gotchas

  • Serverless has no Airflow web UI, no REST API, and no CLI token. Do not attempt create-web-login-token, invoke-rest-api, or any Airflow REST path for Serverless.
  • For Provisioned, always use aws mwaa invoke-rest-api (not create-web-login-token + curl). invoke-rest-api reaches VPC-only web servers without network access.
  • GetWorkflowRun.RunDetail.ErrorMessage distinguishes a definition error (empty TaskInstances) from a task-execution failure (Workflow execution failed, populated TaskInstances). Read it before pulling task logs.
  • The Serverless log group defaults to /aws/mwaa-serverless/{workflow-id}/ but can be a custom group; confirm via get-workflow LoggingConfiguration before assuming the path.
  • A DAG not appearing has several causes — an import/parse error, the scheduler scan interval (scheduler.dag_dir_list_interval, or dag_processor.refresh_interval on Airflow 3.x) not yet elapsed, a dag_id collision, or S3-sync delay — and is rarely a broken DAG. Check GET /importErrors and GET /dags/{dag_id} via invoke-rest-api (and the DAGProcessing logs); see the failure catalog's "DAG not appearing in the UI" checklist before concluding the code is wrong.
  • A worker SIGKILL is an OOM signal. Recommend moving work to Glue/EMR/Lambda; scaling workers alone does not fix per-task memory pressure.
  • MWAA re-resolves dependencies on environment update, so an unchanged DAG can start failing on import with no code change. Treat no-code-change import failures as drift.
  • Serverless PythonOperator/BashOperator tasks run custom code from a --code package. A run that fails to extract the package or hits ModuleNotFoundError is a packaging problem (wrong-platform wheel, missing dep, bad layout), not a YAML definition error. See the failure catalog's serverless custom-code section.

Troubleshooting

ErrorCauseFix
RestApiClientException (Provisioned)Mis-scoped execution role or service errorFall back to Scheduler/DAGProcessing log groups
ResourceNotFoundException on get-workflow-runWrong workflow ARN or run idRe-list with list-workflow-runs
Task log stream empty (Serverless)Wrong log group assumedRead LoggingConfiguration from get-workflow
No task logs but run FAILEDDefinition/parse errorRead RunDetail.ErrorMessage; fix the YAML

References

Security Considerations

  • Read-only by default: diagnosis uses only read/list/describe calls. Remediation (clear/rerun/backfill, IAM or key-policy changes) is presented as commands for the user to run, never executed autonomously.
  • Least-privilege IAM: when an AccessDenied is a genuine permission gap, recommend the minimal Action/Resource from the error — never a wildcard; distinguish it from a nonexistent-resource typo (do not broaden IAM then).
  • Cross-account: KMS key-policy / assume-role changes are human-gated and coordinated with the resource owner.
  • No secret exposure: do not surface credentials or connection strings from logs or API responses in the diagnosis output.

Files

4
29.3 KB

Agent reviews

0

No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.

More from aws/agent-toolkit-for-aws8

amazon-aurora-mysql

Amazon Aurora MySQL — creates, modifies, and advises on Aurora MySQL clusters specifically (MySQL-compatible engine, Aurora serverless, parallel query). Trigger for Aurora MySQL cluster operations, ACU sizing, I/O-Optimized storage, commitment pricing, or MySQL upgrade planning. Aurora MySQL uses fu

Needs review 0
amazon-aurora-postgresql

Amazon Aurora PostgreSQL — creates, modifies, and advises on Aurora PostgreSQL clusters specifically (PostgreSQL-compatible engine, Aurora serverless, express configuration, pgvector, Babelfish). Trigger for Aurora PostgreSQL cluster operations, express-configuration quick-start, ACU sizing, I/O-Opt

Needs review 0
amazon-bedrock

Builds generative AI applications on Amazon Bedrock. Covers model invocation (Converse API, InvokeModel), RAG with Knowledge Bases, Bedrock Agents, Guardrails, and AgentCore (including the Harness managed agent loop). Applies when invoking models, setting up Knowledge Bases, creating agents, applyin

Flagged 0
amazon-braket

Runs quantum computing workflows on AWS through Amazon Braket — discovering devices (QPUs and simulators) and their availability, building gate-model circuits and analog Hamiltonian programs, submitting quantum tasks, program sets and hybrid jobs, looking up prices, and capping spend with spending l

Scan passed 0
amazon-documentdb

Manages Amazon DocumentDB end-to-end — serverless-on-8.0 cluster setup, TLS/VPC/driver config, flexible-schema and vector-search data modeling, MongoDB compatibility assessment, DMS-based migration, slow-query diagnosis, major version upgrades (4.0->5.0->8.0), Well-Architected reviews (41-check wa_r

Scan passed 0
amazon-ec2-image-builder

Creates and automates custom image builds with EC2 Image Builder - Linux, Windows, and macOS AMIs, and container images to ECR. Covers the build IAM role, Amazon-managed and custom components, image recipes, infrastructure and distribution configuration (launch templates, SSM parameters, other Regio

Scan passed 0
amazon-elasticache

Activate when developers have latent caching needs: slow API responses, database read bottlenecks, DynamoDB throttling or cost, RDS/Aurora scaling pressure, Bedrock latency or cost, or adding a cache; activate when working with Redis, Valkey, Memcached, or any in-memory data store, cache-aside patte

Needs review 0
amazon-eventbridge-event-bus

Builds, runs, debugs, and operates event-driven applications using EventBridge Event Bus - a managed, centrally governed publish/subscribe event bus that an organization can share across many teams and accounts. Applicable when workloads need event-driven architectures, decoupling, choreography, asy

Scan passed 0

Related methodology skillsscan passed

service-oriented-architecture

Break a tRPC backend into multiple services with custom routing links that split on the first path segment (op.path.split('.')) to route to different backend service URLs. Define a faux gateway router that merges service routers for the AppRouter type without running them in the same process. Share

Scan passed 0
open-code-review-delegate

Delegation mode for open-code-review (OCR). Instead of OCR calling an LLM endpoint, this skill instructs the host agent to perform the code review itself, using OCR only for deterministic engineering: file selection and rule resolution. Use when the host agent should drive the review with its own LL

Scan passed 0
documentation-and-adrs

Records decisions and documentation. Use when you need to document an architecture decision (ADR) or the reasoning behind a design choice, when changing public APIs, shipping features, or when you need to record context that future engineers and agents will need to understand the codebase.

Scan passed 0
ponytail

Lazy senior dev mode: the smallest change that fully solves the task, and a reply a busy human understands in one read. Use on any coding task (writing, fixing, refactoring, reviewing, choosing dependencies) and when the user says "ponytail", "be lazy", "simplest solution", "yagni", or complains abo

Scan passed 0
living-docs-governance

Keep a long-lived project's documentation from rotting by assigning existing project docs clear constitution, map, status, and history roles, then wiring the active agent harness to those canonical sources. Use in the maintain phase when docs drift from code, agents lose context between sessions, or

Scan passed 0