managed-airflow-dag-authoring
Provides guidance for authoring Apache Airflow DAGs in Managed Service for Apache Airflow (MSAA; formerly Cloud Composer). Covers environment context discovery, Airflow 2 vs 3 compatibility, authoring best practices, and local/remote validation processes. Use when creating or extending an Airflow DA
- 0
- Installs
- —
- Rating
- —
- Success rate
- 1
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 919ffc505e8a4142… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
GCP Managed Airflow DAG Authoring Guide
This skill guides you through authoring and validating Apache Airflow DAGs for Managed Service for Apache Airflow (MSAA; formerly Cloud Composer) environments.
Phase 1: Context Discovery
Before writing any DAG code, you MUST understand the constraints (e.g. version of Airflow) and capabilities of your target environment if user is willing to provide them.
1.1 Identify Target Environment & Access
Determine if you have direct access to the target Managed Airflow environment, local development environment or if you are working offline (only changing local files without validation).
- If environment access is available: Use
gcloudto inspect the environment (see Section 1.3). - If offline: Rely on user provided details.
1.2 Identify Development Environment
Determine if a local development environment is available.
- Check if
composer-devCLI is installed. - Check if a local Python environment with
airflowis available.
1.3 Inspect Target Environment (if available and requested)
Run the following commands to discover version constraints:
-
Get Airflow/Image Version:
gcloud composer environments describe {env_name} \ --location {region} \ --format="value(config.softwareConfig.imageVersion)" -
Get Installed Packages (Versions):
gcloud composer environments describe {env_name} \ --location {region} \ --format="value(config.softwareConfig.pypiPackages)" -
Get DAGs GCS Bucket:
gcloud composer environments describe {env_name} \ --location {region} \ --format="value(config.dagGcsPrefix)"
Phase 2: DAG Authoring Best Practices
2.1 General Airflow Best Practices
- Idempotency: Every task SHOULD be idempotent. Running it multiple times with the same inputs (e.g., execution date) SHOULD produce the same result and not duplicate data.
- No Top-Level Code Execution: Do NOT execute database queries, external API calls, or heavy computations at the top level of the DAG file (outside of tasks/operators). This code runs every few seconds during DAG parsing and will degrade performance.
- Explicit Catchup: Always set
catchup=Falsein the DAG definition unless historical backfilling is explicitly required. - Use Airflow Variables/Connections: Never hardcode credentials or
environment-specific configs. Use
Variable.get()(withdeserialize_json=Trueif applicable) andBaseHook.get_connection(). Access variables via Jinja templates (e.g.,{{ var.value.my_var }}) to avoid database calls during DAG parsing.
2.2 Airflow 2 vs Airflow 3 Compatibility
Use managed-airflow-migrations skill to navigate adjusting the code to specific target Airflow version.
Phase 3: Validation Process
You MUST validate DAGs before concluding your task.
3.1 Local Validation (Offline/Pre-deployment)
3.1.1 Static Analysis & Linting
Use ruff or pylint if available.
ruff check path/to/dag.py
- If targeting Airflow 3, check with Airflow 3 rules if rulesets are available.
3.1.2 Local Dev Environment (composer-dev)
If the user has composer-dev configured:
-
Copy the DAG to the local directory with DAGs:
cp path/to/dag.py $(composer-dev describe {local_env} --format="value(dags_directory)") -
Verify parsing:
composer-dev run-airflow-cmd {local_env} dags list-import-errors
3.2: Target Environment Validation
Only perform these steps if you have GCP access and are authorized to deploy to a target environment.
3.2.1 Deploy to GCS
Upload the DAG to the target environment's GCS bucket:
gcloud storage cp path/to/dag.py gs://{target_bucket}/dags/
3.2.2 Verify via Airflow CLI
Wait 1-2 minutes for the scheduler to parse the file, then run:
-
Check for Import Errors:
gcloud composer environments run {env_name} \ --location {region} \ dags list-import-errors
Pass Criteria: Output should be "No data found" or empty.
-
Verify DAG is Listed:
gcloud composer environments run {env_name} \ --location {region} \ dags list | grep {dag_id}
3.2.3 Monitor Cloud Logging
Check for runtime parsing errors in Cloud Logging:
resource.type="cloud_composer_environment"
resource.labels.environment_name="{env_name}"
log_id("airflow-scheduler")
severity>=ERROR
textPayload:"{dag_file_name}"
Definition of Done
- DAG code adheres to Airflow version constraints of the target environment.
- DAG code follows best practices (no top-level execution, idempotent if possible).
- DAG parses locally without import errors.
- (If environment is available) DAG is deployed to the target environment and verified to have no import errors.
Files
1- SKILL.md
3bd6b3d56c5.7 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from google/skills8
Configures best-practice alerting policies for AI agents using OpenTelemetry (OTel) metrics, generating output as Terraform (.tf) configuration files. Use when analyzing, writing, or deploying alerting policies to monitor agent latency, error rates, token usage, and quality metrics. Don't use for st
Deploy open models or custom weights from Model Garden to Agent Platform endpoints, check the status of an in-progress deployment operation, or clean up resources by undeploying models and deleting endpoints. Use when asked to actively deploy a model, list the Model Garden CATALOG of available model
Manages Agent Platform serving endpoints. Use when you need to create, list, describe, update, or delete serving endpoints for model deployment on Agent Platform. Also use when troubleshooting endpoint permission, quota, or resource busy errors. Don't use for deploying models to endpoints or for run
Measures and improves the quality of AI models and agents on Google Cloud using the Eval Quality Flywheel methodology. Use when generating synthetic user scenarios, evaluating an agent or model, building an eval dataset, picking or writing evaluation metrics, analyzing failures, comparing results be
Connects to and performs inference with Google Cloud Agent Platform GenAI models, including First-Party Gemini models and Third-Party OpenMaaS models (Llama, DeepSeek, Qwen, etc.). Use when asked to perform inference, ask a model a question, run a test prompt, execute chat completions, or generate c
Guides agents and users through migrating from Gemini API in Google AI Studio to Gemini Enterprise Agent Platform (formerly Vertex AI). Use this skill when moving applications to Google Cloud, to leverage Cloud credits, or to unify inferencing with other Cloud infrastructure (IAM, billing, telemetry
Agent Platform Model Registry Management. Use when you need to upload, list, describe, update, or delete machine learning models (and their versions) in the Agent Platform Model Registry. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform models.
Manages and orchestrates prompts in Agent Platform. Use when you need to create, list, retrieve, version, or delete managed prompts in Agent Platform. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform prompts.
Related methodology skillsscan passed
Break a tRPC backend into multiple services with custom routing links that split on the first path segment (op.path.split('.')) to route to different backend service URLs. Define a faux gateway router that merges service routers for the AppRouter type without running them in the same process. Share
Performs AI-powered code review on Git changes using the `ocr` CLI from alibaba/open-code-review. Use when the user asks to review code, review a pull request, review staged/unstaged changes, review a commit, or compare branches for code quality issues. Produces line-level review comments and can au
Guides systematic root-cause debugging. Use when tests fail, builds break, something that worked yesterday broke, behavior doesn't match expectations, or you encounter any unexpected error. Use when you need to figure out what broke and why — a systematic approach to finding and fixing the root caus
Lazy senior dev mode: the smallest change that fully solves the task, and a reply a busy human understands in one read. Use on any coding task (writing, fixing, refactoring, reviewing, choosing dependencies) and when the user says "ponytail", "be lazy", "simplest solution", "yagni", or complains abo
Python testing strategies using pytest, TDD methodology, fixtures, mocking, parametrization, and coverage requirements. Use when writing pytest tests — fixtures, mocks, parametrization, or coverage.