migrating-to-amazon-redshift
Guides an end-to-end data-warehouse migration to Amazon Redshift — discovery, schema/SQL/stored-procedure/macro/script conversion, data migration, validation, performance comparison, and reporting. Source-routed via `references/<source>/`; Teradata (Vantage) is the supported source; additional sourc
- 0
- Installs
- —
- Rating
- —
- Success rate
- 14
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 004792724f3da7f2… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Migrating to Amazon Redshift
What this skill is
This skill is AI guidance, not an execution framework. It is entirely Markdown knowledge (rules, mappings, patterns, best practices) — no executable code. All execution — conversion, the discovery/migration/validation runners, dependencies, and infrastructure — you (the AI) generate at runtime from this knowledge, tailored to the customer's environment.
Principle: knowledge over shipped code → less drift, nothing for the customer to run or depend on, reliable first-time results. Do not look for a pyproject, a tools package, an orchestrator engine, or shipped scripts — there are none by design; you generate execution.
Runtime: this skill works with or without the AWS MCP server — step guidance uses AWS CLI syntax. Running it with the AWS MCP server is recommended for sandboxed execution and audit logging; without it, the AI runs the generated scripts on the host shell (assumes Bash, Python 3, and AWS CLI + credentials). Do not assume MCP-only tools are available.
Source routing
This skill migrates a supported source data warehouse to Amazon Redshift. First identify the
source system, then load that source's knowledge under references/<source>/:
- Teradata (Vantage) →
references/teradata/— supported (all references below). - Other sources (e.g. Snowflake, Oracle) — unsupported; each is added as its own
references/<source>/set when ready.
The workflow is source-agnostic (discovery → convert → migrate → validate → performance → report); only the conversion knowledge is source-specific. Everything below is the Teradata set.
When to use
- Migrating a Teradata system (Vantage) to Amazon Redshift.
- Converting Teradata DDL, SQL, stored procedures, macros, or BTEQ to Redshift/RSQL.
- Assessing Teradata→Redshift migration complexity/effort.
Operating principles
- Discovery is strictly read-only (SELECT-only) on the source. Never change production
state: no DDL/DML, and never enable logging (
BEGIN/REPLACE QUERY LOGGING). If DBQL is empty, mark itunavailableand fall back to always-onDBC.AMPUsageV— seereferences/teradata/discovery-queries.md. - Skill provides knowledge; you generate execution. Read the
references/to reason and convert — apply the rules inreferences/teradata/conversion-rules.mddirectly for conversion, and generate the discovery/migration/validation runners (and the read-only discovery collector fromreferences/teradata/discovery-queries.md) tailored to the environment. - Generate, don't assume a framework. Assume the environment has Bash, Python 3, and AWS
CLI + credentials. Any Python lib a generated script needs (
teradatasql,boto3, …) ispip install-ed on demand by that script / its run-instructions — pin exact versions. Teradata TTU (BTEQ/TPT) is Linux/Windows-only — not macOS; prefer WRITE_NOS +teradatasql(cross-platform, no client) for discovery/extract unless a TTU/Linux host exists. - Credentials: use a read-only Teradata user; prefer IAM roles over IAM users. For
production, reference credentials from AWS Secrets Manager or Systems Manager
Parameter Store. For local development only, a git-ignored
.envfile or profile may be used — never commit it. Never hard-code or echo secrets. In a portable bundle, reference a co-located credentials file and ship acredentials.env.exampletemplate — the real file is git-ignored. - Persist state in files. All generated output goes under a git-ignored
output/in the user's working dir; keepoutput/state.mdcurrent so work is resumable.
Workflow (phases)
Run in order; each phase's result/ feeds the next (see references/teradata/orchestration.md).
- Discovery — inventory the source. →
references/teradata/discovery-queries.md(read-only collection SQL + BTEQ driver template the AI generates) →output/discovery/result/inventory.json - Conversion — schema + code. Apply the conversion rules directly, flag the
manual-rewrite long tail, and fix Redshift errors from the references. →
references/teradata/conversion-rules.md,references/teradata/data-type-mapping.md,references/teradata/architecture-mapping.md,references/teradata/stored-procedure-migration.md,references/teradata/bteq-to-rsql.md,references/teradata/common-errors.md - Data migration — extract → S3 → COPY, restartable. →
references/teradata/data-migration-patterns.md - Validation — counts/aggregates/sampling. →
references/teradata/validation-patterns.md - Performance — baseline vs Redshift; size the target. →
references/teradata/performance.md,references/teradata/sizing.md - Reporting — aggregate all phases. →
references/teradata/reporting.md
Conversion (how the AI applies it)
There is no converter to run — convert by applying the rules in
references/teradata/conversion-rules.md directly (with the type / architecture / stored-procedure /
BTEQ references): apply the deterministic rules to the well-understood bulk, flag the
manual-rewrite constructs with their suggested rewrites, assign a confidence per object,
and fix any Redshift errors using references/teradata/common-errors.md. The reference docs are the
single source of truth; conversion-rules.md includes golden input→output examples to match.
Execution modes (connectivity)
- Connected — your host can reach Teradata/Redshift → run the generated scripts in place.
- Disconnected — it can't → generate a self-contained bundle under
output/<phase>/(script + co-located credentials template + relativeresult/+run-instructions.md); the operator runs it on a reachable host and copiesresult/back. The copied-backresult/is the durable state — read it (+state.md) and continue.
Project-workspace layout (per migration run)
<project-workspace>/
migration-config.yaml # operator-authored: endpoints, scope, strategy
.gitignore # ignores output/
output/ # everything generated (git-ignored)
state.md # progress cursor
discovery/ … result/inventory.json
conversion/ … result/{ddl,sql,procedures,rsql}/ manual_review.json
data_migration/ … result/{extract,load,templates}/ migration_manifest.json
validation/ … result/validation_report.json
performance/ … result/{perf_baseline,perf_compare}.json
reporting/ result/migration_report.md
Security considerations
- No shipped code or dependencies. This skill is text-only — the customer runs nothing from it. Any runner the AI generates MUST pin exact dependency versions, validate/sanitize inputs (file paths, SQL, shell args), and never print or log credentials, secrets, or PII.
- Least privilege + ephemeral credentials. Use a read-only Teradata user for discovery. On AWS
prefer IAM roles over IAM users and IAM auth over username/password. Keep secrets in
AWS Secrets Manager / Parameter Store — never hard-code, echo, or commit them (credentials files
are git-ignored; ship only
*.exampletemplates). - Data in transit / at rest. Use TLS to both engines; stage extracts in an encrypted S3
bucket (SSE) with a least-privilege bucket policy; load via
COPY … IAM_ROLE(not access keys). Enable encryption on the target Redshift cluster. - Blast radius. Discovery is read-only by design. Migration writes to the target — validate against a throwaway / non-production Redshift first, and never point a generated write-path at production without explicit operator confirmation.
- No secret leakage in artifacts. Generated
output/…(manifests, reports,state.md) MUST NOT embed credentials or endpoints beyond what the operator supplies inmigration-config.yaml. - COPY
IAM_ROLEhardening. Scope the role's policy to the specific staging prefix (not bucket-wides3:*), and include condition keys in its trust policy (aws:SourceAccount/aws:SourceArn, orsts:ExternalIdfor cross-account) to prevent confused-deputy assumption — per Redshift IAM-role authorization best practices. - Logging & monitoring. Enable CloudTrail (S3 data events on the staging bucket + Redshift management events), Redshift audit logging (connection/user-activity logs to S3 or CloudWatch), and CloudWatch alarms on COPY failures or unusual staging-bucket access during the migration.
The AWS MCP server (recommended runtime) additionally provides sandboxed execution and audit logging for the generated scripts.
References (specialized knowledge)
| File | Topic |
|---|---|
references/teradata/orchestration.md | phase workflow + state model |
references/teradata/conversion-rules.md | the 72 conversion rules (source of truth) |
references/teradata/data-type-mapping.md | TD→RS type mapping |
references/teradata/architecture-mapping.md | PI→DISTKEY, PPI→SORTKEY, Join Index→MV |
references/teradata/stored-procedure-migration.md | SP → PL/pgSQL |
references/teradata/bteq-to-rsql.md | BTEQ → RSQL |
references/teradata/common-errors.md | common Redshift errors + fixes |
references/teradata/discovery-queries.md | DBC system-view inventory queries |
references/teradata/data-migration-patterns.md | COPY/TPT/micro-batch/checkpoint |
references/teradata/validation-patterns.md | row-count/aggregate/sample compare |
references/teradata/performance.md | representative-query extraction + compare |
references/teradata/sizing.md | RG node type + count from the source profile |
references/teradata/reporting.md | migration status-report generation |
Files
14- SKILL.md
f56a48045510.8 KB - references/teradata/architecture-mapping.md
48f9f7bccc5.8 KB - references/teradata/bteq-to-rsql.md
37648a94b34.1 KB - references/teradata/common-errors.md
83190660693.7 KB - references/teradata/conversion-rules.md
7a354eb2dc16.8 KB - references/teradata/data-migration-patterns.md
e520b3f6e612.9 KB - references/teradata/data-type-mapping.md
65416e191e3.4 KB - references/teradata/discovery-queries.md
d63db3a33d14.7 KB - references/teradata/orchestration.md
e41992814d7.1 KB - references/teradata/performance.md
ac361e09026.6 KB - references/teradata/reporting.md
2a9948d7ef4.2 KB - references/teradata/sizing.md
1800d0fffb7.3 KB - references/teradata/stored-procedure-migration.md
2fce5086ec5.1 KB - references/teradata/validation-patterns.md
4cca23d9ac5.9 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from aws/agent-toolkit-for-aws8
Amazon Aurora MySQL — creates, modifies, and advises on Aurora MySQL clusters specifically (MySQL-compatible engine, Aurora serverless, parallel query). Trigger for Aurora MySQL cluster operations, ACU sizing, I/O-Optimized storage, commitment pricing, or MySQL upgrade planning. Aurora MySQL uses fu
Amazon Aurora PostgreSQL — creates, modifies, and advises on Aurora PostgreSQL clusters specifically (PostgreSQL-compatible engine, Aurora serverless, express configuration, pgvector, Babelfish). Trigger for Aurora PostgreSQL cluster operations, express-configuration quick-start, ACU sizing, I/O-Opt
Builds generative AI applications on Amazon Bedrock. Covers model invocation (Converse API, InvokeModel), RAG with Knowledge Bases, Bedrock Agents, Guardrails, and AgentCore (including the Harness managed agent loop). Applies when invoking models, setting up Knowledge Bases, creating agents, applyin
Runs quantum computing workflows on AWS through Amazon Braket — discovering devices (QPUs and simulators) and their availability, building gate-model circuits and analog Hamiltonian programs, submitting quantum tasks, program sets and hybrid jobs, looking up prices, and capping spend with spending l
Manages Amazon DocumentDB end-to-end — serverless-on-8.0 cluster setup, TLS/VPC/driver config, flexible-schema and vector-search data modeling, MongoDB compatibility assessment, DMS-based migration, slow-query diagnosis, major version upgrades (4.0->5.0->8.0), Well-Architected reviews (41-check wa_r
Creates and automates custom image builds with EC2 Image Builder - Linux, Windows, and macOS AMIs, and container images to ECR. Covers the build IAM role, Amazon-managed and custom components, image recipes, infrastructure and distribution configuration (launch templates, SSM parameters, other Regio
Activate when developers have latent caching needs: slow API responses, database read bottlenecks, DynamoDB throttling or cost, RDS/Aurora scaling pressure, Bedrock latency or cost, or adding a cache; activate when working with Redis, Valkey, Memcached, or any in-memory data store, cache-aside patte
Builds, runs, debugs, and operates event-driven applications using EventBridge Event Bus - a managed, centrally governed publish/subscribe event bus that an organization can share across many teams and accounts. Applicable when workloads need event-driven architectures, decoupling, choreography, asy
Related database skillsscan passed
Build persistent multi-agent operating systems on Claude Code. Covers kernel architecture, specialist agents, slash commands, file-based memory, scheduled automation, and state management without external databases. Use when building a persistent multi-agent system on Claude Code with its own memory
Use when the user wants to provision infrastructure or third-party services using Stripe Projects. Triggers: "I need a database", "set up auth", "add caching", "give me a Postgres", "provision Redis", "I need hosting", "add a vector DB", "get me an API key for X", "get credentials for X", "sign up f
Build and troubleshoot Cloudflare Basin analytics workflows with Basin Pipelines, Basin Catalog, and Basin SQL. Use for streaming data into R2 Iceberg tables, managing catalogs, or querying those tables; also use for requests using the former Data Platform, Pipelines, R2 Data Catalog, or R2 SQL name
Manages deprecation and migration. Use when removing old systems, APIs, or features. Use when migrating users from one implementation to another. Use when migrating a database schema in production, such as renaming or dropping a column without downtime (expand/contract). Use when deciding whether to
Sets up, manages, queries, and configures Cloud Firestore databases (Standard/Enterprise edition), including data modeling, security rules, indexes, and SDK integrations (Web, Python, iOS, Android, Flutter). Use when creating/listing Firestore databases, defining data models/indexes, writing SDK que
Migrates Neo4j driver code and Cypher queries from older versions (4.x, 5.x)