verify-agent-action
Review a proposed AI-agent action or human-approval packet before execution. Use when an agent wants to run a consequential tool, command, deployment, message, purchase, credential operation, or data mutation; when checking whether approval still matches the exact action; or when auditing action evi
- 0
- Installs
- —
- Rating
- —
- Success rate
- 1
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 37dcd2638091acda… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Verify Agent Action
Treat a plausible approval screen as a claim, not proof. Verify the complete decision path before a human or an external enforcement point decides whether to act.
Preserve the safety boundary
- Never execute, approve, sign, send, purchase, deploy, or mutate anything.
- Never convert this review into execution authority.
- Never infer missing evidence, identities, timestamps, or parameters.
- Treat a valid schema, checksum, or signature as insufficient by itself.
- Treat signatures as evidence of attribution and integrity, not factual truth.
- Keep supporting and refuting evidence separate; do not average conflict away.
- Fail closed on a material mismatch. Use
INCONCLUSIVEwhen required evidence is unavailable.
Set this field in every final result:
{"execution_authorized": false}
Collect the review packet
Request only the artifacts needed for the review:
- The original user or system request.
- The exact proposed action:
- operation or tool name
- target resource
- complete parameters
- filesystem and network scope
- maximum execution count
- not-before and expiry times
- The assessment that claims the action is justified.
- The source evidence and policy used by that assessment.
- The approval record, including approver identity, role, action digest, nonce, audience, issue time, expiry, and use count.
- The latest monitoring events and expected heartbeat interval.
- The current trusted time and any prior nonce-use record.
List missing fields before analysis. Do not silently substitute defaults.
Build the exact action identity
Create one normalized action object without dropping fields:
{
"operation": "git.push",
"target": "owner/repository",
"parameters": {
"branch": "fix/example",
"commit": "40-character-sha",
"remote": "origin"
},
"filesystem_scope": [],
"network_scope": ["github.com:443"],
"execution_count": 1,
"not_before": "RFC3339 timestamp",
"expires_at": "RFC3339 timestamp"
}
Use a project-specified canonicalization and digest algorithm when provided. Otherwise, report that cryptographic identity cannot be independently verified; still compare every field structurally.
Never normalize away a security-relevant distinction such as:
- branch, commit, repository, environment, recipient, amount, currency, or host
- recursive, force, overwrite, privileged, destructive, or dry-run flags
- filesystem roots, CIDRs, ports, domains, execution counts, or expiry
Run the six controls
Evaluate every control as PASS, FAIL, INCONCLUSIVE, or NOT_APPLICABLE.
1. Recompute the assessment
- Re-run the declared deterministic evaluator from the declared source inputs when its implementation is available.
- Compare the complete canonical result, not selected fields.
- Mark
FAILif the received result differs from recomputation. - Mark
INCONCLUSIVEwhen only schema validation, an internal checksum, or an unverifiable evaluator claim is available.
2. Match the exact approved action
- Compare the proposed action with the action bound into the approval.
- Compare the complete normalized object and its digest.
- Mark
FAILif any material field changed after approval. - Treat a broad target or scope as a mismatch when the evidence justifies only a narrower action.
3. Reject replay and identity ambiguity
- Verify the nonce is unique and unused.
- Verify subject, audience, issuer, approver role, issue time, not-before time, expiry, and maximum use count.
- Mark
FAILfor a reused nonce, wrong audience, expired approval, future-dated approval, excessive use count, revoked identity, or role mismatch. - Mark
INCONCLUSIVEif no trustworthy replay store or time source exists.
4. Test reviewer independence
Build a dependence table for every reviewer or evaluator:
| Dimension | Compare |
|---|---|
| Model | family, version, fine-tune |
| Provider | account and control plane |
| Prompt | shared template or ancestry |
| Retrieval | overlapping sources and indexes |
| Tools | shared evaluator code and runtime |
| Operator | common owner or approval authority |
Do not count correlated reviewers as independent quorum members. Mark FAIL if
the policy requires independent approval and the remaining independent set is
too small.
5. Preserve evidence and contradiction
- Inventory every evidence identifier referenced by the assessment.
- Confirm each item is present, authenticatable, within its validity window, and relevant to the claim.
- Record support and refutation independently:
| Support | Refutation | Epistemic state |
|---|---|---|
| absent | absent | UNDETERMINED |
| present | absent | SUPPORTED_ONLY |
| absent | present | REFUTED_ONLY |
| present | present | CONFLICTED |
- Mark
FAILif evidence was removed, altered, expired, or concealed in a way that changes the result. - Never convert
CONFLICTEDinto a numeric average that appears safe.
6. Verify lifecycle and monitoring
- Confirm the action is inside its validity window.
- Verify monitoring-event signatures or integrity evidence when available.
- Check sequence numbers, previous-event digests, and expected heartbeat cadence.
- Treat missing, stale, reordered, or broken-chain telemetry as a failure when policy requires continuous monitoring.
- Do not interpret silence as health.
Challenge convenient conclusions
Before producing the final result, attempt these mutations mentally or with project-provided test fixtures:
- Replace a blocked assessment with an allowed result.
- Change one approved target, parameter, scope, amount, or commit.
- Reuse an otherwise valid approval nonce.
- Replace independent reviewers with correlated copies.
- Remove one refuting evidence item.
- Stop the monitoring heartbeat after approval.
If any mutation would pass the reviewed controls, record the affected control
as FAIL; do not merely recommend future hardening.
Determine the review result
Use exactly one result:
ELIGIBLE_FOR_HUMAN_DECISION: all required controls pass.ELIGIBLE_WITH_CONTROLS: no required control fails, and explicit external controls can resolve the listed conditions before execution.BLOCKED: at least one required control fails or the action exceeds the justified scope.INCONCLUSIVE: no required control is proven false, but evidence needed for a safe decision is missing or unverifiable.
ELIGIBLE_FOR_HUMAN_DECISION is not approval. A human authority and a separate
enforcement point remain responsible for any real action.
Report in this format
# Agent Action Review
## Result
- Review result: BLOCKED | INCONCLUSIVE | ELIGIBLE_WITH_CONTROLS |
ELIGIBLE_FOR_HUMAN_DECISION
- Execution authorized: false
- Exact action digest: <verified value or NOT_VERIFIED>
## Action
- Operation:
- Target:
- Material parameters:
- Scope:
- Validity window:
- Maximum uses:
## Control matrix
| Control | Status | Evidence | Reason |
|---|---|---|---|
| Recomputed assessment | PASS/FAIL/INCONCLUSIVE/N/A | ... | ... |
| Exact action binding | ... | ... | ... |
| Replay and identity | ... | ... | ... |
| Reviewer independence | ... | ... | ... |
| Evidence completeness | ... | ... | ... |
| Monitoring freshness | ... | ... | ... |
## Supporting evidence
- ...
## Refuting evidence and defeaters
- ...
## Required next action
- State the smallest concrete step that could change the result.
## Boundaries
- State what this review did not prove.
Lead with the result and the exact reason. Prefer a reproducible blocker over a confidence score.
Files
1- SKILL.md
13138bebda8.0 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from github/awesome-copilot8
Check any AI agent codebase against the OWASP Agentic Security Initiative (ASI) Top 10 risks. Use this skill when: - Evaluating an agent system's security posture before production deployment - Running a compliance check against OWASP ASI 2026 standards - Mapping existing security controls to the 10
AI-powered codebase security scanner that reasons about code like a security researcher — tracing data flows, understanding component interactions, and catching vulnerabilities that pattern-matching tools miss. Use this skill when asked to scan code for security vulnerabilities, find bugs, check for
Use this skill when the user explicitly asks to map, document, or onboard into an existing codebase. Trigger for prompts like "map this codebase", "document this architecture", "onboard me to this repo", or "create codebase docs". Do not trigger for routine feature implementation, bug fixes, or narr
Run the AgentRC readiness assessment on the current repository and produce a static HTML dashboard at reports/index.html. Wraps `npx github:microsoft/agentrc readiness` and hands off rendering to the @ai-readiness-reporter custom agent. Supports policies (--policy) for org-specific scoring. Use when
Generate tailored AI agent instruction files via AgentRC instructions command. Produces .github/copilot-instructions.md (default, recommended for Copilot in VS Code) plus optional per-area .instructions.md files with applyTo globs for monorepos. Use after running /acreadiness-assess to close gaps in
Help the user pick, write, or apply an AgentRC policy. Policies customise readiness scoring by disabling irrelevant checks, overriding impact/level, setting pass-rate thresholds, or chaining org baselines with team overrides. Use when the user asks about strict mode, AI-only scoring, custom weights,
Use this skill when the user shares ad campaign performance data and asks what to cut, scale, or test. Trigger for prompts like "analyze my ad campaigns", "where am I wasting ad spend", "reallocate my ad budget", "which ads are actually working", or "ROAS analysis". Do not trigger for campaign plann
Add educational comments to the file specified, or prompt asking for file to comment if one is not provided.
Related devops skillsscan passed
Operational controls for long-lived or cloud-hosted agent systems — runtime lifecycle (start, pause, stop, restart), observability (logs, metrics, traces), least-privilege safety scopes and kill switches, and rollout/rollback change management with audit logs and success/cost metrics. Use when runni
Land and deploy workflow. (gstack)
Build or maintain Cloudflare Sandbox apps on the stable @cloudflare/sandbox package. Use sandbox-next for preview apps and sandbox-migrate-to-next for stable-to-preview migrations.
Deploy tRPC on AWS Lambda with awsLambdaRequestHandler() from @trpc/server/adapters/aws-lambda for API Gateway v1 (REST, APIGatewayProxyEvent) and v2 (HTTP, APIGatewayProxyEventV2), and Lambda Function URLs. Enable response streaming with awsLambdaStreamingRequestHandler() wrapped in awslambda.strea
Automates CI/CD pipeline setup. Use when setting up or modifying build and deployment pipelines. Use when you need to automate quality gates, configure test runners in CI, or establish deployment strategies.
Deploys and manages full-stack web applications (Next.js, Angular) with Server-Side Rendering (SSR) using Firebase App Hosting. Use when deploying Next.js/Angular apps, configuring apphosting.yaml or firebase.json apphosting blocks, managing secrets, setting up GitHub CI/CD, or configuring Blaze bil