regression-sweep
Re-validates every previously-confirmed finding against its current target — detects drift (patched, mitigated, re-introduced). Cron-driven weekly sweep across the org's validated finding tree.
- 0
- Installs
- —
- Rating
- —
- Success rate
- 2
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 a42e34428a2e8f80… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Regression Sweep
Walk the entire validated/*.json tree, re-fire each finding's poc.py, compare output against the recorded poc_output.txt, and write a weekly drift report. Mounted onto the cloud-agent task #4.
Trigger
Cron weekly (default Mondays 02:00 UTC). May also be invoked ad-hoc after a major patch deployment.
Workflow
- Index validated findings. Glob
validated/*.jsonand resolve each entry'sFINDING_DIR(underfindings/finding-NNN/). - Per finding:
- Re-run
python3 poc.pywith a 60-second timeout. - Capture stdout/stderr into
findings/finding-NNN/evidence/validation/regression-{week}-rerun.txt. - Diff against
findings/finding-NNN/evidence/validation/poc-rerun-output.txt(the validator's original re-run output) using a normalized line-set comparison (strip timestamps, request IDs, ephemeral tokens). - Re-check the finding's CVE via
tools/nvd-lookup.py— has severity changed?
- Re-run
- Classify each finding into one of:
still_valid— re-run matches baseline within tolerance, CVSS unchanged.drift_severity— re-run matches, but CVSS shifted ≥1.0 (NVD re-scored).newly_invalid— re-run output diverges, exploit no longer fires. Likely patched.newly_revalidated— finding had been markedREJECTEDlater, but now fires again. Regression.inconclusive— re-run errored (network, target unreachable). Retry next sweep.
- Write report to
artifacts/regression-{YYYYWww}.json+ human-readableregression-{YYYYWww}.md. Transitions (newly_revalidated,drift_severity,newly_invalid) appear in the report under explicit headers for analyst review.
Output
{OUTPUT_DIR}/
artifacts/
regression-{YYYYWww}.json # machine-readable result
regression-{YYYYWww}.md # human summary
findings/finding-NNN/evidence/validation/
regression-{YYYYWww}-rerun.txt # captured re-run output
regression-{week}.json schema:
{
"week": "2026W19",
"swept_at": "2026-05-13T02:00:00Z",
"counts": {"still_valid": 47, "drift_severity": 2, "newly_invalid": 5, "newly_revalidated": 1, "inconclusive": 3},
"findings": [
{"finding_id": "finding-012", "asset": "asset42", "cve": "CVE-2024-12345",
"verdict": "newly_invalid", "reason": "PoC output diverged: response now 404",
"baseline_cvss": 9.8, "current_cvss": 9.8}
]
}
Rules
- Demonstrate, never disrupt. Re-firing a PoC is observation, not mutation. Every
poc.pyalready satisfied the demonstrate-only constraints at validation time (task-03 Safety section); the sweep simply re-executes the same script — which by contract reads a proof signal and exits. If a PoC at re-run attempts a mutating action it should not have contained originally, the sweep aborts that finding withinconclusiveand emits a stderr WARN — the PoC needs re-validation, not regression scoring. - Bounded per-PoC time. 60-second timeout per
poc.py. Timeouts →inconclusive, notnewly_invalid. Avoids false-positive "patched" claims caused by network blips. - Normalized diff. Strip timestamps (
\d{4}-\d{2}-\d{2}T\d{2}:\d{2}), request-IDs ([a-f0-9]{32,}), and ephemeral session tokens before comparing. Real exploit output is structurally stable. - No new findings. A regression sweep can flip status of existing findings but cannot create new ones. Newly observed vulns belong to the Validation Run task, not this skill.
- Idempotent. Re-running the same week's sweep overwrites the same
regression-{YYYYWww}.json. The per-findingregression-{week}-rerun.txtis timestamped to preserve history. - Cap concurrent re-runs. Max 5 parallel PoC re-runs per sweep to avoid hammering production assets.
References
reference/diff-normalization.md— full normalization rules and per-finding output-stability heuristics.
Files
2- SKILL.md
20bbecde1f4.0 KB - reference/diff-normalization.md
3f9301a8653.0 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from transilienceai/communitytools8
Offensive AI security testing and exploitation framework. Systematically tests LLM applications for OWASP Top 10 vulnerabilities including prompt injection, model extraction, data poisoning, and supply chain attacks. Integrates with pentest workflows to discover and exploit AI-specific threats.
API security testing - GraphQL, REST API, WebSocket, and Web-LLM attack techniques.
Stitches confirmed single-asset findings into multi-hop attack paths across the organization. Builds a graph where nodes are assets and edges are confirmed exploit hops citing the findings that enable them.
Acquire an authenticated session THROUGH MFA/OTP on an in-scope target and emit a reusable session artifact (Playwright storageState + Bearer) so executors can test the post-auth attack surface. Use when the highest-value authenticated classes (BOLA/IDOR/mass-assignment/injection on the real data AP
Authentication security testing - auth bypass, JWT attacks, OAuth flaws, password attacks, 2FA bypass, CAPTCHA bypass, and bot detection evasion.
Smart contract security testing and blockchain CTF exploitation. Covers Solidity vulnerability analysis, EVM storage manipulation, delegatecall attacks, CREATE/CREATE2 address prediction, and common DeFi exploit patterns. Use when analyzing Solidity contracts, solving blockchain challenges, or testi
Client-side vulnerability testing - XSS (reflected/stored/DOM), CSRF, CORS misconfiguration, Clickjacking, DOM-based attacks, and Prototype Pollution.
Cloud and container security testing - AWS, Azure, GCP, Docker, and Kubernetes misconfigurations and exploitation.
Related backend skillsscan passed
MyBatis and MyBatis-Spring patterns for mapper design, XML and annotation SQL, result mapping, dynamic SQL safety, transactions, batching, pagination, and query performance. Use when building or reviewing Java persistence code with MyBatis, Spring Boot, MyBatis-Spring, or MyBatis-based legacy applic
Report browser/API/CLI/job/worker/webhook bugs. (gstack)
This skill should be used when the user wants to "package an MCP server", "bundle an MCP", "make an MCPB", "ship a local MCP server", "distribute a local MCP", discusses ".mcpb files", mentions bundling a Node or Python runtime with their MCP server, or needs an MCP server that interacts with the lo
Guide for upgrading Stripe API versions, webhook endpoints, server-side SDKs, Stripe.js, and mobile SDKs
PostHog integration for server-side Node.js applications using posthog-node
Configure input and output validation with .input() and .output() using Zod, Yup, Superstruct, ArkType, Valibot, Effect, or custom validator functions. Chain multiple .input() calls to merge object schemas. Standard Schema protocol support. Output validation returns INTERNAL_SERVER_ERROR on failure.