pentest-engagement
Run a professional penetration engagement OR a network vulnerability scan from a scope. WEB mode (apex domains / app URLs) — mandatory surface expansion, systematic OWASP attack-class coverage, reversible active exploitation, authoritative validation, Transilience PDF. NETWORK mode (a list of IPs/CI
- 0
- Installs
- —
- Rating
- —
- Success rate
- 2
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 6cde102839b78f0f… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Pentest Engagement
Orchestrates a scoped pentest end-to-end via the pentest-engagement workflow. It is the breadth-complete, coverage-gated counterpart to the flag-shaped htb-solve — same engine (coordinator-loop, now with interleaved per-finding validation built into the loop), but driven by an attack-class coverage matrix instead of a flag, with surface expansion and root-cause severity baked in.
When to use
A real (non-CTF) engagement defined by a scope — either:
- WEB — web / API / cloud apps defined by apex domains / asset URLs, or
- NETWORK — a list of IPs / CIDRs / ranges (e.g. 1500 hosts) to scan for live services and vulnerabilities.
The workflow auto-detects the mode in Setup (engagement_kind): predominantly IPs/CIDRs → network; apex domains / app URLs → web. For HackTheBox/CTF use hackthebox (htb-solve) instead.
Run it
WEB — from a scope file or inline:
Workflow('pentest-engagement', { scope_file: 'projects/pentest/<engagement>-scope.md' })
Workflow('pentest-engagement', { scope: { engagement_name, apex_domains:[], assets:[...], creds_env:[...], roe, business_tier } })
NETWORK — inline IP/CIDR list or a scope file containing one (a plain newline list of IPs/CIDRs is accepted):
Workflow('pentest-engagement', { targets: ['10.0.0.0/24', '192.0.2.0/24', '198.51.100.7'] })
Workflow('pentest-engagement', { scope_file: 'projects/pentest/<engagement>-ips.txt', scan_profile: 'standard' })
Options (shared): maxConcurrent (default = prudent, derived from CPU cores — ~half the cores, capped 2–8; never hundreds/thousands of parallel tasks), dryRun, max_experiments, business_tier, report (default true).
Options (network): scan_profile light (bounded 1-1024 + curated less-common, for large/fast sweeps) / standard (full-range -p- on every reachable host, DEFAULT; two-stage SYN→-sV on found-open ports, host-count-guarded, message-bus + non-443-TLS aware) / full (-p- + bounded UDP); udp (top-50 UDP; off for light/standard, auto-on for full); slice_size (hosts per scan worker, auto ≈64 IP-list / 2 CIDR-heavy); deepen_top (deep-dive the N highest-value hosts, default 10, 0 to skip); geo_vantages (≤2 gcp zones for the 2nd-vantage allowlist re-probe; overrides the US+EU default), auto_provision (default true; false = detect+flag only, no cloud spend). On a source-IP/geo-allowlist signature the workflow auto-provisions a 2nd-geography vantage and re-probes the filtered hosts before concluding "no surface."
Write the scope file per reference/scope-file-format.md. Credentials are referenced by env-var name only and read from the repo .env via python3 tools/env-reader.py — never inline secret values.
Phases (what the workflow does)
- Setup —
env-readercreds, parse scope, classify kind (web|network), read CPU cores → prudent parallel-task cap,OUTPUT_DIR = projects/pentest/<date>_<engagement>/, STARTED Slack (gated). - Expand (WEB, the #1 fix) — MANDATORY CT-log / passive-DNS / origin-discovery across every in-scope apex (
crt.sh,certspotter,subfinder, origin-discovery for CDN/WAF-fronted hosts). Scope = the discovered surface, not the handoff. Builds the per-asset work list + seeds each asset's coverage matrix. Scan (NETWORK, replaces Expand) — slice the IP/CIDR set into machine-prudent batches; one nmap worker per slice (agents scale with slices ≈ dozens, never with IP count) runs the SAME pipeline: host discovery (reachability is unknown) → bounded common+less-common port/service scan → CVE surfacing (nmap --script vulners,nuclei, each CVE-ID enriched viatools/nvd-lookup.py) → writes a uniform per-IP treehosts/<ip>/{recon,host.json,findings}+ a mergedrecon/inventory/. - Assess (single interleaved stage — no separate downstream validation pass) — WEB: each asset →
coordinator-loop(coverage mode), which validates each candidate the instant it is materialized on fresh blind agents (strict per-finding cure/drop loop) before search continues. NETWORK: boundedcoordinator-loopdeep-dives on only thedeepen_tophighest-value hosts, same interleaved per-finding validation (everything else is the uniform tool-scan, not a per-host agent). Coverage-by-VALID: a class is covered only by aVALID/REPAIREDfinding, a justified N/A, or a genuine negative — a class whose candidates were all rejected/dropped stayspendingand search keeps going. - Correlate —
attack-path-stitcher+risk-prioritiseracross all validated findings → ranked org roadmap. - Report (deterministic — no agent authors the report) — JS hands the resolved engagement block + exact commands to ONE finalize runner:
tools/report_data_build.pymerges the namespaced interim finding-JSONs into the canonicalreport_data.json(the sole-owner assembly), then the format-dispatched renderer runs —transilience→ the canonicalgenerate_report.pyPDF skill,custom→custom_report_cmd(or a Markdown fallback). JS then hard-gates:report_dataassembled ∧ (transilience:WROTE∧bytes>0∧ the[assets: …/formats/transilience-report-style]provenance tag), retry-once →BLOCKED. OnlyVALID/REPAIREDfindings appear (drop-entirely —validated/is confirmed-only by construction); REJECTED (false-positives/) and uncured DROPPED (dropped/) never appear and there is no gaps/assurance section. The finalize runner also runsnetwork_coverage_map.py(swept-host tail) +coverage_gate.pyover the whole engagement, writesreports/coverage-matrix.json, and the deliverable includes a deterministic Attack Pattern Coverage section (surface-unit × attack-class). COMPLETE is a hard 100% gate: it requires the report to assemble+render AND the coverage gate to reportcomplete:true(every applicable cell covered) — for BOTH web and network (network additionally requires scan-completion). Any untested applicable cell →INCOMPLETE_coverage/BLOCKED. - Package & deliver — a verified
<report_id>_deliverable.zip(reports/ input/ logs/ artifacts/), a short stats summary (summary.md: agents, findings by severity, elapsed; tokens/cost renderunavailable — no runtime token counter), and the workflow returnsslack_offer: true.
Post-run (main loop, outside the workflow): because coordinators must not call AskUserQuestion, the invoking agent shows the returned summary and asks whether to post the deliverable_zip to Slack; on yes, python3 tools/slack-send.py --channel "$PENTEST_SLACK_CHANNEL_ID" (gated on a successful, COMPLETE engagement).
Report options: report_format transilience (default) | custom; custom_report_cmd (the custom renderer, receives the report_data.json path + reports/ dir); prior_report (a prior PDF or report_data.json — its title/sector/scope are metadata-only, never seeding the work list) + version to mint the cover version + "Supersedes" line.
Determinism (be honest)
The decision layer is a frozen pure-JS computeVerdict (parity-guarded, fixture-pinned), its operands are frozen per engagement via NVD/KEV --cache-dir snapshots (artifacts/nvd-cache/, artifacts/kev-snapshot.json), and the adversarial quorum is raised to 3. Given identical inputs the verdict is provably identical. But a fresh live run is highly reproducible, not 100% — LLM sampling produces the booleans/numbers that feed the verdict, and live-target drift moves the inputs; that ceiling is inherent and stated plainly, not papered over. Byte-identical results are a guarantee 100% only on replay of a frozen evidence set (the Phase-2 replay cache, artifacts/validation-cache/) — that is the sole context in which "same request → same result" is a guarantee rather than a strong tendency.
Boundaries
- Orchestrator only — never run the coordinator loop inline; the coverage/bookkeeping discipline needs the workflow boundary.
- A missing credential is not a global block — the unauthenticated surface is always tested; only a no-reachable-asset scope blocks.
- Reversible own-org/own-tenant writes are authorized by default (create-then-delete is non-destructive); destructive ops, DoS, brute force, and out-of-scope tenants are prohibited (set in RoE).
References
- scope-file-format.md — scope file schema + worked example
- coverage-matrix.md — the canonical attack-class coverage contract (completion gate)
- principles.md — scope-is-the-surface, reversible active exploitation, real-tools-first, root-cause severity
- pentest-report.md — Transilience report structure + §7.1 root-cause severity
- validator-role.md — engagement-validator attack-class coverage check (8)
Files
2- SKILL.md
fe0ef46edd9.6 KB - reference/scope-file-format.md
ca53489df411.0 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from transilienceai/communitytools8
Offensive AI security testing and exploitation framework. Systematically tests LLM applications for OWASP Top 10 vulnerabilities including prompt injection, model extraction, data poisoning, and supply chain attacks. Integrates with pentest workflows to discover and exploit AI-specific threats.
API security testing - GraphQL, REST API, WebSocket, and Web-LLM attack techniques.
Stitches confirmed single-asset findings into multi-hop attack paths across the organization. Builds a graph where nodes are assets and edges are confirmed exploit hops citing the findings that enable them.
Acquire an authenticated session THROUGH MFA/OTP on an in-scope target and emit a reusable session artifact (Playwright storageState + Bearer) so executors can test the post-auth attack surface. Use when the highest-value authenticated classes (BOLA/IDOR/mass-assignment/injection on the real data AP
Authentication security testing - auth bypass, JWT attacks, OAuth flaws, password attacks, 2FA bypass, CAPTCHA bypass, and bot detection evasion.
Smart contract security testing and blockchain CTF exploitation. Covers Solidity vulnerability analysis, EVM storage manipulation, delegatecall attacks, CREATE/CREATE2 address prediction, and common DeFi exploit patterns. Use when analyzing Solidity contracts, solving blockchain challenges, or testi
Client-side vulnerability testing - XSS (reflected/stored/DOM), CSRF, CORS misconfiguration, Clickjacking, DOM-based attacks, and Prototype Pollution.
Cloud and container security testing - AWS, Azure, GCP, Docker, and Kubernetes misconfigurations and exploitation.
Related security skillsscan passed
Security checklist for Solidity AMM contracts, liquidity pools, and swap flows. Covers reentrancy, CEI ordering, donation or inflation attacks, oracle manipulation, slippage, admin controls, and integer math. Use when auditing or writing Solidity AMM, liquidity pool, or swap code.
Security audit: supported static findings; qualified profiles add reproduction and repair candidates. (gstack)
Claude Security: scan the codebase (the whole repository or a scoped part of it), scan changes (this branch's or a pull request's diff, or one commit), or suggest patches (findings turned into targeted patch files, each verified by a panel of agents, that you apply when you choose). Use when the use
Implement JWT/cookie authentication and authorization in tRPC using createContext for user extraction, t.middleware with opts.next({ ctx }) for context narrowing to non-null user, protectedProcedure base pattern, client-side Authorization headers via httpBatchLink headers(), WebSocket connectionPara
Hardens code against vulnerabilities. Use when auditing an input handler for vulnerabilities, when handling user input, authentication, data storage, or external integrations, or when checking a login flow is safe against the OWASP Top Ten. Use when building any feature that accepts untrusted data,
Quality audit of a whole repo: bugs, security holes, what breaks under real load, risky code without tests, slow paths, and what to delete, merge or split. Ranked, each finding explained in plain English. One-shot report, changes nothing. Use for "audit this codebase", "review the whole repo", "find