io.github.jamie7893/keelen

Keelen

Autonomous dev team steered from chat: plain-English requests in, tested merged PRs out.

1.0.0
Version
remote
Transport
43
Tools

Security review

Review passed

Reviewed Jan 1, 2000.

  • tools: 43 tools scanned
  • metadata: scanned

No findings.

Tools (43)

  • list_projects

    List the projects in this workspace, with what each one is building.

  • create_project

    Start building something new: creates a GitHub repo and begins work on it. For a NEW repo. Given an existing repo, import_project(repo_full_name) is the right call — this one would create a second, empty repo beside it (list_github_repos() browses what the workspace can see). Scaffolds a new GitHub repo, a bootstrap-mode project, and submits `build_description` as the project's first Roadmap Request. `name` is a concise GitHub short repo slug (no owner); `project_kind` is REQUIRED and one of library | node_library | python_library | service | cli | web_app | godot_game | roblox_game; `preview_command` is required iff `project_kind == 'web_app'`. `engine` is OPTIONAL — one of claude_code | codex | glm | kimi | grok (defaults to claude_code); codex, glm, kimi, and grok require the workspace to have a matching connected credential. `org` is OPTIONAL — a GitHub organization login to create the repo inside (e.g. your company org)

  • submit_request

    Ask for a change in plain English: a feature, a bug fix, or a new direction. Each call carries ONE feature or intent; a multi-feature ask belongs in separate requests. `text` must be under 16000 characters. Returns {thread_id, status, next_action, poll_after_seconds, next_step}; next_step names the re-check of get_request_status after poll_after_seconds.

  • get_request_status

    Check what happened to a request, and read any questions it asked back. Intake is async (~5min cadence), so the status changes between calls. `next_action` is one of: "wait" (still processing), "answer_questions" (answer_request takes one answer per question), "done" (see generated_roadmap_item_ids), "cancelled" (terminal, no items), "failed" (see intake_failure_reason).

  • set_product_vision

    Say what this product is for, so every run knows what it is building toward. `vision_md` is free-form markdown and must be non-empty. Tenant-scoped: a project not in the caller's workspace 404s. Returns {project_id, product_vision_md, updated_at, next_step}.

  • set_product_goal

    Set the outcome to aim at right now, so work is prioritised toward one thing. `goal_md` is free-form markdown and must be non-empty. Tenant-scoped: a project not in the caller's workspace 404s. Returns {project_id, product_goal_md, updated_at, next_step}.

  • list_roadmap

    See what is planned for a project and in what order. Queued rows also say whether the expand picker would elect them (`electable`) and, when not, why it skips them (`skip_reason`: awaiting_intake_thread | awaiting_clarification | crash_capped | noop_capped | crash_cooldown | noop_cooldown | held_on_open_pr | other). The first reason that applies, in that order, is the one returned. `awaiting_clarification` means the planner asked a planning question about the item and waits for the answer: `list_escalations` returns that card with the item's `roadmap_item_id`. A skipped item ahead of yours does not delay it.

  • reorder_roadmap

    Change what gets built first. `ordered_ids` is the desired front-to-back order of queued roadmap-item ids (list_roadmap returns them). The first id becomes the highest priority — the cadence expands the lowest-priority_int queued item next. Horizon pins still dominate: a pinned-later item stays at the back and a pinned-now item at the front, regardless of position in `ordered_ids`. Ids that are unknown or no longer queued are skipped; duplicates are rejected. Returns {updated, queue, next_step}.

  • cancel_roadmap_item

    Cancel / close a single roadmap item (list_roadmap returns its id). For a DELIVERED or duplicate item that keeps re-parking: once the work has shipped, every expand produces no dev-ready tasks and files a recurring `roadmap_item_parked` escalation that has to be acked. Cancelling drops the item out of the expand queue AND resolves any open expand-lane escalation for it. Business-level replay guard: an already expanded/cancelled item changes no state, and the dashboard can restore a cancelled item. Refuses (409) while the item is actively being expanded (a retry once that iteration ends succeeds). Returns {id, status, changed, next_step}.

  • clear_horizon_pin

    Clear a queued roadmap item's horizon pin (now/next/later → none). A horizon pin dominates the queue sort, so `reorder_roadmap` cannot move a pinned item out of its band — a stale `now` pin on a delivered/duplicate item clogs the front of the queue. This unpins it and reprices the queue so the item follows plain priority order again (and reorder_roadmap can then move it). Only queued items carry a settable pin (in-flight / shipped items derive theirs), so this refuses (422) on a non-queued item — a delivered item is closed with cancel_roadmap_item. Returns {id, previous_pin, horizon_pin, next_step}.

  • rollback_roblox_place

    Roll a Roblox project's place back to a previously-published version. For a `roblox_game` project, re-publishes the RETAINED build artifact for `version_number` (the web Roblox Publishing card lists published versions) — it never rebuilds from source, so rollback is fast + deterministic. This mints a NEW Roblox version pointing at the old build. Unknown version → 404; a version with no retained artifact → 422; a place open in Studio / rate-limited → 409 (a retry can succeed); an invalid or unscoped Open Cloud key → 409 / 403. Returns {version_number, env, published_at, status, next_step}.

  • project_status

    Check how a project is doing: what is in flight, what shipped, what is stuck. Returns lifecycle status, open task count, queued roadmap item count, runs today, verification pass rate, open blockers, pause state, and freshness. Zero `open_tasks` does not mean an empty roadmap: `queued_roadmap_items` counts planned work awaiting expansion, including work held while the scheduler is disabled. list_roadmap returns those items. `glm_peak_paused` is NOT a fault and sets no pause columns: it is the ephemeral GLM peak-hours skip. True means clean PRs hold and runs stop until `glm_peak_resumes_at`. It is an intentional cost gate, so the accurate description is "waiting for off-peak", not a failure. For `web_app` projects the optional `ui_review` block can mint a GitHub installation token and read the default branch's `ui-review.json`, so this status call is not guaranteed to be a purely local read.

  • close_task

    Close a task that should not be built — a duplicate, or work already shipped. For a task on the board that is obsolete: the change already landed in another PR, a sibling task covers it, or the user changed direction. The task is marked `cancelled` and keeps ALL of its history (acceptance criteria, QA steps, iterations) — nothing is deleted. `reason` is REQUIRED and is recorded on the audit trail in one line. `superseded_by_pr_number` (or `superseded_by_task_id`) records verified provenance when the work was genuinely delivered somewhere else, instead of a bare abandon. Closing does NOT claim the content is on the default branch, so any task that declared a dependency on this one keeps waiting; those are delivered or re-planned separately. Refuses with 409 while the task is being worked on by a running iteration (an in-flight iteration must finish, or its machine must be stopped, before the close succeeds). A task in another wo

  • get_product_goal

    Read a project's current product goal (the outcome work is aimed at). The paired setter `set_product_goal` replaces the whole document rather than appending to it, so a write without a prior read silently discards whatever the user already recorded; adding a line means reading the current text, editing it, and setting the full result back. Returns {project_id, product_goal_md, updated_at}. `product_goal_md` is None when no goal has been set. Tenant-scoped: a project not in the caller's workspace 404s.

  • get_product_vision

    Read a project's current product vision (what the product is for). The paired setter `set_product_vision` replaces the whole document rather than appending to it, so a write without a prior read silently discards whatever the user already recorded; adding a line means reading the current text, editing it, and setting the full result back. Returns {project_id, product_vision_md, updated_at}. `product_vision_md` is None when no vision has been set. Tenant-scoped: a project not in the caller's workspace 404s.

  • list_escalations

    See what the work is stuck on and waiting for a human decision about. `project_status` only COUNTS open escalations; this returns each one with its kind, reason, detail_md, recommended_action, and (when task-scoped) the blocked task's title + PR url. A card about a roadmap item carries its `roadmap_item_id` (null for a task-only, project-level or synthetic row), which matches the `id` in `list_roadmap`. A "forever-paused" project with no open `project_pause` row is surfaced as a synthetic `orphan:<project_id>` row. Each row carries a server-derived `task_retry_available` flag and an exact `next_tool`: eligible recovery blocks route to `retry_blocked_task`, roadmap parks to `rearm_roadmap_item`, platform conflicts to their structured resolver, stale work to `replan_task`, and only safe auxiliary cards to `resolve_escalation`.

  • resolve_escalation

    Acknowledge a handled ask only when no blocked task becomes invisible. `decision_md` is a required short note (why/how it was resolved), appended to the escalation's detail_md as an audit trail. Resolving a project_pause / orphan_pause RESUMES the project (clears the pause). A blocked task's final task_block/operator_action cannot be acknowledged: its typed retry, platform-policy resolution, replan, close, or supersede operation is the correct path instead. The acknowledgement itself is replay-guarded (a repeat changes nothing), but the call is not idempotent end to end: resuming a paused project lets queued refinement run, and that refinement can DELETE tasks. Accepts a real escalation UUID or a synthetic `orphan:<project_id>`.

  • retry_blocked_task

    Grant one fresh attempt after fixing a recovery-budget task block. ``escalation_id`` is the task_block id returned by ``list_escalations``. Eligibility, tenant, task, failure class, task status, and competing blockers are all derived server-side. This is distinct from ``resolve_escalation``, which remains acknowledgement-only for task blocks. Business-level replay guard: an already-retried block is not retried again; the call is not idempotent end to end, because an authenticated request also persists credential-use state.

  • rearm_roadmap_item

    Run the dashboard's replay-guarded same-item dependency re-arm. ``escalation_id`` is an ``expand_produced_nothing`` or ``roadmap_item_parked`` escalation id. Keelen verifies delivered structural prerequisites, grants one bounded expand retry, preserves the roadmap item id and ``depends_on_item_ids``, and resolves the matching cards in one transaction. A bounded live GitHub check may be used when the local merge state is stale, so a call can reach out even when the re-arm itself changes nothing.

  • resolve_platform_policy_conflict

    Resolve a platform-policy dead-end with a structured plan decision. ``decision`` is one of ``reuse_existing_evidence``, ``split_task``, ``raise_project_cap``, or ``remove_scenario``. ``raise_project_cap`` also requires ``project_cap`` (6..32). The decision becomes a new task-plan section; the same task lineage is requeued before the escalation resolves. On a project that runs the ui-review evidence lifecycle (``ui-review.json`` is a retained evidence catalog), ``raise_project_cap`` changes the per-run capture budget -- how many scenarios one capture run holds -- not catalog capacity: the catalog has no scenario limit, so no decision is needed to make room in it. On a ``web_app`` task, ``reuse_existing_evidence`` also repairs a frozen theme-only collision in the same transaction: criteria whose evidence specs differ only by a light/dark variant token (for example ``setup-timeline-light`` and ``setup-timeline-dark``) are refrozen to

  • set_ui_review_scenario_cap

    Set this web_app project's ui-review scenario budget, or reset it. ``scenario_cap`` is an integer from 6 through 32, or ``null`` to reset to the platform default of 12. The screenshot cap derives from it and moves with it, so the two can never starve each other. Raising is always allowed, including from a completely full manifest. LOWERING is refused when the default branch already declares more scenarios or screenshots than the smaller budget allows, and is also refused when that manifest cannot be read — both caps are enforced when keelen pushes and not in your CI, so an over-cap manifest fails every push while CI stays green. On a project with the retained evidence catalog (the response carries ``evidence_lifecycle: true``) the same parameter is a CAPTURE-RUN setting: it limits how many scenarios and screenshots one capture run holds, a pull request that needs more is split into capture batches, and the catalog keeps every scena

  • replan_task

    Replace an unretryable blocked task without losing its lineage. Creates a fresh, explicitly planned task under the same roadmap item and attempt lineage, transfers prerequisites and downstream dependents, and cancels the stale task with supersession provenance. The old PR, task, failures, criteria, and evidence remain in the audit history; nothing is marked delivered.

  • control_scheduler

    Start, pause, or resume a project's autonomous work. `action` is one of: "enable" | "disable" | "resume" | "process_now". - enable/disable flip scheduler_enabled (the loop dispatches only enabled, status='active' projects). - resume clears a pause (peak/backoff/manual) so the project dispatches again. - process_now durably prioritizes and immediately attempts the next intake batch, independent of the dev scheduler switch. A gate returns a typed reason and recovery step instead of a spawn promise. - enable/resume can execute pending refinement that DELETES tasks, and process_now can start an intake batch, so the tool is destructive even when the echoed scheduler state looks unchanged. Returns the resulting scheduler state.

  • archive_project

    Archive a project (reversible shelve) — frees a project slot in-tool. Stops any in-flight machine, then flips the project to `archived`: it drops out of the per-tier project cap (freeing a slot for a new project) and the loop stops dispatching it, but the project + its history are kept and can be restored from the dashboard project page. Stopping a machine DESTROYS the running process state — restoring the project later cannot bring that execution back, so this is destructive despite being reversible. A replay on an already-archived project changes no row, but is not idempotent end to end. This is preferable to `delete_project` unless the project must be gone. Owner-scoped (an MCP key is owner-only); a project not in the workspace 404s.

  • delete_project

    Soft-delete a project (the harder option) — frees a slot and hides it. Stops any in-flight machine, then flips the project to `deleted`: it disappears from `list_projects`, drops out of the project cap, and the loop stops dispatching it. Deleting ALSO permanently purges the project's stored review artifacts. The row is retained for audit and re-importing the repo restores that deleted row in place (unlike `archive_project` there is no in-tool restore), but the purged artifacts are gone for good and the stopped workers' process state cannot be recreated. A replay on a deleted project changes no row; the call is not idempotent end to end. Owner-scoped; a project not in the workspace 404s.

  • run_security_review

    Find security problems in a repository: a deep, whole-codebase review. Spawns a one-shot audit that scans the repo across a kind-aware taxonomy (secrets + git history, vulnerable/abandoned deps, injection, SSRF, path traversal, deserialization, crypto, info-leak, plus web authz/session/CORS, library API-misuse, game client-trust, or infra/CI as applicable) and posts findings to the project's Security review for human triage. The findings are reviewed by a human and the ones worth fixing go into the loop as Requests; no fix is applied automatically. Billable; one audit in-flight per project. (Triggering is disabled while the feature is hardened for production: a project not on the operator allowlist — empty by default — returns a message instead of spawning; earlier results stay visible.) Requires a paid plan; a free or trial workspace gets a message telling the user to upgrade.

  • run_legal_exposure_review

    Find where a codebase creates legal exposure (privacy, consent, data handling). Spawns a one-shot review machine that reads the checkout offline and maps what the code DOES onto commonly cited legal obligations, filtered by the project's saved compliance profile (jurisdictions plus eleven product facts). Findings land on the project's Legal page for a human to triage, readable with get_legal_exposure_findings. No change is ever applied automatically. THIS IS NOT LEGAL ADVICE AND IT IS NOT A LEGAL CLEARANCE. The result carries a `disclaimer_md` field that must reach the user before any summary. The review is not exhaustive, so an empty result is never proof that anything is in order. Requires a saved compliance profile (409-shaped refusal without one), a plan that carries the feature, and the project on the operator allowlist. Billable; one review in flight per project. Every refusal comes back as `ok: false` with an actionable

  • get_legal_exposure_findings

    Read a project's legal exposure findings, worst exposure first. Returns each finding as an OBSERVATION plus the obligation commonly cited over that pattern, its citation, the date the citation was last checked, and whether counsel has reviewed the registry entry (`counsel_reviewed`, false today for every entry). `exposure_order` is an ORDER, never a score: there is no grade, no percentage, and no overall state in this output. `not_determinable` lists the checks that could not reach a verdict; omitting them would turn a partial review into a clean answer. The `disclaimer_md` field must reach the user. `limit` defaults to 25 (max 100). `include_all` adds resolved, dismissed, and out-of-scope rows to the default actionable set. Read-only, no compute, no rate limit; it works even while the trigger is closed for the project.

  • run_control_gap_review

    Start a source-bound security control gap review for one project. The review records bounded engineering observations under the user's selected Cyber Essentials or CMMC Level 1 or Level 2 context. It does not determine framework standing, and it does not make changes. The `disclaimer_md` field must reach the user before any observation is summarised. The project needs a saved framework profile, an eligible plan, and a place on the operator allowlist. Billable; one review is allowed in flight per project. Refusals return `ok: false` with a next step and do not start work. Rate-limited per workspace.

  • get_control_gap_findings

    Read a project's security-control observations and review boundaries. The result groups records by evidence class without a total or an overall framework outcome. `not_determinable` and `outside_review_scope` qualify every observation. An empty group does not establish that a control is in place, and `disclaimer_md` must reach the user. `limit` defaults to 25 (max 100). `include_all` includes resolved, dismissed, and out-of-scope records. This read-only tool has no compute quota and remains available after the trigger closes for a project.

  • answer_request

    Answer a thread's clarifying questions (status must be awaiting_answers). `answers` is a list of {"idx": <int from get_request_status>, "answer_md": <str, 1..2000 chars>}. Every question needs exactly one answer. The call flips the thread back to intake_pending; get_request_status reports the subsequent state.

  • refine_request

    Say what is wrong with what a request produced, and have it reworked. Status must be "done". `feedback_md` is 1..1979 chars. Moves the thread to refine_pending; get_request_status reports the revised items. Refinement re-plans the thread's roadmap items and can EDIT OR HARD-DELETE the task rows it generated earlier (the machine patch-tasks endpoint deletes them outright), so it is destructive: up to five refinements per thread each carry that power.

  • signup

    Create a Keelen account (or start agent login) — emails a 6-digit code. UNAUTHENTICATED — the only tool besides verify_email that works before a bearer key is configured. `email` is where the code is sent. Flow: signup(email) -> the user reads the 6-digit code from their inbox -> verify_email(email, code) returns a reveal-once API key, stored as this server's `Authorization: Bearer <api_key>` header in the MCP client config; reconnecting the client and calling get_onboarding_status() then continues setup. The code expires in 15 minutes; a repeat signup resends it. Response is uniform whether or not the email already has an account (enumeration-safe), so signup doubles as agent LOGIN. Rate-limited per IP and per email. The `email` must be explicitly provided and confirmed by the user in chat before this call. It is not inferred from a client profile, the logged-in account, git config, or any other ambient source; a candidate address

  • verify_email

    Redeem the emailed 6-digit code for a reveal-once workspace API key. UNAUTHENTICATED. `email` + `code` must match a code issued by signup(email) within the last 15 minutes (5 attempts max). The returned `api_key` is shown exactly ONCE — store it ONLY in the MCP client config ("Authorization: Bearer <api_key>"), NEVER in a repo or a file you might commit. Reconnecting this server with the header set and calling get_onboarding_status() continues setup. An invalid/expired/consumed code returns a uniform error; a fresh code comes from signup(email).

  • get_onboarding_status

    Check what is set up so far and what to do next to get building. Reports engine_connected / github_connected / project_count / payment_status and a `next_action` string whose value is authoritative, with this tool polled between steps. `next_action` is "call_tool" (the tool named in `next_tool`, with `next_step` describing the call) until onboarding is complete, then "done". (`next_action_detail` echoes the pre-2026-07-29 dict shape and is DEPRECATED — it is removed 2026-10-29; next_action/next_tool are the current fields.) - engine step: the dashboard /login link goes to the user. Engine subscriptions (Claude / Codex / GLM) are connected in the DASHBOARD for security — engine credentials are never asked for or pasted in this chat. - github step: connect_github() returns an install link. - project step: create_project(...) for a new repo, or import_project(repo_full_name) for an existing one. - launch step: get_provisioning

  • open_dashboard

    Get a one-click, pre-authenticated dashboard sign-in link for the owner. The engine-connect step — and any dashboard task (billing card update, a project page) — needs a signed-in browser. Because this server has already authenticated the workspace owner, this mints a single-use magic-link login token and returns a `/login?token=…` deep link: opening it signs the user straight into the dashboard (no email round-trip, no password) and lands them where onboarding left off. The returned `login_url` works once and expires in 15 minutes; a repeat call mints a fresh one.

  • connect_github

    Get the link that connects a GitHub account, so work can reach real repos. Returns an `install_url` that the user opens in a browser to pick the GitHub account/org, approve the install, and land on a "connected" page; get_onboarding_status() then reports github_connected true. The link expires in 10 minutes; a repeat call mints a fresh one.

  • list_github_repos

    List repos the workspace's GitHub connection can see (for import_project). Each entry has full_name, default_branch, private, language, pushed_at. A `full_name` is what import_project(repo_full_name) connects. Reading the list can MINT a short-lived GitHub installation token, so this is not a pure read. Returns 409 if GitHub isn't connected yet — connect_github() is required first. A provider failure is translated into a typed, actionable error instead of passing GitHub's response through: a missing or suspended App installation, a revoked OAuth token, or a non-rate-limit permission denial is a 409 naming connect_github(); a recognized rate limit is a 429; other provider failures are a 502; a timeout or transport failure is a 503.

  • import_project

    Connect an EXISTING GitHub repo as a Keelen project. This is the counterpart of create_project: it is for a repo that already exists, while create_project scaffolds a brand-new one. `repo_full_name` is "owner/repo" — it MUST be visible to the workspace's GitHub connection (list_github_repos() to browse; a non-visible repo 404s). `engine` is OPTIONAL — one of claude_code | codex | glm | kimi | grok (defaults to claude_code); codex/glm/kimi/grok require a matching connected credential. `build_description` is OPTIONAL but STRONGLY recommended — a plain-language "what should Keelen build first?" submitted as the project's first Request so the loop has work; an imported project with no Request sits idle until submit_request(project_id, ...) adds one. `project_kind` is OPTIONAL — one of library | node_library | python_library | service | cli | web_app | godot_game | roblox_game | unknown. Omit it and the kind is auto-detected. It is

  • get_provisioning_status

    Check whether a new project has finished setting up and is ready to build. Returns overall (provisioning | ready | errored), a 7-stage checklist, and a user-facing error_kind when a stage failed. next_action is "wait" with poll_after_seconds (~10s) while provisioning, "wait" with NO poll_after_seconds when errored (fix the reported cause, then re-check — polling again without fixing it will not change the result), and "done" when overall is 'ready' — submit_request(project_id, text) then steers the loop. Reaching 'ready' also PERSISTS the owner's onboarding-completion state, so this is not a pure read. Tenant-scoped: a project not in the caller's workspace 404s.

  • get_billing

    Billing status + a Stripe checkout link when a NEW subscription is needed. `plan` is one of starter | pro | agency (default starter). When the workspace has NO live subscription and needs one (open-signup unpaid, churned, or converting from a free/trial tier), returns a `checkout_url` with next_action "browser" — the user opens it in a browser (the one setup step that can't happen in chat). Compute unlocks automatically once payment completes (a Stripe webhook flips the workspace to active); blocking on it is unnecessary. A past_due workspace gets NO checkout — the fix is a card update in the dashboard billing page (a new checkout would create a second subscription); next_step carries that fix. Subscribed or suspended-with-subscription states return checkout_url=None with an explanatory next_step.

  • project_digest

    Read a compact progress digest for one project (delivery, activity, holds, next decision). Returns the closed MCP projection: delivery counts and the five newest delivered items, a recent loop-activity count, current holds, the single highest-priority open decision, and `more_decisions`. Two steering fields are ALWAYS present: `next_action` is `"call_tool"` with `next_tool="list_escalations"` when a decision exists, otherwise `"none"` with `next_tool: null`. `since` is an optional ISO-8601 timestamp with a timezone; it defaults to 24 hours ago and is clamped to the last seven days. This tool is **off by default and not generally available**: it is registered only when the server was started with `KEELEN_PROGRESS_DIGEST_ENABLED` enabled, and a live disable refuses calls until restart. Tenant-scoped: a project outside the caller's workspace 404s.

  • stop_iteration

    Stop one running iteration after confirmed runtime cleanup. Requires a personal active workspace owner. Workspace keys and Portfolio delegation cannot stop compute. The exact project iteration number binds the runtime. Cleanup failure leaves it active; a busy subject is refused. Confirmed cleanup finalizes unfinished runtime iterations, holds its legacy tasks for owner close or replan and records the owner and reason. It grants no delivery, retry or merge authority. A repeat returns the same receipt.