ZeroWidth Workbench
Build, run and publish Workbench flows, tasks, shims and knowledge bases.
- 1.0.0
- Version
- remote
- Transport
- 55
- Tools
Security review
Partly reviewedReviewed 1h ago.
- tools: 55 tools scanned
- metadata: scanned
- mediumReviewRemote tools take credentials as input
Whatever an agent passes to a remote tool leaves the machine. Never send connection strings, tokens or passwords to a third-party MCP server unless it is the service those credentials belong to.
workbench_flows_revisions_list, workbench_flows_update, workbench_flows_publish
Tools (55)
search_docs
Search ZeroWidth product documentation. Returns matching pages with title, slug, public URL, and a query-relevant snippet. Use this when the user asks about a ZeroWidth product (Compass, Workbench, Caliper, Prism, Ledger, Napkin, zv1), an API behavior, or a policy. No authentication required — the docs corpus is public.
get_doc
Fetch the full Markdown body of a specific docs page by its slug. Use this after `search_docs` when the user needs the complete content of a page. No authentication required.
list_docs
Enumerate all available docs pages, optionally filtered by product (e.g. 'compass', 'legal', 'overview'). Use this to discover what slugs exist before calling `get_doc`. No authentication required.
workbench_flows_revisions_list
The flow's revision history, newest first: published versions carry a `version` label and note; entries with version null are draft autosaves. Use it to answer 'is this published?' (any entry with a version), to find a revision id for caliper_evals_create's flowRevisionId, or to see when the draft last changed.
workbench_flows_update
Changes a flow's name, description, visibility, tags, or archived state — the metadata around the flow, NOT its orchestration (prompts and nodes change through workbench_flows_edit_text). Pass only what changes. Archiving hides the flow from the default gallery but leaves it runnable, scheduled, and shared; pass archived: false to restore. May return `needs_confirmation`.
workbench_flows_publish
Snapshots the flow's current DRAFT as a named published version — the version schedules (workbench_flows_schedule), guest share links (workbench_flows_share_create), public-API runs, and 'published'-stage evals execute. The draft keeps evolving after this; runs on the published side don't change until the next publish. Publishing does NOT run the flow, but any eval set to runOnPublish starts a run (that spends credit — mention it when one exists). Publish only after the user has tested the draft (workbench_flows_run) or asked for it. May return `needs_confirmation`.
workbench_flows_delete
Deletes a flow. Its schedules stop, guest share links die, and evals targeting it can no longer run (their binding shows flowOk: false) — check caliper_flow_performance for evals and workbench_flows_schedules_list before proposing, and prefer workbench_flows_update with archived: true when the user just wants it out of the way. May return `needs_confirmation`.
workbench_flows_fork
Copies a PUBLIC flow from another workspace (a template) into this one as a new draft the user owns. Identify it by its flowUuid (the id on Workbench's public template pages and in `workbench_flows_get` output); optionally pin which published version to copy. Flows already in this workspace can't be forked — open them instead. Counts toward the plan's flow cap. May return `needs_confirmation`.
workbench_flows_schedules_list
Every schedule on one flow: cadence, whether it's paused (`enabled: false`), the next fire time, delivery, and the last run's status. Read this before workbench_flows_schedule (don't create a duplicate) and to get the scheduleId for workbench_flows_schedules_update / workbench_flows_schedules_delete / workbench_flows_schedules_run_now.
workbench_flows_schedules_update
Changes one schedule: `enabled: false` pauses it (configuration kept), `enabled: true` resumes and recomputes the next fire time from now; `interval` / `scheduleTime` / `scheduleTimezone` / `scheduleDayOfWeek` change the cadence; `label`, `delivery`, and `input` can change too. Pass only what changes. Get the scheduleId from workbench_flows_schedules_list. May return `needs_confirmation`.
workbench_flows_schedules_delete
Removes a schedule for good. Prefer workbench_flows_schedules_update with enabled: false when the user might want it back. May return `needs_confirmation`.
workbench_flows_schedules_run_now
Runs one iteration of a schedule immediately through the real scheduled pipeline (same input, same delivery — the email or channel it normally posts to) without moving its cadence. Works on paused schedules. Use it to test a schedule the user just set up. Spends workspace credit and delivers for real, so it sits behind the approval gate — may return `needs_confirmation`. The run is fire-and-forget: check workbench_flows_schedules_list for lastStatus, or workbench_executions_get for the trace.
workbench_flows_shares_list
Every guest link on one flow, newest first: label, whether it's still live, the shareUrl (live links only — dead ones have no URL), expiry, and how many guests and messages it has seen. Use it to answer 'who has access', to recover a link the user lost, and to get the shareId for workbench_flows_shares_revoke.
workbench_flows_shares_revoke
Kills one guest link immediately — anyone holding it gets a closed page from then on. Conversations already had stay in the flow's history. Get the shareId from workbench_flows_shares_list. May return `needs_confirmation`.
workbench_flows_run
Execute a flow and return its outputs synchronously. Runs the draft by default; pass source:"published" (optionally a version) to run the live published version. Spends workspace LLM budget, bounded by the token's cost cap. `input` is the flow's input envelope, e.g. {"kind":"chat","messages":[…]} or {"kind":"form","values":{…}}.
workbench_flows_schedule
Set a specific built flow to run automatically on a cadence — the deterministic counterpart to a zv1 routine. Each run executes the flow's latest PUBLISHED revision with the fixed `input` envelope and routes the output per `delivery`. USE THIS when the user wants a flow they've built to run on a schedule ("run my digest flow every morning"). Do NOT create a routine for this — a routine runs a free-form instruction, not a built flow. The schedule runs the PUBLISHED flow, so the flow must be published (check workbench_flows_revisions_list; publish with workbench_flows_publish if it isn't). `input` is fixed for every run, so the flow itself should fetch anything time-varying at run time. Manage existing schedules with workbench_flows_schedules_list / workbench_flows_schedules_update.
workbench_flows_list
Lists every flow in the active workspace the caller can see. Returns summaries (id, name, visibility, updatedAt) — fetch one with `workbench_flows_get` for the full orchestration body.
workbench_executions_get
The flow's stack traces: recent executions with per-node timelines — what each node received, produced, how long it took, and the exact error when one failed. USE THIS when a run misbehaves instead of guessing: read the failing node's inputs/error, then propose a fix (workbench_flows_edit_text) grounded in what actually happened. Pass executionId to inspect one run, or just flowId for the most recent runs. Node inputs/outputs are truncated for transport — the full record is in the flow's dev drawer.
workbench_flows_get
One flow. Default `view: summary` lists its nodes and links so you can pick one; `view: node` with a nodeId returns that node's full settings and prompt — read THAT before workbench_flows_edit_text, and copy `find` text from it character-for-character, never from memory. `view: full` returns the whole orchestration body.
workbench_flow_authoring_guide
READ THIS FIRST before writing or editing any raw flow orchestration JSON. Covers the document shape, the settings-vs-input-ports rule (temperature, max_tokens, tools, response_format are PORTS, not settings — constants reach ports via value nodes), exact link format, plugin links, and three complete worked examples.
workbench_node_catalog_get
The ground truth for what nodes exist and what their ports actually are — verify against this instead of recalling. Search by keyword/category for summaries; pass `slug` for one node's full detail (inputs, outputs, settings). Inputs are PORTS fed by links; settings live on the node — see `workbench_flow_authoring_guide`. The catalog is the WORKSPACE'S: model nodes the workspace's inference policy forbids come back with `allowed: false` — never author with those; pick an allowed model. Pass `flowId` to include the flow's pinned imports as nodes.
workbench_flows_scaffold
Creates a runnable first-draft flow from a spec you author: pick the simplest pattern that fits (classifier for read-and-bucket, structurer for transform/extract/draft-for-review, agent for genuinely conversational), write a production-quality system prompt grounded in what the user told you, and mark anything stubbed with [STUB: ...] markers plus stubNotes. The draft opens in Workbench's simple editor at /w/<workspace>/flows/<id> — give the user that path. May return `needs_confirmation`; show the user what you're proposing and wait for their approval, then re-call with the approvalId. When the draft comes from a Compass change, pass its opportunityId so the flow and the change point at each other (Compass shows 'Open in Workbench'; the flow shows where it came from). Then: workbench_flows_run to try it, workbench_flows_publish when it's ready for schedules and share links.
workbench_flows_edit_text
Applies ONE precise text replacement to a flow's DRAFT orchestration — Edit-tool semantics: `find` must be the EXACT current text, copied character-for-character from workbench_flows_get, and must occur exactly once anywhere in the flow (system prompts, node settings, metadata, a node's type). `find` is text inside ONE value, not JSON: to change a model, find the model slug alone (`anthropic-claude-haiku-4-5`) and replace it with another `llm` slug from workbench_node_catalog_get. It can't add or remove nodes or links. Zero or multiple matches return an error instead of guessing. This is the improvement primitive: check receipts first with caliper_flow_performance, cite the run id in `note`, apply the edit after approval, then re-run the eval with caliper_evals_run and report the score delta — never claim improvement without the before/after. Edits land on the draft only; the published version changes when someone publishes (workbench_flows_publish, after the user has seen the result).
workbench_flows_share_create
Creates a no-account guest link where anyone can chat with the flow's latest PUBLISHED revision at workbench's /s/<token> page — the fastest way to put a working flow in a stakeholder's hands ('here, try it'). The flow must have a published version (workbench_flows_revisions_list shows it; publish with workbench_flows_publish if not). Conversations are capped per guest; the link can expire, and you can list links with workbench_flows_shares_list and kill one with workbench_flows_shares_revoke. May return `needs_confirmation` — say who the link is for and wait.
workbench_kb_get
One knowledge base: name, recipe (docs / tabular / graph), status, document and chunk counts, embedding model, and ingestion settings. Read it to confirm a KB has content before pointing a flow at it. Get the kbId from workbench_kb_list.
workbench_kb_documents_list
Every document ingested into the KB's current content: id, name, type, size, chunk count. Empty for a KB with nothing ingested yet (or one whose ingestion is still running — check workbench_kb_ingestion_status). Use it before workbench_kb_add_documents to avoid re-adding a document, and to get the documentId for workbench_kb_document_delete.
workbench_kb_ingestion_status
The state of one ingestion run started by workbench_kb_add_documents: PENDING / RUNNING / DONE / FAILED, chunks written, and the error when it failed. Read it when the user asks whether their documents are in yet, or before a search that needs them — don't poll in a loop; ingestion of a few documents takes under a minute. Pass the runId the add-documents call returned.
workbench_kb_delete
Deletes a knowledge base and everything in it. Flows that search it will find nothing — check which agents use it (the user knows; the KB page in Workbench lists them) and say so before proposing. May return `needs_confirmation`.
workbench_kb_document_delete
Removes one ingested document and all its chunks; searches stop returning it at once. Get the documentId from workbench_kb_documents_list. To replace a document, remove it then workbench_kb_add_documents the new version. May return `needs_confirmation`.
workbench_kb_list
Lists every knowledge base in the active workspace the caller can see. Returns summaries (id, name, description, status, doc/chunk counts, embedding model). Use an id with `workbench_kb_search` to retrieve content.
workbench_kb_search
Retrieves the most relevant chunks from one knowledge base — the core RAG primitive. `mode:"semantic"` (default) embeds the query and ranks by meaning (spends a small amount of workspace credit); `mode:"keyword"` is a free case-insensitive substring match. Each hit carries the chunk text, its document, and a score (cosine similarity 0-1 for semantic, match count for keyword). Get a kbId from `workbench_kb_list`.
workbench_kb_create
Creates an empty knowledge base. Recipe 'docs' (default) for reference material an agent searches at runtime; 'tabular' for spreadsheet-style data queried with SQL; 'graph' for entity/relationship extraction. After creating, add content with workbench_kb_add_documents, then attach the KB to an agent in the Workbench editor (or tell the user to). May return `needs_confirmation`.
workbench_kb_add_documents
Adds documents to a knowledge base and starts ingestion (chunk + embed; graph KBs also extract entities; tabular KBs load CSV as tables). Text goes inline (Markdown, plain text, CSV — encoding utf8); binary files (PDF, .docx) go base64-encoded with their mimeType. Up to 50 sources per call, ~10 MB each. Ingestion runs in the BACKGROUND — the result carries a runId; check it with workbench_kb_ingestion_status when the user asks, don't poll. workbench_kb_search works once it finishes. Content from <attached-file> blocks is ideal source material. Check workbench_kb_documents_list first so you don't add a document twice. Embedding spends workspace inference credit, so this sits behind the approval gate: may return `needs_confirmation`.
workbench_shim_list
Every shim the user can see: id, name, the question it answers, how many answers and examples it has, and the published version's held-out accuracy (null until something is published). A shim is a small decision model that runs inside an app with no model call. Start here to get a shimId.
workbench_shim_get
One shim in full: the question, where its inputs come from, every answer with its description and examples, what the shim still needs before it can build, and the newest build's report — held-out accuracy, recall per answer, which answers get mistaken for which, and the compiler's notes. Read it before writing examples: match the setting and the existing examples' register, and put new examples where recall is low or the compiler says an answer is thin.
workbench_shim_create
Creates a shim from a name, the question it answers, and its answers. Add examples afterwards with workbench_shim_add_examples — three or more per answer and it builds on its own. May return `needs_confirmation`.
workbench_shim_add_answer
Adds one answer (a label and, optionally, when it applies). The shim needs three examples of it before it builds again with the new answer. May return `needs_confirmation`.
workbench_shim_add_examples
Adds example inputs to one answer. Write inputs a real person would actually type in the shim's setting — varied in length, register and specifics, each one clearly this answer and not another — never paraphrases of the answer's name. Read the shim first so new examples fit alongside the existing ones and land where recall is low. Exact duplicates are skipped. The shim rebuilds on its own afterwards; read it again for the new report. May return `needs_confirmation`.
comments_list
Lists the comment threads on one workspace entity (open first, then resolved) with authors and timestamps. Read this before weighing in on contested work — the threads are where disagreement lives before it becomes a decision.
comments_create
Posts a comment on a workspace entity — a new thread, or a reply when rootId is given. Use it to leave findings where the discussion already lives (an eval result on the flow being debated, a summary on a long thread). Mention people via mentionedUserIds (from workspace member ids) to ring their notification bell; never mention someone who didn't ask to be pulled in.
comments_resolve
Sets a comment thread's resolved state (rootId = the thread's root comment id). Resolve ONLY when the human asked or the thread's question is demonstrably settled — and say what settled it in a reply first. Reopening is for new evidence.
workbench_tasks_get
One Task in full: its sentence, status, the latest saved plan (the clauses — which run flows, which zv1 handles, who decides at each gate) with the flow and people names the plan refers to, and its recent runs. Read this before editing a task's sentence (workbench_tasks_compile with taskId) and when the user asks what a task does. Get the taskId from workbench_tasks_list.
workbench_tasks_update
Changes a task's name — the gallery label only. The sentence and plan are versioned and change through workbench_tasks_compile + workbench_tasks_freeze, never here. May return `needs_confirmation`.
workbench_tasks_delete
Deletes a task. Runs still in flight are stopped (a gate waiting on someone is closed), and its schedule trigger stops firing; past runs stay in history. Check workbench_tasks_get for live runs and say so before proposing. May return `needs_confirmation`.
workbench_tasks_list
Lists the workspace's Tasks — sentence-orchestrated processes — with their sentence, status, and recent runs (id, status, who a waiting run is on). Use this to find a task the user names ('claims intake') and to see what's in flight; workbench_tasks_get reads one in full.
workbench_tasks_create
Creates a DRAFT Task from a name + sentence. Use this once the user's sentence feels settled after you've refined it together — then compile it. The sentence should be THEIR words: trigger first ('When a claim arrives…'), then steps, branches, and who decides what.
workbench_tasks_compile
Runs the compiler over a sentence and returns the PROPOSED plan (or repair-grade problems — fix by refining the sentence with the user, not by guessing). Nothing persists. Present the proposal in plain terms: which clauses run flows, which you (zv1) will handle, who decides, what the compiler noted. Then iterate or freeze. Compile may take ~15s.
workbench_tasks_freeze
Freezes a reviewed plan as the Task's next version and activates it. Pass the EXACT plan object a compile returned — never hand-edit it (change the sentence and recompile instead). Freeze ONLY after the user has seen the proposal and said go; echo what you're freezing in one line first. In-flight runs keep their pinned versions. After freezing, offer to start a first run.
workbench_tasks_start
Kicks off a run of a Task. When the task expects an input document (its trigger has an input), pass the user's material as `input` — pasted text, extracted attachment text, or JSON. Cite the run id back and tell the user what happens next (the run may immediately be waiting on someone).
workbench_tasks_run_view
The run's current state rendered as its sentence: what ran (with timings), which path each gate took, what's waiting on whom and for how long, what's coming up. Use whenever the user asks where something stands, and after acting on a gate.
workbench_tasks_act
Records THIS user's verdict on a run that is waiting on them — the labels come from the gate (workbench_tasks_run_view shows them via the plan; a wrong label errors listing the real options). Echo the verdict + note back to the user in one line BEFORE calling; that echo is the confirmation. Refuses when the gate is assigned to someone else.
search_workspace
Finds entities across every tool by name in one call — Workbench flows, Compass pages, Caliper datasets, evals, rubrics, reviews, specs and sources (apps sending agent traces), Ledger entries, Napkin sketches and decks. Use it FIRST when the user names something without saying where it lives ('the onboarding flow', 'that invoice page'); reach for a tool's own list only when you already know the tool. Each hit carries its id, kind, and workspace-relative path, so the id feeds the matching *_get tool and the path makes a link. Results only include what the user can see, and only kinds this token may read.
entity_tags_get
Returns the tags on a batch of entities of one kind — the labels galleries organize by. Ids come from the kind's list/get tool or from search_workspace. Use it before entity_tags_set so you replace the full set knowingly, and to answer 'what is this filed under'. Entities the user can't see are omitted.
entity_tags_browse
Without a tag: every tag in use across the workspace with how many entities carry it, most-used first — the vocabulary the team already organizes by. With a tag: everything filed under it across every tool, each with its kind, id, title, and path. Use it to reuse existing labels instead of inventing near-duplicates, and to answer 'show me everything about X' when X is a label.
entity_tags_set
Replaces the FULL tag set on one entity (an empty list clears it). Read the current tags with entity_tags_get first and pass the merged list — this is not additive. Tags are lowercase letters, numbers, spaces, and hyphens; prefer labels already in use (entity_tags_browse) so the workspace's vocabulary stays small. The id comes from the kind's list/get tool or search_workspace.