ZeroWidth
Search, write and run work across Compass, Workbench, Caliper, Ledger, Prism and Napkin.
- 1.0.0
- Version
- remote
- Transport
- 200
- Tools
Security review
Review passedReviewed 1d ago.
- tools: 200 tools scanned
- metadata: scanned
No findings.
Tools (200)
search_docs
Search ZeroWidth product documentation. Returns matching pages with title, slug, public URL, and a query-relevant snippet. Use this when the user asks about a ZeroWidth product (Compass, Workbench, Caliper, Prism, Ledger, Napkin, zv1), an API behavior, or a policy. No authentication required — the docs corpus is public.
get_doc
Fetch the full Markdown body of a specific docs page by its slug. Use this after `search_docs` when the user needs the complete content of a page. No authentication required.
list_docs
Enumerate all available docs pages, optionally filtered by product (e.g. 'compass', 'legal', 'overview'). Use this to discover what slugs exist before calling `get_doc`. No authentication required.
list_posts
Lists ZeroWidth's published Distributed Cognition essays (newest first): title, slug, excerpt, byline, public URL. Use this when the user asks what's been published, or when a workspace question might already have an essay behind it — then fetch the full text with `get_post`. No authentication required.
get_post
Fetch the full Markdown body of one Distributed Cognition essay by slug (from `list_posts`). Link the public URL when citing it to the user. No authentication required.
workbench_flows_revisions_list
The flow's revision history, newest first: published versions carry a `version` label and note; entries with version null are draft autosaves. Use it to answer 'is this published?' (any entry with a version), to find a revision id for caliper_evals_create's flowRevisionId, or to see when the draft last changed.
workbench_flows_update
Changes a flow's name, description, visibility, tags, or archived state — the metadata around the flow, NOT its orchestration (prompts and nodes change through workbench_flows_edit_text). Pass only what changes. Archiving hides the flow from the default gallery but leaves it runnable, scheduled, and shared; pass archived: false to restore. May return `needs_confirmation`.
workbench_flows_publish
Snapshots the flow's current DRAFT as a named published version — the version schedules (workbench_flows_schedule), guest share links (workbench_flows_share_create), public-API runs, and 'published'-stage evals execute. The draft keeps evolving after this; runs on the published side don't change until the next publish. Publishing does NOT run the flow, but any eval set to runOnPublish starts a run (that spends credit — mention it when one exists). Publish only after the user has tested the draft (workbench_flows_run) or asked for it. May return `needs_confirmation`.
workbench_flows_delete
Deletes a flow. Its schedules stop, guest share links die, and evals targeting it can no longer run (their binding shows flowOk: false) — check caliper_flow_performance for evals and workbench_flows_schedules_list before proposing, and prefer workbench_flows_update with archived: true when the user just wants it out of the way. May return `needs_confirmation`.
workbench_flows_fork
Copies a PUBLIC flow from another workspace (a template) into this one as a new draft the user owns. Identify it by its flowUuid (the id on Workbench's public template pages and in `workbench_flows_get` output); optionally pin which published version to copy. Flows already in this workspace can't be forked — open them instead. Counts toward the plan's flow cap. May return `needs_confirmation`.
workbench_flows_schedules_list
Every schedule on one flow: cadence, whether it's paused (`enabled: false`), the next fire time, delivery, and the last run's status. Read this before workbench_flows_schedule (don't create a duplicate) and to get the scheduleId for workbench_flows_schedules_update / workbench_flows_schedules_delete / workbench_flows_schedules_run_now.
workbench_flows_schedules_update
Changes one schedule: `enabled: false` pauses it (configuration kept), `enabled: true` resumes and recomputes the next fire time from now; `interval` / `scheduleTime` / `scheduleTimezone` / `scheduleDayOfWeek` change the cadence; `label`, `delivery`, and `input` can change too. Pass only what changes. Get the scheduleId from workbench_flows_schedules_list. May return `needs_confirmation`.
workbench_flows_schedules_delete
Removes a schedule for good. Prefer workbench_flows_schedules_update with enabled: false when the user might want it back. May return `needs_confirmation`.
workbench_flows_schedules_run_now
Runs one iteration of a schedule immediately through the real scheduled pipeline (same input, same delivery — the email or channel it normally posts to) without moving its cadence. Works on paused schedules. Use it to test a schedule the user just set up. Spends workspace credit and delivers for real, so it sits behind the approval gate — may return `needs_confirmation`. The run is fire-and-forget: check workbench_flows_schedules_list for lastStatus, or workbench_executions_get for the trace.
workbench_flows_shares_list
Every guest link on one flow, newest first: label, whether it's still live, the shareUrl (live links only — dead ones have no URL), expiry, and how many guests and messages it has seen. Use it to answer 'who has access', to recover a link the user lost, and to get the shareId for workbench_flows_shares_revoke.
workbench_flows_shares_revoke
Kills one guest link immediately — anyone holding it gets a closed page from then on. Conversations already had stay in the flow's history. Get the shareId from workbench_flows_shares_list. May return `needs_confirmation`.
workbench_flows_run
Execute a flow and return its outputs synchronously. Runs the draft by default; pass source:"published" (optionally a version) to run the live published version. Spends workspace LLM budget, bounded by the token's cost cap. `input` is the flow's input envelope, e.g. {"kind":"chat","messages":[…]} or {"kind":"form","values":{…}}.
workbench_flows_schedule
Set a specific built flow to run automatically on a cadence — the deterministic counterpart to a zv1 routine. Each run executes the flow's latest PUBLISHED revision with the fixed `input` envelope and routes the output per `delivery`. USE THIS when the user wants a flow they've built to run on a schedule ("run my digest flow every morning"). Do NOT create a routine for this — a routine runs a free-form instruction, not a built flow. The schedule runs the PUBLISHED flow, so the flow must be published (check workbench_flows_revisions_list; publish with workbench_flows_publish if it isn't). `input` is fixed for every run, so the flow itself should fetch anything time-varying at run time. Manage existing schedules with workbench_flows_schedules_list / workbench_flows_schedules_update.
workbench_flows_list
Lists every flow in the active workspace the caller can see. Returns summaries (id, name, visibility, updatedAt) — fetch one with `workbench_flows_get` for the full orchestration body.
workbench_executions_get
The flow's stack traces: recent executions with per-node timelines — what each node received, produced, how long it took, and the exact error when one failed. USE THIS when a run misbehaves instead of guessing: read the failing node's inputs/error, then propose a fix (workbench_flows_edit_text) grounded in what actually happened. Pass executionId to inspect one run, or just flowId for the most recent runs. Node inputs/outputs are truncated for transport — the full record is in the flow's dev drawer.
workbench_flows_get
One flow. Default `view: summary` lists its nodes and links so you can pick one; `view: node` with a nodeId returns that node's full settings and prompt — read THAT before workbench_flows_edit_text, and copy `find` text from it character-for-character, never from memory. `view: full` returns the whole orchestration body.
workbench_flow_authoring_guide
READ THIS FIRST before writing or editing any raw flow orchestration JSON. Covers the document shape, the settings-vs-input-ports rule (temperature, max_tokens, tools, response_format are PORTS, not settings — constants reach ports via value nodes), exact link format, plugin links, and three complete worked examples.
workbench_node_catalog_get
The ground truth for what nodes exist and what their ports actually are — verify against this instead of recalling. Search by keyword/category for summaries; pass `slug` for one node's full detail (inputs, outputs, settings). Inputs are PORTS fed by links; settings live on the node — see `workbench_flow_authoring_guide`. The catalog is the WORKSPACE'S: model nodes the workspace's inference policy forbids come back with `allowed: false` — never author with those; pick an allowed model. Pass `flowId` to include the flow's pinned imports as nodes.
workbench_flows_scaffold
Creates a runnable first-draft flow from a spec you author: pick the simplest pattern that fits (classifier for read-and-bucket, structurer for transform/extract/draft-for-review, agent for genuinely conversational), write a production-quality system prompt grounded in what the user told you, and mark anything stubbed with [STUB: ...] markers plus stubNotes. The draft opens in Workbench's simple editor at /w/<workspace>/flows/<id> — give the user that path. May return `needs_confirmation`; show the user what you're proposing and wait for their approval, then re-call with the approvalId. When the draft comes from a Compass change, pass its opportunityId so the flow and the change point at each other (Compass shows 'Open in Workbench'; the flow shows where it came from). Then: workbench_flows_run to try it, workbench_flows_publish when it's ready for schedules and share links.
workbench_flows_edit_text
Applies ONE precise text replacement to a flow's DRAFT orchestration — Edit-tool semantics: `find` must be the EXACT current text, copied character-for-character from workbench_flows_get, and must occur exactly once anywhere in the flow (system prompts, node settings, metadata, a node's type). `find` is text inside ONE value, not JSON: to change a model, find the model slug alone (`anthropic-claude-haiku-4-5`) and replace it with another `llm` slug from workbench_node_catalog_get. It can't add or remove nodes or links. Zero or multiple matches return an error instead of guessing. This is the improvement primitive: check receipts first with caliper_flow_performance, cite the run id in `note`, apply the edit after approval, then re-run the eval with caliper_evals_run and report the score delta — never claim improvement without the before/after. Edits land on the draft only; the published version changes when someone publishes (workbench_flows_publish, after the user has seen the result).
workbench_flows_share_create
Creates a no-account guest link where anyone can chat with the flow's latest PUBLISHED revision at workbench's /s/<token> page — the fastest way to put a working flow in a stakeholder's hands ('here, try it'). The flow must have a published version (workbench_flows_revisions_list shows it; publish with workbench_flows_publish if not). Conversations are capped per guest; the link can expire, and you can list links with workbench_flows_shares_list and kill one with workbench_flows_shares_revoke. May return `needs_confirmation` — say who the link is for and wait.
workbench_kb_get
One knowledge base: name, recipe (docs / tabular / graph), status, document and chunk counts, embedding model, and ingestion settings. Read it to confirm a KB has content before pointing a flow at it. Get the kbId from workbench_kb_list.
workbench_kb_documents_list
Every document ingested into the KB's current content: id, name, type, size, chunk count. Empty for a KB with nothing ingested yet (or one whose ingestion is still running — check workbench_kb_ingestion_status). Use it before workbench_kb_add_documents to avoid re-adding a document, and to get the documentId for workbench_kb_document_delete.
workbench_kb_ingestion_status
The state of one ingestion run started by workbench_kb_add_documents: PENDING / RUNNING / DONE / FAILED, chunks written, and the error when it failed. Read it when the user asks whether their documents are in yet, or before a search that needs them — don't poll in a loop; ingestion of a few documents takes under a minute. Pass the runId the add-documents call returned.
workbench_kb_delete
Deletes a knowledge base and everything in it. Flows that search it will find nothing — check which agents use it (the user knows; the KB page in Workbench lists them) and say so before proposing. May return `needs_confirmation`.
workbench_kb_document_delete
Removes one ingested document and all its chunks; searches stop returning it at once. Get the documentId from workbench_kb_documents_list. To replace a document, remove it then workbench_kb_add_documents the new version. May return `needs_confirmation`.
workbench_kb_list
Lists every knowledge base in the active workspace the caller can see. Returns summaries (id, name, description, status, doc/chunk counts, embedding model). Use an id with `workbench_kb_search` to retrieve content.
workbench_kb_search
Retrieves the most relevant chunks from one knowledge base — the core RAG primitive. `mode:"semantic"` (default) embeds the query and ranks by meaning (spends a small amount of workspace credit); `mode:"keyword"` is a free case-insensitive substring match. Each hit carries the chunk text, its document, and a score (cosine similarity 0-1 for semantic, match count for keyword). Get a kbId from `workbench_kb_list`.
workbench_kb_create
Creates an empty knowledge base. Recipe 'docs' (default) for reference material an agent searches at runtime; 'tabular' for spreadsheet-style data queried with SQL; 'graph' for entity/relationship extraction. After creating, add content with workbench_kb_add_documents, then attach the KB to an agent in the Workbench editor (or tell the user to). May return `needs_confirmation`.
workbench_kb_add_documents
Adds documents to a knowledge base and starts ingestion (chunk + embed; graph KBs also extract entities; tabular KBs load CSV as tables). Text goes inline (Markdown, plain text, CSV — encoding utf8); binary files (PDF, .docx) go base64-encoded with their mimeType. Up to 50 sources per call, ~10 MB each. Ingestion runs in the BACKGROUND — the result carries a runId; check it with workbench_kb_ingestion_status when the user asks, don't poll. workbench_kb_search works once it finishes. Content from <attached-file> blocks is ideal source material. Check workbench_kb_documents_list first so you don't add a document twice. Embedding spends workspace inference credit, so this sits behind the approval gate: may return `needs_confirmation`.
workbench_shim_list
Every shim the user can see: id, name, the question it answers, how many answers and examples it has, and the published version's held-out accuracy (null until something is published). A shim is a small decision model that runs inside an app with no model call. Start here to get a shimId.
workbench_shim_get
One shim in full: the question, where its inputs come from, every answer with its description and examples, what the shim still needs before it can build, and the newest build's report — held-out accuracy, recall per answer, which answers get mistaken for which, and the compiler's notes. Read it before writing examples: match the setting and the existing examples' register, and put new examples where recall is low or the compiler says an answer is thin.
workbench_shim_create
Creates a shim from a name, the question it answers, and its answers. Add examples afterwards with workbench_shim_add_examples — three or more per answer and it builds on its own. May return `needs_confirmation`.
workbench_shim_add_answer
Adds one answer (a label and, optionally, when it applies). The shim needs three examples of it before it builds again with the new answer. May return `needs_confirmation`.
workbench_shim_add_examples
Adds example inputs to one answer. Write inputs a real person would actually type in the shim's setting — varied in length, register and specifics, each one clearly this answer and not another — never paraphrases of the answer's name. Read the shim first so new examples fit alongside the existing ones and land where recall is low. Exact duplicates are skipped. The shim rebuilds on its own afterwards; read it again for the new report. May return `needs_confirmation`.
caliper_datasets_get
One dataset with a page of its items — read this BEFORE editing items (caliper_datasets_update_item needs the item id) and before extending coverage, so you don't add cases that already exist. Items come `limit` at a time (default 25) from `offset`; `total` is the full count. Long fields are cut at ~800 characters with a truncation marker. Get the datasetId from caliper_datasets_list.
caliper_datasets_update
Changes a dataset's title, description, or visibility. Pass only what changes; items are untouched (use caliper_datasets_add_items / caliper_datasets_update_item / caliper_datasets_remove_item for those). May return `needs_confirmation`.
caliper_datasets_delete
Deletes a dataset. Evals bound to it stop being runnable (their binding shows datasetOk: false), and reviews/specs over its items lose their source — check caliper_evals_list for evals that reference it and say so before proposing. May return `needs_confirmation`.
caliper_datasets_update_item
Rewrites one item's content, keeping its id (ratings and eval results keyed to it stay attached). The item's KIND can't change — a Q&A item stays Q&A, a sequence stays a sequence — so pass the same shape you read from caliper_datasets_get. Chat and raw items can't be edited from here. May return `needs_confirmation`.
caliper_datasets_remove_item
Drops one item. Ratings and run results keyed to it are orphaned (kept in history, gone from the dataset). A dataset must keep at least one item. May return `needs_confirmation`.
caliper_datasets_generate
Creates a NEW dataset of model-written items — Q&A by default, or scripted sequences / simulated people via `shape` — pass flowId and the generator reads the flow's prompt, mode, and schema to write realistic cases for THAT flow; `description` adds guidance (or stands alone when there's no flow). Use this when the user wants test cases fast and has none; prefer caliper_datasets_create with hand-written items when real scenarios are already in hand (Compass pages, a transcript). This SPENDS workspace inference credit (one generator call), so it sits behind the approval gate: say so and expect `needs_confirmation`. Returns the dataset summary; read the items with caliper_datasets_get and tell the user to review them before trusting an eval built on them.
caliper_datasets_generate_items
Appends model-written items to an EXISTING dataset, in the style of what's already there (existing items are the few-shot examples; the output shape matches theirs unless overridden). Use to widen coverage when the user says 'more like these' or 'add edge cases'; write them by hand with caliper_datasets_add_items when the scenarios are known. SPENDS workspace inference credit, so it sits behind the approval gate — say so. Existing items and their ratings are untouched.
caliper_rubrics_list
Lists the rubrics in the active workspace (id, name, description, criterion count). Check here BEFORE caliper_rubrics_create — reuse an existing rubric's id in caliper_evals_create when one already scores the same job. Read a rubric's criteria with caliper_rubrics_get.
caliper_rubrics_get
One rubric with its criteria — each criterion's name, what the judge looks for, and the score scale. Read this to explain a score (which criterion slipped and what it asks for) or before caliper_rubrics_update. Get the rubricId from caliper_rubrics_list or an eval's rubricId.
caliper_rubrics_update
Changes a rubric's title, description, visibility, or replaces its criteria wholesale (pass the FULL list — criteria get fresh ids). Evals snapshot the rubric when they're created, so an existing eval keeps scoring with the criteria it started with; say so, and offer to create a new eval when the criteria change materially. May return `needs_confirmation`.
caliper_rubrics_delete
Deletes a rubric. Existing evals keep their snapshot of it and keep running; nothing new can bind to it. May return `needs_confirmation`.
caliper_evals_get
One eval: what it targets (Workbench flow + stage, or external), its dataset and rubric ids, the rubric criteria it scores with (the snapshot taken at creation), schedule + runOnPublish + regression settings, run counts, latest score, and `binding` health (datasetOk / flowOk false = a run would fail at resolution — say so before proposing caliper_evals_run). Get the evalId from caliper_evals_list or caliper_flow_performance.
caliper_evals_runs_item_execution
What the flow actually DID on one run item: every step in order (what it said, which tools it called with what arguments, what came back) and the final answer, plus status, duration, and cost. Read this when a low score needs explaining beyond the judge's reasoning — a wrong tool call or an empty tool result is usually the cause, and the fix is different from a prompt fix. Pass the run item's `id` (NOT `itemId`) from caliper_evals_runs_get. Items whose outputs were supplied from outside (external evals) have no trace and return not_found. Steps are capped for transport; the run page in Caliper has the full record.
caliper_evals_update
Changes an eval's title, description, visibility, schedule (interval + time anchor), runOnPublish (queue a run whenever the flow publishes), or regressionThreshold (score drop vs the previous run that triggers an alert; null = off). Pass only what changes. The dataset, rubric snapshot, and target are fixed at creation — create a new eval to change those. Scheduled and on-publish runs spend credit on their own, so state that plainly when turning them on. May return `needs_confirmation`.
caliper_evals_delete
Deletes an eval and stops its schedule. Run history is kept but no longer reachable from the eval. May return `needs_confirmation`.
caliper_evals_runs_cancel
Stops a run that is still PENDING / RUNNING / SCORING — the brake on a run that's spending more than expected or was started by mistake. The queue stops at once; an item already handed to the flow finishes on its own timeout. The run settles as FAILED with a 'cancelled' reason and keeps the items it completed. A run that already finished returns run_not_live. May return `needs_confirmation`.
caliper_sources_list
Apps sending their agent's traces to Caliper — from their own code, OpenTelemetry, or a published Workbench flow — busiest first: name, how it sends, traces in the last 14 days (per day), and the tools its agent calls most with the share of traces using each. Start here when the user asks how their agent behaves on real traffic; then caliper_sources_over_time for one source.
caliper_sources_over_time
One source day by day: traces, failures, tool calls, lookups that found nothing, tokens, cost, and speed (meanMs; p50/p95 as 'answered within' bucket edges); each tool with its calls and the share of each day's traces that used it; the documents lookups landed on most; models used. Use it to answer 'how often is it calling web search and is that changing' — compare the first and last weeks and name the day it moved. Follow with caliper_traces_list on that day or tool to show the conversations behind it.
caliper_traces_list
The traces behind a point on a source's chart, newest first, 50 a page: the question and answer (cut short), tools called, time, tokens, cost, status. Narrow by day, tool, document, or failures; page with `before` (the nextBefore from the last page). Open one with caliper_traces_get.
caliper_traces_get
One trace: what came in, what went out, and every step in order — model calls (model, tokens, cost, reasoning), tool calls (arguments and result), lookups (query and documents with scores), anything else — each with timing and status. Long values are cut at 2,000 characters. Use it to explain why one conversation went the way it did.
caliper_source_feeds_list
Datasets this source keeps adding to as conversations arrive: which dataset, for what (review, eval, spec), the pick it matches, sampling, how many have landed, and whether it's still running or why it stopped.
caliper_sources_to_dataset
Saves the conversations behind a pick (a day, a tool, a document, failures — the same narrowing as caliper_traces_list) into a new dataset (newDatasetName) or an existing one (datasetId), newest first up to `limit`. `purpose` decides what each item keeps: review = the agent's reply, tool calls included, for people to rate; eval = the reply becomes the expected answer; spec = only the questions, for people to answer. Set keepAdding to make it live: new matching conversations keep arriving (every one, or 1 in 10 / 1 in 100), up to 1,000. Then offer the next step — caliper_evals_create, or a review in Caliper. May return `needs_confirmation`.
caliper_source_feeds_stop
Stops a source from adding new conversations to a dataset. What it already added stays. Get the feedId from caliper_source_feeds_list. May return `needs_confirmation`.
caliper_starters_list
The starter sets Caliper ships: prompt injection, system-prompt leaking, over-refusal, personal-data handling, bias under ambiguity. Each is a small original dataset (single messages AND multi-turn build-ups: false memory, fabricated earlier turns, slow escalation) paired with an anchored rubric. Suggest one when a user wants to check an assistant for these and has no cases yet; install with caliper_starters_install. Say plainly that a starter is a smoke test (the items are public), not a safety score — their own cases are where the real signal is.
caliper_starters_install
Copies one starter into the workspace as a new dataset and a new rubric the user owns and can edit. Returns both ids; the next step is caliper_evals_create binding them to the flow (or target 'external'). Needs caliper:datasets:write and caliper:evals:write.
caliper_datasets_list
Lists datasets in the active workspace. Returns summaries (id, name, item count); items are not included.
caliper_flow_performance
How a Workbench flow is ACTUALLY doing, with receipts: every Caliper eval targeting the flow, recent runs with scores, the latest run decomposed into per-criterion averages, the score delta vs the previous run, and the worst-scoring items WITH the judge's reasoning. Use this BEFORE claiming a flow works or proposing changes — and cite the runId + scores when you do. The worst items are diagnostic: failures clustered around missing company facts suggest a knowledge gap (consider proposing a Compass interview with the workflow owner) rather than a prompt problem.
caliper_evals_runs_list
Recent runs for one eval, newest first: status (PENDING/RUNNING/SCORING/DONE/FAILED), overall score once DONE, label, and timestamps. THE CHECK-BACK for caliper_evals_run: when the user asks how the run went, read this — cite the run id and score, and compare against the PREVIOUS run's score for the delta. Still don't poll in a loop; check when the user asks or when reporting.
caliper_evals_runs_get
One run: status, overall score, and by default the ten lowest-scoring items with their per-criterion scores and the judge's reasoning (`items: all` for every item, `none` for totals only). THE ANSWER to 'why did the score drop' — read this, then name the criterion that slipped and quote the reasoning on the lowest items. Works for runs Caliper ran and for runs submitted from CI (triggeredBy 'ci', usually labeled with a pull request number). Still no polling loops; read it when the user asks.
caliper_evals_list
Lists evals in the active workspace with their target config (which Workbench flow, which dataset/rubric). Use to find the evalId for caliper_evals_run.
caliper_evals_run
Queues a new run of an eval — inference over the dataset, then LLM-judge scoring. THE VERIFY STEP of the improvement loop: after an approved workbench_flows_edit_text, run the eval again and report the score delta vs the previous run. Runs take a while — but you're brought back into THIS conversation automatically with the scores the moment it finishes, so tell the user it's queued and that you'll follow up here; never poll or ask them to check back. Costs workspace LLM budget, so it sits behind the approval gate: may return `needs_confirmation`.
caliper_datasets_create
Creates a dataset of test items — THE FIRST STEP of setting up evaluation for a flow. Three item shapes: Q&A (input + optional expectedOutput, the golden answer); SEQUENCE (turns: 2-20 scripted user messages the model answers one at a time with its own earlier replies in front of it, + expectedResponse for the final reply, optional expectedBehavior for the whole conversation); SIMULATED (goal + optional persona/disposition/strategy/maxTurns + expectedBehavior; a platform flow plays a person adaptively, Caliper-run evals only; disposition is a preset id like genuine, pressure, confused, impatient, or free text). Use sequences and simulated items for the slow attacks and for real customers with real needs: a model that holds on message one often folds on message ten. Write good inputs from real usage: the Compass pages the flow was built from are the best source of realistic scenarios.
caliper_datasets_add_items
Appends items to an existing dataset — use this to grow coverage (new edge cases, scenarios from a completed interview) instead of creating a parallel dataset. Same three shapes as caliper_datasets_create (Q&A, sequence, simulated). Existing items and their ratings are untouched.
caliper_rubrics_create
Creates the scoring rubric an eval's LLM judge uses — 1-10 criteria, each scored on a numeric scale (default 1-5), judged pass/fail (kind pass_fail), or checked in code with no judge (kind check: contains, not_contains, matches a regex, valid_json, max_chars, equals_expected). Write criteria about the FLOW'S JOB (accuracy to source material, tone, refusal behavior), not generic 'quality'.
caliper_evals_create
Binds a dataset + rubric to something under test as a repeatable eval — THE LAST SETUP STEP before scoring. Two targets: a Workbench flow (pass flowId; Caliper runs inference itself, then caliper_evals_run scores it), or 'external' (target: 'external'; the user's own model runs elsewhere and their script submits outputs through the public API, usually from CI — see the docs guide 'Run evals in CI'). For an external eval, hand back the eval id and tell the user to create a workspace API key with the ci_evals preset in their workspace settings; keys can't be minted from here. flowStage 'draft' evals the live draft (pre-publish); 'published' (default) evals the latest published revision at run time, or one pinned with flowRevisionId. Reuse an existing rubric from caliper_rubrics_list when one already scores this job.
compass_pages_list
Lists pages in the active workspace's Compass compendium. Optional `type` filter narrows to one node type (WORKFLOW, PERSON, SYSTEM, PAIN_POINT, DOCUMENT, VALUE, PRIORITY, AUDIENCE, OFFERING). Returns summaries, 50 at a time (`limit` / `offset`, `total` and `nextOffset` in the result) — fetch one with `compass_pages_get` for the full body. When you know what you're looking for, `compass_pages_search` is the better first call.
compass_pages_search
Keyword search over page titles, bodies, and tags (case-insensitive substring match), paginated. The fast way to check whether a subject already has a page before creating or linking. Returns summaries with a 300-char body snippet — fetch full text with `compass_pages_get`.
compass_pages_get
Fetches one Compass page by id, including its full Markdown body, header fields, and any attached Napkin sketches — view an attached sketch's actual drawing with `napkin_boards_view` (its boardId) before discussing it.
compass_pages_create
Creates a typed page in the workspace's Compass compendium. Check `compass_pages_list` first — don't create a page for a subject the map already has; link to it instead. May return `needs_confirmation` — if so, tell the user what you're proposing, wait for their approval, then re-call with the approvalId. Only record what the human has actually told you.
compass_pages_update
Edits an existing page's title, description, body, tags, header fields, or visibility — use this to FIX what you (or an extraction) got wrong instead of creating a duplicate. Fetch the current page with `compass_pages_get` first and preserve what the human wrote; title / body / tags replace the field wholesale, while `headerFields` MERGES over the current ones (send only the keys you're changing — 'set the owner' leaves status alone). Page ids come from compass_pages_list. May return `needs_confirmation` — tell the user what you're changing, wait for approval, then re-call with the approvalId.
compass_links_create
Creates a typed, directed edge between two pages — the knowledge graph's connective tissue. Canonical directions: PERSON owns WORKFLOW/SYSTEM, PERSON involved_in WORKFLOW, WORKFLOW uses SYSTEM, SYSTEM uses SYSTEM, PAIN_POINT affects WORKFLOW/SYSTEM/PERSON, DOCUMENT documents anything, relates_to as fallback. Idempotent on (from, to, kind) — re-creating an existing edge returns it. May return `needs_confirmation`; tell the user what you're proposing, wait for their approval, then re-call with the approvalId.
compass_page_links_list
All typed edges touching one page, both directions, each hydrated with the other endpoint's page summary. Use this to understand a subject's neighborhood before adding to it.
compass_interviews_create
Mints a stakeholder-interview invite: a no-account guest link where the person talks to an interviewer agent briefed by your focusPrompt, and the transcript flows back into Compass as reviewable draft pages. THE KNOWLEDGE-GAP MOVE: when caliper_flow_performance shows failures clustered on missing company facts (the judge says the flow invented a policy, missed a rule, didn't know who owns something), the fix is usually not a prompt edit — it's asking the human who actually knows. Write a focusPrompt that names the SPECIFIC gaps (cite the eval run id), pick the owner of the relevant workflow as interviewee when the map knows one, and hand the user the invite link to forward. May return `needs_confirmation` — tell the user who you want to interview and why, then wait.
compass_interview_targets
Ranks the PEOPLE the map says know about a page (workflow, system, pain point): `owns` links first, then `involved_in`, then weaker edges. Each target carries contact email + role from their PERSON page and their latest interview (skip someone who just gave one — interview fatigue is real). Empty result = the map doesn't know an owner: ASK THE USER who runs this, create the PERSON page + owns link from the answer, and the map gets smarter. Use before compass_interviews_create to pick the interviewee.
compass_interview_invite
Emails the guest link for an existing interview, framed as coming from the REQUESTING USER (their name signs it; replies go to them). Author `message` in their voice — short, human, says why THEIR knowledge matters and that it takes ~15 minutes, no account needed. The approval card shows the exact subject + message before anything sends. Limit: the invite plus ONE reminder; a third ask is the user's conversation to have. Completion arrives as a notification with draft-page counts — don't poll.
compass_gaps_list
Lists the workspace's open questions — the gap registry: what the map doesn't know yet. Defaults to OPEN gaps. This is where you keep your head — check it before asking the user something you may already have flagged, and compose interview briefs from a subject's open questions.
compass_gaps_create
Records something you don't yet know — the gap registry is your working memory, so use it liberally while mapping or interviewing. Give one clear question, WHY it matters (what answering it unblocks), and how you'd resolve it (ask the user / interview a specific person / connect a source). No approval needed — noting your own uncertainty isn't acting on the user's behalf. Attach it to a subject (subjectType/subjectId) when it's ABOUT a specific page or person.
compass_gaps_resolve
Closes a gap once you've learned the answer (ANSWERED, with the resolution) or decided it doesn't matter (DISMISSED). Keep the registry honest — resolve gaps as their answers land (from an interview, a doc, or the user) so it always reflects what's still unknown. Gaps you raised close without a card; a question a person wrote needs their approval.
compass_opportunities_list
The workspace's changes — fixes, chores, new steps, experiments. Each carries `statusName` (the workspace's own word for where it is) and `statusCategory` (TRIAGE undecided, BACKLOG not started, ACTIVE in progress, DONE, CANCELED dropped), plus the older lane key in `status`. `experiment: true` marks a change measured against the Ledger: `ledgerEntryId` null means no expectation registered; `ledger.verdict` carries confirmed/missed after settlement. Scores (value/feasibility/risk) and a workflow anchor are optional. Filter by lane key in `status` or by workflowPageId; compass_statuses_list has the status names.
compass_opportunities_propose
File an AI-PROPOSED opportunity into the workspace's review queue — the scouting verb (the opportunity-scout routine's main move). Unlike compass_opportunities_create this needs NO approval: the proposal itself is the human gate — it lands in the Compass Inbox and the cockpit's Needs-you for accept/dismiss. Check compass_opportunities_list first so you never duplicate an idea. Anchor to a workflow page when one fits; score value/feasibility/risk 1–5.
compass_opportunities_create
Record a change — a fix, a chore, a new step, or an experiment worth trying. Anchor it to a workflow page when one fits (compass_pages_list); leave unanchored otherwise. It lands in the board's first status unless you name another (compass_statuses_list). Set `experiment` only when the user wants it measured against the Ledger. May return `needs_confirmation` — summarize and wait for approval.
compass_opportunities_register_expectation
The honesty mechanism: write what the experiment is expected to change BEFORE evidence exists. Creates a Ledger decision entry and links it to the opportunity — never backfill an expectation to match an outcome. Bind a metric (ledger_metrics_list) + comparator + target when the expectation is measurable; readings then land on the entry as evidence automatically. One expectation per opportunity — revise by superseding in Ledger. May return `needs_confirmation`.
compass_opportunities_mark_implemented
The measurement window's boundary: when the user says the experiment's change actually shipped / went live / rolled out, record the landing date. Readings before it are baseline; after it, evidence of effect. Recorded once — it cannot move afterward, so confirm the date. Attaches to the linked Ledger entry as evidence. May return `needs_confirmation`.
compass_map_view
Renders the workspace's visual map to an image and returns it so you can SEE it the way the user does: an isometric drawing where every page is a structure whose shape is its type — people are figures, audiences are crowds, workflows are gears lying on the ground, systems are database drums, offerings are price tags, pain points are warning signs, documents are standing sheets of paper, values are shields, priorities are flags — sized by how many other pages connect to them, placed near what they link to, with pages connected to nothing parked to one side, and routes drawn between linked pages. A coloured ring on the ground around a page shows the changes touching it by stage (blue undecided, amber not started, green in progress, violet done; thicker = more), with a key in the top-left corner. Use this when the user asks about the shape of their map, where the problems are, what connects to what, or anything spatial. The text part counts the pages of each type, the links, and how many
compass_pages_delete
Moves a page to the trash (soft delete — its connections and anchored opportunities go with it, and compass_pages_restore brings the whole set back). Use this to REMOVE a page you or an extraction created wrongly, or one the user says no longer belongs on the map; to fix a wrong title or body, use compass_pages_update instead. Page ids come from compass_pages_list / compass_pages_search. May return `needs_confirmation` — name the page and wait for approval.
compass_pages_restore
Puts a trashed page back on the map, together with exactly the connections and opportunities its delete took. Use when the user wants a deleted page back (the pageId from the earlier compass_pages_delete, or from the Compass trash). Fails with not_found when the page is already live or never existed. May return `needs_confirmation`.
compass_links_update
Edits an existing edge between two pages: its `kind` (how the two relate) and/or its `note`. Use to CORRECT a connection typed wrongly — 'Dana doesn't own billing, she's involved in it'. Link ids come from compass_page_links_list. Fails with conflict when the new reading already exists between the same two pages (delete this one instead). May return `needs_confirmation`.
compass_links_delete
Deletes one edge from the graph (the pages stay). Use when a connection is simply wrong — the system isn't used by that workflow, the person left the team. Link ids come from compass_page_links_list. Re-creating the same edge later revives it. May return `needs_confirmation`.
compass_opportunities_update
Corrects an opportunity's content: label, description, priority, what done means, the 1–5 value / feasibility / risk scores, the owner (a PERSON page), notes, tags, workflow anchor, external link, visibility. Fields you omit are untouched; the lane is NOT here — move it with compass_opportunities_set_status. Opportunity ids come from compass_opportunities_list. May return `needs_confirmation` — say what you're changing and wait.
compass_opportunities_delete
Removes an opportunity from the pipeline (soft delete). For an AI proposal the user doesn't want, this is the dismiss verb; for a captured experiment that was a duplicate or a mistake, the remove verb. To conclude a real experiment without evidence, prefer compass_opportunities_set_status REJECTED — that keeps the record. Ids come from compass_opportunities_list. May return `needs_confirmation`.
compass_opportunities_accept
The review verb for the Inbox: promotes an AI-PROPOSED opportunity (from compass_opportunities_propose or the opportunity scout) into the workspace's own pipeline. compass_opportunities_set_status does NOT do this — a proposal stays in the Inbox until accepted. Optional `status` decides its lane in the same step (BACKLOG to park it, EXPERIMENTING to start it); omitted, it lands in NEW. Fails with not_proposed when the row isn't an AI proposal. Ids come from compass_inbox_list / compass_opportunities_list. May return `needs_confirmation` — name the proposal and wait.
compass_interviews_list
Every stakeholder interview in the workspace, newest first: who was invited, status (INVITED / IN_PROGRESS / COMPLETED), focus, message count, whether the transcript was already reviewed (documentPageId set), and the invite URL while the link is still live. Check this before inviting someone again — interview fatigue is real — and to answer 'who have we already asked?'. Read one in full with compass_interviews_get.
compass_interviews_get
One interview in full — metadata plus the transcript of what the guest actually said. Read this before summarizing an interview or drafting pages from it: the transcript is the source you work from, and your compass skill has the rules for typing and splitting what's in it. Ids come from compass_interviews_list or compass_inbox_list.
compass_interviews_revoke
Kills an interview's invite link — the guest's next visit sees that the link is no longer active. Use when an invite went to the wrong person, the user changed their mind, or the link leaked. Idempotent. Ids come from compass_interviews_list. May return `needs_confirmation`.
compass_inbox_list
The Compass Inbox in one read: finished interviews awaiting review (transcript in, not yet turned into pages), AI-proposed opportunities awaiting accept / dismiss, and the caller's pending approval cards for Compass writes. THE place to answer 'what needs me?' for the map. Next moves: compass_interviews_get to read a transcript, then propose pages from it with compass_pages_create; compass_opportunities_accept or compass_opportunities_delete for a proposal.
compass_gaps_update
Edits an open question's wording, rationale, or suggested resolution — for sharpening a vague question or fixing one you phrased badly. Closing a gap is compass_gaps_resolve, not this. Gap ids come from compass_gaps_list. May return `needs_confirmation`.
compass_statuses_list
The statuses this workspace's changes move through, in board order — the names people use. Each sits in a group: Undecided (TRIAGE), Not started (BACKLOG), In progress (ACTIVE), Done (DONE), Dropped (CANCELED). Groups are stable across workspaces; names aren't, so use the names when talking to people and the groups when reasoning about progress. Move a change with compass_opportunities_set_status.
compass_opportunities_set_status
Move a change to one of the workspace's statuses, by name (compass_statuses_list has them — e.g. "In review"). The old lane keys (NEW / QUALIFYING / BACKLOG / EXPERIMENTING / SETTLED / REJECTED) still work and land on that lane's default status. Only changes marked as experiments follow the Ledger: when an experiment enters an In progress status with no registered expectation, offer compass_opportunities_register_expectation; when it enters Done with its Ledger entry still open, offer ledger_entries_settle. Other changes just move. May return `needs_confirmation`.
compass_changes_upsert
Create or update up to 100 changes in one call, each matched by `url` — the GitHub issue, Linear or Jira ticket, PR or doc it mirrors. A url already linked to a change updates that change (label, description, status, tags, experiment, workflow — omitted fields are left alone); any other url creates a change with that link. Run it again with the same urls to keep Compass in step; nothing duplicates. Statuses are by name (compass_statuses_list). Returns one result per item — created / updated with the change's key, or failed with why; a failed item doesn't stop the rest. Summarize what you're about to bring in before calling. May return `needs_confirmation` — one approval covers the whole batch.
prism_fields_list
Lists Prism fields (idea-exploration canvases) in the active workspace: name, origin idea, node count. Fetch one with `prism_fields_get` for its nodes.
prism_fields_get
One field's full tree: every node (id, parent, axis, description, pinned, color, whether it has an image / desk research) plus the AI-clustered groups. The origin node has parentId=null; children sit along named axes. Use node ids with the expand / research / study tools.
research_findings_list
Cross-instrument findings for one subject (today: a Prism node) — attached feedback studies with response counts + theme counts, and whether desk research exists. (Interviews are study-level now — read them per study with prism_interviews_list.) THE PLACE TO LOOK before proposing new research: cite what already exists.
prism_fields_create
Creates a field from an origin idea — the seed of an exploration canvas. Counts against the workspace's field quota. May return `needs_confirmation` — tell the user what you're proposing and wait for approval.
prism_nodes_expand
THE core Prism gesture: generate variations of a node along a named semantic axis ('more visceral', 'for the skeptic', 'stripped to essentials'…) in one of four directions. Runs generation against the workspace's inference credit. May return `needs_confirmation`.
prism_nodes_update
Sets pinned state and/or the color tag on a node. Pins are the cross-field shortlist. May return `needs_confirmation`.
prism_nodes_research
Web-sourced secondary research on a node's idea — market context with citations, stored on the node. depth 'thorough' runs three angled passes (market, evidence, shifts) at ~3× the credit. `focus` steers what it goes after; without one it answers a generic brief off the node's own description, so pass it whenever the user has said what they actually want to know. Runs against the workspace's inference credit. May return `needs_confirmation`.
prism_studies_list
Feedback studies attached to one Prism node: title, question count, response count, open/closed.
prism_studies_get
One study's questions, response count, open/closed state, and the STORED analysis (summary + themes with strength; insights.activeRun set = an analysis is running right now — wait for it before citing). participantUrl is the live share link to forward when the study is open and one has been minted. Cite theme strengths when reporting results.
prism_studies_draft
Generates a short study (5-8 questions probing the node's problem space — respondents never see the idea itself) and creates it as a DRAFT. The user reviews, previews, and opens it for responses from the study page (or via prism_studies_update publish). Runs generation against inference credit. May return `needs_confirmation`.
prism_studies_draft_standalone
Generates and creates a DRAFT study that isn't tied to any canvas node — brand tracking, workspace-level research, anything you can describe in a sentence. It appears on Prism's Studies page; the user previews and opens it for responses there (or via prism_studies_update publish). Runs generation against inference credit. May return `needs_confirmation`.
prism_studies_draft_from_interviews
Qual-first design: drafts a survey GROUNDED in the study's completed interviews — the recurring claims become measurable questions, the tensions become the choices, in the interviewees' own words. Replaces the study's current questions with the draft (nothing fields until publish). Needs at least one completed interview (see prism_interviews_list). An optional steer biases the instrument ('focus on pricing'). Runs generation against inference credit. May return `needs_confirmation`.
prism_studies_update
publish: true takes a draft live (one-way — the participant link starts working). closed toggles whether a live study accepts responses. Also edits the study: title, the `questions` instrument (whole replacement — read prism_studies_get first; LOCKED once responses exist, the server refuses with `invalid` and the answer is a new study), `adaptiveFollowUps` (AI probes on curious answers), the closing chat (`endChatEnabled` + `endChatFocus`), `completionRedirectUrl` for a BYO panel (null clears), and `autoAnalyzeTarget` (null = manual only). Study ids come from prism_studies_list_workspace. May return `needs_confirmation`.
prism_studies_results
Everything the Results tab shows, as numbers you can trust without counting raw rows: fielding funnel (opens → starts → completes, median completion time, sources, per-question drop-off), per-question aggregates (scale distributions + means + NPS on 0–10, choice/multi counts, rank first-place + mean ranks, MaxDiff set scores, word frequencies, text samples), and quality flags (speeders, duplicate participant ids). ALWAYS use this for quantitative questions about a study — never tally answers yourself.
prism_research_digest
A cross-study digest: every study with responses — title, origin (standalone / canvas node / tracker wave), response count, latest analysis summary, top themes, and which insights were kept in Ledger. The starting point for 'what have we learned about X?' — follow up with prism_studies_get (insights) or prism_studies_results on the studies that matter.
prism_studies_report
An AI research analyst writes the report's narrative layer — executive summary, key findings with their numbers, recommendations — from the study's computed record (funnel, per-question aggregates, themes, quotes). Stored on the study; the printable report at the study page renders it above the charts. Runs against inference credit. May return `needs_confirmation`.
prism_series_list
Recurring research programs (ADR 0031): name, cadence, enabled state, next wave time, wave count, and the metric + score rule when the series posts readings. Waves themselves are ordinary studies.
prism_series_launch_wave
CLOSES the current wave (no more responses; it scores and posts to Ledger if bound) and fields the next one from the series' frozen instrument, readying its participant link. Fielding spends credit for follow-ups, closing chats, and analysis, and closing a live wave can't be undone — say both plainly before proposing. Manual launches don't move the schedule clock. May return `needs_confirmation`.
prism_interviews_list
The study's one-on-one qual conversations: guest, status, focus, and — once completed — the structured memo (summary, key claims, tensions, verbatim quotes). Raw transcript included only while no memo exists. inviteUrl is the link to forward for interviews still waiting on their guest.
prism_studies_segments
The stored k-means segmentation over the study's numeric answers: named segments with size, share, and the distinguishing features (segment mean vs overall). Null when none computed yet — prism_studies_segments_compute discovers them (needs 30+ responses and 2+ numeric questions).
prism_studies_cut
Server-computed banner cut with significance: a target question (scale/number → group means + NPS on 0–10; choice/multi → per-option shares) split by acquisition source or by any single-choice question, each group tested against its complement at 95% (Welch t / two-proportion z). vsRest says higher/lower/not_significant — or not_tested when either side is under the 30-response floor; never present not_tested as 'no difference'. Use this for every 'does X differ by Y?' question instead of eyeballing raw rows.
prism_studies_analyze
Kicks the theme analysis (incremental — reads only responses and interviews since the last run; the previous themes carry forward). Returns immediately with the run's phase; the run continues server-side, so wait a moment and re-read prism_studies_get — findings are fresh once its insights stop reporting an active run. upToDate: true means there was nothing new (a healthy no-op, not an error). Runs against inference credit. May return `needs_confirmation`.
prism_interviews_create
Creates a one-on-one AI-led interview on a study and returns the invite link to forward — the guest needs no account. The transcript stays on the study and joins the next analysis run (it never lands in Compass). focusPrompt briefs the interviewer on what to dig into; omitted, it explores the study's territory. May return `needs_confirmation`.
prism_insights_promote
Promotes an analysis theme (or, with no themeId, the analysis summary) into the workspace's Ledger as a belief, carrying the supporting verbatims as evidence with links back to the study. THE step that turns a finding into workspace memory — use when the user says an insight matters. May return `needs_confirmation`.
prism_studies_segments_compute
Runs k-means over the study's numeric answers (silhouette-picked k) and names the discovered segments from their computed profiles. Needs 30+ responses and 2+ numeric questions; recompute overwrites. Runs naming against inference credit. May return `needs_confirmation`.
prism_fields_update
Edits a field's name (null reverts to the origin idea), description, or visibility. Nodes are untouched. Field ids come from prism_fields_list. May return `needs_confirmation`.
prism_fields_delete
Removes a field and its whole canvas (soft delete; frees the workspace's field quota). Use when the user is done with an exploration or one was created by mistake — studies attached to its nodes are not deleted. Field ids come from prism_fields_list. May return `needs_confirmation` — name the field and wait.
prism_studies_list_workspace
All studies regardless of origin — standalone, canvas-node, and series waves — newest first: title, question count, response count, draft / live / closed state, and the field / node / series it hangs off. The starting point for 'which studies do we have?'; prism_studies_list is the per-node view, prism_research_digest the findings view. Read one with prism_studies_get.
prism_studies_delete
Permanently deletes a study with its responses, interviews, analysis, and share links. Right for a draft that won't be fielded or a duplicate; for a live study the user just wants to stop, prefer prism_studies_update closed: true — that keeps the data. Study ids come from prism_studies_list_workspace. May return `needs_confirmation` — name the study and its response count, then wait.
prism_series_create
Creates a research program that fields the SAME instrument on a schedule — brand tracking, a weekly pulse — each wave an ordinary study. The questions freeze once the first wave fields (that's the trendline), so get them right with the user first. Optionally bind a Ledger metric + score rule so every wave posts a reading (a `scorer` needs a `metricId`). Waves launch on schedule, or on demand with prism_series_launch_wave. May return `needs_confirmation`.
prism_series_update
Edits a series: `enabled: false` pauses the schedule (true resumes), interval / scheduleTime / scheduleTimezone / dayOfWeek move the cadence, name and follow-up / auto-analyze settings change any time. The instrument (questions, scorer) is editable only while no wave has fielded — after that the server refuses with `invalid`, and the answer is a new series. Series ids come from prism_series_list. May return `needs_confirmation`.
ledger_entries_list
The workspace's memory: decisions (changes with pre-registered expectations), lessons (distilled beliefs), and observations (captured facts). CONSULT BEFORE ACTING — before proposing a flow change, prompt edit, or process decision, filter by the Compass page it touches (compassPageId) and check whether prior attempts exist and how they settled. Filter status=open for unsettled expectations awaiting evidence (each carries a derived `lapsed` flag — true when its deadline has passed; surface lapsed ones when the user asks what needs attention); kind=lesson for what the team already believes.
ledger_entries_get
One entry's full anatomy — the six fields, attached evidence (Caliper runs, measurements, observations), and the supersede chain (what replaced it, or what it replaced). Cite entry ids when telling the user about prior related decisions.
ledger_entries_create
Records a memory entry. Every entry is a TITLE (`summary`: one short plain sentence) over a BODY (`rationale`: the detail, markdown welcome) — never put the detail in the title. Three kinds: `decision` — a change being made now; PRE-REGISTRATION IS THE POINT, so `prediction` (what we expect) must be written NOW, before any evidence exists, and never backfilled to match an outcome. `lesson` — a distilled belief the team already holds (`lesson` text required, no prediction). `observation` — a durable fact worth remembering (no prediction): something already true, never something planned. Ideas, pitches, backlog items, and upcoming work are NOT entries — a dated piece of work that carries out a decision is a plan item (ledger_plan_add on that decision), and a running list you keep across runs belongs in a Napkin doc or sheet. When a user states something durable about their business in conversation, offer to capture it as a lesson or observation. When a user shares MEETING NOTES, propose
ledger_entries_settle
The ritual moment: evidence has landed, the expectation closes, the lesson is written. Only propose settlement when attached evidence actually answers the prediction — check `ledger_entries_get` first and cite the evidence in the lesson. `lesson` (what we now believe) is required and permanent; settlement happens exactly once. May return `needs_confirmation`.
ledger_metrics_list
The numbers the workspace watches — each with unit, latest reading, and how many open decisions are bound to it. Consult when a user mentions a number that sounds like a tracked metric, and before recording a reading.
ledger_metrics_create
Create a metric — a number the workspace watches (triage time, weekly signups, cost per run). Check ledger_metrics_list first; names are unique per workspace. When a user says they want to track or measure something, offer this. May return `needs_confirmation`.
ledger_metrics_record_reading
Record one observation of a tracked metric. `metricId` accepts a metric id OR its snake_case slug from ledger_metrics_list; an unknown one is a not_found error — check ledger_metrics_list, create it with ledger_metrics_create, or pass `createIfMissing: true` to mint a "measure" metric at that slug in the same call (an "event" metric when the reading carries `labels`). The reading automatically lands as evidence on every open decision whose prediction is bound to this metric — so when a user reports a number ("triage is down to 12 minutes"), offer to record it. For event-kind metrics, omit `value` to count one occurrence. May return `needs_confirmation`.
ledger_metrics_get
One metric's definition (unit, kind, direction, target, cadence) plus its readings newest first and the open decisions bound to it. Use before answering 'how is X trending?' or before recording a reading against it. `metricId` accepts the id or the snake_case slug from ledger_metrics_list. For an event metric whose readings carry labels, `labels` lists each label and its values, largest total first. Pass `slice` to get the series for part of it ("new users in DE"), or `by` to split the series by one label ("new users by country"); either returns `series`, summed per `bucket`.
ledger_entries_update
Corrects an entry in place — its title (summary), body (rationale), the pre-registered prediction of an OPEN decision (null clears it), or its facet (filing category; editable even after settlement). Works on open decisions, and on lessons, observations, and beliefs (they carry no expectation, so fixing their wording is fine any time). Use to fix a typo, swapped fields, or a misrecorded detail. Never edits kind or lesson. Fails with conflict on a settled decision — its claim is corrected by a human superseding it in Ledger — and on superseded or retracted entries. Entry ids come from ledger_entries_list. May return `needs_confirmation`.
ledger_entries_retract
Takes an entry out of the curated ledger — THE remedy when you recorded something wrong (a decision that wasn't made, a duplicate, a fact the user corrects). Reversible from Ledger and fully audited, so it is safe to offer as soon as the user says 'that's not right'. Not for overturning a settled claim the team once believed — a human supersedes that. Fails with conflict if already retracted. Ids come from ledger_entries_list. May return `needs_confirmation`.
ledger_entries_add_evidence
Records what happened against an open decision — a manual observation the user reports ('the pilot team says triage feels faster'), an implementation note, or a reference to a Caliper run. Evidence is what settlement later reads, so attach it as it arrives and cite it in the lesson. Metric readings attach themselves via ledger_metrics_record_reading — don't duplicate them here. Fails with conflict on superseded entries. Ids come from ledger_entries_list. May return `needs_confirmation`.
ledger_metrics_update
Corrects a metric's name, unit, description, kind (measure / event), direction (which way is good), target, cadence, icon, level, or the metrics it drives (its place in the metric tree). Readings are untouched. Only what you pass changes. The slug is not editable here — external writers address metrics by slug. Fails with conflict when a rename collides with a live metric. `metricId` is the id from ledger_metrics_list. May return `needs_confirmation`.
ledger_metrics_archive
Takes a metric out of the gallery (`archived: true`) or puts it back (`false`). Readings stay, and entries that settled against it still read correctly — this is the cleanup for a metric minted once and abandoned, or one the workspace stopped watching. `metricId` accepts the id or slug from ledger_metrics_list. May return `needs_confirmation`.
ledger_plan_list
Pass `entryId` for one decision's plan items, or `from` + `to` (YYYY-MM-DD, at most about a year apart) for the calendar: decision spans (recorded day → expectation deadline) plus every dated item inside the window, including standalone dates with no decision. Filter the calendar by `ownerUserId` to answer 'what's mine this week' or to find tomorrow's items to draft. Never use it to compare or rank people's output.
ledger_plan_add
Adds a plan item to an OPEN decision (the work that carries it out: a post, a launch step, an email). Omit entryId for a standalone date (a holiday, an event you're only watching). Items are all-day unless you pass startTime with timeZone (and optionally endTime), for a webinar, a scheduled post, or a launch at noon. If the date costs money or time and comes with an expectation, record it as a decision with ledger_entries_create instead. Use repeatWeeklyUntil for a weekly cadence: it writes one row per week (max 60), each movable on its own. Fails with conflict once the decision is settled. May return `needs_confirmation`.
ledger_plan_update
Edits one plan item while its decision is open: move it (dueOn), rename it, change the owner (null clears), attach the link, or set status (planned | done | skipped). Mark done only when the user says it shipped; a link alone doesn't mean done. Fails with conflict once the decision is settled. May return `needs_confirmation`.
ledger_plan_delete
Removes one plan item while its decision is open, for an item added by mistake or work that's no longer planned. If the work was planned and then dropped, prefer ledger_plan_update with status 'skipped' so settlement can see it. Fails with conflict once the decision is settled. May return `needs_confirmation`.
ledger_metric_presets_list
Ready-made metrics a business can start tracking, each with a unit, a cadence, which direction is good, a business surface (facet), and cross-cutting tags. Reach for this when a workspace has few or no metrics, when someone asks what they should be measuring, or when a decision needs a number to settle against and none exists. Filter by `facet` (where it lives in the business) or `tag` (what kind of number it is). Suggest a SMALL set — three to six that fit what you know about this business — and say in one sentence why each one, rather than listing the catalog. Adopt with ledger_metric_presets_adopt. These are starting points: a workspace renames and retargets them freely afterwards.
ledger_metric_presets_adopt
Creates the named presets as real metrics the workspace owns, tagged from the catalog. Safe to repeat: a preset the workspace already has comes back `already_present` rather than creating a second series, and a workspace's own renames and targets are never overwritten. Adopt only what the user agreed to — a metric nobody reads is noise on the Metrics tab, and eight thoughtful ones beat forty. Tell them the metrics start empty and the next step is a feed (ledger_feeds_create) or a first reading. May return `needs_confirmation`.
ledger_metric_starters_list
Starter trees by shape of business or team (subscription software, services firm, online store, sales team, service operation). Each lists its metrics with a level (outcome / driver / activity) and the links between them (which number moves which). Reach for this before ledger_metric_presets_list when a workspace has no metrics yet or asks how its numbers fit together: pick the starter that matches what you know about the business, describe its tree in a sentence or two, and offer to adopt it with ledger_metric_starters_adopt, dropping any metric that doesn't fit.
ledger_metric_starters_adopt
Creates a starter's metrics at their levels and links them. Pass `slugs` to keep only some of its metrics; links are made only where both ends exist. Safe to repeat and safe after a different starter: metrics the workspace already has are left as they are and get linked into the tree. Adopt only what the user agreed to. Tell them the metrics start empty and the next step is connecting a source (ledger_feeds_create) or a first reading. May return `needs_confirmation`.
ledger_entries_draft
Writes up to 50 entries as DRAFTS: proposals that stay out of the record until a person confirms each one in Ledger's Drafts view. Use it for entries you found rather than were told — decisions in meeting notes, or a team's past changes read from its tracker, pull requests or launch posts (set `fromHistory: true`). For history: take each claim from what the source said AT THE TIME, set `landedAt` to when it shipped, attach the source URL, and leave `rollout` as full unless the source says otherwise. These read as low confidence because they were written down after the fact; say so plainly rather than overstating them. Tell the person how many drafts are waiting and that they review them under Decisions → Drafts. May return `needs_confirmation`.
ledger_feed_sources_list
The connected servers this workspace exposes to you, with the id a feed needs. Their tools appear to you namespaced as `ext__<name>__<tool>` — call one directly to see what it returns before proposing a feed. A server the workspace has switched off for you is not listed and cannot be fed from here.
ledger_feeds_list
The standing instructions for how metrics' numbers arrive — which tool each one calls, how often, and how the last run went. Check before proposing a feed so you don't duplicate one, and consult when a user asks why a metric is stale: a feed with a failing last run is usually the answer.
ledger_feeds_preview
Call a connected server's tool once and see what a mapping would pull out of the response — nothing is written and no feed is created. Send no `mapping` for a first look: you get a sample of the response plus the paths that hold numbers. Then send a mapping to confirm it finds the readings you expect. Always do this before ledger_feeds_create; proposing a feed whose mapping you haven't seen work is how a metric fills up with the wrong number. May return `needs_confirmation`.
ledger_feeds_create
Freeze a tool call as a metric's standing source: it runs on the cadence you give and records what it finds, with no model involved. `metric` takes an id or a snake_case slug; an unknown slug starts tracking that metric. Confirm the mapping with ledger_feeds_preview first. Tell the user they can backfill history afterwards with ledger_feeds_backfill. May return `needs_confirmation`.
ledger_feeds_update
Change a feed's arguments, mapping, cadence, or pause it. Reach for this when a feed's last run reports `empty` (the response shape moved, so the mapping needs a new path) or `error`. Changing the mapping changes what the number means, so say what you're changing and why. May return `needs_confirmation`.
ledger_feeds_run
Run a feed immediately over its own window and record what it finds. Use it right after creating one to prove the mapping works — a run that comes back `empty` names the path that missed, which is what you fix with ledger_feeds_update. Does not move the feed's schedule. Re-running is safe: readings dedupe per time bucket. May return `needs_confirmation`.
ledger_feeds_backfill
Run the same frozen instruction over a past range, so a metric has history instead of starting the day it was set up. Offer this whenever you create a feed — a chart with a year behind it is worth far more than one that begins today, and an expectation can be judged against what normal looked like. Works when the source returns a series; a source that only ever reports 'right now' will write one point. Safe to repeat: overlapping ranges dedupe. Keep ranges to a few hundred points; a run that fills up says so. May return `needs_confirmation`.
ledger_feeds_delete
Remove a feed. The readings it already wrote stay — the series is the record and outlives the instruction. Prefer pausing with ledger_feeds_update when the user might want it back. May return `needs_confirmation`.
comments_list
Lists the comment threads on one workspace entity (open first, then resolved) with authors and timestamps. Read this before weighing in on contested work — the threads are where disagreement lives before it becomes a decision.
comments_create
Posts a comment on a workspace entity — a new thread, or a reply when rootId is given. Use it to leave findings where the discussion already lives (an eval result on the flow being debated, a summary on a long thread). Mention people via mentionedUserIds (from workspace member ids) to ring their notification bell; never mention someone who didn't ask to be pulled in.
comments_resolve
Sets a comment thread's resolved state (rootId = the thread's root comment id). Resolve ONLY when the human asked or the thread's question is demonstrably settled — and say what settled it in a reply first. Reopening is for new evidence.
napkin_boards_list
Lists sketch boards in the active workspace — name, id, kind (sketch / deck), shape count, last activity. Boards are SPATIAL canvases (drawing, diagrams); the markdown DOCUMENTS are docs — napkin_docs_list. Archived boards are hidden unless `includeArchived`; `q` narrows by name. When the user mentions a sketch, drawing, whiteboard, or napkin, find it here, then LOOK at it with `napkin_boards_view` before discussing its contents. When referring the user to a sketch in your reply, put `[[board:ID]]` alone on its own line — it renders as a clickable card with a live thumbnail.
napkin_boards_get
Returns a board's shapes as data — each with its id, kind, position, size, and text — so you can TARGET one to change or remove with napkin_draw (`updates` / `deletes` take these ids and absolute board coordinates). Pen strokes come back as a point count, not the points. For what the board LOOKS like, use napkin_boards_view instead. Board ids come from napkin_boards_list.
napkin_boards_view
Renders the sketch board to an image and returns it so you can SEE what's drawn — layout, arrows, handwriting-style strokes, sticky notes — not just data about it. Use this before answering any question about a sketch's contents, and cite the board when you do — `[[board:ID]]` alone on its own line embeds the sketch card in your reply.
napkin_boards_update
Edits a board's details: name, description, visibility, or `archived` (true takes it out of the gallery; false brings it back — the recovery move for a board you created by mistake or the user no longer wants). Board ids come from napkin_boards_list. Boards are working material: no approval card, every change attributed.
napkin_deck_outline
Returns a Napkin deck as a markdown outline — one section per slide in presentation order, with slide titles, text content (bold + bullets preserved), empty layout stubs still awaiting content, and a summary of drawn marks. This is the cheap way to know what a deck SAYS; use `napkin_slide_view` when you need to see how a specific slide LOOKS. Works on any board that has slides. Cite the deck with `[[board:ID]]` alone on its own line.
napkin_slide_view
Renders a single slide of a Napkin deck to an image — exactly what present mode shows, cropped to the slide. Navigate by 1-based `slide` number in presentation order (get the map from `napkin_deck_outline` first; the result echoes `slideIndex`/`slideCount` so you can step through a deck slide by slide). Use this to check visual layout, drawings, and images that the text outline can't carry.
napkin_boards_create
Creates a blank Napkin board — a sketch (free canvas) or a deck (slides). NOT for written documents: 'draft/write a doc, note, memo' is napkin_docs_create. Use when the user asks for one, or proactively when a sketch/deck would carry the conversation better than words; hand it over with `[[board:ID]]` alone on its own line. For a deck with content, prefer `napkin_deck_write` (one call, whole deck).
napkin_deck_write
Creates a deck from a markdown outline — the SAME dialect `napkin_deck_outline` reads back, so any deck you've read shows the format. Decks are SLIDES for presenting; a prose document ('write this up', 'draft a doc') is napkin_docs_create instead. Rules: `## Heading` starts a slide (heading becomes the slide title and its biggest text block); optional trailing `[layout: title | section | title-body | title-lead | two-column | three-column | comparison | statement | quote | closing | blank]`; body lines fill the layout's remaining text areas in order, split into areas by a line containing only `---`; `- ` bullets are kept as bullets. Slides default to title-body (or title, when bodyless). Layout text areas you don't fill stay as visible prompts for the user. Example: ## Q3 Review [layout: title] What happened, what's next ## Revenue [layout: title-body] - up 40% QoQ - churn flat ## Bets [layout: two-column] Double down on decks --- Sunset legacy plans After writing, cite the deck with
napkin_slide_add
Appends one slide to an existing deck. Pick a layout (title | section | title-body | title-lead | two-column | three-column | comparison | statement | quote | closing | blank), give the title, and optionally `content` — markdown for the layout's remaining text areas, split by `---` lines (two-column/comparison take one block per column). Check the deck first with `napkin_deck_outline`.
napkin_slide_fill
Fills a slide's empty layout text areas by NAME — the names `napkin_deck_outline` surfaces as `[empty text stub: "…"]`. Pass `texts` mapping those exact names to content (markdown `- ` bullets welcome). Read the outline first to see which stubs a slide still has open.
napkin_draw
Draws shapes, strokes, and labels on a board — diagram on a sketch, annotate a deck slide — and, via `updates` / `deletes`, changes or removes shapes that are already there (ids from napkin_boards_get; the fix for something you drew wrong). For NEW elements never compute absolute board coordinates: with `slide`, coordinates are slide-local — (0, 0) top-left to (960, 540) bottom-right of that slide. Without `slide` (sketches), (0, 0) is the top-left of the existing content — the same area `napkin_boards_view` renders, so place by what you saw there; on an empty board just start at (0, 0). Example — circle a slide's title and margin-note it: {"boardId": "…", "slide": 2, "elements": [ {"type": "ellipse", "x": 60, "y": 40, "w": 400, "h": 90, "color": "#dc2626"}, {"type": "arrow", "x": 560, "y": 140, "x2": 470, "y2": 90, "color": "#dc2626"}, {"type": "text", "x": 575, "y": 130, "w": 260, "text": "tighten this claim", "color": "#dc2626"} ]} After drawing, cite the board with `[[board
napkin_deck_set_theme
Paints every slide of a deck with a theme — slide background, heading and body typefaces, and text colors — in one call, the same as picking it in the deck's theme menu. Font sizes are kept. Prefer "brand" (the workspace's own brand kit) unless the user asks for something else. Built-in themes: plain (Helvetica on white. Gets out of the way.); editorial (Garamond headings on cream. Reads like a document.); stage (Helvetica, light on dark, for a projected room.); blueprint (Consolas headings on mist. For technical decks.); warm (Rounded on amber. Softer than it sounds.); notebook (Marker headings. Keeps the sketchbook feel.). For one-off colors or sizes on a single shape, use napkin_draw's `updates` instead. Check the result with napkin_slide_view.
napkin_slide_update
Changes one slide as a whole: move it to another position in the deck (`moveTo`, 1-based — the other slides shift around it), hide it from the show without deleting it (`skipped`), show or hide its slide number (`pageNumber`), or replace its speaker notes (`notes`, markdown). Slide numbers come from napkin_deck_outline. To change what's ON the slide, use napkin_slide_fill or napkin_draw.
napkin_brand_list
Lists the workspace's brand kits: id, name, description, and which one is the workspace's brand (`isDefault`). Most workspaces have one. Use napkin_brand_get to read a kit.
napkin_brand_get
Reads a brand kit — the workspace's brand unless you name another. READ THIS BEFORE writing or designing anything for the workspace: a deck, a doc, an interface, copy. Returns `brief` (the whole kit as markdown: every section — overview, voice, colors, typography, imagery, motion, and whatever else the team wrote — with the reason behind each rule; follow the reasons when a case isn't covered), `css` (the `--brand-*` variables; interfaces already link it as brand.css, so style with var(--brand-accent) etc. rather than hex values), and `tokens` (the resolved values Napkin uses). Before any kit exists it returns values from the workspace's colors with `kit: null`.
napkin_brand_check
Finds words the brand says to avoid in a piece of writing, with what to say instead and why. Pass `text`, or a deck (`boardId`) or doc (`docId`) to check its words. Run it on your own drafts before handing them back.
napkin_brand_apply
Restyles a piece in the brand — the workspace's brand unless you name a kit. `deck` (a board id): every slide gets the brand's background, typefaces, and text colors; sizes are kept. `interface` (an interface id): rewrites its brand.css, so anything styled with var(--brand-*) updates. `accent` makes one of the kit's colors the piece's accent — use it when the piece is about one product or campaign that has its own color in the brand (read the kit's Colors section to see which). Docs need nothing — they read the brand when exported. Check a deck afterwards with napkin_slide_view.
napkin_brand_draft
Creates a new brand kit, or edits one that is NOT the workspace's brand. Use it when someone asks you to put their brand together (from their site, a deck, or what they tell you) or to try a variation. You can't change the workspace's brand itself — make a draft and tell the user they can choose it in Napkin's Brand section. A kit is a document of sections, each with `title`, `kind` (overview, voice, colors, typography, logo, imagery, motion, layout, examples, custom), markdown `body`, and `rules` ({rule, why}); colors sections hold `swatches` ({name, value: 6-digit hex, note}), typography sections hold `faces` ({name: what it's for, family: a font name or sans / serif / mono / rounded / marker, note}), voice sections hold `use` and `avoid` ({term, instead, why}), and any section can hold `assets` ({fileId: an image already in the workspace's files, name, note, backdrop: hex}) — logo versions, example images — and `links` ({url, title, note}). An examples section holds work that gets t
napkin_brand_fonts
Gets the font files for a kit's typefaces from Google Fonts and keeps them in the kit, so decks, slide pictures and the canvas draw the brand's real typefaces instead of a stand-in. Only typefaces that name a family (like Poppins) and don't have files yet are fetched. Works on any kit, including the workspace's brand: it adds the files for the typefaces the kit already names and changes nothing else. A typeface Google doesn't have is reported back; the user can upload its files in the kit's typography section.
napkin_brand_examples
Shows you the pictures in a kit's Examples sections — screenshots of work that gets the brand right — with each one's note on what makes it good, plus the example links. Look before you design a deck, page or interface in the brand, and hold your work to the same bar: the layouts, density, type sizes and use of color you see there.
napkin_brand_add_files
Adds pictures to a brand kit section — logo versions, product marks, example screenshots, imagery. Each file comes from a public https `url` (fetched by the server) or, for a small file that isn't online, base64 `data` with a `filename`. PNG, JPEG, WebP, GIF and SVG, up to 20 MB each. Every file becomes a workspace file and an asset on the section, appended after the ones already there. Give each a `name` and a `note` saying what it's for; on logo sections set `backdrop` to the hex of the ground it's made for (#ffffff for a dark logo, #000000 for a reversed one), which is how pages pick the right version. Same rule as napkin_brand_draft: you can change any kit except the workspace's brand. Files that fail are reported and the rest still land.
napkin_brand_view
Look at the pictures in a brand kit — logos, product marks, footage stills, layout references, examples — before you use them. Without `section` or `names` it lists every picture by section with its note, so you know what exists. With `section` (a section title, like Imagery) or `names` (pictures' names as the brief lists them) it shows you those pictures, up to 8 at a time. Look before you build a design around a picture: whether a photo has room for text, which part of it carries the color, and whether it suits texture inside type are things you can only tell by seeing it.
napkin_deck_export
Exports a deck's slides as PNG files at each slide's real pixels — a square post at 1080×1080, a story at 1080×1920, a leaderboard at 728×90 — ready to upload to an ad platform or a social post. One slide comes back as a PNG; several as a .zip named by deck, slide and size. Hidden slides are left out. The file is saved to the workspace's files; give the user `downloadUrl` (it opens for anyone signed in to the workspace) and say it's also in their files. Slide numbers and the workspace mark aren't drawn. For a PDF or PowerPoint, the user exports from the deck's File menu.
napkin_layouts_list
Lists the slide layouts napkin_deck_compose can build from — the brand kit's own first, then the built-in ones — with when to use each and the slots it takes (`needs` must be filled; `takes` are optional). Every layout adapts to every size, fits its text, and uses the brand's colors, faces and logos; the logo fills itself in. Pick by what the slide has to say: one line (statement), a number (stat), a quote, a list (points, steps), an event, a carousel (cover, inner, closing), a picture (image-top, image-full, split, corner, product-shot), or a display ad (banner).
napkin_deck_compose
Builds a designed deck: each slide is a layout with its slots filled, or HTML and CSS you write; a real browser lays it out, and it becomes an ordinary Napkin deck people can edit. Use this whenever a deck should look designed — it's how to get layout, big numbers, color blocks, and pictures right. LAYOUTS FIRST - For ads, posts, carousels and title slides, start from a layout (napkin_layouts_list): give `layout` and `slots` instead of `html`. Layouts adapt to every size, fit their text, keep to each size's safe zones, and use the brand's own logos. A slide made from a layout remembers it, so when someone resizes it, it's laid out again instead of scaled. - A set of ads is one slide with `sizes`: ["1:1","4:5","9:16","1.91:1"] — each size gets its own arrangement. - Write HTML only for something no layout does. HOW TO WRITE A SLIDE - `html` is the inside of one slide: a `.slide` box 960×540 px (the deck's own size if it has one). Margins are reset; lay out with flexbox or grid, paddin
napkin_docs_list
Lists the workspace's Napkin docs (quick markdown documents — drafts, notes, working text; the LINEAR surface, distinct from boards which are drawing canvases) newest first, with excerpts. Archived docs are hidden unless `includeArchived`; `q` narrows by title. Fetch one with napkin_docs_get. When referring the user to a doc in your reply, put `[[doc:ID]]` alone on its own line — it renders as a clickable card. NEVER cite a doc as [[board:ID]]; boards are sketches.
napkin_docs_get
Fetch one Napkin doc's full markdown body by id (from napkin_docs_list). Read before editing — napkin_docs_update replaces or appends against the current body.