ai.zerowidth/caliper

ZeroWidth Caliper

Build datasets and rubrics, then run evals that score your AI in Caliper.

1.0.0
Version
remote
Transport
45
Tools

Security review

Review passed

Reviewed 1d ago.

  • tools: 45 tools scanned
  • metadata: scanned

No findings.

Tools (45)

  • search_docs

    Search ZeroWidth product documentation. Returns matching pages with title, slug, public URL, and a query-relevant snippet. Use this when the user asks about a ZeroWidth product (Compass, Workbench, Caliper, Prism, Ledger, Napkin, zv1), an API behavior, or a policy. No authentication required — the docs corpus is public.

  • get_doc

    Fetch the full Markdown body of a specific docs page by its slug. Use this after `search_docs` when the user needs the complete content of a page. No authentication required.

  • list_docs

    Enumerate all available docs pages, optionally filtered by product (e.g. 'compass', 'legal', 'overview'). Use this to discover what slugs exist before calling `get_doc`. No authentication required.

  • caliper_datasets_get

    One dataset with a page of its items — read this BEFORE editing items (caliper_datasets_update_item needs the item id) and before extending coverage, so you don't add cases that already exist. Items come `limit` at a time (default 25) from `offset`; `total` is the full count. Long fields are cut at ~800 characters with a truncation marker. Get the datasetId from caliper_datasets_list.

  • caliper_datasets_update

    Changes a dataset's title, description, or visibility. Pass only what changes; items are untouched (use caliper_datasets_add_items / caliper_datasets_update_item / caliper_datasets_remove_item for those). May return `needs_confirmation`.

  • caliper_datasets_delete

    Deletes a dataset. Evals bound to it stop being runnable (their binding shows datasetOk: false), and reviews/specs over its items lose their source — check caliper_evals_list for evals that reference it and say so before proposing. May return `needs_confirmation`.

  • caliper_datasets_update_item

    Rewrites one item's content, keeping its id (ratings and eval results keyed to it stay attached). The item's KIND can't change — a Q&A item stays Q&A, a sequence stays a sequence — so pass the same shape you read from caliper_datasets_get. Chat and raw items can't be edited from here. May return `needs_confirmation`.

  • caliper_datasets_remove_item

    Drops one item. Ratings and run results keyed to it are orphaned (kept in history, gone from the dataset). A dataset must keep at least one item. May return `needs_confirmation`.

  • caliper_datasets_generate

    Creates a NEW dataset of model-written items — Q&A by default, or scripted sequences / simulated people via `shape` — pass flowId and the generator reads the flow's prompt, mode, and schema to write realistic cases for THAT flow; `description` adds guidance (or stands alone when there's no flow). Use this when the user wants test cases fast and has none; prefer caliper_datasets_create with hand-written items when real scenarios are already in hand (Compass pages, a transcript). This SPENDS workspace inference credit (one generator call), so it sits behind the approval gate: say so and expect `needs_confirmation`. Returns the dataset summary; read the items with caliper_datasets_get and tell the user to review them before trusting an eval built on them.

  • caliper_datasets_generate_items

    Appends model-written items to an EXISTING dataset, in the style of what's already there (existing items are the few-shot examples; the output shape matches theirs unless overridden). Use to widen coverage when the user says 'more like these' or 'add edge cases'; write them by hand with caliper_datasets_add_items when the scenarios are known. SPENDS workspace inference credit, so it sits behind the approval gate — say so. Existing items and their ratings are untouched.

  • caliper_rubrics_list

    Lists the rubrics in the active workspace (id, name, description, criterion count). Check here BEFORE caliper_rubrics_create — reuse an existing rubric's id in caliper_evals_create when one already scores the same job. Read a rubric's criteria with caliper_rubrics_get.

  • caliper_rubrics_get

    One rubric with its criteria — each criterion's name, what the judge looks for, and the score scale. Read this to explain a score (which criterion slipped and what it asks for) or before caliper_rubrics_update. Get the rubricId from caliper_rubrics_list or an eval's rubricId.

  • caliper_rubrics_update

    Changes a rubric's title, description, visibility, or replaces its criteria wholesale (pass the FULL list — criteria get fresh ids). Evals snapshot the rubric when they're created, so an existing eval keeps scoring with the criteria it started with; say so, and offer to create a new eval when the criteria change materially. May return `needs_confirmation`.

  • caliper_rubrics_delete

    Deletes a rubric. Existing evals keep their snapshot of it and keep running; nothing new can bind to it. May return `needs_confirmation`.

  • caliper_evals_get

    One eval: what it targets (Workbench flow + stage, or external), its dataset and rubric ids, the rubric criteria it scores with (the snapshot taken at creation), schedule + runOnPublish + regression settings, run counts, latest score, and `binding` health (datasetOk / flowOk false = a run would fail at resolution — say so before proposing caliper_evals_run). Get the evalId from caliper_evals_list or caliper_flow_performance.

  • caliper_evals_runs_item_execution

    What the flow actually DID on one run item: every step in order (what it said, which tools it called with what arguments, what came back) and the final answer, plus status, duration, and cost. Read this when a low score needs explaining beyond the judge's reasoning — a wrong tool call or an empty tool result is usually the cause, and the fix is different from a prompt fix. Pass the run item's `id` (NOT `itemId`) from caliper_evals_runs_get. Items whose outputs were supplied from outside (external evals) have no trace and return not_found. Steps are capped for transport; the run page in Caliper has the full record.

  • caliper_evals_update

    Changes an eval's title, description, visibility, schedule (interval + time anchor), runOnPublish (queue a run whenever the flow publishes), or regressionThreshold (score drop vs the previous run that triggers an alert; null = off). Pass only what changes. The dataset, rubric snapshot, and target are fixed at creation — create a new eval to change those. Scheduled and on-publish runs spend credit on their own, so state that plainly when turning them on. May return `needs_confirmation`.

  • caliper_evals_delete

    Deletes an eval and stops its schedule. Run history is kept but no longer reachable from the eval. May return `needs_confirmation`.

  • caliper_evals_runs_cancel

    Stops a run that is still PENDING / RUNNING / SCORING — the brake on a run that's spending more than expected or was started by mistake. The queue stops at once; an item already handed to the flow finishes on its own timeout. The run settles as FAILED with a 'cancelled' reason and keeps the items it completed. A run that already finished returns run_not_live. May return `needs_confirmation`.

  • caliper_sources_list

    Apps sending their agent's traces to Caliper — from their own code, OpenTelemetry, or a published Workbench flow — busiest first: name, how it sends, traces in the last 14 days (per day), and the tools its agent calls most with the share of traces using each. Start here when the user asks how their agent behaves on real traffic; then caliper_sources_over_time for one source.

  • caliper_sources_over_time

    One source day by day: traces, failures, tool calls, lookups that found nothing, tokens, cost, and speed (meanMs; p50/p95 as 'answered within' bucket edges); each tool with its calls and the share of each day's traces that used it; the documents lookups landed on most; models used. Use it to answer 'how often is it calling web search and is that changing' — compare the first and last weeks and name the day it moved. Follow with caliper_traces_list on that day or tool to show the conversations behind it.

  • caliper_traces_list

    The traces behind a point on a source's chart, newest first, 50 a page: the question and answer (cut short), tools called, time, tokens, cost, status. Narrow by day, tool, document, or failures; page with `before` (the nextBefore from the last page). Open one with caliper_traces_get.

  • caliper_traces_get

    One trace: what came in, what went out, and every step in order — model calls (model, tokens, cost, reasoning), tool calls (arguments and result), lookups (query and documents with scores), anything else — each with timing and status. Long values are cut at 2,000 characters. Use it to explain why one conversation went the way it did.

  • caliper_source_feeds_list

    Datasets this source keeps adding to as conversations arrive: which dataset, for what (review, eval, spec), the pick it matches, sampling, how many have landed, and whether it's still running or why it stopped.

  • caliper_sources_to_dataset

    Saves the conversations behind a pick (a day, a tool, a document, failures — the same narrowing as caliper_traces_list) into a new dataset (newDatasetName) or an existing one (datasetId), newest first up to `limit`. `purpose` decides what each item keeps: review = the agent's reply, tool calls included, for people to rate; eval = the reply becomes the expected answer; spec = only the questions, for people to answer. Set keepAdding to make it live: new matching conversations keep arriving (every one, or 1 in 10 / 1 in 100), up to 1,000. Then offer the next step — caliper_evals_create, or a review in Caliper. May return `needs_confirmation`.

  • caliper_source_feeds_stop

    Stops a source from adding new conversations to a dataset. What it already added stays. Get the feedId from caliper_source_feeds_list. May return `needs_confirmation`.

  • caliper_starters_list

    The starter sets Caliper ships: prompt injection, system-prompt leaking, over-refusal, personal-data handling, bias under ambiguity. Each is a small original dataset (single messages AND multi-turn build-ups: false memory, fabricated earlier turns, slow escalation) paired with an anchored rubric. Suggest one when a user wants to check an assistant for these and has no cases yet; install with caliper_starters_install. Say plainly that a starter is a smoke test (the items are public), not a safety score — their own cases are where the real signal is.

  • caliper_starters_install

    Copies one starter into the workspace as a new dataset and a new rubric the user owns and can edit. Returns both ids; the next step is caliper_evals_create binding them to the flow (or target 'external'). Needs caliper:datasets:write and caliper:evals:write.

  • caliper_datasets_list

    Lists datasets in the active workspace. Returns summaries (id, name, item count); items are not included.

  • caliper_flow_performance

    How a Workbench flow is ACTUALLY doing, with receipts: every Caliper eval targeting the flow, recent runs with scores, the latest run decomposed into per-criterion averages, the score delta vs the previous run, and the worst-scoring items WITH the judge's reasoning. Use this BEFORE claiming a flow works or proposing changes — and cite the runId + scores when you do. The worst items are diagnostic: failures clustered around missing company facts suggest a knowledge gap (consider proposing a Compass interview with the workflow owner) rather than a prompt problem.

  • caliper_evals_runs_list

    Recent runs for one eval, newest first: status (PENDING/RUNNING/SCORING/DONE/FAILED), overall score once DONE, label, and timestamps. THE CHECK-BACK for caliper_evals_run: when the user asks how the run went, read this — cite the run id and score, and compare against the PREVIOUS run's score for the delta. Still don't poll in a loop; check when the user asks or when reporting.

  • caliper_evals_runs_get

    One run: status, overall score, and by default the ten lowest-scoring items with their per-criterion scores and the judge's reasoning (`items: all` for every item, `none` for totals only). THE ANSWER to 'why did the score drop' — read this, then name the criterion that slipped and quote the reasoning on the lowest items. Works for runs Caliper ran and for runs submitted from CI (triggeredBy 'ci', usually labeled with a pull request number). Still no polling loops; read it when the user asks.

  • caliper_evals_list

    Lists evals in the active workspace with their target config (which Workbench flow, which dataset/rubric). Use to find the evalId for caliper_evals_run.

  • caliper_evals_run

    Queues a new run of an eval — inference over the dataset, then LLM-judge scoring. THE VERIFY STEP of the improvement loop: after an approved workbench_flows_edit_text, run the eval again and report the score delta vs the previous run. Runs take a while — but you're brought back into THIS conversation automatically with the scores the moment it finishes, so tell the user it's queued and that you'll follow up here; never poll or ask them to check back. Costs workspace LLM budget, so it sits behind the approval gate: may return `needs_confirmation`.

  • caliper_datasets_create

    Creates a dataset of test items — THE FIRST STEP of setting up evaluation for a flow. Three item shapes: Q&A (input + optional expectedOutput, the golden answer); SEQUENCE (turns: 2-20 scripted user messages the model answers one at a time with its own earlier replies in front of it, + expectedResponse for the final reply, optional expectedBehavior for the whole conversation); SIMULATED (goal + optional persona/disposition/strategy/maxTurns + expectedBehavior; a platform flow plays a person adaptively, Caliper-run evals only; disposition is a preset id like genuine, pressure, confused, impatient, or free text). Use sequences and simulated items for the slow attacks and for real customers with real needs: a model that holds on message one often folds on message ten. Write good inputs from real usage: the Compass pages the flow was built from are the best source of realistic scenarios.

  • caliper_datasets_add_items

    Appends items to an existing dataset — use this to grow coverage (new edge cases, scenarios from a completed interview) instead of creating a parallel dataset. Same three shapes as caliper_datasets_create (Q&A, sequence, simulated). Existing items and their ratings are untouched.

  • caliper_rubrics_create

    Creates the scoring rubric an eval's LLM judge uses — 1-10 criteria, each scored on a numeric scale (default 1-5), judged pass/fail (kind pass_fail), or checked in code with no judge (kind check: contains, not_contains, matches a regex, valid_json, max_chars, equals_expected). Write criteria about the FLOW'S JOB (accuracy to source material, tone, refusal behavior), not generic 'quality'.

  • caliper_evals_create

    Binds a dataset + rubric to something under test as a repeatable eval — THE LAST SETUP STEP before scoring. Two targets: a Workbench flow (pass flowId; Caliper runs inference itself, then caliper_evals_run scores it), or 'external' (target: 'external'; the user's own model runs elsewhere and their script submits outputs through the public API, usually from CI — see the docs guide 'Run evals in CI'). For an external eval, hand back the eval id and tell the user to create a workspace API key with the ci_evals preset in their workspace settings; keys can't be minted from here. flowStage 'draft' evals the live draft (pre-publish); 'published' (default) evals the latest published revision at run time, or one pinned with flowRevisionId. Reuse an existing rubric from caliper_rubrics_list when one already scores this job.

  • comments_list

    Lists the comment threads on one workspace entity (open first, then resolved) with authors and timestamps. Read this before weighing in on contested work — the threads are where disagreement lives before it becomes a decision.

  • comments_create

    Posts a comment on a workspace entity — a new thread, or a reply when rootId is given. Use it to leave findings where the discussion already lives (an eval result on the flow being debated, a summary on a long thread). Mention people via mentionedUserIds (from workspace member ids) to ring their notification bell; never mention someone who didn't ask to be pulled in.

  • comments_resolve

    Sets a comment thread's resolved state (rootId = the thread's root comment id). Resolve ONLY when the human asked or the thread's question is demonstrably settled — and say what settled it in a reply first. Reopening is for new evidence.

  • search_workspace

    Finds entities across every tool by name in one call — Workbench flows, Compass pages, Caliper datasets, evals, rubrics, reviews, specs and sources (apps sending agent traces), Ledger entries, Napkin sketches and decks. Use it FIRST when the user names something without saying where it lives ('the onboarding flow', 'that invoice page'); reach for a tool's own list only when you already know the tool. Each hit carries its id, kind, and workspace-relative path, so the id feeds the matching *_get tool and the path makes a link. Results only include what the user can see, and only kinds this token may read.

  • entity_tags_get

    Returns the tags on a batch of entities of one kind — the labels galleries organize by. Ids come from the kind's list/get tool or from search_workspace. Use it before entity_tags_set so you replace the full set knowingly, and to answer 'what is this filed under'. Entities the user can't see are omitted.

  • entity_tags_browse

    Without a tag: every tag in use across the workspace with how many entities carry it, most-used first — the vocabulary the team already organizes by. With a tag: everything filed under it across every tool, each with its kind, id, title, and path. Use it to reuse existing labels instead of inventing near-duplicates, and to answer 'show me everything about X' when X is a label.

  • entity_tags_set

    Replaces the FULL tag set on one entity (an empty list clears it). Read the current tags with entity_tags_get first and pass the merged list — this is not additive. Tags are lowercase letters, numbers, spaces, and hyphens; prefer labels already in use (entity_tags_browse) so the workspace's vocabulary stays small. The id comes from the kind's list/get tool or search_workspace.