agentcheck
Synthetic checks, nightly regression replay and model-drift alerts for AI agents
- 0.1.1
- Version
- remote
- Transport
- 11
- Tools
Security review
Review passedReviewed 23h ago.
- tools: 11 tools scanned
- metadata: scanned
No findings.
Tools (11)
agentcheck_get_status
Current status of a public monitored target: overall state, uptime over 24h/7d/30d, last check time, last nightly scores, open incidents and the badge/status URLs. Use the owner (GitHub login) and target slug from the status page URL https://agentwares-agentcheck.vercel.app/<owner>/<slug>. No API key needed.
agentcheck_get_pricing
Machine-readable pricing for agentcheck: tiers with monthly USD price, target limits, check interval and features, plus per-run add-ons. Same data as /pricing.json. No API key needed.
agentcheck_model_retirements
Is this model retiring? Pass model IDs as your code or config names them (gpt-5-mini, claude-sonnet-4-5-20250929, gemini-2.5-flash, openai/o3-mini). For each, returns every announced shutdown on OpenAI's, Anthropic's and Google's own deprecation pages (OpenAI API, Claude API, Gemini API, Vertex AI): status, shutdown date, recommended replacement, the replacement's list-price ratio per token, the API changes the provider documents, the source URL with the date it was read, and a page with the details. retiring: false means no shutdown for that ID is on those pages. Language models only. No API key needed.
agentcheck_api_sunsets
When does the API version this code pins stop working, and what happens then? Pass provider and version pairs as your code pins them (shopify 2025-10, meta v21.0, google-ads v22, hubspot v1, twilio-flex 2.16.2, twilio-notify v1, keap xml-rpc, exchange-online ews, stripe 2024-06-20). For each, returns every dated end the provider's own versioning, changelog or retirement page lists: status, end date (and whether the provider gives only a month), what happens after it (error, silent fall-forward to another version, app delisting, or unsupported), whether the failure is silent, the replacement version, the source URL with the date it was read, and a page with the details. Where the provider's own pages give different dates for the same end, other_dates lists each one with its page and date_note says which the counts run to. A pair with no entry comes back with the reason: older than every listed version, no end date published yet, or a provider that publishes none (Stripe). Covers Shopify
agentcheck_statuspage_preview
Read a public Atlassian Statuspage page once: its name, overall status, components (id, name, status, group) and unresolved incidents, from the page's own /api/v2/summary.json. Pass any address on the page: status.example.com, https://status.example.com/history or the JSON URL. Use it before POST /api/v1/statuspage, which turns the components you give a URL into a live agentcheck status page checked by agentcheck itself, with no signup; the answer carries that request's body. Same read as /statuspage/preview: only public addresses, 5-second timeout, 512 KB cap. Private pages cannot be read. No API key needed.
agentcheck_create_target
Enroll something to monitor: an http endpoint (JSON or OpenAI-style chat), a remote MCP server (Streamable HTTP url) or an A2A agent (origin with /.well-known/agent-card.json). Pass `checks` to create checks in the same call (POST /api/v1/probe proposes three). Returns the target id, the public status page, the badge SVG URL and a README snippet. The first check runs on the next scheduled tick, within five minutes; call agentcheck_run_now to run immediately. Free tier: 1 target, hourly; Starter+: 5-minute checks. Requires an API key.
agentcheck_add_check
Add a check to one of your targets. A check runs an input (http path/prompt, mcp tool call, a2a message) on a schedule and judges the answer with a golden: exact, contains, regex and json_schema cost nothing; rubric and baseline use the LLM judge (baseline = same outcome as the last known-good answer). Returns the check id. Requires an API key.
agentcheck_run_now
Run every check of one of your targets immediately (outside the schedule) and return pass/fail per check with latency, judge cost and any incident opened or closed. Use it right after enrolling, or to confirm a fix. Requires an API key.
agentcheck_record
Record one production interaction with your agent (the prompt or messages, the final answer, the tools it called, optionally the model) as a trace on a target. Call it from your agent or from a proxy in front of it after each task; promote a good trace with agentcheck_promote_trace to replay it nightly and catch regressions. Requires an API key.
agentcheck_promote_trace
Turn an imported or recorded trace into a replayable check whose golden is the recorded outcome (tool sequence + final-answer rubric). The check runs daily and in the nightly replay (Pro). Returns the check. Requires an API key.
agentcheck_list_incidents
Open and recent incidents across your targets (or one target): when they opened/closed, the failing check and the cause. Requires an API key.