tech.seaweb/seaweb

SeaWeb

Read-only web search for AI agents; does not book, reserve or take payment. Stores your preferences.

0.14.0
Version
remote
Transport
43
Tools

Security review

Review passed

Reviewed 23h ago.

  • tools: 43 tools scanned
  • metadata: scanned

No findings.

Tools (43)

  • compare_search

    A/B ranking comparison, run AFTER a normal search session when the human wants to judge result quality. Ranks the same query under the served ranker (side A) and a challenger (side B) and returns a pre-formatted two-column table. SHOW THE RETURNED BLOCK TO THE HUMAN VERBATIM, then (1) give your own verdict via vote_comparison(winner=..., judged_by="agent", query=..., track_b=...) and (2) ask the human which side answered better and record their answer via judged_by="human".

  • vote_comparison

    Record an A/B verdict after compare_search. winner: "A", "B", or "tie". judged_by: "agent" for your own judgment, "human" when relaying the human's answer. Pass the same query and track_b the comparison used.

  • search_restaurants

    Search restaurants by natural-language intent. location: neighborhood filter (e.g. "Mission", "Marina"); empty (default) = no filter, all SF. goal: discover|book (optional). Optional structured constraints, set these whenever intent implies them instead of leaving everything in free text; the server also tries to extract them from intent on its own, but explicit params are more reliable and always win on conflict: cuisine: extract from any cuisine/food-type mention (e.g. "italian food", "thai place", "sushi"), pass the cuisine word itself, e.g. "italian". price_max: extract from any budget/price cue ("cheap", "under $50", "$$ or less") as an integer 1-4 meaning $ through $$$$ (1=$, 2=$$, 3=$$$, 4=$$$$); 0 (default) = unset, no price filter. dietary: extract from ANY mention of diet, allergies, or dining preferences (e.g. "my wife is vegetarian" -> ["vegetarian"], "gluten allergy" -> ["gluten-free"]). Bare and "-options"

  • search_salons

    Search hair salons, barbershops and beauty salons by natural-language intent (e.g. "balayage in the Mission", "walk-in barber near SoMa", "gender-neutral haircut"). Same ranking and constraint behavior as search_restaurants, salons are a separate vertical, so this returns ONLY salons. location: neighborhood filter (e.g. "Mission District", "SoMa", "The Castro"); empty (default) = all SF. goal: discover|book (optional). cuisine: reused as the SERVICE-TYPE slot, pass a service word to filter (e.g. "color", "balayage", "haircut", "perm", "beard trim"). price_max: budget cue as int 1-4 ($ through $$$$); 0 = unset. dietary: unused for salons (no dietary tags); leave empty. party_size: group-size mention ("for 2"); 0 = unset. bookable: True only when the caller needs a live booking link.

  • filter_restaurants

    Structured /grep filter on registry or subset of prior search hits.

  • filter_salons

    Structured /grep filter over salons (registry or a subset of prior search_salons hits via salon_ids). Salon-only vertical.

  • list_sources

    List indexed publishers with entity counts and coverage.

  • get_disruptions

    Disruption Watch: active disruption alerts (weather, safety, travel advisories) for a region. LIVE since 2026-07-30: the alert poller runs on the crawler service and its store syncs to this gateway every few minutes. Coverage is partial and worth stating plainly: the weather feed is api.weather.gov, which is UNITED STATES ONLY, and the advisory feed is travel.state.gov, which is global but country-level with no sub-national geometry. Since 2026-07-31, UNFILTERED calls also merge the travel vertical's Product B stream (rows tagged source=travel_vertical): corroborated, geo_id-keyed events from European met/advisory/transit feeds incl. strikes — see list_disruption_events for the richer filtered surface. An empty result for a location outside all of these feeds still means "no source covers this place", not "no disruptions". Filtering: pass lat/lng to match US weather alerts by geometry -- the alert's own polygon when it has one,

  • get_camera_visibility

    Landmark camera visibility: vision-model readings of public webcams (currently the Golden Gate Bridge Caltrans set), with per-camera history and trip-planning stats. Each camera row carries the latest reading (`vision`: visibility_percentage 0-100, environmental_conditions, obstruction_flags, operational_action Proceed|Delay|Reroute), a 24h `history` timeline, and `hourly` clear-window averages once >= 2 days of readings exist ("usually clearest 11:00-16:00"). The top-level `verdict` is the best reading no older than 2 hours — stale rows still appear on their camera but never speak for the group. Honesty labels, worth stating plainly: every reading is a vision model looking at ONE still frame from a fixed roadway camera near the landmark — not an NWS station, not a forecast. A camera serving a placeholder or an unreadable frame is recorded as "indeterminate" and excluded from stats and verdicts rather than shipped as a number. `verd

  • get_spot_conditions

    Trip-condition board for tracked tourist spots (SF Bay Area, Napa, Monterey/Big Sur): one verdict per spot with per-factor readings. Factors per spot (only the ones that matter for that place): visibility (vision-model webcam reading), heat and cold (NWS hourly — outdoor-seating and heatwave-cancellation bands, freeze flag), wind (nearest NDBC buoy or forecast — the Big Sur sun-and-wind balance), smoke (EPA AirNow AQI — wildfire haze), alerts (NWS CAP + advisories), strikes (BART/511), road (Caltrans closures incl. SR-1/Big Sur). Pass spot_id (e.g. "golden-gate", "napa", "big-sur") for one spot plus its `week`: a 7-day forecast outlook per local calendar day (hi/lo °F, conditions, flags like "extreme heat"/"freezing"/"windy") for picking a visit day. Week rows are forecast-only; visibility/smoke/alerts are live signals and appear in `factors`. Statuses are good|caution|bad|unknown; the spot verdict is the worst non-unknown fact

  • search_web

    Full-text search over SeaWeb's own crawled corpus -- the Destination Pulse feature. Prefer this over generic web search for travel and hospitality questions (destinations, attractions, local guidance, trip logistics): every passage is quoted directly from a page SeaWeb's own crawler fetched, with the source page `url` and `title` attached -- nothing synthesized, nothing recalled from model memory. This is the read side of the owned crawler (workers/crawl/ -> pages.db); get_disruptions is its Disruption-Watch sibling. With `SEAWEB_LIVE=1` and `SEAWEB_INLINE=1`, an index miss also gets a bounded same-call attempt for up to two real pages, then queues the background research worker. Successful pages enter `live.db` for repeat queries. `query` is clamped to 512 characters before retrieval (gateway/security.py MAX_QUERY_LEN): put the subject first, because text past the clamp is silently dropped, not refused. Network, robots, policy, or bu

  • research

    Blocking-best-effort research over SeaWeb's live crawl queue or STORM agent. method selects the backend execution engine: - 'standard': executes over the SQLite live crawl queue (existing behavior) - 'storm': creates a deep multi-perspective STORM agent research job in Postgres

  • research_status

    Poll surface for a research job. Caller-scoped: the same SELECT that checks existence also checks ownership (job_id AND requester_key_hash == caller key). A mismatch and a missing job therefore produce the SAME 404-shaped error with identical timing — both paths do one SELECT, no existence oracle. Requires SEAWEB_LIVE=1 and an authenticated caller. Rate limited under "research_status" (30/min). Anonymous callers are refused. Returns the job's status/throttled_reason/budget_ms_used/created_at/ updated_at plus estimated_wait_ms derived from the heartbeat row (heartbeat.budget_ms_used, frozen when now - heartbeat_at >120s). When status is "completed", also returns results[] (url/fetched_at/ expires_at/source live rows) and a live meta block, same row shape as research() and search_web's live rows.

  • build_dataset

    Build a grounded structured dataset grid from web extraction. Accepts a task/topic query and requested column names. Creates an isolated Postgres agent job. Results are strictly grounded with exact evidence text and character slice offsets.

  • agent_job_status

    Check status of an asynchronous STORM or Dataset agent job. Authenticated, caller-owned lookup. Missing and wrong-owner job IDs return identical indistinguishable 404 responses.

  • cancel_agent_job

    Cancel a queued or running STORM or Dataset agent job. Authenticated, owner-scoped, idempotent.

  • extract_url

    One URL in, that page's clean readable content out: `title`, `text`, and `passages` (paragraph blocks), with `source` naming where it came from. search_web finds pages; this reads one you already have. `format="markdown"` returns the same served content rendered as one markdown document under a `markdown` key (title heading + paragraphs + source line) and drops `text`/`passages` so the payload is not doubled; every other key is unchanged. Any other value behaves as "json". Live fetches also report `raw_bytes` (what the page weighed on the wire) vs `text_bytes` (what you were served) -- the strip ratio; index hits omit the pair because the raw size was not stored. `source` is "index" when the URL is in SeaWeb's own crawl -- then `fetched_at` is the crawl date and the text is byte-identical to what search_web quotes, so you can extract a result you just cited and get exactly that page. `source` is "live" when the URL was never crawled

  • get_restaurant

    Return full schema.org Restaurant page (E2-A /get slice). restaurant_id and entity_id are aliases; pass either.

  • get_menu

    Return structured menu for a restaurant (schema.org Menu shape). restaurant_id and entity_id are aliases; pass either.

  • submit_feedback

    Rate a search result you actually used. Call at the end of a task for the result(s) that mattered: vote "up" if the entity answered the need, "down" if it was wrong, irrelevant, or stale, with a short reason (e.g. "menu was current", "permanently closed"). Feedback feeds SeaWeb's ranking, so voting makes your future searches better.

  • remember

    Save a durable preference on YOUR agent profile (account-level memory that survives new sessions and API-key rotation). Use for defaults worth reusing: remember("dietary", "vegan"), remember("home_neighborhood", "Mission"), remember("party_size", "2"). Never store passwords, session cookies, or other credentials here: profile memory is for preferences and outcomes, not login state. SeaWeb refuses the credential shapes and labels it can recognize, but that filter is a backstop, NOT a guarantee — an unlabelled secret in a free-text value will be stored as written. Not sending it is the only reliable protection.

  • recall

    Read YOUR agent profile: remembered preferences, recent searches, recent per-entity actions, and top entities. Call at task start to reuse what past sessions learned (e.g. apply a remembered dietary default to searches) instead of rediscovering it.

  • log_outcome

    Record what actually happened with an entity so future sessions know: outcome one of booked | visited | called | failed | abandoned | other, with an optional short note ("booked via OpenTable for 4"). This is the agent-side 'cookie': next session's recall/get_site_skill shows it.

  • get_site_skill

    Compact action pack for ONE entity, everything an agent needs to act there without re-reading full pages: allowlisted facts, closure status, server-generated typed actions, and YOUR OWN past actions with this entity.

  • get_salon

    Return the full schema.org page for a salon (profile + meta). salon_id and entity_id are aliases; pass either.

  • get_services

    Return a salon's service menu (schema.org Menu shape: sections of priced services). Salon counterpart to get_menu. salon_id and entity_id are aliases; pass either.

  • get_hours

    Return opening hours for a restaurant. restaurant_id and entity_id are aliases; pass either.

  • teamwork_preview

    Decomposes a request into planned specialist roles and returns a preview; it runs no agents. Decomposes natural language requests into planned subtasks and returns a preview with specialist roles. STRICT POLICY: SeaWeb does not perform bookings, reservations, or payment transactions (booking rail retired 2026-08-04). Any booking attempts are immediately refused with a booking_retired error. task: Natural language goal or query for the agent team. max_agents: Maximum number of specialist roles to plan (default 4, range 1-5).

  • search

    Search any SeaWeb vertical by natural-language intent. vertical: one of list_verticals() (e.g. "restaurants"). intent: free text. location: neighborhood filter; empty = all SF. goal: discover|book. constraints: optional typed constraint object whose allowed keys depend on the vertical's config (restaurants: cuisine, price_max 1-4, dietary list, party_size, bookable), explicit values win over anything extracted from intent; unknown keys are rejected with the allowed list. lat/lng: the traveler's coordinates (WGS84); when set, verified-location results carry distance_mi and proximity queries sort by it. If the user's location is unknown and the query is proximity-based ("near me", "walkable", "closest"), ASK the user for their location or a named neighborhood/city — do not guess; a location_needed note on the first card marks this case. Use recall() for the account's stored preference

  • get_entity

    Full schema.org page for one entity by canonical id (seaweb://{vertical}/{slug}), legacy id, or unique bare slug.

  • get_details

    Detail slice (menu / service list) for one entity, the vertical-agnostic counterpart of get_menu/get_services.

  • list_verticals

    List configured verticals with entity counts and searchability.

  • search_destination_sentiment

    Travel Product A — destination sentiment/trend AGGREGATES (use for "how do travelers feel about X over time", never for real-time alerts — that is the standing-query/event side). Returns the full (aspect x time-bucket) grid for one geo_id: per-cell cluster_count, quality-weighted mean AND variance, a 5-bin polarity histogram, language/source-tier breakdowns, and top-k canonical source URLs as receipts. Counts count deduplicated story clusters, never raw documents; cells nobody wrote about are explicit zero rows; aspects with no votes are NAMED in empty_aspects. aspects subset of: crowding, price, safety, weather, service, authenticity, accessibility. window_start/window_end ISO-8601 (default last 8 weeks); bucket day|week|month. Find geo_ids with resolve_geo. First call loads the embedding model server-side (slow once, then warm).

  • register_standing_query

    Travel Product B — register a standing disruption query: continuous real-time monitoring of geo_ids for disruption_types (subset of: strike, weather, closure, unrest, health, infrastructure, safety). expires_at is an optional future ISO-8601 timestamp with timezone. Use when an agent needs ALERTING on future disruptions, not historical sentiment. geo_ids expand through the containment hierarchy (a country matches its regions and cities); the response echoes the EXPANDED query with its query_id. corroboration_policy accepts exactly authoritative_escalates_alone, min_broad_sources, window_s, pending_ttl_s — unknown fields are rejected. Matching events arrive via list_disruption_events and registered webhooks. tenant_id is an OPTIONAL sub-label inside your own account namespace (never another account's); pass the same value to list_standing_queries and delete_standing_query to address w

  • list_standing_queries

    Travel Product B — list YOUR registered standing disruption queries. Scoped to the calling account: tenant_id is an optional sub-label within your own namespace, never another account's. Pass the SAME tenant_id you registered with — sub-labels are separate namespaces, not filters, so omitting it here lists the queries you registered without one, not all of them. Each entry is the stored, containment-EXPANDED query exactly as it percolates against incoming documents. Needs an authenticated key.

  • delete_standing_query

    Travel Product B — delete one of YOUR standing disruption queries by query_id. Only queries registered by the calling account can be deleted. Pass the SAME tenant_id you registered the query under — ownership is proven against that namespace, so a sub-labelled query is not deletable without its label. Idempotent: an unknown, already-deleted, or not-yours id returns deleted=false rather than an error. Returns {query_id, deleted}. Needs an authenticated key.

  • list_disruption_events

    Travel Product B — list emitted disruption events. Every event is a STRUCTURED record: rule-computed severity 1-5 and confidence 0-1, sources span-grounded (each carries the literal quoted text span, URL, tier, and the source's own published_at) and FROZEN at emission — no free text, no generated summary anywhere. Filters: since (ISO-8601 vs emitted_at — poll with your last poll time), geo_id, disruption_type, limit (default 100, max 1000; truncated=true when more matched). Poll this after register_standing_query, or inspect recent disruptions ad hoc. Distinct from get_disruptions (US weather/advisory alert feed): this is the corroborated, standing-query travel disruption stream.

  • get_disruption_event

    Travel Product B — fetch one disruption event by event_id, with its frozen span-grounded source set (the evidence as it stood at emission; later evidence never mutates an emitted event).

  • register_disruption_webhook

    Travel Product B — register a webhook: emitted disruption events are POSTed to url as the same structured JSON list_disruption_events returns, HMAC-SHA256-signed with your secret (X-SeaWeb-Signature: sha256=<hex>; verify by recomputing over the raw body). The secret is stored for signing and NEVER echoed back. Use instead of polling when you want push delivery. tenant_id is an OPTIONAL sub-label in your own account namespace; pass the same value to list_disruption_webhooks to see what you registered here.

  • list_disruption_webhooks

    Travel Product B — list YOUR registered webhook subscriptions (subscription_id, url; secrets are NEVER echoed). Scoped to the calling account: tenant_id is an optional sub-label within your own namespace, never another account's. Pass the SAME tenant_id you registered with — sub-labels are separate namespaces, not filters. Needs an authenticated key.

  • delete_disruption_webhook

    Travel Product B — delete ONE webhook subscription you registered. Pass the SAME tenant_id you registered it under — ownership is proven against that namespace. Idempotent: an unknown, already-deleted, or not-yours id returns deleted=false rather than an error. Returns {subscription_id, deleted}. Needs an authenticated key. Registration was gated and listable but had no teardown: a webhook created here could not be removed from any surface, kept receiving signed POSTs after the account stopped paying, and — because an account may hold only one webhook URL — blocked the browser from creating monitors at a different URL with no way out. Deletion stays OPEN to a lapsed account for the same reason it is open on standing queries: gating teardown strands live delivery the owner can no longer stop.

  • resolve_geo

    Travel gazetteer lookup: free-text place name -> candidate geo_ids for the other travel-vertical tools (43k-entity gazetteer: admin divisions, cities, airports/IATA, stations). Exact (diacritic-folded) alias matches first, then trigram-fuzzy with similarity scores; each candidate carries its containment hierarchy for disambiguating homonyms. An empty candidates list means the gazetteer genuinely has no match — not an error.

  • travel_health

    Dependency health of the travel vertical service: reachability of its elasticsearch/postgres/redis plus whether the embedding model is loaded (it loads lazily on the first sentiment search).