Stipple — Document Verification & Extraction
Document forensics: tamper/AI checks, fields, tables, identity, screening, tenders, citations.
- 0.3.1
- Version
- remote
- Transport
- 16
- Tools
Security review
Review passedReviewed Jan 1, 2000.
- tools: 16 tools scanned
- metadata: scanned
No findings.
Tools (16)
verify_document
Forensically inspect a document (PDF or image) for authenticity: tampering signs, AI-generation indicators, arithmetic reconciliation (financial docs), and provenance. USE THIS WHEN someone shares a payslip, bank statement, invoice, receipt, ID, certificate, or contract and asks: is this genuine / real / authentic? has it been edited, doctored, or photoshopped? can I trust this file? (For "did an AI *write* this prose" use `detect_ai_text`; for "are this report's citations real" use `verify_references`. Both are available in this canonical suite.) Provide the document ONE way: `url` (a public http(s) link — fetched server-side, the cheapest call: no need to download or encode anything) OR `bytes_b64` (inline base64, plus `filename` so PDF-vs-image routing is right). Returns the headline result — `risk_band` (low/medium/high/insufficient/error), `inspection_quality` (coverage, orthogonal to risk), `recommended_action`, a `summary`, the
check_document
Cheap cache-check: has this exact document already been inspected? Hash the file yourself (sha256, lowercase hex) and call this before verify_document to skip a redundant (paid) inspection. Returns {cached, warrant_id, permalink}.
get_warrant
Retrieve a stored warrant by id (e.g. 'warrant_<hex>') — the full bundle as JSON, or a human-readable Markdown report when as_markdown=True. USE THIS WHEN you have a warrant_id from an earlier verify_document / check_document call and need the FULL evidence — every signal that fired, per-page findings, provenance — rather than the summary the original call returned. Use as_markdown=True to get a report you can show a human verbatim.
submit_feedback
Record thumbs up/down on a warrant's rating (the engine's precision-flywheel label source). verdict must be 'up' or 'down'; note is optional free text. USE THIS WHEN the ground truth became known after a verify_document call — e.g. the document was later confirmed genuine or fraudulent — so the engine learns from the outcome. Tell it what happened; it sharpens future inspections for everyone.
extract_fields
Extract structured FIELDS from a document (PDF or image) with a vision model. USE THIS WHEN you need specific values OUT of a document — a payslip's gross/net, an invoice's total/ABN, a form's checkboxes, a table's cells — rather than a yes/no about the document. (For "is this genuine?" use verify_document; "what kind of document is this?" is `options={"classify": true}` right here.) Say WHAT to pull, four ways: - `fields`: an ad-hoc list — names like ["gross_pay","abn"], or objects {"name":..., "type":"text|amount|date|boolean", "description":...}. THE general case: ask for exactly the fields your task needs. Use type "boolean" for a checkbox/tickbox. `"question"` works instead of `"description"` if you would rather just ask: {"name":"customer_name", "question":"What is the customer name?"}. - `template`: a named preset — "payslip", "tax_invoice", "bank_statement", "receipt". - NEITHER: AUTO — the document is clas
verify_identity
Run an Australian identity check over a SET of identity documents. A vision model reads each document (which ID it is, which fields it shows — name/photo/address/signature — and its issue date); a deterministic engine then tallies them against a scheme and reports whether identity is established, and exactly what's still missing if not. USE THIS WHEN someone needs to verify a person's identity from their documents — KYC / onboarding / "do these documents satisfy the 100-point check?" Pass ALL the person's documents together (a passport alone is 70 points; the check needs >= 100). `documents` is a list, each item ONE of: {"url": "https://..."} (public link, fetched server-side) or {"bytes_b64": "...", "filename": "passport.pdf"} (inline). Up to 10. `scheme`: "afp_100_point" (points, default) or "austrac_safe_harbour" (category combinations). Returns `{established, points/target or satisfied_path, documents[] (per-document: type, fields show
check_pack
Check whether a SET of documents satisfies a checklist — completeness, cheaply. USE THIS WHEN you have an application / onboarding pack and need "do we have the required documents, and what's still missing?" Each document is CLASSIFIED (one cheap page-1 read — never full field extraction or multi-page), then matched against the checklist's required slots. (For "is a document genuine?" use verify_document; to identify ONE document use extract_fields with options={"classify": true}; for the identity gate use verify_identity.) Define the checklist ONE of two ways: - `scheme`: a named preset — "income_proof", "lending_prequal", "rental_application". - `requirements`: an ad-hoc checklist — a list of document-type names like ["payslip","bank_statement"], or objects {"key":..., "accepts":[types], "optional":bool}. `documents` is a list (up to 12), each ONE of: {"url": "https://..."} (public link, fetched server-side) or {"bytes_b64": "...
screen_adverse_media
Screen a person or organisation for ADVERSE MEDIA and SANCTIONS exposure (KYC/AML). PEP lists are not screened: `sanctions.flags.pep` is always false and `sanctions.note` says so. USE THIS WHEN onboarding or due-diligence asks: does this subject appear in negative news (fraud, money laundering, bribery, sanctions, trafficking, enforcement action), or on a sanctions list? Pairs naturally after verify_identity. Identify the subject ONE of two ways: pass `name` (plus any of `dob` as YYYY-MM-DD, `country`, `aliases`, `employer`, `role` — these sharpen matching and cut same-name false positives), OR pass an identity document via `url`/`bytes_b64` (+`filename`) and the subject is read from it. Returns `{subject, sanctions{...}, adverse_media{...}, risk_flag, headline, limitations}`: sanctions candidates are corroboration-gated (a name-only hit is `possible`, NEVER confirmed — one common name matches several different people); media hits are entity-d
detect_ai_text
Estimate the PROBABILITY that a document's text was AI-GENERATED (LLM-written prose). USE THIS WHEN someone shares prose — an essay, cover letter, article, review, application, or report (or a link to one) — and asks: did an AI / ChatGPT write this? is this human-written? detect AI text. Provide the document ONE way: `text` (pasted markdown/plain prose), `url` (a public http(s) link to a page or PDF — fetched server-side, the cheapest call), OR `bytes_b64` (a base64 PDF/file, plus `filename` for routing). Returns `{probability, lean, tells, reasoning, applicable}`. HONEST SCOPE: the probability is the model's CONFIDENCE, not a calibrated truth — it can false-flag templated/coached or non-native-English writing. It works on PROSE only: for a form/table/numeric document (payslip, statement) it returns `applicable: false` and abstains, because AI-text detection false-positives badly there — use `verify_document` (the authenticity engine)
verify_references
Fact-check a document's REFERENCES and CLAIMS — built for AI-generated reports whose citations must be checked before they're trusted. USE THIS WHEN someone shares a report, article, whitepaper, or deep-research export (or a link to one) and asks: is this accurate / legit? are these citations real? fact-check this. did the AI make this up? Also use it proactively before relying on any AI-written document. Provide the document ONE way: `url` (a public http(s) link to a PDF or web page — fetched server-side, the cheapest call: no need to download or encode anything), `text` (pasted markdown/plain prose), OR `bytes_b64` (a base64 PDF; URLs are read from the PDF's link annotations, so they're exact). Default (fast): provenance (is it a ChatGPT deep-research export?), citation resolution (live / archived / dead, papers matched against arXiv/Crossref to catch 'real ID, wrong paper'), and internal MATH (recompute the doc's own arithmetic). Set `de
check_source_overlap
Check whether text OVERLAPS text published on the public web — a plagiarism-style check: does this text appear elsewhere? was this copied? find the source of this text. Provide the document ONE way: `text` (pasted prose), `url` (a public http(s) link — fetched server-side; that page and its host are excluded from matches), OR `bytes_b64` (a base64 PDF/.docx/text file, plus `filename` for routing). Returns two evidence tiers, never mixed: `matches` are EXACT/near-verbatim overlaps confirmed against the fetched source page — each carries the quoted text from both sides, the source URL, and char spans for highlighting. `possible_paraphrases` are model JUDGEMENTS (reworded overlap), clearly labelled, never quotes, and alone they cap the overlap band at "low". `overlap_band` summarises: none | low | notable | high. HONEST SCOPE: this searches the PUBLIC WEB within capped queries — it is not an academic-database check, absence of matches is neve
find_tenders
Search open tenders across Australia; New Zealand rows only when you ask for them (jurisdiction=NZ). FREE, within the weekly cap. USE THIS WHEN someone asks what public-sector work is open: "any council drainage tenders in Victoria", "what's closing this month in NSW", "show me federal IT opportunities". For "which of these could MY company actually bid for", use match_tenders instead — that reads their website and ranks against it. `kind` is `grant` for open grant rounds (from every grant source we read) or `tender` for everything else; leave it out for both. Each result says which it is in `kind`. `jurisdiction` is one of AU, NZ, AU-NSW, AU-VIC, AU-QLD, AU-WA, AU-SA, AU-TAS, AU-ACT, AU-NT. `tier` is federal, national, state, council, university or health. `closing_before` is an ISO date. `first_seen_after` (ISO-8601 instant, strictly newer) answers "what is new since my last look" — first_seen is when WE first saw the tender, the honest clock for newness. There is deliberately no `
match_tenders
Rank open tenders against what a company actually does. Free, inside the weekly cap. USE THIS WHEN someone asks which opportunities suit a specific business: "what could we bid for", "is there anything for a civil contractor in Victoria", "find work for acme.com.au". Give `company_url` — a plain domain is fine, we resolve it — and we read their site, build a capability profile, and score the shortlist against it. `example` runs a built-in profile (civil, it, facilities) with no site read, for demonstrating the shape of the answer. Returns `{profile, matched, shown, withheld, withheld_reason, matches[], degraded, score_means, coverage}`. Each match has `score`, `band`, `why[]` — the company's own stated capabilities this tender needs — `gaps[]`, things the tender asks for that their website does not mention, and `checks[]` (each a `label`), computed: outside the states asked for, or prequalification named in the tender's text (a heuristic,
tender_sources
Every source we search, what it is allowed to do, and what the last run returned. FREE. USE THIS WHEN someone asks where the data comes from, whether a particular portal is covered, or why a search came back empty. It is the honesty surface: it names sources behind login walls, sources whose robots.txt refuses us, and sources that returned nothing on the last run and why. Returns `{sources[], coverage}` — per source: id, tag, name, refresh mode, jurisdiction, tier, how it is accessed, what its robots.txt says, how many tenders we hold from it, and its status on the most recent run; its URL (`source_url`) with an API key or a signed-in account only. Snapshot sources include their observed date and are not presented as nightly feeds.
buyer_awards
What a buyer has awarded, what is ending, and what they plan. FREE. USE THIS WHEN someone asks about a specific buyer before a bid: "who holds Transport for NSW's work", "what is ending soon at Queensland Health", "what does this agency usually pay". Give `buyer` (the organisation name as published) or `buyer_key` (from a tender's buyer, or a previous answer). Returns `{buyer, expiring[], planned[], recent_awards[], top_suppliers[], open_tenders[], computed_at, sources}`: the nightly rollup (awards in the window, value quartiles as published, median response window), contracts ending within 12 months with the incumbent, planned procurements with their quarter and spend band, the suppliers who win from them (name and share), and open tenders under the same name. ANONYMOUS CALLERS SEE COUNTS, VALUES, DATES AND BUYERS; supplier and incumbent names are withheld and `withheld_reason` says so. Relay that sentence as it is. Links to other sites - an award's notice (`url`), a plan's evidenc
find_signals
Signals: what may be tendered before it is. FREE. USE THIS WHEN someone asks what is coming: "which contracts in Queensland end in the next six months", "what is planned for ICT next quarter", "what is expiring for this buyer". `kind` is one of contract_expiry (a contract ending, with its incumbent), planned_procurement (a buyer's stated plan with its quarter and spend band as published) or recurring_tender (derived from our own history, labelled `derived`). `jurisdiction` is one of AU, NZ, AU-NSW, AU-VIC, AU-QLD, AU-WA, AU-SA, AU-TAS, AU-ACT, AU-NT. `window_before` is an ISO date: signals whose window starts on or before it. `q` searches the subject, buyer and incumbent. Returns `{total, results[], computed_at, sources}`. Each signal carries `confidence` (`published` or `derived` - a vocabulary, not a score), its window (never invented: an expiry's window IS the contract's end date; a planned row with no parseable quarter has none), `evidence_ref`, and `evidence_url` with an API key