io.github.77777R7/octocrawl

octocrawl

Web scraper for agents: blocked, empty and wrong pages reported as such, with an Evidence Record.

0.4.1
Version
remote + npm
Transport
2
Tools

Security review

Partly reviewed

Reviewed 1h ago. Tool definitions changed on Oct 11, 2026.

  • tools: 2 tools scanned
  • metadata: scanned
  • packages: 1 checked
  • mediumReviewRemote tools take credentials as input

    Whatever an agent passes to a remote tool leaves the machine. Never send connection strings, tokens or passwords to a third-party MCP server unless it is the service those credentials belong to.

    scrape
  • mediumReviewTool definitions changed after an earlier review

    A server that changes its tool descriptions after being approved ("rug pull") can slip new instructions to agents. Re-check what changed before trusting it.

    changed 2026-10-11

Tools (2)

  • scrape

    Fetch one URL through the Octocrawl coverage ladder. Compact by default; set debug=true for the full audit. The result's warnings name what its content cannot vouch for: robots_overridden (robots.txt disallows the URL; a local server fetched it because you named it), or client_rendered_suspected when the HTTP page looks like a shell its scripts fill in and the browser rung found nothing better. Its agentHints, when present, say what to change next time (a login wall, a robots.txt rule, a gate, a wait). metadata.scrapeId names the call's record for get_scrape.

  • map

    List a site's URLs without fetching each page: the start URL, the links on its page (read on the http lane alone; no browser) and the entries of the sitemaps the site declares (robots.txt Sitemap: lines, else /sitemap.xml), inside one deadline. Every URL is in the crawl's scope (the start host and its www twin, the start URL's path subtree, assets left out, similar URLs folded) and allowed by its host's robots.txt unless ignoreRobotsTxt is set; what was left out is counted. A title is never fetched: the start page's own, an anchor's text or a sitemap's news title. At the deadline the answer is what was found, status partial (failed when nothing), stoppedBy timeout. Compact by default ({ id, status, stoppedBy, links: [{ url, title?, description?, robots? }], warning?, agentHints?, counts }; robots only on a link robots.txt keeps out, under ignoreRobotsTxt); debug=true returns the full map with each link's evidence (via, sitemapFile, lastmod, robots), the sources read and the refusals. O