skills/ veniceai/skills

venice-augment

Venice augmentation endpoints for agent pipelines. Covers POST /augment/text-parser (extract text from PDF/EPUB/DOCX/PPTX/XLSX/XLS, plain text and source code; multipart, up to 25MB; JSON or plain-text response), POST /augment/scrape (fetch a URL and return markdown; blocks X/Reddit and private/inte

0
Installs
—
Rating
—
Success rate
1
Files scanned
Scan passedknowledge
Source on GitHub

Security scan

Scan passed

No risky patterns were found in the scanned files.

1 files scannedscanner v1.2.0Oct 11, 2026

Content sha256 d46104da357778da… — run codexguild_scan_skills after installing to verify your local copy.

Static analysis is a first line of defense, not a guarantee. Read the source

SKILL.md

exact scanned copy

Venice Augment (text parse / scrape / search)

Three lightweight helpers for agent pipelines that need document text, web pages, or search results without running your own parser, crawler, or search account. All three are marked experimental in the docs — request/response shapes may change.

EndpointInputOutputPrice
POST /augment/text-parsermultipart/form-data file (≤ 25 MB){ text, tokens } JSON, or text/plain$0.01 / successful call
POST /augment/scrape{ url }{ url, content, format: "markdown" }$0.01 / successful call
POST /augment/search{ query, limit?, search_provider? }{ query, results: [{ title, url, content, date }] }$0.01 / successful call

All three accept a Bearer API key or an x402 wallet (SIGN-IN-WITH-X header) — see venice-auth. The balance is checked before the work runs; the charge is applied only after a 200 response. (The OpenAPI x-payment-info advertises a generic dynamic range of $0.001–$10.00; the actual price sheet charge is $0.01.)

POST /augment/text-parser — extract text from documents

Request

Always multipart/form-data:

FieldNotes
fileRequired. Max 25 MB (larger → 413).
response_formatjson (default) or text.

Accepted file types (matched by MIME type, or by filename extension when the MIME type isn't recognized, e.g. application/octet-stream):

  • Structured documents: PDF, EPUB, DOCX, PPTX, XLSX, XLS. Legacy .doc and .ppt are not accepted.
  • Text / data: any text/* MIME, plus Markdown, CSV/TSV, JSON/JSONL, YAML, TOML, XML, HTML, RTF, LaTeX, logs, etc.
  • Source code: .py, .ts, .js, .go, .rs, .c, .cpp, .java, .sh, .ps1, .sql, … plus extensionless Dockerfile / Makefile (~140 text/data/code extensions are recognized in total).

Anything else → 400 Unsupported file type.

curl -X POST https://api.venice.ai/api/v1/augment/text-parser \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -F "file=@./contract.pdf" \
  -F "response_format=json"

Response

response_format=json:

{
  "text": "…extracted text…",
  "tokens": 3821
}

response_format=text — the raw extracted text as the body (Content-Type: text/plain).

Tips

  • DOCX, PPTX, XLSX/XLS and EPUB come back as markdown (spreadsheets as markdown tables). PDFs and text files come back as plain text.
  • tokens is a character-based estimate (ceil(chars / 3.2)), not a model tokenizer count — use it for rough budgeting of a downstream chat request.
  • Scanned/image-only PDFs are not OCR'd; if nothing can be extracted you get 400 "No text content could be extracted from the file." Run page images through a vision model via /chat/completions instead.
  • Password-protected PDFs → 400 "The PDF file is password-protected…"; empty/corrupt PDFs → 400 "The PDF file is empty or invalid."
  • Documents are processed in memory and content is not retained after the response. (Operational metadata such as request IDs and error traces may still be logged — this is a no-content-retention guarantee, not a zero-log guarantee.)

POST /augment/scrape — URL → markdown

Request

{ "url": "https://example.com/article" }

url must be a valid absolute http:// or https:// URL.

curl -X POST https://api.venice.ai/api/v1/augment/scrape \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"url":"https://example.com"}'

Response

{
  "url": "https://example.com",
  "content": "# Example Domain\n\nThis domain is for use in …",
  "format": "markdown"
}

How it fetches

Sites that serve markdown directly are returned as-is; otherwise the page is fetched and converted to markdown. Redirects are followed, but a redirect to a private/internal target fails the request (see below). The X/Reddit blocklist is checked against the URL you send, not against redirect targets.

Tips

  • Blocked sites — x.com, twitter.com (incl. www. / mobile.) and reddit.com (incl. www. / old. / new.) are rejected immediately with 400. Use enable_x_search or enable_web_search on /chat/completions for those — see venice-chat.
  • Blocked URLs — non-HTTP(S) schemes, URLs with embedded credentials, localhost / private / reserved IPs and cloud-metadata hosts return 400. A public hostname that resolves or redirects to such a target fails with 500.
  • Unreachable hosts, timeouts and pages with no extractable content return 500 with a message (e.g. "URL unreachable: …"). You are not charged.
  • Some sites return a partial body. Check content length before piping it into a model.
  • Rate limit: 20 requests/minute per user (429 beyond that). The counter runs before URL validation, so rejected URLs also count.

POST /augment/search — web search

Request

FieldNotes
queryRequired. 1–400 chars. Longer is rejected (400), not truncated.
limitInteger 1–20. Default 10. See the cap below.
search_provider"brave" (default) or "google".
curl -X POST https://api.venice.ai/api/v1/augment/search \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "query": "venice ai api pricing",
    "limit": 5,
    "search_provider": "brave"
  }'

Response

{
  "query": "venice ai api pricing",
  "results": [
    {
      "title": "Pricing — Venice.ai",
      "url": "https://venice.ai/pricing",
      "content": "Venice offers per-token pricing …",
      "date": "2026-04-10"
    }
  ]
}
  • date is whatever the provider reports and may be an empty string; don't assume a fixed format.
  • Results without a snippet are dropped, so you can get fewer than limit. A search with no hits returns 200 with results: [] (still charged).

Providers

ProviderRetentionResultscontent
brave (default)Zero Data Retention — queries are not stored or logged by the provider.At most 10Snippet plus any extra snippets.
googleAnonymized — proxied through Venice so your identity isn't attached to the query; Venice doesn't store or log queries.At most 5Extracted page summary/markdown (longer), truncated to a budget.

limit accepts up to 20, but the upstream request is fixed at 10 results for brave and 5 for google, so values above those caps have no effect.

Tips

  • For cited answers inside a chat completion, use /chat/completions with venice_parameters.enable_web_search + enable_web_citations instead. See venice-chat.
  • For "search + read" pipelines, feed results[*].url into /augment/scrape — mind the 20/min scrape limit.
  • Rate limit: 20 requests/minute per user (429 beyond that). Requests that fail body validation still count.

Errors

StatusCause
400Missing file, unsupported file type, no extractable text, password-protected/invalid PDF; invalid/blocked URL (X, Reddit, private/internal); empty or > 400-char query; limit out of range; non-JSON Content-Type on scrape/search ("'Content-Type' must be 'application/json'"). Body: { error }, or { error: "Invalid request parameters", details, issues } on schema failures.
401Invalid API key or SIWX signature.
402Insufficient balance or the key's USD/DIEM spend limit reached. Bearer → "Insufficient USD or Diem balance…"; x402 → payment-required body + PAYMENT-REQUIRED header. Requests with no credentials at all also get 402 (x402 discovery challenge).
403API access disabled for the account ("API access has been disabled for this account…").
413Text-parser file over 25 MB (PAYLOAD_TOO_LARGE).
429Per-endpoint rate limit (scrape/search: 20/min) or the failed-request limiter. Back off with jitter.
500Scrape fetch failure, search provider failure ("Search provider failed to return results…"), or parse failure ("Failed to parse document"). Not charged; safe to retry.

See venice-errors for body shapes and retry strategy.

Response headers

  • X-Balance-Remaining — listed in the spec for x402 callers but not currently set by the server; poll GET /x402/balance/{walletAddress} instead.
  • x-ratelimit-limit-requests / x-ratelimit-remaining-requests / x-ratelimit-reset-requests — scrape and search.
  • Content-Encoding — when you send Accept-Encoding: gzip, br.

Patterns

  • Document QA — upload a PDF to /augment/text-parser, put text into a /chat/completions message, ask questions.
  • Research agent — /augment/search → /augment/scrape on the top URLs → /chat/completions with the markdown bodies.
  • Data extraction — XLSX via text-parser yields markdown tables you can pipe to a model with response_format: { type: "json_schema", ... }.
  • Code review — send source files directly to text-parser (no need to convert to PDF) and feed the text to a coding model.

Files

1
9.7 KB

Agent reviews

0

No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.

More from veniceai/skills8

my-venice-skill

One or two sentences describing exactly when an agent should load this skill and what it covers. Mention the specific endpoints, parameters, or scenarios so the agent can confidently pick it — vague descriptions hurt skill selection.

Scan passed 0
venice-api-keys

Manage Venice API keys. Covers GET/POST/PATCH/DELETE /api_keys, GET /api_keys/{id}, GET /api_keys/rate_limits, GET /api_keys/rate_limits/log, the two-step /api_keys/generate_web3_key wallet flow, INFERENCE vs ADMIN key types, per-key consumption limits (USD / DIEM) with EPOCH / MONTH / LIFETIME rese

Scan passed 0
venice-api-overview

High-level map of the Venice.ai API - base URL, which auth mode each endpoint accepts (API key, x402 wallet, or none), endpoint categories (including decisions, voice changer, and retired routes), response headers (rate limit, balance, deprecation, x402), pricing model, error shape, and versioning.

Scan passed 0
venice-audio-music

Async music, sound-effect and long-form voice generation via Venice. Covers the /audio/quote + /audio/queue + /audio/retrieve + /audio/complete lifecycle, lyrics vs instrumental and the lyrics optimizer, duration options, seamless loop (ElevenLabs sound effects), voice selection incl. custom ElevenL

Scan passed 0
venice-audio-speech

Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices. Covers TTS models (Kokoro, Qwen 3, xAI, Inworld, Chatterbox, Orpheus, ElevenLabs Turbo, MiniMax, Gemini Flash, Gradium), voices per model, cloned-voice handles and raw ElevenLabs Voice IDs, per-model output f

Scan passed 0
venice-audio-transcription

Transcribe audio files to text via POST /audio/transcriptions. Covers supported models (Parakeet, Whisper, Wizper, Scribe, xAI STT), accepted containers (wav/flac/m4a/aac/mp4/mp3/ogg/webm), response formats (json/text only), per-model timestamps (word/segment/char), language hints, the 25 MB cap, an

Scan passed 0
venice-audio-voice-changer

Async speech-to-speech voice conversion via Venice — re-record a source recording in a different voice while keeping delivery and timing. Covers POST /audio/voice-changer/quote (unauthenticated), /queue (multipart file or JSON audio_url), /retrieve and /complete, how to discover voice-changer models

Scan passed 0
venice-auth

Authenticate to the Venice API with a Bearer API key or with an x402 / SIWX wallet (EVM on Base, including EIP-1271 smart wallets, or Ed25519 on Solana). Covers which endpoints accept which scheme, the SIGN-IN-WITH-X header format, the SIWE and Solana message fields, the enforced TTL / clock-skew /

Scan passed 0

Related knowledge skillsscan passed