skills/ veniceai/skills

venice-responses

Use Venice's Alpha POST /responses endpoint - an OpenAI-compatible, stateless Responses API with typed output blocks (reasoning, message, function_call, web_search_call). Covers request shape, input items, tools (function, web_search, x_search), reasoning controls, incomplete responses, streaming ev

0
Installs
—
Rating
—
Success rate
1
Files scanned
Scan passedai-ml
Source on GitHub

Security scan

Scan passed

No risky patterns were found in the scanned files.

1 files scannedscanner v1.2.0Oct 10, 2026

Content sha256 d600addb1b5f9d72… — run codexguild_scan_skills after installing to verify your local copy.

Static analysis is a first line of defense, not a guarantee. Read the source

SKILL.md

exact scanned copy

Venice Responses API (Alpha)

POST /api/v1/responses is Venice's OpenAI-compatible Responses endpoint. It returns a typed output array instead of a single message.content string — useful for agents that need to separate reasoning, messages, tool calls, and web-search events. Internally the request is translated to a chat completion, so model support matches venice-chat.

Alpha. The spec labels it Alpha (and its description still says "Alpha testers only"), but access is no longer restricted: any Bearer API key or x402 wallet can call it. Schemas may still change.

Use when

  • A client library expects the OpenAI Responses shape (output[] with type: "reasoning" | "message" | "function_call" | "web_search_call").
  • You want reasoning, message, and tool-call output cleanly separated.
  • You want SSE streaming with typed events.

Otherwise use venice-chat — it has structured output, audio/video/file inputs, E2EE, sampling controls, and every venice_parameters field.

Limitations vs /chat/completions

LimitationDetail
StatelessNothing is stored. Send the full history each call. previous_response_id, store, background are ignored.
No E2EEE2EE-capable models return 400 unless venice_parameters.enable_e2ee: false (TEE-only mode). For encrypted inference use /chat/completions.
Text + image input onlyinput_text / input_image (and text / image_url parts). No audio, video, or file parts.
No structured outputtext.format / response_format are dropped. Use /chat/completions.
Subset of venice_parameterscharacter_slug, enable_e2ee, enable_web_search, enable_web_scraping, enable_web_citations, include_venice_system_prompt, include_search_results_in_stream. Other keys (strip_thinking_response, disable_thinking, enable_x_search, return_search_results_as_documents) are silently dropped.
No model feature suffixesmodel: "zai-org-glm-5-1:enable_web_search=on" resolves the model but ignores the suffix.
Few generation controlsOnly temperature, top_p, max_output_tokens.

Authentication

Same as the rest of the API — Authorization: Bearer <key> or SIGN-IN-WITH-X: <SIWX> for x402 wallets. See venice-auth.

Minimal request

curl https://api.venice.ai/api/v1/responses \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "zai-org-glm-5-1",
    "input": "Explain why the sky is blue in one paragraph."
  }'

OpenAI SDK (Python):

import os
from openai import OpenAI
client = OpenAI(api_key=os.environ["VENICE_API_KEY"], base_url="https://api.venice.ai/api/v1")
resp = client.responses.create(model="zai-org-glm-5-1", input="Explain why the sky is blue.")
print(resp.output_text)

Request fields

FieldNotes
modelRequired. Model ID, trait, or compatibility mapping.
inputRequired. A string, or an array of input items (below).
max_output_tokensPositive integer. Mapped to max_tokens; above the model's model_spec.maxCompletionTokens → 400 on models with an enforced cap.
temperature (0–2), top_p (0–1)Sampling.
reasoning.effortnone | minimal | low | medium | high | xhigh | max (per-model support: model_spec.capabilities.reasoningEffortOptions). reasoning may be null.
reasoning.enabledfalse disables reasoning on supported models and suppresses reasoning blocks. Ignored when an effort is set.
reasoning.summaryauto | concise | detailed. Accepted but not forwarded.
toolsSee Tools.
tool_choice"auto" | "none" | "required" | {"type":"function","function":{"name":"..."}}. Dropped when no function tools remain.
web_searchBoolean. true forces web search on (same as a {"type":"web_search"} tool).
includeOnly "reasoning.encrypted_content" has an effect (adds encrypted_content to reasoning blocks when the provider returns it).
streamBoolean. SSE with typed events.
anon_user_idOptional end-user id: 1–128 printable ASCII characters, no `
fallbacksUp to 10 {model} entries. Anthropic beta refusal fallback for Claude Fable 5; forwarded only on direct Anthropic routes.
venice_parametersSubset listed above. Example: {"character_slug":"alan-watts","enable_web_search":"auto"}.

The body is permissive: other fields (instructions, metadata, parallel_tool_calls, n, stop, seed, prompt_cache_key, store, previous_response_id, background, text, user) are accepted without error but never reach inference (user still splits the error budget per value). Put system instructions in the input array instead of instructions.

Input items

ItemShapeHandling
Message{role, content} or {type:"message", role, content}role: user / assistant / system / developer (developer becomes system). content is a string or an array of parts.
Function call{type:"function_call", call_id, name, arguments}Replayed as an assistant tool call.
Function output{type:"function_call_output", call_id, output}output may be a string, array, object, number, boolean, or null. input_image parts inside an array output are forwarded to the model as images.
Reasoning{type:"reasoning", ...}Accepted but discarded — reasoning is not carried between turns.
Item reference{type:"item_reference", id}Accepted but discarded (nothing is stored to reference).

Content parts: input_text, output_text (to replay assistant output), and input_image. input_image.image_url may be a URL string (OpenAI Responses style) or {url, detail}; detail (auto / low / high) may also sit on the part. Messages without type additionally accept Chat-style text and image_url parts. Image URLs get the same validation as on /chat/completions (public, no redirects, ≥ 64 px); failures → 400. Use a vision model: message images are not capability-checked on this endpoint (images inside a function_call_output on a non-vision model do return 400).

Tools

ToolEffect
{"type":"function","function":{name, description, parameters, strict}}Function calling. The flat OpenAI form {"type":"function","name":...,"parameters":...} is also accepted. Use a model with supportsFunctionCalling (not pre-checked on this endpoint, unlike chat).
{"type":"web_search"}Forces Venice web search on (not auto). search_context_size / user_location are accepted but ignored.
{"type":"x_search", ...}xAI native web + X search on models with supportsXSearch (Grok); ignored on other models. Optional filters: allowed_x_handles / excluded_x_handles (≤ 10 each), from_date, to_date, enable_image_understanding, enable_video_understanding.
code_interpreter, file_search, computer_use_preview, othersAccepted and dropped. Unknown tool types that carry a name are treated as function tools.

Response shape

{
  "id": "resp_chatcmpl-abc123",
  "object": "response",
  "created_at": 1735689600,
  "model": "zai-org-glm-5-1",
  "status": "completed",
  "output": [
    {"type": "reasoning", "id": "rs_1", "summary": ["I considered Rayleigh scattering..."]},
    {"type": "web_search_call", "id": "ws_1", "status": "completed"},
    {"type": "function_call", "id": "fc_1", "call_id": "call_abc", "name": "get_weather",
     "arguments": "{\"city\":\"Paris\"}", "status": "completed"},
    {"type": "message", "id": "msg_1", "status": "completed", "role": "assistant",
     "content": [{"type": "output_text", "text": "The sky is blue because... ^1^",
       "annotations": [{"type": "url_citation", "url": "https://example.com/rayleigh",
         "title": "Rayleigh scattering", "start_index": 27, "end_index": 30}]}]}
  ],
  "usage": {
    "input_tokens": 20,
    "input_tokens_details": {"cached_tokens": 8},
    "output_tokens": 80,
    "output_tokens_details": {"reasoning_tokens": 40},
    "total_tokens": 100
  }
}
  • Non-streamed output order: reasoning → web_search_call → function_call(s) → message. The message block is omitted when the model returned only tool calls with no text.
  • status is completed or incomplete. Errors before or during inference come back as HTTP errors (streaming uses response.failed).
  • Incomplete responses keep their output and usage. When generation stops on max_output_tokens or a content filter, status: "incomplete", incomplete_details: {"reason": "max_output_tokens" | "content_filter"}, and message / function_call blocks carry status: "incomplete".
  • input_tokens_details appears only when cached tokens are non-zero; output_tokens_details only when the provider reports reasoning tokens. There is no cost field (unlike /chat/completions).

Output block types

typePurpose
reasoningReasoning from thinking models. summary[] holds text; encrypted_content appears only if you sent include: ["reasoning.encrypted_content"] and the provider returned encrypted reasoning. Sending it back in input has no effect.
messageMain text. content[].type === "output_text" with annotations[].
function_callTool call: name, JSON-string arguments, call_id. Answer with a function_call_output item with the same call_id.
web_search_callMarker that Venice web search ran.

url_citation annotations are built only when the text contains single-index ^n^ markers — set venice_parameters.enable_web_citations: true to get them. Each annotation spans the marker itself; multi-index markers such as ^1,3^ are not annotated.

Streaming

With stream: true, events are event: <type> + data: {...} pairs; payloads carry type (equal to the event name) and an increasing sequence_number — except response.web_search.done, whose payload has type: "web_search_call", id, status, results and no sequence_number. Typical flow:

event: response.created                      # status: in_progress
event: response.web_search.done              # Venice-specific; only when search ran, carries results[{index,url,title,snippet}]
event: response.output_item.added            # item.type = reasoning
event: response.reasoning.delta
event: response.output_item.added            # item.type = message
event: response.content_part.added
event: response.output_text.delta            # repeated
event: response.output_item.added            # item.type = function_call
event: response.function_call_arguments.delta
event: response.output_item.done             # reasoning, then message (after content_part.done), then each function_call
event: response.completed                    # or response.incomplete, with the full response
data: [DONE]
  • On an upstream failure you get response.failed (response.status: "failed", response.error: {code, message}) followed by data: [DONE].
  • The final response.completed / response.incomplete payload differs slightly from a non-streamed response: annotations are always empty and function calls come after the message.
  • include_search_results_in_stream has no effect here; search results always arrive in response.web_search.done.

Errors

StatusWhen
400Invalid body, E2EE-capable model without enable_e2ee: false, invalid image, unsupported reasoning.effort for the model, context too long, max_output_tokens over the cap
401Invalid API key or SIWX sign-in; also a model that requires a paid subscription
402No credentials at all (x402 discovery body — not 401), insufficient balance, or API-key spend limit. x402: PAYMENT_REQUIRED body with topUpInstructions + siwxChallenge and a PAYMENT-REQUIRED header (see venice-x402)
403Model blocked by the key's modelPrivacy, region, or provider restriction
404Unknown model
422Content-policy violation on an input image
429Rate limited
500 / 503Inference failed (upstream overloads and timeouts also surface as 500 here, not 429 / 504; retry with backoff) / model offline

The spec lists an X-Balance-Remaining header on x402 200 responses, but the server does not currently set it — poll GET /x402/balance/{walletAddress} instead. See venice-errors.

Migration notes (from /chat/completions)

  • messages → input (the same role/content objects work; system prompts go in as role: "system" or "developer" items).
  • max_tokens → max_output_tokens; reasoning_effort → reasoning.effort.
  • Tool results → function_call_output items keyed by call_id.
  • venice_parameters.character_slug, enable_web_search, enable_web_citations, enable_web_scraping, include_venice_system_prompt → pass inside venice_parameters (not as model suffixes).
  • enable_x_search → add an {"type":"x_search"} tool instead.
  • strip_thinking_response / disable_thinking → use reasoning.enabled: false.
  • Structured output, audio / video / file inputs, seed / stop / n, logprobs, prompt-cache routing, and full E2EE → stay on /chat/completions.

Gotchas

  • Unknown character_slug is not rejected here (chat returns 404). The request runs without the character, without your system messages, and without the Venice system prompt or web search. Validate slugs first with GET /characters/{slug} (venice-characters; Bearer key only — wallet callers can't, and should use /chat/completions, which returns 404 for unknown slugs).
  • Reasoning items in input are discarded; there is no cross-turn reasoning carry-over on this endpoint.
  • tool_choice objects must be {"type":"function","function":{"name":...}}; the flat {"type":"function","name":...} form fails validation.
  • Stateless: previous_response_id is silently ignored, so omitting history silently loses context.

Files

1
14.5 KB

Agent reviews

0

No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.

More from veniceai/skills8

my-venice-skill

One or two sentences describing exactly when an agent should load this skill and what it covers. Mention the specific endpoints, parameters, or scenarios so the agent can confidently pick it — vague descriptions hurt skill selection.

Scan passed 0
venice-api-keys

Manage Venice API keys. Covers GET/POST/PATCH/DELETE /api_keys, GET /api_keys/{id}, GET /api_keys/rate_limits, GET /api_keys/rate_limits/log, the two-step /api_keys/generate_web3_key wallet flow, INFERENCE vs ADMIN key types, per-key consumption limits (USD / DIEM) with EPOCH / MONTH / LIFETIME rese

Scan passed 0
venice-api-overview

High-level map of the Venice.ai API - base URL, which auth mode each endpoint accepts (API key, x402 wallet, or none), endpoint categories (including decisions, voice changer, and retired routes), response headers (rate limit, balance, deprecation, x402), pricing model, error shape, and versioning.

Scan passed 0
venice-audio-music

Async music, sound-effect and long-form voice generation via Venice. Covers the /audio/quote + /audio/queue + /audio/retrieve + /audio/complete lifecycle, lyrics vs instrumental and the lyrics optimizer, duration options, seamless loop (ElevenLabs sound effects), voice selection incl. custom ElevenL

Scan passed 0
venice-audio-speech

Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices. Covers TTS models (Kokoro, Qwen 3, xAI, Inworld, Chatterbox, Orpheus, ElevenLabs Turbo, MiniMax, Gemini Flash, Gradium), voices per model, cloned-voice handles and raw ElevenLabs Voice IDs, per-model output f

Scan passed 0
venice-audio-transcription

Transcribe audio files to text via POST /audio/transcriptions. Covers supported models (Parakeet, Whisper, Wizper, Scribe, xAI STT), accepted containers (wav/flac/m4a/aac/mp4/mp3/ogg/webm), response formats (json/text only), per-model timestamps (word/segment/char), language hints, the 25 MB cap, an

Scan passed 0
venice-audio-voice-changer

Async speech-to-speech voice conversion via Venice — re-record a source recording in a different voice while keeping delivery and timing. Covers POST /audio/voice-changer/quote (unauthenticated), /queue (multipart file or JSON audio_url), /retrieve and /complete, how to discover voice-changer models

Scan passed 0
venice-augment

Venice augmentation endpoints for agent pipelines. Covers POST /augment/text-parser (extract text from PDF/EPUB/DOCX/PPTX/XLSX/XLS, plain text and source code; multipart, up to 25MB; JSON or plain-text response), POST /augment/scrape (fetch a URL and return markdown; blocks X/Reddit and private/inte

Scan passed 0

Related ai-ml skillsscan passed

ai-regression-testing

Regression testing strategies for AI-assisted development. Sandbox-mode API testing without database dependencies, automated bug-check workflows, and patterns to catch AI blind spots where the same model writes and reviews code. Use when adding regression coverage to AI-assisted code, or when the sa

Scan passed 0
amazon-workspaces-agent-access

Connects AI agents to remote Windows desktop applications on Amazon WorkSpaces Applications (AppStream 2.0) through the managed Agent Access MCP server, and guides reliable desktop automation. Covers connecting an agent to the MCP endpoint (SigV4, streaming URL, and Active Directory SAML/Domain Join

Scan passed 0
finetuning-technique

Selects a fine-tuning technique (SFT, DPO, RLVR, or RLAIF) for the user's use case and validates it against the selected model's available recipes. Use when the user has decided to finetune and needs to choose a technique, or when the technique needs to be validated against a model. Requires a base

Scan passed 0
building-livekit-agents

Builds voice and chat AI agents with LiveKit Agents and LiveKit Cloud. Use when the user asks to "build a voice agent", "create a LiveKit agent", "add voice AI to my app", "implement handoffs", "structure an agent workflow", "my agent is slow / too chatty", "it says it booked but nothing was saved",

Scan passed 0
neo4j-vector-index-skill

Create and manage Neo4j vector indexes, run vector similarity search (ANN/kNN),

Scan passed 0
agent-platform-model-registry

Agent Platform Model Registry Management. Use when you need to upload, list, describe, update, or delete machine learning models (and their versions) in the Agent Platform Model Registry. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform models.

Scan passed 0