venice-audio-transcription
Transcribe audio files to text via POST /audio/transcriptions. Covers supported models (Parakeet, Whisper, Wizper, Scribe, xAI STT), accepted containers (wav/flac/m4a/aac/mp4/mp3/ogg/webm), response formats (json/text only), per-model timestamps (word/segment/char), language hints, the 25 MB cap, an
- 0
- Installs
- —
- Rating
- —
- Success rate
- 1
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 8d31afee05ea1ede… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Venice Transcription (/audio/transcriptions)
POST /api/v1/audio/transcriptions takes an audio file and returns text. It's OpenAI-compatible with multipart/form-data — the OpenAI SDK's audio.transcriptions.create() works unchanged.
| Method | Path | Auth | Notes |
|---|---|---|---|
POST | /api/v1/audio/transcriptions | Bearer key or x402 (SIWX) | multipart/form-data, file field file, max 25 MB. Billed per second of audio. |
Use when
- You need STT (speech-to-text) for voice notes, meetings, podcasts, short audio.
- You need word/segment timestamps for subtitles or chapters.
- You want to pick between Venice-hosted Parakeet, Whisper-family models, ElevenLabs Scribe, or xAI STT.
For video, there is no transcription endpoint any more — POST /video/transcriptions is retired and returns 410. Extract the audio track and send it here, or ask a video-capable chat model via venice-chat.
Minimal request
curl https://api.venice.ai/api/v1/audio/transcriptions \
-H "Authorization: Bearer $VENICE_API_KEY" \
-F "file=@./meeting.m4a" \
-F "model=nvidia/parakeet-tdt-0.6b-v3" \
-F "response_format=json" \
-F "timestamps=false"
{ "text": "Alright everyone, let's kick off the meeting...", "duration": 184.2 }
With timestamps=true, the JSON also carries a timestamps object (see below).
Request (multipart/form-data)
Only the fields below are read; anything else in the form is ignored.
| Field | Type | Default | Notes |
|---|---|---|---|
file | binary | — | Required. Real file part (no base64). Accepted: wav/wave, flac, m4a, aac, mp4, mp3, ogg/oga, webm. Checked by extension/MIME and then by binary signature. Max 25 MB. |
model | string | — | Send it. The OpenAPI schema lists nvidia/parakeet-tdt-0.6b-v3 as default, but that default is never applied: omitting model returns 400 "model is required". |
response_format | json / text | json | Only these two. text returns a text/plain body with just the transcript. |
timestamps | bool (true/false as form string) | false | Adds timestamps to the JSON response. |
language | string | — | ISO 639-1 hint (en, ja, …). Forwarded by Whisper, Wizper, Scribe and xAI STT; ignored by Parakeet (auto-detects). |
Response
{
"text": "…",
"duration": 184.2,
"timestamps": {
"word": [{ "word": "Alright", "start": 0.12, "end": 0.48 }],
"segment": [{ "text": "Alright everyone…", "start": 0.12, "end": 4.9 }],
"char": [{ "char": "A", "start": 0.12, "end": 0.15 }]
}
}
duration (seconds) and timestamps are optional. Which timestamp arrays appear depends on the model:
| Model | Timestamp granularity |
|---|---|
openai/whisper-large-v3 | segment + word |
fal-ai/wizper | segment |
elevenlabs/scribe-v2 | word |
stt-xai-v1 | word |
nvidia/parakeet-tdt-0.6b-v3 | may include segment, word and/or char |
Models
All five are in the live GET /models?type=asr list. Price is model_spec.pricing.per_audio_second.usd.
| Model ID | Privacy | Notes |
|---|---|---|
nvidia/parakeet-tdt-0.6b-v3 | private | Venice-hosted, fast. Ignores language. |
openai/whisper-large-v3 | private | Multilingual; language hint; segment + word timestamps. |
fal-ai/wizper | private | Whisper v3 variant; language hint; segment timestamps. |
elevenlabs/scribe-v2 | anonymized | language hint; word timestamps. |
stt-xai-v1 | anonymized | language hint; word timestamps. |
A key with modelPrivacy: PRIVATE_ONLY gets 403 on the anonymized ones (PRIVATE_TEXT keys are not restricted here). Failed transcriptions are not charged.
OpenAI SDK
import OpenAI from 'openai'
import fs from 'node:fs'
const client = new OpenAI({
apiKey: process.env.VENICE_API_KEY,
baseURL: 'https://api.venice.ai/api/v1',
})
const out = await client.audio.transcriptions.create({
file: fs.createReadStream('meeting.m4a'),
model: 'openai/whisper-large-v3',
response_format: 'json',
language: 'en',
// @ts-expect-error — Venice-specific extra, passes through multipart
timestamps: true,
})
console.log(out.text)
Long files
There's no server-side chunking, and uploads are capped at 25 MB. Split long recordings client-side (on silence, or fixed segments), transcribe each chunk, then concatenate with offset timestamps.
ffmpeg -i long.mp3 -f segment -segment_time 600 -c copy chunk_%03d.mp3
Errors
| Code | Meaning |
|---|---|
400 | Missing model, bad params (e.g. response_format not json/text), no file part (including a JSON body instead of multipart → "No audio file provided"), unsupported extension/MIME, or unrecognized binary signature. |
401 | Authentication failed. |
402 | Insufficient balance. Bearer → {"error":"Insufficient USD or Diem balance…"}, or `"API key USD |
403 | A PRIVATE_ONLY key calling an anonymized model, or region restriction. |
404 | Unknown model. |
413 | File larger than 25 MB ({"code":"PAYLOAD_TOO_LARGE","error":"File exceeds the maximum allowed size of 25 MB."}). |
422 | Upstream provider couldn't process the audio (zero-length, silent, corrupt, unsupported format or language, provider-side refusal). No suggested_prompt. |
429 | Rate limited. |
500 | Inference failure. |
502 | Temporary upstream ASR failure — {"error":"Audio transcription failed due to a temporary upstream error. Please retry."} (no code field). Retry with backoff. |
503 | Model temporarily offline — retry with jitter. |
See venice-errors for body shapes and retry strategy.
Gotchas
- Always send
model— the documented default never applies. filemust be a real multipart file part. JSON + base64 is not supported.- There is no
verbose_json,srtorvtt. For subtitles, useresponse_format=json+timestamps=trueand render the timings yourself.textdrops timestamps entirely. - Check which granularity your model returns before building on
timestamps.wordvstimestamps.segment. - A file with a valid extension but a non-audio binary signature is rejected; re-encode to a standard profile (e.g. MP3 44.1 kHz or 16 kHz).
- On
429, back off; throttle big batches rather than firing everything in parallel.
Files
1- SKILL.md
479b0fa5146.8 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from veniceai/skills8
One or two sentences describing exactly when an agent should load this skill and what it covers. Mention the specific endpoints, parameters, or scenarios so the agent can confidently pick it — vague descriptions hurt skill selection.
Manage Venice API keys. Covers GET/POST/PATCH/DELETE /api_keys, GET /api_keys/{id}, GET /api_keys/rate_limits, GET /api_keys/rate_limits/log, the two-step /api_keys/generate_web3_key wallet flow, INFERENCE vs ADMIN key types, per-key consumption limits (USD / DIEM) with EPOCH / MONTH / LIFETIME rese
High-level map of the Venice.ai API - base URL, which auth mode each endpoint accepts (API key, x402 wallet, or none), endpoint categories (including decisions, voice changer, and retired routes), response headers (rate limit, balance, deprecation, x402), pricing model, error shape, and versioning.
Async music, sound-effect and long-form voice generation via Venice. Covers the /audio/quote + /audio/queue + /audio/retrieve + /audio/complete lifecycle, lyrics vs instrumental and the lyrics optimizer, duration options, seamless loop (ElevenLabs sound effects), voice selection incl. custom ElevenL
Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices. Covers TTS models (Kokoro, Qwen 3, xAI, Inworld, Chatterbox, Orpheus, ElevenLabs Turbo, MiniMax, Gemini Flash, Gradium), voices per model, cloned-voice handles and raw ElevenLabs Voice IDs, per-model output f
Async speech-to-speech voice conversion via Venice — re-record a source recording in a different voice while keeping delivery and timing. Covers POST /audio/voice-changer/quote (unauthenticated), /queue (multipart file or JSON audio_url), /retrieve and /complete, how to discover voice-changer models
Venice augmentation endpoints for agent pipelines. Covers POST /augment/text-parser (extract text from PDF/EPUB/DOCX/PPTX/XLSX/XLS, plain text and source code; multipart, up to 25MB; JSON or plain-text response), POST /augment/scrape (fetch a URL and return markdown; blocks X/Reddit and private/inte
Authenticate to the Venice API with a Bearer API key or with an x402 / SIWX wallet (EVM on Base, including EIP-1271 smart wallets, or Ed25519 on Solana). Covers which endpoints accept which scheme, the SIGN-IN-WITH-X header format, the SIWE and Solana message fields, the enforced TTL / clock-skew /
Related tooling skillsscan passed
Open an executable and its argument array in a visible terminal window through a reusable, shell-free launch plan with dry-run, JSON, capability detection, detached fallback, and standalone recovery modes. Use when Codex needs to open an interactive CLI, SSH session, local development process, sandb
Web performance regression detection. (gstack)
Audit and improve CLAUDE.md files in repositories. Use when user asks to check, audit, update, improve, or fix CLAUDE.md files. Scans for all CLAUDE.md files, evaluates quality against templates, outputs quality report, then makes targeted updates. Also use when the user mentions "CLAUDE.md maintena
Helps you build and check a color system for your project. It generates palettes, names semantic tokens, converts between formats and measures contrast.
Creates a new Angular app using the Angular CLI. This skill should be used whenever a user wants to create a new Angular application and contains important guidelines for how to effectively create a modern Angular application.
Audit, diagnose, or optimize website loading and interaction performance, Core Web Vitals, and Lighthouse performance scores.