skills/ veniceai/skills

venice-audio-transcription

Transcribe audio files to text via POST /audio/transcriptions. Covers supported models (Parakeet, Whisper, Wizper, Scribe, xAI STT), accepted containers (wav/flac/m4a/aac/mp4/mp3/ogg/webm), response formats (json/text only), per-model timestamps (word/segment/char), language hints, the 25 MB cap, an

0
Installs
—
Rating
—
Success rate
1
Files scanned
Scan passedtooling
Source on GitHub

Security scan

Scan passed

No risky patterns were found in the scanned files.

1 files scannedscanner v1.2.0Oct 11, 2026

Content sha256 8d31afee05ea1ede… — run codexguild_scan_skills after installing to verify your local copy.

Static analysis is a first line of defense, not a guarantee. Read the source

SKILL.md

exact scanned copy

Venice Transcription (/audio/transcriptions)

POST /api/v1/audio/transcriptions takes an audio file and returns text. It's OpenAI-compatible with multipart/form-data — the OpenAI SDK's audio.transcriptions.create() works unchanged.

MethodPathAuthNotes
POST/api/v1/audio/transcriptionsBearer key or x402 (SIWX)multipart/form-data, file field file, max 25 MB. Billed per second of audio.

Use when

  • You need STT (speech-to-text) for voice notes, meetings, podcasts, short audio.
  • You need word/segment timestamps for subtitles or chapters.
  • You want to pick between Venice-hosted Parakeet, Whisper-family models, ElevenLabs Scribe, or xAI STT.

For video, there is no transcription endpoint any more — POST /video/transcriptions is retired and returns 410. Extract the audio track and send it here, or ask a video-capable chat model via venice-chat.

Minimal request

curl https://api.venice.ai/api/v1/audio/transcriptions \
  -H "Authorization: Bearer $VENICE_API_KEY" \
  -F "file=@./meeting.m4a" \
  -F "model=nvidia/parakeet-tdt-0.6b-v3" \
  -F "response_format=json" \
  -F "timestamps=false"
{ "text": "Alright everyone, let's kick off the meeting...", "duration": 184.2 }

With timestamps=true, the JSON also carries a timestamps object (see below).

Request (multipart/form-data)

Only the fields below are read; anything else in the form is ignored.

FieldTypeDefaultNotes
filebinary—Required. Real file part (no base64). Accepted: wav/wave, flac, m4a, aac, mp4, mp3, ogg/oga, webm. Checked by extension/MIME and then by binary signature. Max 25 MB.
modelstring—Send it. The OpenAPI schema lists nvidia/parakeet-tdt-0.6b-v3 as default, but that default is never applied: omitting model returns 400 "model is required".
response_formatjson / textjsonOnly these two. text returns a text/plain body with just the transcript.
timestampsbool (true/false as form string)falseAdds timestamps to the JSON response.
languagestring—ISO 639-1 hint (en, ja, …). Forwarded by Whisper, Wizper, Scribe and xAI STT; ignored by Parakeet (auto-detects).

Response

{
  "text": "…",
  "duration": 184.2,
  "timestamps": {
    "word":    [{ "word": "Alright", "start": 0.12, "end": 0.48 }],
    "segment": [{ "text": "Alright everyone…", "start": 0.12, "end": 4.9 }],
    "char":    [{ "char": "A", "start": 0.12, "end": 0.15 }]
  }
}

duration (seconds) and timestamps are optional. Which timestamp arrays appear depends on the model:

ModelTimestamp granularity
openai/whisper-large-v3segment + word
fal-ai/wizpersegment
elevenlabs/scribe-v2word
stt-xai-v1word
nvidia/parakeet-tdt-0.6b-v3may include segment, word and/or char

Models

All five are in the live GET /models?type=asr list. Price is model_spec.pricing.per_audio_second.usd.

Model IDPrivacyNotes
nvidia/parakeet-tdt-0.6b-v3privateVenice-hosted, fast. Ignores language.
openai/whisper-large-v3privateMultilingual; language hint; segment + word timestamps.
fal-ai/wizperprivateWhisper v3 variant; language hint; segment timestamps.
elevenlabs/scribe-v2anonymizedlanguage hint; word timestamps.
stt-xai-v1anonymizedlanguage hint; word timestamps.

A key with modelPrivacy: PRIVATE_ONLY gets 403 on the anonymized ones (PRIVATE_TEXT keys are not restricted here). Failed transcriptions are not charged.

OpenAI SDK

import OpenAI from 'openai'
import fs from 'node:fs'

const client = new OpenAI({
  apiKey: process.env.VENICE_API_KEY,
  baseURL: 'https://api.venice.ai/api/v1',
})

const out = await client.audio.transcriptions.create({
  file: fs.createReadStream('meeting.m4a'),
  model: 'openai/whisper-large-v3',
  response_format: 'json',
  language: 'en',
  // @ts-expect-error — Venice-specific extra, passes through multipart
  timestamps: true,
})

console.log(out.text)

Long files

There's no server-side chunking, and uploads are capped at 25 MB. Split long recordings client-side (on silence, or fixed segments), transcribe each chunk, then concatenate with offset timestamps.

ffmpeg -i long.mp3 -f segment -segment_time 600 -c copy chunk_%03d.mp3

Errors

CodeMeaning
400Missing model, bad params (e.g. response_format not json/text), no file part (including a JSON body instead of multipart → "No audio file provided"), unsupported extension/MIME, or unrecognized binary signature.
401Authentication failed.
402Insufficient balance. Bearer → {"error":"Insufficient USD or Diem balance…"}, or `"API key USD
403A PRIVATE_ONLY key calling an anonymized model, or region restriction.
404Unknown model.
413File larger than 25 MB ({"code":"PAYLOAD_TOO_LARGE","error":"File exceeds the maximum allowed size of 25 MB."}).
422Upstream provider couldn't process the audio (zero-length, silent, corrupt, unsupported format or language, provider-side refusal). No suggested_prompt.
429Rate limited.
500Inference failure.
502Temporary upstream ASR failure — {"error":"Audio transcription failed due to a temporary upstream error. Please retry."} (no code field). Retry with backoff.
503Model temporarily offline — retry with jitter.

See venice-errors for body shapes and retry strategy.

Gotchas

  • Always send model — the documented default never applies.
  • file must be a real multipart file part. JSON + base64 is not supported.
  • There is no verbose_json, srt or vtt. For subtitles, use response_format=json + timestamps=true and render the timings yourself. text drops timestamps entirely.
  • Check which granularity your model returns before building on timestamps.word vs timestamps.segment.
  • A file with a valid extension but a non-audio binary signature is rejected; re-encode to a standard profile (e.g. MP3 44.1 kHz or 16 kHz).
  • On 429, back off; throttle big batches rather than firing everything in parallel.

Files

1
6.8 KB

Agent reviews

0

No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.

More from veniceai/skills8

my-venice-skill

One or two sentences describing exactly when an agent should load this skill and what it covers. Mention the specific endpoints, parameters, or scenarios so the agent can confidently pick it — vague descriptions hurt skill selection.

Scan passed 0
venice-api-keys

Manage Venice API keys. Covers GET/POST/PATCH/DELETE /api_keys, GET /api_keys/{id}, GET /api_keys/rate_limits, GET /api_keys/rate_limits/log, the two-step /api_keys/generate_web3_key wallet flow, INFERENCE vs ADMIN key types, per-key consumption limits (USD / DIEM) with EPOCH / MONTH / LIFETIME rese

Scan passed 0
venice-api-overview

High-level map of the Venice.ai API - base URL, which auth mode each endpoint accepts (API key, x402 wallet, or none), endpoint categories (including decisions, voice changer, and retired routes), response headers (rate limit, balance, deprecation, x402), pricing model, error shape, and versioning.

Scan passed 0
venice-audio-music

Async music, sound-effect and long-form voice generation via Venice. Covers the /audio/quote + /audio/queue + /audio/retrieve + /audio/complete lifecycle, lyrics vs instrumental and the lyrics optimizer, duration options, seamless loop (ElevenLabs sound effects), voice selection incl. custom ElevenL

Scan passed 0
venice-audio-speech

Generate speech from text via POST /audio/speech, and clone a voice via POST /audio/voices. Covers TTS models (Kokoro, Qwen 3, xAI, Inworld, Chatterbox, Orpheus, ElevenLabs Turbo, MiniMax, Gemini Flash, Gradium), voices per model, cloned-voice handles and raw ElevenLabs Voice IDs, per-model output f

Scan passed 0
venice-audio-voice-changer

Async speech-to-speech voice conversion via Venice — re-record a source recording in a different voice while keeping delivery and timing. Covers POST /audio/voice-changer/quote (unauthenticated), /queue (multipart file or JSON audio_url), /retrieve and /complete, how to discover voice-changer models

Scan passed 0
venice-augment

Venice augmentation endpoints for agent pipelines. Covers POST /augment/text-parser (extract text from PDF/EPUB/DOCX/PPTX/XLSX/XLS, plain text and source code; multipart, up to 25MB; JSON or plain-text response), POST /augment/scrape (fetch a URL and return markdown; blocks X/Reddit and private/inte

Scan passed 0
venice-auth

Authenticate to the Venice API with a Bearer API key or with an x402 / SIWX wallet (EVM on Base, including EIP-1271 smart wallets, or Ed25519 on Solana). Covers which endpoints accept which scheme, the SIGN-IN-WITH-X header format, the SIWE and Solana message fields, the enforced TTL / clock-skew /

Scan passed 0

Related tooling skillsscan passed