building-livekit-agents
Builds voice and chat AI agents with LiveKit Agents and LiveKit Cloud. Use when the user asks to "build a voice agent", "create a LiveKit agent", "add voice AI to my app", "implement handoffs", "structure an agent workflow", "my agent is slow / too chatty", "it says it booked but nothing was saved",
- 0
- Installs
- —
- Rating
- —
- Success rate
- 2
- Files scanned
Security scan
Scan passedNo risky patterns were found in the scanned files.
Content sha256 42b7ff0eb1ce77c5… — run codexguild_scan_skills after installing to verify your local copy.
Static analysis is a first line of defense, not a guarantee. Read the source
SKILL.md
Building LiveKit agents
This skill covers how to structure a voice agent. It has no API specifics, because those change;
get them from reading-livekit-docs.
It assumes LiveKit Cloud, the recommended path: managed infrastructure, plus LiveKit Inference for models so you don't manage per-provider API keys.
Where the agent runs and which LiveKit the project uses are separate questions. An agent the user self-hosts (on their own servers instead of LiveKit Cloud's agent hosting) still connects to LiveKit Cloud and can still use LiveKit Inference. Inference is a LiveKit Cloud feature, so it's only off the table when the project runs on LiveKit OSS. The architecture advice applies either way; on LiveKit OSS, models come from each provider's own plugin and API keys.
Before you write code
- Load
reading-livekit-docsand look up the APIs you're about to use. Don't write LiveKit code from memory. - Confirm the project is connected to a LiveKit Cloud project (or a LiveKit OSS server):
LIVEKIT_URL,LIVEKIT_API_KEY,LIVEKIT_API_SECRET, usually in.env. The CLI can set these up. - Decide the workflow shape before writing the first agent class (see "Structure" below). Splitting a monolith into handoffs later is much more work than starting with two agents.
- Plan how you'll verify it. Decide now whether you'll use
debugging-livekit-agents(drive a real conversation),testing-livekit-agents(assert on turns), or both, because it affects how you factor the code.
How voice changes the requirements
A voice agent is more than a chat agent with a speaker attached. These constraints drive most design decisions:
Latency. Users expect a reply within a few hundred milliseconds. Context size, tool count, whether a tool call sits on the critical path, and whether responses stream all add to or save from that budget. Plan for network stalls and provider timeouts too; they happen routinely.
Context size. A 10,000-token system prompt with 50 tool definitions feels sluggish on any model, because the model re-reads all of it every turn. Give each phase only the tools it can reach and the instructions it needs.
Listening. Users can't skim or scroll back, and they'll talk over the agent. Long replies are a bug, silence sounds broken, and interruptions are normal.
Structure: handoffs and tasks
The usual failure is one agent that does everything. It collects every tool, instruction, and piece of state until it's slow and unreliable, and by that point splitting it is a rewrite.
Handoffs transfer control from one agent to another. Put them at natural conversation boundaries, like greeting → intake → resolution, or general support → billing specialist. Each agent then carries only its own tools and instructions. Choose a boundary where the context can be summarized for the next agent. If the next agent needs everything the previous one had, the boundary is in the wrong place.
Tasks are tightly scoped prompts aimed at one outcome. Use them for discrete operations that don't need a full agent, or where a focused prompt works better than a general one.
If you can't say in one sentence what an agent is responsible for, split it.
Tools
- Tool descriptions drive behavior. When an agent calls the wrong tool or calls one at the wrong time, check the description before blaming the model. The most common cause is a description that doesn't say when to use the tool.
- Keep tools off the critical path where you can. Users hear every tool call they wait on as latency.
- Plan for tool failure. Decide what the agent says when a backend is down or returns nothing. An agent that makes up an answer when a tool fails is very hard to catch later.
The model interprets; your code owns the state
The model reads the conversation and proposes actions. Application code owns the records, the permission checks, the state transitions, and every external effect. Most agents that "work in the demo and fail in production" have that line blurred somewhere.
- Never classify intent with code. Approval, refusal, correction, cancellation, "next Tuesday" — the runtime model interprets those. A regex, a keyword list, or a phrase whitelist will be wrong in ways you never test, and adding one as a "conservative" second gate has the same defect. Validate structure in code (typed dates, enums, required fields); leave meaning to the model.
- A tool call is the model's interpretation, not proof it was right. Keep message provenance, version checks, ordering, and business rules in code, where they can be checked.
- Tools return facts, not sentences. Compact data, outcomes, and actionable errors; the model chooses the wording. Script exact text only when the task mandates a verbatim disclosure.
- Follow the user, not a form. Accept facts the caller volunteers together, ask only for what's missing or ambiguous, and never demand ritual wording ("say yes to confirm") after a clear answer.
- One authoritative state object per session, and keep model-supplied facts separate from trusted identity, the clock, ids, and receipts.
Make every change mean exactly one thing
The costliest agent bugs are mutations that did more or less than the caller meant: "no note for him" clearing the whole list, a correction that also reset a confirmed field, a re-stated value that invalidated an approval. Before writing a mutating tool, state its target, what changes, and what must stay the same — then pair it with the nearest request that must do something different.
The rules in short: omission preserves; missing, empty, unknown, and cleared are four different
things; collections get application-issued ids; validate before applying; a scoped negative never
clears a collection; unchanged values are no-ops. When a task requires review before an effect,
approval is a later real user message for that version, delivery is tracked at the speech
boundary, and success is published only after the write commits. The full treatment — including
closing, output ownership, and how text and audio input take different hook paths — is in
references/state-and-effects.md. Read it before building anything that books, edits, confirms, or
ends calls.
Start with a failing complete-path test
Before expanding the tool surface or polishing the persona, pick one ordinary user goal and drive it through the real agent to its required effect — the booking exists, the record changed, the call ended. Write the expected result from the user's request, not from the application's own export. Then pair it with the first guard that must refuse, because a test that rejects everything proves nothing about the guard. Keep that pair green while you add everything else.
Verify before you call it done
Prompt changes break agent behavior as easily as code changes do, and trying it once by hand doesn't count as verification.
- While building, drive conversations with
debugging-livekit-agents. It runs your agent locally, lets you send turns as text or speech, and shows the tool calls behind each reply. - Before you call it done, write tests with
testing-livekit-agents. At minimum, cover the core behavior the user asked for, tool invocation with correct arguments if there are tools, and one failure path. - Before shipping a change to a live agent, run simulations with
writing-livekit-scenariosandrunning-livekit-simulations.
If the user asks for no tests, build without them, mention once that you'd recommend them before production, and move on.
Common mistakes
- Starting with one agent "just for now." You're deciding the structure up front; the implementation can still be simple.
- Putting off latency. It compounds, and it only gets more expensive to fix.
- Copying an example you don't understand. An example shows one pattern. Pasted whole, it brings extra context and components you can't explain.
- Assuming your model knowledge is current. It isn't. See
reading-livekit-docs. - Shipping on manual testing alone. Prompt edits change behavior without any visible error, and tests let you find out before users do.
Related skills
- Facts, APIs, changelogs:
reading-livekit-docs - Drive a live conversation while building:
debugging-livekit-agents - Turn-level tests:
testing-livekit-agents - Whole-conversation testing:
writing-livekit-scenarios,running-livekit-simulations - State, approvals, commits, delivery, closing:
references/state-and-effects.md - Deploying it and keeping it healthy in production:
operating-livekit-agents
Files
2- SKILL.md
881967fbed9.4 KB - references/state-and-effects.md
71ee0b4dda8.2 KB
Agent reviews
0No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.
More from livekit/agent-skills6
Drives a live multi-turn conversation with a LiveKit agent running locally, using `lk agent debugger`. Use when the user says "test my agent", "try my agent", "does this work", or "why did it call that tool", after editing an agent to check how it behaves, or when an agent on a speech-to-speech mode
Deploys and operates a LiveKit agent in production: shipping a version to LiveKit Cloud and rolling it back, secrets and configuration, the worker process model and prewarming, safe async inside worker processes, provider timeouts and degradation, graceful shutdown, SDK upgrades, and observability.
Looks up current LiveKit facts (API signatures, CLI flags, config options, model and provider support, SDK changelogs, pricing) from the docs instead of answering from memory. Use whenever a question touches LiveKit specifics: "does LiveKit support X", "what changed in agents 1.8", "what are the arg
Runs LiveKit agent simulations and acts on the results. Use when the user says "run my simulations", "regression test my agent before deploying", "run the scenarios", "use lk agent simulate", "did my agent pass", "why did this scenario fail", "run simulations in CI", "test the audio pipeline", "chec
Writes turn-level tests for a LiveKit agent in the user''s normal test suite: pytest (Python) or Vitest (Node.js). Use when the user asks to "write tests for my agent", "add a test for this tool", "test the handoff", "pin this bug", "why does my agent test fail", or after building or changing agent
Creates and maintains the scenarios a LiveKit agent simulation runs, and wires the agent to consume them. Use when the user asks "what should I test", "generate simulation scenarios", "write scenarios for my agent", "add a scenario for X", "organize my scenario files", "my simulations are flaky", "t
Related ai-ml skillsscan passed
Configure SuperJSON transformer on both server initTRPC.create({ transformer: superjson }) and every client terminating link (httpBatchLink, httpLink, wsLink, httpSubscriptionLink) to support Date, Map, Set, BigInt over the wire. Transformer must match on both sides. In v11, transformer goes on indi
Unified media generation via fal.ai MCP — image, video, and audio. Covers text-to-image (Nano Banana), text/image-to-video (Seedance, Kling, Veo 3), text-to-speech (CSM-1B), and video-to-audio (ThinkSound). Use when the user wants to generate images, videos, or audio with AI.
Store and query vector embeddings using Amazon S3 Vectors, a cost-effective long-term vector storage service with its own API namespace (s3vectors). Triggers on: create S3 vector bucket, vector index, store embeddings, semantic search, RAG vector storage, similarity search, vector database, migrate
Generates python code that evaluates SageMaker models. Supports two evaluation types: LLM-as-Judge and Custom Scorer. Use when the user says "evaluate my model", "run a benchmark", "test model performance", "how did my model perform", "compare models", or other similar requests.
Use Neo4j GenAI Plugin ai.text.* functions and procedures for in-Cypher
Guides agents to interactively discover customer requirements for live, bidirectional multi-agent AI systems that process continuous streams of multimodal data for real-time technical guidance and safety monitoring. Generates a custom Google Cloud solution that uses opinionated best practices and ar