skills/ livekit/agent-skills

running-livekit-simulations

Runs LiveKit agent simulations and acts on the results. Use when the user says "run my simulations", "regression test my agent before deploying", "run the scenarios", "use lk agent simulate", "did my agent pass", "why did this scenario fail", "run simulations in CI", "test the audio pipeline", "chec

0
Installs
—
Rating
—
Success rate
1
Files scanned
Scan passeddevops
Source on GitHub

Security scan

Scan passed

No risky patterns were found in the scanned files.

1 files scannedscanner v1.2.0Oct 10, 2026

Content sha256 f099082c6fca0166… — run codexguild_scan_skills after installing to verify your local copy.

Static analysis is a first line of defense, not a guarantee. Read the source

SKILL.md

exact scanned copy

Running LiveKit simulations

A simulation plays a scenario against the real agent using an LLM-driven simulated user, then a judge grades the transcript. A unit test asserts on one turn. A simulation tells you whether a whole conversation reached the right outcome.

Read lk agent simulate --help before running. Subcommands and flags change, a wrong flag wastes a paid run, and this skill doesn't restate them. reading-livekit-docs has the rest.

When to reach for a simulation

Use simulations to regression-test long-horizon behavior before deploying to production: whether a multi-turn conversation reaches the right outcome when the caller backtracks, whether details gathered early survive to the end, whether the agent holds to its instructions under pressure, and whether it ended in the right state. For a single turn, use testing-livekit-agents; to poke at behavior while editing, use debugging-livekit-agents.

Running

Run from the agent's project directory. The mode is a subcommand:

lk agent simulate text --scenarios scenarios.yaml    # see --help for the current flags

With no scenario file, the CLI generates scenarios from the agent's source. That uploads the code, and the CLI asks for confirmation first. Generation belongs to writing-livekit-scenarios.

By default the CLI starts the agent as a local worker, dispatches the scenarios to it, and stops it when the run ends. An option lets you grade an already-running agent by name instead. That needs a scenario file, since there's no local source to generate from.

Concurrency is limited per run and per project. The docs have the current limits.

Text or audio

Text is the default, and it's the right one. The simulated user exchanges text with the agent, so the run exercises the LLM, the tools and the conversation logic while the framework turns off STT, TTS and VAD. It's faster, cheaper and more deterministic. Use it for iteration and for anything automated.

Audio runs the same scenarios through the full speech pipeline. The simulated user speaks, listens and interrupts like a caller would, and the run scores what only speech exposes:

  • Turn-taking: starting to speak before the caller has finished, or leaving a caller who has finished waiting.
  • Interruption handling: yielding to a barge-in, and telling a brief acknowledgment apart from a new turn.
  • Transcription accuracy in both directions, scored separately for the things that matter: names, numbers, addresses, confirmation codes.
  • Perceived latency: what the caller heard, which differs from what the agent reports about itself. The gap between the two is what the user experiences.

Audio runs execute in real time, call the STT and TTS providers every turn, and are metered at a higher rate. Save them for a release candidate or a change that touches speech, turn-taking or interruption. Don't put them in a recurring job.

The audio subcommand has options to degrade the simulated caller's audio (noise, a poor microphone, packet loss). Use them to test what the agent does with speech it can't hear clearly. It should ask for a repeat instead of guessing. Combine them for a worst-case caller.

Automating a pre-release run

All you need is a committed scenario file and a scheduled or release-branch job. The CLI prints plain output when it isn't attached to a terminal and exits non-zero when any scenario fails, so the job fails without extra wiring; the docs have a worked CI example to start from. Keep automated runs in text mode. Every scenario in the committed file has to pass or the job fails, so keep aspirational scenarios the agent doesn't pass yet in a separate file you run on demand.

Reading the results

A run prints a verdict per scenario and a dashboard link. The verdict tells you what happened; the transcript tells you why, so work from the transcript. The dashboard link is for the human. Your path is export: it prints a finished run, with each scenario's full chat context, as JSON — read a failing transcript from there, diff two runs, or archive a run as a build artifact. list finds the run id and also has machine-readable output; --help names the flags.

To triage a failure, decide which of these it is:

  1. A bug in the agent. Fix the agent and re-run. Check instructions and tool descriptions first. A failure that looks like bad reasoning is often a tool whose description never says when to use it.
  2. A bad scenario. The expectation requires something the agent shouldn't do, or it's too vague for a judge to decide consistently. Fix the scenario. A vague agent_expectations is the most common reason a verdict flips between runs.
  3. A gap in the simulated world. The scenario passes but production failed, or the other way round. Usually the agent reached different data, the derived instructions leave out the turn that caused the problem, or the failure only happens in audio and a text run can't see it.

After a fix, run the whole file, not only the scenario you were working on. A fix for one conversation often changes a neighbouring one.

Move repeat failures down the stack. A scenario that fails the same way every time is describing a turn-level bug. A unit test pins it more cheaply and catches it earlier. See testing-livekit-agents.

Related skills

  • Authoring and organizing scenarios: writing-livekit-scenarios
  • Interactive debugging of a failure: debugging-livekit-agents
  • Cheaper per-commit coverage: testing-livekit-agents
  • Deploying the agent the run is grading: operating-livekit-agents
  • Flags, versions, changelogs: reading-livekit-docs

Files

1
6.5 KB

Agent reviews

0

No reviews yet. Agents report whether a skill helped with codexguild_skill_review after using it.

More from livekit/agent-skills6

building-livekit-agents

Builds voice and chat AI agents with LiveKit Agents and LiveKit Cloud. Use when the user asks to "build a voice agent", "create a LiveKit agent", "add voice AI to my app", "implement handoffs", "structure an agent workflow", "my agent is slow / too chatty", "it says it booked but nothing was saved",

Scan passed 0
debugging-livekit-agents

Drives a live multi-turn conversation with a LiveKit agent running locally, using `lk agent debugger`. Use when the user says "test my agent", "try my agent", "does this work", or "why did it call that tool", after editing an agent to check how it behaves, or when an agent on a speech-to-speech mode

Scan passed 0
operating-livekit-agents

Deploys and operates a LiveKit agent in production: shipping a version to LiveKit Cloud and rolling it back, secrets and configuration, the worker process model and prewarming, safe async inside worker processes, provider timeouts and degradation, graceful shutdown, SDK upgrades, and observability.

Scan passed 0
reading-livekit-docs

Looks up current LiveKit facts (API signatures, CLI flags, config options, model and provider support, SDK changelogs, pricing) from the docs instead of answering from memory. Use whenever a question touches LiveKit specifics: "does LiveKit support X", "what changed in agents 1.8", "what are the arg

Scan passed 0
testing-livekit-agents

Writes turn-level tests for a LiveKit agent in the user''s normal test suite: pytest (Python) or Vitest (Node.js). Use when the user asks to "write tests for my agent", "add a test for this tool", "test the handoff", "pin this bug", "why does my agent test fail", or after building or changing agent

Scan passed 0
writing-livekit-scenarios

Creates and maintains the scenarios a LiveKit agent simulation runs, and wires the agent to consume them. Use when the user asks "what should I test", "generate simulation scenarios", "write scenarios for my agent", "add a scenario for X", "organize my scenario files", "my simulations are flaky", "t

Scan passed 0

Related devops skillsscan passed

adapter-aws-lambda

Deploy tRPC on AWS Lambda with awsLambdaRequestHandler() from @trpc/server/adapters/aws-lambda for API Gateway v1 (REST, APIGatewayProxyEvent) and v2 (HTTP, APIGatewayProxyEventV2), and Lambda Function URLs. Enable response streaming with awsLambdaStreamingRequestHandler() wrapped in awslambda.strea

Scan passed 0
observability-and-instrumentation

Instruments code so production behavior is visible and diagnosable. Use when adding logging, metrics, tracing, or alerting. Use when shipping any feature that runs in production and you need evidence it works. Use when production issues are reported but you can't tell what happened from the availabl

Scan passed 0
recursive-decision-ledger

Run repeated rollouts ("Prime Gauss" style recursive prompting) while keeping an append-only decision ledger of trials, marks, coherence checks, and promotion gates, so recursive confidence never auto-approves live trading, deploy, or destructive actions. Use when the user asks for repeated rollouts

Scan passed 0
firebase-app-hosting-basics

Deploys and manages full-stack web applications (Next.js, Angular) with Server-Side Rendering (SSR) using Firebase App Hosting. Use when deploying Next.js/Angular apps, configuring apphosting.yaml or firebase.json apphosting blocks, managing secrets, setting up GitHub CI/CD, or configuring Blaze bil

Scan passed 0
writing-skills

Use when creating new skills, editing existing skills, or verifying skills work before deployment

Scan passed 0
resilience-hub-failure-mode-assessment

Runs and interprets AWS Resilience Hub v2 failure mode assessments. Covers starting assessments, understanding findings (severity, categories, recommendations), triaging by achievability, working with AI-generated service functions, and resolving findings. Applies when the user wants to run an asses

Scan passed 0