Skip to content

Integrating AI Avatars with LiveKit: A Developer Guide to the Spatius LiveKit Plugin

If you’re building a voice AI application with LiveKit Agents, adding a real-time avatar face is one of the highest-leverage improvements you can make to perceived engagement. The technical barrier is lower than it looks — if your voice pipeline already runs on LiveKit, the Spatius LiveKit Plugin drops in as a single dependency with minimal configuration.

This guide covers how the integration works, how to set it up, and what to watch for in production.

Why LiveKit + AI Avatar Is a Natural Pairing

LiveKit is a popular open-source WebRTC framework used by teams building real-time voice and video applications. LiveKit Agents — the Python framework for building AI voice agents — handles the ASR, LLM, and TTS orchestration that powers conversational AI. Spatius handles what LiveKit Agents doesn’t: driving an avatar’s face in real time from the audio stream that the voice agent produces.

The integration works cleanly because of how Spatius is architectured. Spatius doesn’t replace or intercept your voice pipeline — it observes the audio your TTS outputs and keeps the avatar’s Motion data in sync. Your LLM, your voice models, your session logic: all of that stays unchanged. Spatius adds a face.

Why not just stream video from the cloud? For the same reason you’d choose LiveKit over a hosted video conferencing product: control, latency, and cost. Cloud-rendered avatar video requires substantially more bandwidth than Spatius’s client-rendered path. The documented Spatius Motion-data rate is approximately 10–15 KB/s because only driving data is transmitted, not rendered video frames. The avatar renders in the browser using WebGL or WebGPU.

Architecture Overview

User Speech
    ↓
[LiveKit Room] → ASR → LLM → TTS (your stack)
                                    ↓
                            Spatius Plugin
                                    ↓
                    [Cloud: Motion Server]
                                    ↓
                          Motion data (~10–15 KB/s)
                                    ↓
                        [Browser: 3DGS avatar renders locally]

The Spatius LiveKit Plugin hooks into the TTS audio output in your agent and transmits it to Spatius Motion Server, which generates Motion data. Motion data is streamed to the client SDK in the browser, which renders the 3DGS avatar model in real time using WebGL or WebGPU. The avatar’s lip sync and expression are driven by Motion data, not by video streaming.

Important scope note: the LiveKit Plugin currently supports Web only. For iOS and Android deployments, use Direct Mode or Backend Mode with the platform-native SDKs (AvatarKit.xcframework for iOS, Gradle ai.spatius:avatarkit for Android).

Prerequisites

Before starting:

  • A LiveKit Cloud account or self-hosted LiveKit server
  • An existing LiveKit Agents voice pipeline (Python)
  • A Spatius account — the free tier includes 500 credits/month (~50 minutes), which is enough for development and testing
  • Node.js environment for the browser client (the Spatius client SDK is distributed as @spatius/avatarkit-rtc on npm)

For the packaged LiveKit client path, the browser needs the Avatar ID plus short-lived LiveKit credentials (server URL, room name, and token). Do not expose a Spatius API Key, LiveKit API secret, or Agora App Certificate in browser code; this path does not use a Spatius Session Token.

Server Side: Python Setup

Install the Spatius LiveKit plugin for Python:

pip install livekit-plugins-spatius

In your LiveKit Agents worker, create an AvatarSession after the voice pipeline is ready:

from livekit.plugins.spatius import AvatarSession

avatar = AvatarSession()
await avatar.start(session, room=ctx.room)

The plugin receives the agent’s audio, forwards it to Spatius Motion Server, and publishes synchronized audio and motion data for the client-side AvatarKit renderer.

The RTC adapter declares livekit-client ^2.15.9 as its supported peer range. It is a compatibility range, not an exact patch-level pin:

{
  "dependencies": {
    "@spatius/avatarkit": "latest",
    "@spatius/avatarkit-rtc": "latest",
    "livekit-client": "^2.15.9"
  }
}

Browser Client Setup

Install the client SDK:

npm install @spatius/avatarkit @spatius/avatarkit-rtc livekit-client

Initialize the avatar in your browser application:

import { AvatarSDK, AvatarManager, AvatarView, DrivingServiceMode } from '@spatius/avatarkit';
import { AvatarPlayer, LiveKitProvider } from '@spatius/avatarkit-rtc';

await AvatarSDK.initialize('your_app_id', {
  drivingServiceMode: DrivingServiceMode.rtc,
});

const avatar = await AvatarManager.load('your_avatar_id');
const avatarView = new AvatarView(avatar, document.getElementById('avatar-container'));
const player = new AvatarPlayer(new LiveKitProvider(), avatarView);

await player.connect({
  url: livekitServerUrl,
  token: livekitToken,
  roomName: livekitRoomName,
});
player.startPublishing();

The SDK handles WebGL/WebGPU initialization, avatar loading, and the Motion data rendering loop. The avatar view renders inside the container element you provide.

Demo Repository

A full working demo — including server-side Python agent and browser client — is available at:

The demo uses Direct Mode by default. The LiveKit Plugin setup is documented in the demo’s README under the livekit-plugin branch.

Full API reference: docs.spatius.ai

Latency: LiveKit Plugin vs Direct Mode

The LiveKit Plugin exists specifically to reduce end-to-end avatar latency for web deployments.

Direct Mode connects the client directly to Motion Server with a Session Token. Its observed latency depends on the selected region, audio format, network, and the rest of your voice stack; measure the path you plan to ship.

LiveKit Agents fits into the LiveKit-based flow your agent is already managing. This reduces integration overhead for web deployments that already use LiveKit Agents. Spatius supports us-west, ap-northeast, and cn-beijing, with automatic selection when no region is configured.

If you’re already using LiveKit Agents and targeting web, the LiveKit Plugin is the recommended path. If you need mobile support (iOS/Android) or prefer a simpler integration without LiveKit, Direct Mode is the right choice.

Rendering on the Browser: What’s Happening

The browser SDK renders the avatar using WebGL or WebGPU (the SDK prefers WebGPU when available and falls back to WebGL). Avatar assets are loaded by the SDK and rendered locally; exact asset size and rendering performance depend on the avatar and target device.

Motion data is approximately 10–15 KB/s according to the developer docs. Rather than streaming rendered video frames from the cloud, the SDK streams motion data and renders locally.

For browser deployments targeting shared or public kiosks, this architecture also has a privacy benefit: the avatar renders locally, the avatar layer receives the AI-generated audio used for facial driving, while transcripts, prompts, and other sensitive user data can remain in your own voice stack. The 3DGS model is cached after first load.

Handling the Avatar Session Token

The Session Token is short-lived and must be generated server-side. Don’t embed your API Key in client-side code.

Your server generates a Session Token by calling the Spatius API with your App ID and API Key, and returns only the token to the browser. The browser uses the token with the AvatarKit SDK. This is a standard pattern; the demo repository includes server-side token generation examples in Python and Go.

For Go server SDK: github.com/spatius-ai/spatius-sdk-go
For Python SDK: pip install spatius — github.com/spatius-ai/spatius-sdk-python

Common Issues

Avatar renders but lip sync is out of sync: check the configured session audio format and sample rate. Web and native SDKs support mono PCM16 and Opus; the LiveKit plugin uses Ogg Opus by default. The SDK does not auto-resample, so confirm that the source format matches the session configuration.

WebSocket connection fails: check the selected integration path, credentials, and region. Supported regions are us-west, ap-northeast, and cn-beijing; automatic region selection is available when unset.

Avatar doesn’t appear on iOS or Android: the LiveKit Plugin is Web-only. For mobile, switch to Direct Mode with the platform-native SDKs. Note that iOS rendering requires a real device — the iOS Simulator doesn’t support Metal rendering.

livekit-client version mismatch: if you see connection errors or audio issues after installing dependencies, run npm list livekit-client and confirm it satisfies the adapter’s ^2.15.9 peer range.

Try It Before You Integrate

The fastest path to evaluating Spatius before integrating the SDK is the Playground — a live session with a Spatius avatar running in your browser, no signup required: spatius.ai/playground

For a broader overview of how to evaluate any real-time avatar platform before committing to an integration: Avatar SDK Demo: How to Test a Real-Time AI Avatar Before You Commit to a Platform


Give your agent a face that responds.

Start building