Skip to content

OpenAI Realtime API Tutorial: How to Add a Real-Time AI Avatar

Microphone audio flowing into a real-time AI avatar in an OpenAI Realtime API architecture

The cleanest way to add an avatar to an OpenAI Realtime API voice agent is to keep OpenAI responsible for the conversation and send only the assistant’s output audio to a separate avatar layer. Spatius can turn that speech audio into motion data while AvatarKit renders the avatar locally. The agent, tools, prompts, and turn-taking remain in your OpenAI application.

Independent guide: Spatius is not affiliated with, endorsed by, or an official integration partner of OpenAI. This article describes an interoperability pattern based on public APIs and documentation, not a packaged native connector.

Last verified: September 22, 2026.

What each system should own

The OpenAI Realtime API supports live audio conversations, session state, interruptions, and tool calls. It is the voice-agent layer. It does not need to become the avatar renderer.

Spatius Direct Mode starts at the opposite boundary: the application already has approved avatar speech audio, and AvatarKit sends that audio to Motion Server, receives motion data, and renders the avatar locally.

ResponsibilityOpenAI Realtime API applicationSpatius
User microphone and conversation sessionOwnsDoes not own
Agent instructions, tools, permissions, and business logicOwnsDoes not own
Assistant speech audioProducesReceives as avatar-driving input
Facial motionDoes not need to ownMotion Server generates motion data
Avatar renderingDoes not ownAvatarKit renders locally on the client
Product UI, analytics, and human handoffOwnsDoes not own

This separation avoids rebuilding a working voice agent merely to add a face.

Reference architecture

The central rule is simple: the same assistant output audio that the user hears should drive the avatar.

User microphone
    ↓
OpenAI Realtime session
    ├─ conversation state, tools, interruptions
    └─ assistant output audio
              ↓
      application audio handoff
              ↓
      Spatius Motion Server
              ↓ motion data
      AvatarKit on the client
              ↓
      locally rendered avatar

Do not send the user’s microphone audio to the avatar layer. The avatar should move only for the assistant’s approved speech. Prompts, transcripts, tool arguments, customer records, and internal agent state do not need to cross the avatar boundary.

Choose Direct Mode or Backend Mode

Spatius documents multiple integration paths. The right path depends on where your application can access the assistant audio.

Audio ownershipRecommended Spatius pathArchitecture consequence
Browser or mobile client receives the assistant audioDirect ModeClient sends avatar speech audio to Motion Server and renders with AvatarKit
Trusted backend receives and controls the assistant audioBackend ModeBackend connects through a Spatius Server SDK and delivers audio/motion payloads to clients
Existing LiveKit Agents applicationDocumented LiveKit integrationUse the packaged LiveKit path rather than inventing a parallel connector

For a browser Realtime application, OpenAI recommends WebRTC for client connections. The generated voice arrives as remote media. Direct Mode is the natural starting point when the client can access the output audio in the format required by the current Spatius audio documentation.

If your OpenAI session runs over a trusted server connection and the backend owns output audio chunks, Backend Mode creates a clearer security boundary. It also makes the backend responsible for delivery, recovery, and synchronization.

Step 1: establish the Realtime voice session

Follow OpenAI’s current Realtime getting-started guide before adding the avatar. For browser and mobile clients, keep permanent project credentials on the backend and use the current short-lived client authorization flow documented by OpenAI.

At this stage, confirm four behaviors without an avatar:

  1. the user can speak and hear the assistant;
  2. tool calls and business rules work;
  3. the user can interrupt the assistant;
  4. session close and reconnect behavior are observable.

Adding an avatar before the voice path is stable makes failures harder to classify.

Step 2: initialize the avatar independently

Create the Spatius session on a separate product boundary. Keep the Spatius API Key on your backend and return only the short-lived Session Token and public identifiers required by the client. The Direct Mode client guide documents the lifecycle: initialize, load and mount the avatar, connect, send avatar speech audio, then close and dispose.

The avatar should be ready before the first assistant speech begins. Otherwise the user may hear the answer while the character is still loading.

Step 3: hand off assistant audio once

Treat OpenAI’s assistant output as the single source of truth. Your application needs one controlled handoff that makes the same speech available to playback and avatar motion.

on assistant audio:
  preserve chunk order and timestamps
  play or route the approved output once
  send the matching avatar-speech audio to Spatius

on assistant response complete:
  finalize the avatar-speech turn

on user interruption:
  stop assistant playback
  stop or clear the corresponding avatar turn

This is intentionally pseudocode. The exact audio events, media-track access, encoding, and AvatarKit calls depend on the current OpenAI connection method and Spatius SDK version. Use the respective official references instead of copying unverified method names from a blog.

The common failure is duplicate playback: one path plays the OpenAI remote track while a second path plays a copied buffer. Design one authoritative speaker path and use the other branch only for motion generation.

Step 4: synchronize interruption and completion

A voice agent feels broken if the user interrupts but the avatar keeps talking. OpenAI’s Realtime conversation guide describes the session events around speech and response completion. Your avatar integration should map the same lifecycle into visual playback.

Test these transitions explicitly:

  • normal response completion;
  • user barge-in during the first audio chunk;
  • user barge-in near the end of a response;
  • tool call followed by speech;
  • network interruption while speech is active;
  • reconnect after the avatar or voice session fails independently.

Do not infer visual state from transcripts. Audio playback state is the relevant source because the avatar must match what the user actually hears.

Step 5: test the two layers separately

Measure the voice and avatar layers independently before reporting an end-to-end number.

MeasurementVoice-agent layerAvatar layer
Time to first assistant audioOpenAI session and applicationNot included
Audio gaps or turn errorsOpenAI/application transportNot included
Audio-to-motion delayNot includedSpatius path
Client render frame stabilityNot includedAvatarKit and target device
Interruption correctnessShared product behaviorMust stop both layers together
Total costOpenAI usage and application infrastructureSpatius plan and avatar usage

OpenAI publishes model pricing in its current model documentation; Spatius publishes separate avatar pricing. Do not present their combined cost as one universal per-minute rate without a workload model.

When this architecture is a good fit

Use this pattern when the OpenAI Realtime agent is already the product’s conversational system and the team wants to add a visual presence without replacing prompts, tools, permissions, analytics, or voice behavior.

It is less suitable when the team wants one vendor to own the complete agent and cloud-rendered video experience. In that case, a bundled avatar platform may reduce integration work, although it also changes vendor ownership and operating economics.

For the broader decision, see how to add an avatar to an existing SaaS AI agent and the real-time avatar pricing comparison.

OpenAI Realtime avatar FAQ

Is this an official OpenAI and Spatius integration?

No. This is an independent interoperability guide based on public interfaces. Spatius does not claim an OpenAI partnership or packaged OpenAI connector.

Does Spatius replace the OpenAI Realtime API?

No. OpenAI owns the live conversational session, assistant behavior, tools, and generated speech. Spatius uses approved assistant speech audio to generate avatar motion and render the character locally.

Should the browser use WebRTC or WebSocket for OpenAI Realtime?

OpenAI recommends WebRTC for browser and mobile client connections. Server-side architectures may use the currently documented server connection options. Choose the connection before selecting Direct or Backend Mode.

What data must be sent to Spatius?

The avatar layer needs the assistant speech audio and the identifiers and credentials required for the Spatius session. Prompts, user transcripts, tool arguments, and internal agent state do not need to be sent merely to animate the avatar.

What should a proof of concept verify?

Verify audio format, ordered streaming, avatar readiness, normal completion, interruption, reconnects, target-device rendering, and separate cost and latency measurements for the agent and avatar layers.

Add a face without replacing the agent

The OpenAI Realtime API can remain the voice and reasoning system. Spatius can remain the visual embodiment layer. Keeping that boundary explicit makes the architecture easier to test, secure, and change.

Bring your Realtime voice-agent architecture and target client. We will help you evaluate the Spatius Direct or Backend Mode without replacing the agent stack. Request a demo, or ,或Review Spatius pricing.。

Give your agent a face that responds.

Start building