Skip to content

Deepgram Voice Agent API: How to Add a Real-Time AI Avatar

Deepgram Voice Agent audio driving a locally rendered Spatius real-time avatar

Deepgram can run the listening, reasoning, and speaking loop for a voice agent; Spatius can add the visual layer. Route Deepgram’s streamed agent audio to both playback and the Spatius Motion Server, then render the returned motion data with AvatarKit. Your application still owns session state, tools, interruption policy, and recovery.

Independent guide: Spatius is not affiliated with, endorsed by, or an official integration partner of Deepgram. This article describes an interoperability pattern based on public APIs and documentation, not a packaged native connector.

Last verified: September 30, 2026.

What each system should own

Deepgram’s Voice Agent API combines speech-to-text, LLM orchestration, and text-to-speech over one WebSocket. It receives user audio, produces transcripts and agent events, and returns synthesized agent speech as binary audio.

Spatius Direct Mode starts after the agent has approved what it will say. AvatarKit sends that speech audio to Motion Server, receives motion data, and renders the avatar locally.

ResponsibilityDeepgram Voice Agent applicationSpatius
User speech and turn detectionOwnsDoes not own
LLM, prompts, tools, and business rulesOwnsDoes not own
Agent speechProduces output audioReceives approved avatar speech audio
Facial motionDoes not need to ownMotion Server generates motion data
Avatar renderingDoes not ownAvatarKit renders on the client
Product UI, fallback, analytics, and handoffApplication ownsDoes not own

This page is intentionally different from the existing Deepgram speech-to-text integration. That integration uses Deepgram as one listening component. This guide starts with Deepgram’s complete Voice Agent API and routes its final output speech into a separate avatar layer.

Reference architecture

The same agent audio should drive what the user hears and what the avatar presents.

User microphone
    ↓
Deepgram Voice Agent WebSocket
    ├─ STT, LLM orchestration, tools, turn events
    └─ binary agent output audio
              ↓
      ordered application audio queue
          ├─ user playback
          └─ Spatius Motion Server
                    ↓ motion data
            AvatarKit on the client
                    ↓
            locally rendered avatar

Do not synthesize a second response for the avatar or animate from the final transcript alone. A second TTS path can change pronunciation, pauses, or duration. The audio the user actually hears is the correct synchronization source.

For the wider product boundary, see how to add an avatar to an existing SaaS agent and the Spatius integration catalog.

Step 1: establish the Deepgram agent connection

Deepgram documents one Voice Agent WebSocket at wss://agent.deepgram.com/v1/agent/converse. Its message-flow guide requires the client to wait for Welcome, send a Settings message, and wait for SettingsApplied before streaming microphone audio.

The settings define input and output formats plus the Listen, Think, and Speak providers. Deepgram’s current example uses raw linear16 output at 24 kHz. Treat that as an example, not a universal Spatius requirement. Confirm the audio formats supported by the current SDKs before implementation.

Get the voice agent working without an avatar first. Verify normal turns, tools, interruption, session close, and reconnect behavior so later failures can be assigned to the correct layer.

Step 2: fan out output audio once

Deepgram returns agent speech as ordered binary audio frames. Put those frames through one controlled queue, then branch them to playback and the avatar adapter.

on Deepgram agent audio:
  preserve order and timing
  route the approved speech to user playback
  send the matching speech to the Spatius avatar path

on agent audio complete:
  finalize the avatar-speech turn

If playback needs resampling, buffering, or a media container, perform that in a dedicated branch. Do not let one consumer mutate the source stream for the other. Keep enough sequence and session information to discard stale audio after an interruption or reconnect.

The exact AvatarKit methods depend on the current SDK. Use the official Spatius audio concepts and lifecycle documentation rather than copying unverified method names from a blog.

Step 3: coordinate barge-in across audio and motion

Deepgram emits UserStartedSpeaking when the user begins talking over the agent. Its official guide tells clients to stop audio playback immediately. A visual agent must also stop presenting the cancelled response.

On UserStartedSpeakingRequired product action
Audio playerStop and discard queued agent audio
Avatar adapterStop forwarding cancelled speech and clear pending work
Avatar UIReturn to listening or idle rather than finishing stale motion
Tool workflowCancel only work the product policy marks as cancellable
LogsPreserve the interrupted turn and event timing

Stopping the speaker does not automatically clear the avatar. Your application owns the coordinated state change. Use the product patterns in when users should be able to interrupt an AI avatar.

Step 4: preserve an audio-only fallback

The voice agent should remain usable if avatar initialization, asset loading, motion delivery, or client rendering fails. Stop stale animation, surface a clear status, and continue in audio-only mode when the workflow allows it.

This fallback also gives observability a useful boundary: Deepgram may be healthy while the avatar path is degraded, or the avatar may be ready while the voice session has failed. Record Deepgram’s server events alongside the playback queue, avatar session, and client rendering state.

Step 5: test the complete turn

Measure from the end of user speech to the first understandable agent audio and the first correct avatar movement. Test a long response, interruption during the first audio chunk, interruption near completion, tool success, tool failure, reconnect, and avatar-only failure.

Do not publish one latency number unless the region, device, network, models, voice, sample size, and measurement boundary are documented. Deepgram’s Voice Agent product page and Spatius pricing describe different parts of the stack; total cost must include both plus the application infrastructure.

When this architecture is a good fit

Use this pattern when Deepgram should own the unified voice loop, your product should retain tools and user experience, and the avatar should remain a replaceable presentation layer.

Choose a bundled cloud-video avatar agent when the shortest path to a hosted persona matters more than preserving a modular voice-agent architecture. Choose the component approach when stack control, client rendering, device reach, or the ability to change voice and avatar providers matters more.

Deepgram real-time avatar FAQ

Does Deepgram Voice Agent API include an avatar?

No. It provides the real-time voice-agent pipeline. A visual layer such as Spatius must consume the agent’s final speech audio and render the avatar separately.

Can the same Deepgram audio drive playback and lip sync?

Yes. Route the same ordered agent-audio stream to both consumers. Avoid generating a second TTS response for the avatar because timing and pronunciation can diverge.

Who owns interruption behavior?

Deepgram detects user speech and emits events, but the application must stop playback, clear cancelled avatar work, update UI state, and apply its tool-cancellation policy.

Is this the same as a Deepgram STT integration?

No. A Deepgram STT integration uses Deepgram for transcription inside a multi-provider pipeline. This architecture uses the unified Voice Agent API for listening, reasoning orchestration, and speaking, then adds Spatius as the visual output layer.

Does Spatius need the user microphone stream?

The avatar path needs the approved agent speech it must present, not the user’s raw microphone audio. Keep user audio and product data inside the systems that actually require them.

Add a visual layer to your Deepgram agent

Bring one Deepgram Voice Agent conversation, the chosen audio format, target devices, interruption requirements, and expected concurrency. We will help you evaluate the avatar boundary without replacing the voice agent. Request a demo, or ,或Explore integrations.。

Give your agent a face that responds.

Start building