Deepgram can run the listening, reasoning, and speaking loop for a voice agent; Spatius can add the visual layer. Route Deepgram’s streamed agent audio to both playback and the Spatius Motion Server, then render the returned motion data with AvatarKit. Your application still owns session state, tools, interruption policy, and recovery.
Independent guide: Spatius is not affiliated with, endorsed by, or an official integration partner of Deepgram. This article describes an interoperability pattern based on public APIs and documentation, not a packaged native connector.
Last verified: September 30, 2026.
What each system should own
Deepgram’s Voice Agent API combines speech-to-text, LLM orchestration, and text-to-speech over one WebSocket. It receives user audio, produces transcripts and agent events, and returns synthesized agent speech as binary audio.
Spatius Direct Mode starts after the agent has approved what it will say. AvatarKit sends that speech audio to Motion Server, receives motion data, and renders the avatar locally.
| Responsibility | Deepgram Voice Agent application | Spatius |
|---|---|---|
| User speech and turn detection | Owns | Does not own |
| LLM, prompts, tools, and business rules | Owns | Does not own |
| Agent speech | Produces output audio | Receives approved avatar speech audio |
| Facial motion | Does not need to own | Motion Server generates motion data |
| Avatar rendering | Does not own | AvatarKit renders on the client |
| Product UI, fallback, analytics, and handoff | Application owns | Does not own |
This page is intentionally different from the existing Deepgram speech-to-text integration. That integration uses Deepgram as one listening component. This guide starts with Deepgram’s complete Voice Agent API and routes its final output speech into a separate avatar layer.
Reference architecture
The same agent audio should drive what the user hears and what the avatar presents.
User microphone
↓
Deepgram Voice Agent WebSocket
├─ STT, LLM orchestration, tools, turn events
└─ binary agent output audio
↓
ordered application audio queue
├─ user playback
└─ Spatius Motion Server
↓ motion data
AvatarKit on the client
↓
locally rendered avatar
Do not synthesize a second response for the avatar or animate from the final transcript alone. A second TTS path can change pronunciation, pauses, or duration. The audio the user actually hears is the correct synchronization source.
For the wider product boundary, see how to add an avatar to an existing SaaS agent and the Spatius integration catalog.
Step 1: establish the Deepgram agent connection
Deepgram documents one Voice Agent WebSocket at wss://agent.deepgram.com/v1/agent/converse. Its message-flow guide requires the client to wait for Welcome, send a Settings message, and wait for SettingsApplied before streaming microphone audio.
The settings define input and output formats plus the Listen, Think, and Speak providers. Deepgram’s current example uses raw linear16 output at 24 kHz. Treat that as an example, not a universal Spatius requirement. Confirm the audio formats supported by the current SDKs before implementation.
Get the voice agent working without an avatar first. Verify normal turns, tools, interruption, session close, and reconnect behavior so later failures can be assigned to the correct layer.
Step 2: fan out output audio once
Deepgram returns agent speech as ordered binary audio frames. Put those frames through one controlled queue, then branch them to playback and the avatar adapter.
on Deepgram agent audio:
preserve order and timing
route the approved speech to user playback
send the matching speech to the Spatius avatar path
on agent audio complete:
finalize the avatar-speech turn
If playback needs resampling, buffering, or a media container, perform that in a dedicated branch. Do not let one consumer mutate the source stream for the other. Keep enough sequence and session information to discard stale audio after an interruption or reconnect.
The exact AvatarKit methods depend on the current SDK. Use the official Spatius audio concepts and lifecycle documentation rather than copying unverified method names from a blog.
Step 3: coordinate barge-in across audio and motion
Deepgram emits UserStartedSpeaking when the user begins talking over the agent. Its official guide tells clients to stop audio playback immediately. A visual agent must also stop presenting the cancelled response.
On UserStartedSpeaking | Required product action |
|---|---|
| Audio player | Stop and discard queued agent audio |
| Avatar adapter | Stop forwarding cancelled speech and clear pending work |
| Avatar UI | Return to listening or idle rather than finishing stale motion |
| Tool workflow | Cancel only work the product policy marks as cancellable |
| Logs | Preserve the interrupted turn and event timing |
Stopping the speaker does not automatically clear the avatar. Your application owns the coordinated state change. Use the product patterns in when users should be able to interrupt an AI avatar.
Step 4: preserve an audio-only fallback
The voice agent should remain usable if avatar initialization, asset loading, motion delivery, or client rendering fails. Stop stale animation, surface a clear status, and continue in audio-only mode when the workflow allows it.
This fallback also gives observability a useful boundary: Deepgram may be healthy while the avatar path is degraded, or the avatar may be ready while the voice session has failed. Record Deepgram’s server events alongside the playback queue, avatar session, and client rendering state.
Step 5: test the complete turn
Measure from the end of user speech to the first understandable agent audio and the first correct avatar movement. Test a long response, interruption during the first audio chunk, interruption near completion, tool success, tool failure, reconnect, and avatar-only failure.
Do not publish one latency number unless the region, device, network, models, voice, sample size, and measurement boundary are documented. Deepgram’s Voice Agent product page and Spatius pricing describe different parts of the stack; total cost must include both plus the application infrastructure.
When this architecture is a good fit
Use this pattern when Deepgram should own the unified voice loop, your product should retain tools and user experience, and the avatar should remain a replaceable presentation layer.
Choose a bundled cloud-video avatar agent when the shortest path to a hosted persona matters more than preserving a modular voice-agent architecture. Choose the component approach when stack control, client rendering, device reach, or the ability to change voice and avatar providers matters more.
Deepgram real-time avatar FAQ
Does Deepgram Voice Agent API include an avatar?
No. It provides the real-time voice-agent pipeline. A visual layer such as Spatius must consume the agent’s final speech audio and render the avatar separately.
Can the same Deepgram audio drive playback and lip sync?
Yes. Route the same ordered agent-audio stream to both consumers. Avoid generating a second TTS response for the avatar because timing and pronunciation can diverge.
Who owns interruption behavior?
Deepgram detects user speech and emits events, but the application must stop playback, clear cancelled avatar work, update UI state, and apply its tool-cancellation policy.
Is this the same as a Deepgram STT integration?
No. A Deepgram STT integration uses Deepgram for transcription inside a multi-provider pipeline. This architecture uses the unified Voice Agent API for listening, reasoning orchestration, and speaking, then adds Spatius as the visual output layer.
Does Spatius need the user microphone stream?
The avatar path needs the approved agent speech it must present, not the user’s raw microphone audio. Keep user audio and product data inside the systems that actually require them.
Add a visual layer to your Deepgram agent
Bring one Deepgram Voice Agent conversation, the chosen audio format, target devices, interruption requirements, and expected concurrency. We will help you evaluate the avatar boundary without replacing the voice agent. Request a demo, or ,或Explore integrations.。