Inworld AI is a real-time AI infrastructure provider for consumer applications. Its stack includes text-to-speech, speech-to-text, speech-to-speech, LLM routing, and inference. It is not primarily a visual avatar renderer. Teams can use Inworld for the voice and reasoning pipeline, then add a separate visual avatar layer when the product needs a face.
Last verified: September 15, 2026. Confirm live pricing before purchase.
What is Inworld AI in 2026?
Inworld’s current company definition describes an inference provider for real-time AI at consumer scale. It supplies speech models, access to language models, and APIs for live conversational experiences.
That differs from the older perception of Inworld as an AI NPC builder. Gaming remains part of its footprint, but the current homepage emphasizes infrastructure for companions, education, wellness, agentic work, and games. Buyers should assess speech, routing, inference, and unit economics—not rely on an old NPC review.
What does the Inworld stack include?
| Layer | Current Inworld product | What it handles |
|---|---|---|
| Voice output | Realtime TTS-2 and TTS-2 Flash | Streams generated speech, voice design, and cloning |
| Voice input | Realtime STT | Streams transcription with turn and acoustic signals |
| Live conversation | Realtime API | Connects speech in and speech out over WebSocket or WebRTC |
| Reasoning access | Realtime Router | Routes requests across 220+ models and supports conditional routing |
| Compute | Realtime inference and dedicated GPUs | Runs selected models on managed infrastructure |
Inworld says its Realtime API supports full-duplex audio, tool calling, configurable turn-taking, and provider-agnostic model choice. The Realtime Router follows the OpenAI SDK pattern and charges third-party LLM usage at provider rates without an added percentage markup. Teams should still verify model availability, region, rate limits, and feature support for their account.
Inworld AI pricing
Inworld pricing uses dollar-denominated credits across products. A paid plan adds credits equal to the monthly plan price, and usage is deducted at that tier’s TTS, STT, and LLM rates.
| Plan | Monthly commitment | Realtime TTS-2 | TTS-2 Flash | Realtime STT |
|---|---|---|---|---|
| On-Demand | Start free | $25 / 1M characters | $15 / 1M | $0.15/hour |
| Creator | $25 | $20 / 1M | $10 / 1M | $0.10/hour |
| Builder | $100 | $17.50 / 1M | $9 / 1M | $0.10/hour |
| Developer | $300 | $15 / 1M | $8 / 1M | $0.10/hour |
| Growth | $1,500 | $12.50 / 1M | $7 / 1M | $0.10/hour |
| Enterprise | Custom | As low as $5 / 1M | Sub-$5 / 1M | Custom |
LLMs are listed “at cost,” but the underlying model rate still varies. Inworld also offers annual billing, volume commitments, and dedicated compute. Estimate the complete session: incoming transcription, language-model tokens, outgoing speech, retries, idle time, and any visual layer.
Who is Inworld AI best for?
Inworld fits teams with substantial real-time voice usage that want speech, LLM routing, and inference under one relationship. It also suits model experimentation, failover, and consumer-scale unit economics.
It is a weaker fit if the only need is prerecorded avatar video or if the requirement is a complete photorealistic avatar renderer. Those are different product categories. The AI avatar API versus video-generation API guide explains the output difference.
How do you add a visual avatar to Inworld?
A visual avatar should sit after the conversational intelligence and speech layers. Inworld can receive speech, route the request to an LLM, and synthesize the response audio. A dedicated avatar layer can consume that audio, calculate synchronized facial motion, and render the digital human in the client.
Spatius is designed for this separated architecture: audio goes to the Motion Server, compact motion data returns, and the user’s device renders the pixels. This lets the product keep its Inworld voice or model choices while evaluating avatar minutes, concurrency, bandwidth, and device support independently. It is an architectural pairing to prototype, not a claim of a prebuilt Inworld connector.
Read where the avatar fits in an agent architecture before assigning ownership. If the transport layer is LiveKit, the Spatius LiveKit integration guide shows the documented audio-to-avatar pattern.
What should a technical buyer test?
- Measure time to first transcript, first audio, and first visible motion separately.
- Test interruptions, background noise, reconnects, and long sessions.
- Confirm which LLMs and regions are available at the intended tier.
- Model TTS characters, STT hours, LLM tokens, and visual-avatar minutes.
- Check data retention, consent, voice-cloning rights, rate limits, and support terms.
- Test on the browsers, phones, kiosks, or embedded devices customers use.
Inworld AI FAQ
Is Inworld AI only for games and NPCs?
No. Its current positioning covers real-time AI infrastructure for consumer apps, including companions, education, health and wellness, agentic work, and games.
Does Inworld provide a visual AI avatar?
Its current core offer focuses on speech, real-time conversation, LLM routing, and inference. A product needing a rendered digital human should evaluate a separate visual avatar layer.
How much does Inworld AI cost?
On-Demand starts free. Paid monthly tiers are Creator at $25, Builder at $100, Developer at $300, Growth at $1,500, and custom Enterprise. Usage consumes plan credits at tier-specific rates.
Can Inworld work with an existing LLM?
Yes. Inworld’s Router and Realtime API are model-agnostic and list access to OpenAI, Anthropic, Google, and many other providers. Verify the required model and region in the current directory.
Add the visual layer without replacing the voice stack
Inworld can be the real-time voice and reasoning infrastructure while Spatius provides the interactive face.
Give your existing voice AI stack a client-rendered, real-time visual interface. Request a demo, or ,或Review Spatius pricing.。