Text-to-speech
Text-to-speech (TTS) turns an assistant’s reply into audio. Open Settings → Providers → Text-to-Speech as owner/admin, enable it, and choose a provider. This is separate from your chat and speech-to-text models.
Choose a provider
| Option | Where processing runs | Voices |
|---|---|---|
| OpenAI-compatible | At the speech endpoint you configure | The endpoint’s supported voice names |
| ElevenLabs | Cloud | Your available voices and voice clones |
| Piper | On the OE server | Downloadable voice catalog; some voices have multiple speakers |
| KittenTTS | On the OE server, CPU | Eight preset voices |
| Pocket TTS | On the OE server, CPU | Local voices and voice cloning from a clip |
Cloud TTS receives the reply text you ask it to speak. Local engines download models during installation, then synthesize locally. The installer buttons use Linux systemd user services and are not supported inside the single-container Docker deployment. A separately hosted compatible endpoint is another option.
Set up remote speech
Select OpenAI-compatible and enter the full speech URL, key, supported model, and default voice. For ElevenLabs, enter its key and choose a voice available to that account. Save, then use the speech preview to check the result. ElevenLabs also has a Speech pace control.
Use the endpoint’s current model and voice choices; an example voice from another service is not guaranteed to work on it.
Set up Piper
- Select Piper and choose an Initial voice to download.
- Choose Install Piper and wait for the service status to confirm it runs.
- Use the voice catalog to download additional voices. Newly installed voices become available without restarting the service.
- Preview a voice. Use Speech pace to adjust delivery; it saves when you finish moving the slider.
Voice slots show the installed voice names. Multi-speaker voices also have a speaker-ID choice. Older numeric-only choices refer to the LibriTTS-R voice.
Set up KittenTTS or Pocket TTS
Select the provider and use its install button. Wait for the service status, then choose and preview a voice. KittenTTS supplies presets. Pocket TTS also lets you upload a reference voice clip through its voice controls; use audio you have permission to use. Review the upload result and test before assigning it to a device.
Choose voices for wake words
Open Voice devices → Voice config. Pick the voice for each wake-word slot, then choose Save & Push to send the saved settings to the paired devices. A slot’s voice overrides the provider’s global default. Offline devices receive the saved configuration when they reconnect.
Test and troubleshoot
First use the provider/voice preview, then test one paired device. If preview works but the device is silent, check its volume, connection, slot, and Headphone / line-out mode. Use Voice diagnostics to separate speech generation from device playback.
For a remote 401/403, check the key and access to the selected voice. For a local install error, read the progress panel and confirm that the host supports systemd user services. For choppy playback, check device connection/audio health before changing the speech model. Hardware and voice choice affect latency; use the measured timings from your own installation.