Speech-to-text
Speech-to-text (STT) converts a voice recording into the text your assistant reads. Choose either local Faster-Whisper or an OpenAI-compatible transcription service. This setting is shared by the installation’s voice devices.
What you need
Use an owner/admin account and open Settings → Providers → Speech-to-Text. For local installation, use a Linux host with systemd user services. The single-container Docker setup does not support the local installer; use a separate transcription service instead.
Local Faster-Whisper
- Enable Speech-to-Text and select Local — Faster-Whisper.
- If it is not installed, choose CPU or GPU. The GPU profile requires an NVIDIA GPU and driver; use CPU for other hardware.
- Review the disk and memory estimates shown for that profile.
- Choose Save & install and wait for the progress panel to finish.
- Confirm the service status. On supported multi-GPU systems, STT GPU lets you select a GPU; changing it restarts the speech service.
The local service runs on the OE server and needs no provider API key. Its model is downloaded during installation. Existing remote credentials are kept when you change modes, so you can switch back later.
Remote or separately hosted API
- Select Remote API.
- Enter the service’s full transcription endpoint, ending in
/v1/audio/transcriptionsor the equivalent path supplied by the provider. - Enter its API key and supported model name, then choose Save.
- Enable the Speech-to-Text provider if it is disabled.
This mode also supports a compatible service on your LAN. The current remote path requires a nonempty key field; for a server that ignores authentication, use a placeholder such as local. For a server that checks keys, use its real credential. Use an address reachable from OE, not the browser’s localhost.
Remote STT sends the recording to that endpoint. Other voice stages have their own privacy settings; see Which model does what.
Test it
Open Voice devices → Voice diagnostics → Start device check. Choose a paired device, then say its assigned wake word followed by “What is two plus two?” Confirm that What OE heard matches your request. Response timing separates speech recognition from the agent and audio preparation.
Check this microphone checks the browser microphone locally; it does not call STT or test a voice satellite. See Voice diagnostics.
Common problems
- No recognized speech: check device connectivity and microphone capture in Voice diagnostics before changing models.
- Wrong or empty transcript: test in a quieter position and check when recording starts and ends. A wake-word detection alone does not prove the request was captured clearly.
- 401/403: check the endpoint’s key and access permissions.
- 429: inspect the provider’s quota or rate limit.
- Local install fails: read the install panel’s error; confirm systemd is available and, for GPU, the NVIDIA driver is installed.
- Slow response: compare the measured stages. CPU transcription speed varies with hardware; a slow chat model or TTS provider needs a separate fix.