Skip to main content
The speech-to-speech model is a realtime duplex audio model. A member speaks, the model replies in speech in the same session, and it can call the same workspace tools as answers. It is the call path in Observatory, not file transcription. File audio uses speech-to-text. Playback of a written reply uses text-to-speech.

Parameters

The production realtime model accepts speech, text, and images, and returns speech and text.

Configuration