Skip to main content
The speech-to-text model turns spoken audio into transcript text. Convert on the Middleware API uses it for audio files and for the soundtrack of video. The transcript then follows the same split, shape, and extract path as any other source.

Parameters

The production ASR model is language-model speech recognition. It is multilingual. Language can be pinned per ingest.

Configuration