Models
Speech-to-text
Transcription of audio and video into convert
The speech-to-text model turns spoken audio into transcript text. Convert on the Middleware API uses it for audio files and for the soundtrack of video. The transcript then follows the same split, shape, and extract path as any other source.