Skip to main content
Home Skills Speech Recognition

Speech Recognition

Audio Processing

Transcribe speech to text, voice recognition

speech recognitionspeech to texttranscribeASRvoice recognitiondictationaudio transcriptionwhisper
0
Curated Registries
11
Available Models

Available Models & Agents

Browse 11 models and agents with speech recognition capabilities

Search all

ibm-granite/granite-speech-4.1-2b

ibm-granite

Granite Speech 4.1 2B is a compact and efficient speech-language model, specifically designed for multilingual automatic speech recognition (ASR) and bidirectional automatic speech translation (AST) for English, French, German, Spanish, Portuguese and Jap

model

twangodev/qwenasr

twangodev

Serve QwenASR speech recognition and alignment.

model

awerks/whisperx

awerks

Fast automatic speech recognition (70x realtime with large-v2) with word-level timestamps and speaker diarization.

model

nateraw/whisper-large-v3

nateraw

Whisper is a general-purpose speech recognition model.

model

erium/whisperx

erium

Automatic Speech Recognition with Word-level Timestamps & Diarization

model

ibm-granite/granite-speech-3.3-8b

ibm-granite

Granite-speech-3.3-8b is a compact and efficient speech-language model, specifically designed for automatic speech recognition (ASR) and automatic speech translation (AST).

model

subformer/meta-omnilingual-asr-7b

subformer

Omnilingual ASR 7B by Meta (Unofficial) - Automatic speech recognition supporting 1,693 languages with best-in-class accuracy. Meta's recommended variant for mission-critical transcription requiring maximum quality.

model

@cf/openai/whisper

@cf

Whisper is a general-purpose speech recognition model. It is trained on a large dataset of diverse audio and is also a multitasking model that can perform multilingual speech recognition, speech translation, and language identification.

model

@cf/deepgram/flux

@cf

Flux is the first conversational speech recognition model built specifically for voice agents.

model

@cf/openai/whisper-tiny-en

@cf

Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data, Whisper models demonstrate a strong ability to generalize to many datasets and domains without the need for fine-tuning. This is the English-only version of the Whisper Tiny model which was trained on the task of speech recognition.

model

@cf/openai/whisper-large-v3-turbo

@cf

Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation.

model