Model Collection
View the full Granite Speech collection on Hugging Face
Speech Demo
Try Granite Speech in action
WebGPU Demo
Run Granite Speech in your browser
Replicate
Deploy Granite Speech on Replicate
Overview
The Granite Speech 4.1 model family provides compact and efficient speech-language models for multilingual automatic speech recognition (ASR) and automatic speech translation (AST), supporting English, French, German, Spanish, Portuguese, and Japanese. All models are trained on 174,000 hours of audio from public corpora and tailored synthetic datasets.Model Variants
The Granite Speech 4.1 suite includes three specialized variants:- granite-speech-4.1-2b: Balanced ASR and AST capabilities with improved punctuation and capitalization across all supported languages
- granite-speech-4.1-2b-plus: Speech-to-text model with speaker-attributed ASR, timestamps, and keyword-prompted ASR for enhanced recognition of names, acronyms, and technical jargon
- granite-speech-4.1-2b-nar: Non-autoregressive variant (NLE architecture) optimized for fast and accurate ASR with significantly lower latency