
Now loading...
Grok has announced the launch of Grok Voice Transcribe 2.0, its latest model designed for converting speech to text. This upgrade boasts significant improvements, making it one of the most precise transcription models on the market today, achieving double the accuracy of its predecessor, Grok Voice Transcribe 1.0, while maintaining the same pricing structure.
The new transcription model is built on the advanced audio framework already in use with Grok Voice. This framework supports a variety of applications, processing tens of thousands of customer-support calls daily and transcribing millions of hours of video content. It has also been integrated into consumer products, such as the voice assistant feature in Tesla vehicles. Grok Voice Transcribe 2.0 has been trained on an extensive and diverse dataset, including real-world noisy and multilingual audio captured from a variety of environments, and has undergone rigorous post-training refinement.
Compared to other transcription solutions, Grok Voice Transcribe 2.0 excels in challenging audio conditions characterized by issues like unstable phone lines, overlapping voices, regional accents, and the reading of phone numbers or emails. This model is specifically designed to handle these complexities effectively.
On the public Artificial Analysis leaderboard, this latest version holds the top position for accuracy across 32 different streaming transcription models.
In addition to public rankings, Grok Voice Transcribe 2.0 has been evaluated against internal benchmarks drawn from live production data, which includes telephony audio from customer support interactions and multilingual voice commands. This model outperformed its predecessor across all tested scenarios, particularly excelling in telephony applications.
Enhancements in the model are especially marked in its multilingual capabilities, automatically detecting numerous languages and accommodating transitions between them without requiring separate processing. The word error rate in the handling of short phrases has improved dramatically, dropping from 20.6% to 6.8%.
Grok Voice Transcribe 2.0 offers advanced features that allow for easier integration with existing Speech-to-Text API setups without requiring any additional coding. These features encompass both batch and real-time streaming capabilities, precise word-level timestamps, speaker diarization, and support for up to eight independent channels. It also allows for the inclusion of specific terminology relevant to various domains and provides enhanced text formatting for recognizable data types.
Notably, Atlassian has found Grok Voice Transcribe 2.0 to deliver significantly higher accuracy than their previous solution for transcribing screen recordings made with Atlassian Loom. This accuracy can enhance workflows that use AI, enabling a seamless transition from context capture to action execution.
Grok Voice Transcribe 2.0 is priced identically to its predecessor, maintaining a cost of $0.10 per hour for batch transcriptions and $0.20 per hour for real-time streaming, with features like diarization and timestamps included at no extra charge.
The new transcription model will soon be the default option in the Speech-to-Text API, with Grok Voice Transcribe 1.0 scheduled to be phased out in the coming weeks. Users wishing to continue using the previous version can temporarily retain it by specifying grok-voice-transcribe-1.0.
