SpaceXAI has released Grok Voice Transcribe 2.0, a speech-to-text model it says is twice as accurate as version 1.0 at the same price. It uses the audio foundation behind Grok Voice, which handles tens of thousands of support calls each day, transcribes millions of hours of video narration, and powers the Grok assistant in Tesla vehicles. Training draws on noisy, multilingual audio recorded in varied environments and refined through post-training.
The model targets unreliable phone lines, competing voices, local accents, and spoken phone numbers or email addresses. SpaceXAI says it ranks first among 32 streaming models on the public Artificial Analysis accuracy leaderboard. Internal production-derived tests also showed gains on English support calls, Grok conversations, spoken credentials, and short commands in 19 languages. Word error rate on the multilingual short-phrase set fell from 20.6% to 6.8%, while the model led every tested system on telephony audio.
Introducing Grok Voice Transcribe 2.0. It’s the world’s most accurate speech transcription model. pic.twitter.com/7iXVmAEFW2
— SpaceXAI (@SpaceXAI) September 18, 2026
Grok Voice Transcribe 2.0 supports dozens of languages, detects them automatically, and follows language switches within one recording. Its API handles batch jobs, files, URLs, and real-time streams. It offers word-level timestamps and confidence scores, diarization at no extra cost, separate transcription for up to eight channels, and biasing for up to 100 specialized terms. It can also format numbers, dates, currencies, phone numbers, and email addresses, remove filler words, and detect turns for voice agents.
Existing API integrations receive the update without code changes. The model will soon become the default, with version 1.0 due for deprecation in the coming weeks. Teams can temporarily pin grok-voice-transcribe-1.0. Pricing remains $0.10 per hour for batch transcription and $0.20 per hour for streaming, including diarization, timestamps, and key terms.