Unverified. We have not found corroboration in our collected sources.未確認の情報です。収集した情報では裏付けが取れていません。
AI·THE EXPLAINERニュースを理解する
SpaceXAI reports higher accuracy for Grok Voice Transcribe 2.0SpaceXAI、音声認識モデルGrok Voice Transcribe 2を発表
SpaceXAI released Grok Voice Transcribe 2.0, a speech-to-text model it says is twice as accurate as version 1.0 at the same price.SpaceXAIが音声認識モデルGrok Voice Transcribe 2.0をリリースしたと報じられました。精度が従来の2倍で価格は据え置き、という評価は同社によるものです。
THE SOURCE TRAIL情報の根拠
Based on linked publications. No first-party document or classified news organization is linked in this record.掲載元の記事をもとにしています。この記録には当事者・公的資料や、分類済みの報道機関へのリンクがありません。
Inspect the sources ↓出典を確かめる ↓Editorial correction · 2026-09-22編集上の訂正 · 2026-09-22
Attributed the transcription accuracy comparison to the company.音声認識精度の比較は同社の評価であることを明記しました。
What happened何が起きたか
SpaceXAI released Grok Voice Transcribe 2.0, a speech-to-text model built on the audio foundation behind Grok Voice, which handles tens of thousands of support calls daily, transcribes millions of hours of video narration, and powers the Grok assistant in Tesla vehicles. SpaceXAI says it is twice as accurate as version 1.0 at the same price and ranks first among 32 streaming models on the public Artificial Analysis accuracy leaderboard. Internal tests showed word error rate on a multilingual short-phrase set fell from 20.6% to 6.8%, with the model leading on telephony audio. It supports dozens of languages with automatic detection, mid-recording language switching, word-level timestamps, confidence scores, free diarization, up to eight separate channels, and biasing for up to 100 specialized terms.SpaceXAIは音声認識モデル「Grok Voice Transcribe 2.0」を公開した。これはGrok Voiceの音声基盤を活用しており、同基盤は1日に数万件のサポート通話を処理し、数百万時間分の動画ナレーションを文字起こしし、Tesla車両内のGrokアシスタントも支えている。SpaceXAIは価格据え置きでバージョン1.0の2倍の精度を実現したとし、公開の精度リーダーボードArtificial Analysisでストリーミングモデル32種中1位になったとしている。社内テストでは多言語の短いフレーズにおける単語誤り率が20.6%から6.8%に低下し、電話音声のテストでも他モデルを上回った。数十言語に対応し、自動言語検出、録音途中での言語切り替え、単語ごとのタイムスタンプと信頼度スコア、無料の話者分離、最大8チャンネルの個別文字起こし、最大100の専門用語のバイアス設定に対応する。
Why it matters何が変わるのか
The upgrade targets common transcription failure points like unreliable phone lines, competing voices, accents, and spoken credentials, which affects businesses using Grok Voice for support calls and voice agents. Existing API integrations get the update automatically without code changes.雑音の多い電話回線、複数話者の重なり、訛り、口頭で伝えられる認証情報など、文字起こしが失敗しやすい場面に対応しており、Grok Voiceをサポート通話や音声エージェントに使う企業に影響する。既存のAPI連携はコード変更なしで自動的にアップデートされる。
What’s nextこれからの動き
Version 2.0 will soon become the default, with version 1.0 due for deprecation in the coming weeks; teams can temporarily pin grok-voice-transcribe-1.0. Pricing stays at $0.10 per hour for batch transcription and $0.20 per hour for streaming.バージョン2.0は近くデフォルトとなり、バージョン1.0は数週間以内に廃止予定。チームは一時的にgrok-voice-transcribe-1.0を固定して使い続けることができる。価格はバッチ文字起こしが1時間0.10ドル、ストリーミングが1時間0.20ドルのまま据え置かれる。
A little context · terms in this story理解の手がかり · この記事の用語
- AI agent
- Software that uses an AI model and tools to carry out steps toward a task. How independently it can act depends on the product and the permissions you give it.AIモデルとツールを使い、目的に向けて複数の手順を実行するソフトウェア。自律的に動ける範囲は製品や与える権限によって異なります。
- API
- An interface that lets one piece of software use another service. An AI API lets a developer send requests to a model from their own application.ソフトウェアから別のサービスを利用するための窓口。AIのAPIなら、自分のアプリからモデルに処理を依頼できます。
Read the sources原典を読む
These are the source links saved with this story. Older records do not identify which texts were used in the summary.この記事に保存されている出典です。過去の記事には、どの本文を要約に使用したかの記録がありません。
More in AIAIのほかの記事
SpaceXAI releases Grok 4.7 for coding and knowledge workSpaceXAI、コーディングと知識労働向けにGrok 4.7を公開
3 linked sites関連3サイト
Up to 10 references参照画像は最大10枚
Generate + edit生成と編集を統合
Native RGBA透明背景に対応
SIGNAL explanatory diagram · Not a product screenshotSIGNAL解説図 · 実際の製品画面ではありません
Alibaba releases open-weight Qwen-Image-2.1 image modelアリババ、画像生成・編集モデルQwen-Image-2.1を公開
4 linked sites関連4サイト
Z.ai open-sources ZCode after covert code-upload backlashZ.aiのZCode、無断アップロード発覚で公開
7 linked sites関連7サイト