← Speech-to-Text

Google

Gemini 3.5 Transcribe Live

Measured 2026-07-03

* Conditional

Vendor-direct only -- not wired to the Speko gateway, so this row measures the model and not the Speko product. Finalize and first token are measured on the pinned us-east4 vantage (n=30).

WER · batch

not measured

WER · stream

not measured

WER · conversational

13.1%

The batch sibling `gemini-3.5-transcribe` is listed in the Gemini API and returns ZERO output tokens on every request shape tried, while gemini-3.5-flash transcribes the same bytes in the same body — so only the Live model is measurable. 4.9% substitutions puts its hearing third on the board; it loses on 6.9% deletions. Announced 2026-08-26.

Finalize

1.22s

p50–p90 · us-east4 · n=30 · vendor-direct

Time to first token

1.88s

p50–p90 · us-east4 · n=30 · vendor-direct

Cost / min

~$0.009

ESTIMATE, not a published per-minute rate. Google prices this model per token -- $3.50/1M audio in, $21.00/1M text out -- and the ~$0.005/min + ~$0.004/min figures on the pricing page are Google's own estimate at 25 audio tokens/sec input and 175 text tokens/min output, for a blended ~$0.009/min. Speech density moves it, so read it as an order of magnitude rather than a rate card. Every other price on this column is the vendor's published per-minute rate.

Share this result

Embed the live badge

Speko stt rank
Markdown
[![Speko stt rank](https://benchmarks.speko.ai/badge/stt/google-gemini-3-5-transcribe-live.svg)](https://benchmarks.speko.ai/stt/google-gemini-3-5-transcribe-live)
HTML
<a href="https://benchmarks.speko.ai/stt/google-gemini-3-5-transcribe-live"><img src="https://benchmarks.speko.ai/badge/stt/google-gemini-3-5-transcribe-live.svg" alt="Speko stt rank"></a>
URL
https://benchmarks.speko.ai/stt/google-gemini-3-5-transcribe-live