Gemini 3.5 Transcribe Live
Measured 2026-07-03
* Conditional
Vendor-direct only -- not wired to the Speko gateway, so this row measures the model and not the Speko product. Finalize and first token are measured on the pinned us-east4 vantage (n=30).
WER · batch
—
not measured
WER · stream
—
not measured
WER · conversational
13.1%
The batch sibling `gemini-3.5-transcribe` is listed in the Gemini API and returns ZERO output tokens on every request shape tried, while gemini-3.5-flash transcribes the same bytes in the same body — so only the Live model is measurable. 4.9% substitutions puts its hearing third on the board; it loses on 6.9% deletions. Announced 2026-08-26.
Finalize
1.22s
p50–p90 · us-east4 · n=30 · vendor-direct
Time to first token
1.88s
p50–p90 · us-east4 · n=30 · vendor-direct
Cost / min
~$0.009
ESTIMATE, not a published per-minute rate. Google prices this model per token -- $3.50/1M audio in, $21.00/1M text out -- and the ~$0.005/min + ~$0.004/min figures on the pricing page are Google's own estimate at 25 audio tokens/sec input and 175 text tokens/min output, for a blended ~$0.009/min. Speech density moves it, so read it as an order of magnitude rather than a rate card. Every other price on this column is the vendor's published per-minute rate.
Share this result
Embed the live badge
[](https://benchmarks.speko.ai/stt/google-gemini-3-5-transcribe-live)
<a href="https://benchmarks.speko.ai/stt/google-gemini-3-5-transcribe-live"><img src="https://benchmarks.speko.ai/badge/stt/google-gemini-3-5-transcribe-live.svg" alt="Speko stt rank"></a>
https://benchmarks.speko.ai/stt/google-gemini-3-5-transcribe-live