# Speko Benchmarks > Independent, reproducible voice-AI benchmarks: STT, TTS, LLM, S2S, turn-taking, and true cost-per-solve across the whole call. Published numbers are measured, never marketed. Latest run: 2026-07-03 (run 07-03) - 110 models · 42 integrations ## Boards - [Speech-to-Text](https://benchmarks.speko.ai/stt): Word error rate, latency and cost for every speech-to-text model we measure. (measured 2026-07-03) - [Code-switching](https://benchmarks.speko.ai/stt-codeswitch): Two languages in one utterance — error rate and whether the second language survives, per pair, over the live stream. - [Text-to-Speech](https://benchmarks.speko.ai/tts): Naturalness, drift, latency, robustness and cost for every text-to-speech model we measure. (measured 2026-07-03) - [LLM](https://benchmarks.speko.ai/llm): First-token latency, task completion, fabrication and cost for voice-agent LLMs. (measured 2026-08-25) - [S2S](https://benchmarks.speko.ai/s2s): Native voice-to-voice, measured across six concierge scenarios. (measured 2026-07-19) - [True cost-per-solve · whole stack](https://benchmarks.speko.ai/stacks): Cost per grounded solve across a full STT+LLM+TTS stack. (measured 2026-07-03) - [Turn-taking · end-of-turn detection](https://benchmarks.speko.ai/turntaking): End-of-turn detection, ranked on 200 real human clips. - [Open models](https://benchmarks.speko.ai/open): Every open-weights voice model in one place: measured on our boards where we have rows, credible sourced rankings where we do not — with... ## Model pages Every system has a permalink at https://benchmarks.speko.ai// with its metrics, caveats, a rank where the board ranks on a single axis, and an embeddable SVG badge at https://benchmarks.speko.ai/badge//.svg. Examples: - https://benchmarks.speko.ai/stt-codeswitch/soniox-stt-rt-v5 - https://benchmarks.speko.ai/stacks/cheapest - https://benchmarks.speko.ai/turntaking/smart-turn-v3-2-pipecat-whisper-tiny ## Data - Any board as markdown: append `.md` to its URL, e.g. https://benchmarks.speko.ai/tts.md - [Full dataset (JSON)](https://benchmarks.speko.ai/data.json): the exact object the boards render from - [Full corpus (llms-full.txt)](https://benchmarks.speko.ai/llms-full.txt): every blog post plus the current board data in one plain-text file for LLM context - [Methodology](https://benchmarks.speko.ai/#methodology) ## Speko platform - [Speko](https://speko.ai): the voice AI gateway these benchmarks feed - route STT/LLM/TTS per call through one API - [Docs](https://docs.speko.dev) ## Blog - [RSS feed](https://benchmarks.speko.ai/rss.xml) - [Gradium's New Default Reaches the Top Group](https://benchmarks.speko.ai/blog/the-alias-moved) - [$4 and 103 Milliseconds](https://benchmarks.speko.ai/blog/four-dollars-and-103ms) - [Eleven v3 Can Take Calls Now](https://benchmarks.speko.ai/blog/expressive-can-take-calls) - [The Open-Weights Model That Doesn't Go Quiet](https://benchmarks.speko.ai/blog/the-silence-after-hello) - [Turn-Taking for Voice Agents: What Actually Works in 2026](https://benchmarks.speko.ai/blog/turn-taking-for-voice-agents) - [Deepgram Flux TTS Is the Fastest Voice in Our Naturalness Band](https://benchmarks.speko.ai/blog/deepgram-flux-tts) - [It Kept Answering. It Just Stopped Speaking.](https://benchmarks.speko.ai/blog/it-stopped-speaking) - [Your Refund Sounds Exactly Like a Charge](https://benchmarks.speko.ai/blog/the-minus-sign) - [Code-Switching Speech Recognition: Which STT Models Survive Two Languages in One Sentence](https://benchmarks.speko.ai/blog/code-switching-stt-benchmark) - [How Speko Benchmarks STT](https://benchmarks.speko.ai/blog/how-speko-benchmarks-stt) - [Voice Agent Turn-Taking, Measured: Barge-In, Backchannels and End-of-Turn Detection](https://benchmarks.speko.ai/blog/voice-agent-turn-taking-benchmark) - [Our Tamil Column Read 33% Error. The Transcripts Were Fine.](https://benchmarks.speko.ai/blog/wer-vs-cer) - [Bland's New Voice Is Among the Most Natural English We Have Measured](https://benchmarks.speko.ai/blog/bland-joins-the-top-group) - [Mini Is a Price Tier. It Is Not a Latency Tier.](https://benchmarks.speko.ai/blog/mini-is-not-a-latency-tier) - [We Ran the Blind Test. 'Most Natural' Is an Eight-Way Tie.](https://benchmarks.speko.ai/blog/most-natural-is-a-tie) - [The truth about voice AI benchmarks (and how to run your own)](https://benchmarks.speko.ai/blog/truth-about-voice-ai-benchmarks) - [What a Voice Agent Actually Hears](https://benchmarks.speko.ai/blog/what-a-voice-agent-hears) - [Your System Prompt Has a Budget](https://benchmarks.speko.ai/blog/guardrail-budget) - [Both Models Scored 100%. Then We Stress-Tested Them.](https://benchmarks.speko.ai/blog/both-scored-100-stress-test) - [The .99 Problem: The Cheapest Streaming TTS Drops Your Cents](https://benchmarks.speko.ai/blog/the-99-problem) - [gpt-realtime-2.1-mini Gets You Most of the Flagship for a Third of the Price. The Catch Is Latency.](https://benchmarks.speko.ai/blog/gpt-realtime-2-1-mini-best-value) - [gpt-realtime Is the Fastest S2S Model. It Took a Full Conversation to See It.](https://benchmarks.speko.ai/blog/gpt-realtime-2-1-mini-fastest) - [Pulse: Fast to the First Word, Fast to the Last One](https://benchmarks.speko.ai/blog/pulse-stt) - [Speechify's pitch tag swaps the speaker](https://benchmarks.speko.ai/blog/the-pitch-trap) - [The Pause Is the Hard Part: How We Stopped Cutting Callers Off](https://benchmarks.speko.ai/blog/how-speko-uses-smart-turn) - [LiveKit Inference for Voice Agents: Fast First Token, and the Goodbye Most Models Skip](https://benchmarks.speko.ai/blog/livekit-inference-for-voice-agents) - [Everyone Measures the Clip. Nobody Measures the Call.](https://benchmarks.speko.ai/blog/true-cost-per-solve) - [Your Voice Agent Booked the Wrong Name. WER Said It Was 95% Accurate.](https://benchmarks.speko.ai/blog/stt-tool-call-corruption) - [Fast or Natural? Cartesia Sonic-3.5 Refuses to Pick.](https://benchmarks.speko.ai/blog/cartesia-fast-and-natural) - [Your Voice Agent's LLM Speaks Spanish. That Doesn't Mean It Follows the Rules in Spanish.](https://benchmarks.speko.ai/blog/llm-multilingual-reliability) - [How We Benchmark TTS: Gate, Profile, Rank](https://benchmarks.speko.ai/blog/benchmarking-tts-end-to-end) - [Ranking TTS Nativeness with Content-Matched FAD](https://benchmarks.speko.ai/blog/ranking-tts-nativeness) - [How Speko Benchmarks TTS: A Gate, Then a Profile](https://benchmarks.speko.ai/blog/how-speko-benchmarks-tts) - [Semantic Was Supposed to Be the Smart One. It Lost.](https://benchmarks.speko.ai/blog/the-backchannel-trap) - [How Anglicized Is Your TTS? Measuring Phonological Authenticity Across 7 Providers](https://benchmarks.speko.ai/blog/anglicization-pattern) - [The Cartesia Drift: 10% of Voices Hold the Line for a Minute. The Other 90% Don't.](https://benchmarks.speko.ai/blog/cartesia-drift) - [Artificial Analysis Ranks Gemini 3.1 Flash TTS #2. We Asked It for Ten Minutes.](https://benchmarks.speko.ai/blog/10-minute-tts-benchmark) - [Vendors Say They Support 99 Languages. They Don't.](https://benchmarks.speko.ai/blog/the-asterisk-on-supported-languages) - [Speech-to-Speech Got Smart. It Still Can't Replace the Cascade.](https://benchmarks.speko.ai/blog/can-s2s-replace-the-cascade) - [We Tried to Break Four Voice Agents with a Cough. We Failed.](https://benchmarks.speko.ai/blog/designing-barge-in)