Turn-taking · end-of-turn detection

01
Smart Turn v3.2
Pipecat · Whisper-tiny
audio · prosody
94.0%
89.0%
1.0%One false cutoff in 100 incomplete clips.
32msRuns in parallel with STT — no transcript wait.49 p90
02
LiveKit Intl
text · transcript
87.0%
86.0%
12.0%
29msWaits for the STT transcript (~0–125ms) before running.57 p90
03
Turnsense
text · transcript
86.0%
88.0%
16.0%
127msPads every input to 256 tokens regardless of length — slowest as-shipped.193 p90
04
LiveKit EN
text · transcript
82.0%
87.0%
23.0%
5ms24 p90
05
VAD + silence timer
audio · energy
46.9%
100%
100%A silence timer ends every turn — the floor semantic detection has to beat.
~1ms
ModelAccuracy syntheticFalse-cutoff synthetic
Turnsensetext98.8%0.0%
LiveKit Intltext90.6%6.2%
LiveKit ENtext79.2%26.2%
Smart Turnaudio62.4%57.7%
Acoustic cueComplete ENDIncomplete WAITWhat it means
Final F0 slope-119.7 Hz/s+42.9 Hz/scomplete falls, incomplete rises/holds
Final energy slope-14.9 dB/s-2.5 dB/scomplete tapers off